Make agents use tools and reusable skills more reliably, Make RAG retrieval and knowledge-base QA more reliable
Today tracks: make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable, make RAG retrieval and knowledge-base QA more reliable.
This issue fetched and deduplicated 405 candidate papers from the 2026-09-03 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Two Truths and A Lie? Benchmarking Off-the-Shelf LLMs for Requirements Quality Assessment: Performance, False Alarms, and Misses🔗
- 2Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning🔗
- 3Rethinking On-Policy Distillation of Large Language Models II: One Training Example🔗
- 4SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center🔗
- 5Hardware-Aware FP4 FlashAttention-4🔗
- 6A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable, make RAG retrieval and knowledge-base QA more reliable. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, evaluation, benchmark, eval to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Requirements engineering (RE) governs the quality of everything downstream in systems engineering (SE); defective requirements that survive review cycles propagate into design rework, schedule delays, and cost overruns
Keywordsagentevaluationbenchmarkeval
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, alignment, vision-language, video to frame the video generation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos given only an ordered set of event-level captions per video
Keywordsragalignmentvision-languagevideo
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, alignment, post-training, training to frame the training and post-training task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher
Keywordsragalignmentpost-trainingtraining
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, deployment, code, data to frame the data engineering task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Large language model (LLM) agents are increasingly proposed as autonomous SOC analysts, but two limitations make them unreliable at enterprise scale: a finite context window cannot hold a multi-thousand-host authentication graph, and free-form generation offer
Keywordsagentdeploymentcodedata
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around inference, throughput, quantization, systems to frame the systems and deployment task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Blackwell's 4-bit floating-point (FP4) tensor cores do not automatically make attention faster because softmax conversion and on-chip dependencies dominate once its matrix products shrink
Keywordsinferencethroughputquantizationsystems
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, open-source, tool, agents to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Modern AI agent harnesses expose lifecycle hooks that bind shell commands to runtime events such as session start, tool calls, and file edits
Keywordsagentopen-sourcetoolagents
Code/DataCheck the source paper
Other papers worth tracking
The Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMs: Covers a concrete video generation signal; useful as a follow-up candidate.
Compressing Streaming Neural Audio Encoders via Latent-Space Distillation: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Subspace Inference Enables Efficient Active Reward Learning from Preferences: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
InSituMeasure: Probing Situated Measurement Grounding in Industrial Scenes with Multimodal Large Language Models: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Fairness Evaluation of Edge-AI Implementation for Cleft Lip and Palate Speech ASR: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
RARF: Region-Aware Rectified Flows for 3D Brain MRI Inpainting: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
RealCADBench: Benchmarking Parametric CAD Modeling from Industrial Design Intents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
MINERVA: How Small Can a Manipulation Policy Be and Still Solve LIBERO?: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Counterfactual Routing Using Integer Programming with Constraint Generation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
SignSeek: Learning Transferable Representations for Sign Dictionary Retrieval: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking Preservation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Lost in Reordering: Structural Sensitivity of Multilingual LLMs under Semantics-Preserving Perturbations: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Towards a Statistical Understanding of Mixture-of-Experts: Covers a concrete video generation signal; useful as a follow-up candidate.
BRIDGE: An Open-Source Humanoid Platform via Morphology-Control Co-Design for Physical AI: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
When Retrieval Helps: Selective Retrieval for Single-Turn Mental-Health QA: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.