Make RAG retrieval and knowledge-base QA more reliable, Make agents use tools and reusable skills more reliably
Today tracks: make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 431 candidate papers from the 2026-09-16 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1A Zeroth-Order Paradigm for LLM Preference Alignment🔗
- 2PULSE: Unlocking Practical Image Compression on Single-Thread CPU🔗
- 3Collective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery🔗
- 4WetRobo: A Reproducible Robot Kit for Coding Agents in Biological Laboratories🔗
- 5FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory🔗
- 6Risk-Aware World Modeling with Flow-Guided Occupancy Evolution for Selective Trajectory Planning in Automated Driving🔗
What is worth tracking today
Today’s high-signal papers point to: make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, alignment, fine-tuning, memory to frame the training and post-training task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency
Keywordsragalignmentfine-tuningmemory
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, latency, compression, code to frame the code intelligence task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Despite recent progress in learned image compression, existing methods remain computationally expensive on resource-constrained hardware, particularly CPUs
Keywordsagentlatencycompressioncode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, deployment, evaluation, benchmark to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: How does a multi-agent system evolve from a local deviation into collective loss of control?
Keywordsagentdeploymentevaluationbenchmark
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, code, vision-language, robotics to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Automating biological research requires general-purpose, reproducible robot systems that allow individual wet-lab researchers to delegate robot tasks without performing teleoperation or neural-network training
Keywordsagentcodevision-languagerobotics
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around benchmark, code, vision-language, memory to frame the multimodal models task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: State-of-the-art vision-language-action models (VLA) for autonomous driving face critical limitations: excessive parameter counts, inefficient high-resolution image processing, and lack of temporal memory
Keywordsbenchmarkcodevision-languagememory
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around rag, evaluation, temporal, motion to frame the video generation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Safe motion planning in automated driving requires anticipating evolving traffic risks and deciding when to revise the current planned trajectory
Keywordsragevaluationtemporalmotion
Code/DataCheck the source paper
Other papers worth tracking
CapMap-MS-TTA: 3rd Place Solution for the MUMU Track of the 8th LSVOS Challenge at ECCV 2026: Covers a concrete training and post-training signal; useful as a follow-up candidate.
OmniRisk: Omnidirectional Trajectory-Risk Learning for Agile Quadrotor Dynamic Avoidance: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Visual Autoregressive Priors for RAW-to-sRGB Image Signal Processing: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs: Covers a concrete data engineering signal; useful as a follow-up candidate.
WAVE-Go: World-Model Navigation with Adaptive Execution for Wheel-Legged Robots: Covers a concrete code intelligence signal; useful as a follow-up candidate.
Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
SURF: Subtractive Updates for Recommender Forgetting: Covers a concrete video generation signal; useful as a follow-up candidate.
RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reasoning through Evolution: Automatic Meta-path Discovery for LLM-based Fake News Detection: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Hardware-Free Robotics Laboratories in Mixed Reality: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Occluded Gait Recognition with Mixture of Experts: An Action Detection Perspective: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
HiLNO: A Hierarchical Latent Neural Operator with Multi-Scale Supervision for PDEs on General Geometries: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
Where Should Agents Live? Energy-Memory Characterization of Agentic AI for the Edge-Cloud Continuum: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
I code or AI code: A comparative evaluation of AI-rated scores in classroom observations: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
${M}^2$Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models: Covers a concrete multimodal models signal; useful as a follow-up candidate.
CPR: Combining global composing, local performing and full-sequence refining in piano rendering with continuous autoregressive modelling: Covers a concrete vision and image generation signal; useful as a follow-up candidate.
Colla-Q: Toward Collaborative Experts in MoE Quantization via Minimax Precision Balancing: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Designing Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Aligned Consensus Teaching for Label-Efficient Oriented Object Detection in Weakly-Aligned Visible-Infrared Imagery: Covers a concrete multimodal models signal; useful as a follow-up candidate.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.