Make agents use tools and reusable skills more reliably, Make RAG retrieval and knowledge-base QA more reliable
Today tracks: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable.
This issue fetched and deduplicated 379 candidate papers from the 2026-09-18 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1An Interpretable Memory Decision Controller for LLM Agents Based on Three-Signal Complementarity: Decoupling Confidence and Consistency🔗
- 2TrialAtlas: Multi-Agent Research Organization for Clinical Trial Design and Optimization🔗
- 3MIST: Multimodal Survival Prediction with Genomic-Guided Histology Attention🔗
- 4CIBuzzBench: A Benchmark for Cross-Lingual Understanding of Chinese Internet Buzzwords🔗
- 5Scaling Vision-Language Reward Learning for Robot Manipulation in Parallel Simulation🔗
- 6TERMon: Detecting Persistent Behavioral Threats in Edge AI via Hardware-Native Ternary Runtime Monitor🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, retrieval, evaluation to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Memory systems for large language models have focused predominantly on efficient retrieval, whereas the decision of whether retrieved memories should be trusted has received comparatively little attention
Keywordsagentragretrievalevaluation
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, evaluation, memory, agents to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Nearly 90% of drugs entering clinical development ultimately fail, despite billions of dollars in investment
Keywordsagentevaluationmemoryagents
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around alignment, evaluation, code, multimodal to frame the multimodal models task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Multimodal survival models can combine complementary prognostic information from whole-slide images and genomic profiles, but effective fusion remains challenging amid external cohort shift and computational complexity
Keywordsalignmentevaluationcodemultimodal
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, safety, evaluation, benchmark to frame the benchmarks and evaluation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Chinese social media has generated a vast and continually evolving lexicon of internet buzzwords whose meanings are often non-literal and deeply rooted in local cultural and pragmatic contexts
Keywordsragsafetyevaluationbenchmark
Code/DataCheck the source paper
strengthen multimodal understanding of charts, documents, and visual evidence
Signalthis paper targets the concrete research problem behind strengthen multimodal understanding of charts, documents, and visual evidence. It uses the title, abstract, and public signals around latency, evaluation, code, vision-language to frame the training and post-training task, data, or evaluation flow to improve strengthen multimodal understanding of charts, documents, and visual evidence. The main claim is the title, abstract, and public signals indicate: Vision-language models (VLMs) can replace human annotators in preference-based reward learning, but sequential API requests and single-environment data collection make training slow and costly
Keywordslatencyevaluationcodevision-language
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around inference, serving, latency, safety to frame the systems and deployment task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Edge AI accelerators are increasingly deployed in safety-critical environments, where model outputs may control physical actuators, make access-control decisions, or trigger alarms
Keywordsinferenceservinglatencysafety
Code/DataCheck the source paper
Other papers worth tracking
Risk-Aware Occupancy for Safety-Oriented End-to-End Autonomous Driving: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
SIRA: Reasoning-Aware Surgical Instrument Segmentation via Query-Anchored Alignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Robust Structureless Monocular Visual Inertial Initialization Exploiting Line Features and Vanishing Points: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Learning to Move Cities: Deep Meta-Models and Reinforcement Policies for Calibration and Control in Urban Networks: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
VIRGA: Virtual-Agent-Intermediated Riemannian Geometry for Active-Sensing Air-Ground Coordination: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Samsone: A Family of Open Small Audio Language Models for On-Device Inference: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
Driving on Registers, Reasoning on Risk: Risk-Aware Occupancy for Register-Based End-to-End Autonomous Driving: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
PSEE: Progressive Sensor Event Expansion for Point-Supervised Temporal Action Localization: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Hallucination-R1: Robustness-Oriented Paraphrase Generation for Factual Consistency: Covers a concrete training and post-training signal; useful as a follow-up candidate.
AirSplan: Risk-Aware Motion Planning for Quadrotors in Cluttered 3D Gaussian Splats: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
BrainWideBench: Benchmarking large-scale pretraining and across-animal transfer in multi-region neural recordings: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Available Guardrails: Certifying Selective Prediction across ML Systems: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
DiaVLo: Diagnosing Behaviours of Vision-Language Models: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Info3R: Information-Adaptive Test-Time Training for 3D Reconstruction: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Catena: A Comprehensive Software Suite for Large-Scale Connectomics: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Benchmarking the Explanatory Quality of Open-Weight Vision-Language Models in Face Recognition: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.