Make agents use tools and reusable skills more reliably, Make RAG retrieval and knowledge-base QA more reliable, Improve model reasoning, planning, and verification
Today tracks: make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 305 candidate papers from the 2026-07-21 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents🔗
- 2Point Ladder Tuning: Parameter-Efficient Hierarchical Adaptation for 3D Point Cloud Understanding🔗
- 3Cross-Agent Campaign Attribution: Linking Asynchronous Attacks Across LLM Agents🔗
- 4CASE: Causal Alignment and Structural Enforcement for Improving Chain-of-Thought Faithfulness🔗
- 5Wave2Body: Rethinking mmWave Human Pose Estimation as Radar-to-Body Token Translation🔗
- 6The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, deployment, benchmark, code to frame the code intelligence task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer
Keywordsagentdeploymentbenchmarkcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, code, fine-tuning, memory to frame the training and post-training task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Fine-tuning pre-trained point-cloud backbones typically updates all parameters, resulting in substantial computation and memory overhead
Keywordsragcodefine-tuningmemory
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, deployment, evaluation to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: LLM-agent defenses are typically evaluated one session at a time
Keywordsagentragdeploymentevaluation
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around rag, inference, alignment, benchmark to frame the reasoning and planning task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Chain-of-thought (CoT) reasoning is widely used to improve both the performance and interpretability of large language models (LLMs), yet the generated reasoning may not faithfully support the final answer
Keywordsraginferencealignmentbenchmark
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, alignment, code to frame the data engineering task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Millimeter-wave (mmWave) radar enables privacy-friendly human sensing, but its sparse point clouds are physical measurements of view-dependent electromagnetic reflections and only indirectly characterize body articulation
Keywordsraginferencealignmentcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around workflow, retrieval, deployment, safety to frame the safety and alignment task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios
Keywordsworkflowretrievaldeploymentsafety
Code/DataCheck the source paper
Other papers worth tracking
Cognitive Dual-Process Planning for Autonomous Driving with Structured Scene Knowledge and Verifiable Reasoning-Action Consistency: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
Delineate Anything v2: A Global Foundation Model for Field Delineation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Beyond Noisy Signals: Dual-Level Denoising for Multi-modal Sequential Recommendation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
Inference-Time Steering for Cross-Lingual Factual Consistency in LLMs: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing: Covers a concrete vision and image generation signal; useful as a follow-up candidate.
CoGoal3D: Collaborative 3D Object Detection with 3D-Aware Fusion and Refinement: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Measuring Reward-Seeking via Contrastive Belief Updates: Covers a concrete training and post-training signal; useful as a follow-up candidate.
From a Multilingual Streaming ASR Backbone to Kenyan-Language Systems: Data-Centric Adaptation of Nemotron 3.5 for Kikuyu, Dholuo, and Kalenjin: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Spaghetti Architect: A Contamination-Resistant, By-Construction-Labelled, Multi-Language Code Dataset Generator: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
BRIDGE: Bottleneck-Aware Regulator-Set Inference and Diagnosis for Cooperative Gene Regulatory Recovery: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
OmniReasoner: Thinking with Long Audio-Video via Native Tool Use: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Agents in the Wild: Where Research Meets Deployment: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
CircuitKIT : Circuit Discovery, Evaluation, and Application Toolkit for Mechanistic Interpretability: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.