Make agents use tools and reusable skills more reliably, Make RAG retrieval and knowledge-base QA more reliable, Improve code generation, execution feedback, and automated repair
Today tracks: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable.
This issue fetched and deduplicated 281 candidate papers from the 2026-08-14 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1CoDS: Robust Collaborative Perception via Expert-driven Detection and BEV Segmentation🔗
- 2Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead🔗
- 3Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons🔗
- 4CORAL: Curriculum-Optimized Reward Adaptation for LiDAR-Based Goal-Directed Urban Driving🔗
- 5Envs-FORGE: Frontier-Optimized Reward-Grounded Environment Synthesis for Agent RL🔗
- 6HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, code, segmentation, detection to frame the vision and image generation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Collaborative perception breaks through single-view limitations via multi-agent information exchange
Keywordsagentcodesegmentationdetection
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, deployment, evaluation, benchmark to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Nanbeige4.2-3B is a 3B-parameter agentic model built around a Looped Transformer (LT) that reuses one stack of layers for a second forward pass, adding effective depth without additional parameters
Keywordsagentdeploymentevaluationbenchmark
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, deployment, alignment to frame the safety and alignment task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Neuron- and path-level interventions offer the finest-grained route to defending large language models (LLMs) against jailbreak attacks, yet existing methods fall short of this promise, i.e., they often compromise model utility significantly
Keywordsraginferencedeploymentalignment
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around safety, evaluation, code, navigation to frame the robotics and embodied ai task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Reinforcement learning is promising for autonomous urban driving, but long-horizon goal-directed navigation asks a policy to acquire several competing behaviors at once--reaching a distant goal, tracking a route, avoiding obstacles, obeying signals--and a fixe
Keywordssafetyevaluationcodenavigation
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, code, agents to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Reinforcement learning (RL) for terminal agents needs executable training environments with reliable rewards and useful difficulty
Keywordsagentragcodeagents
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, retrieval, serving, alignment to frame the multimodal models task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Existing multimodal RAG methods often flatten structured documents into isolated text and image units, weakening the source organization and local text-image logic needed for faithful evidence selection and placement
Keywordsragretrievalservingalignment
Code/DataCheck the source paper
Other papers worth tracking
Accelerating Large-scale Bundle Adjustment for LiDAR Mapping via Parallel Computing: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Content Based Video Narration of Gameplay with Vision Language Models: Covers a concrete video generation signal; useful as a follow-up candidate.
CMCNet: Aligning Ultrasound Image Embeddings with Textual TI-RADS Representations for Fine-Grained Thyroid Classification: Covers a concrete multimodal models signal; useful as a follow-up candidate.
Leading-Silence Augmentation and Multi-Stage Synthetic Supervision for the Second MLC-SLM Challenge: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
PILOT: Privileged Imitation Learning for End-to-End Motion Planning of Autonomous UAVs under Partial Observability: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing: Covers a concrete training and post-training signal; useful as a follow-up candidate.
CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
AgentRewind: Recoverable Execution for Long-Horizon LLM Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
A Survey of Large Models in Sports: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
APTER: Adaptive Post-Training with Expert-Grounded Rubrics: Covers a concrete training and post-training signal; useful as a follow-up candidate.
MINT: A Universal Zero-Shot Predictor for Transaction Data: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Can Language Models Understand mmWave Data? Benchmarking Large Language Models for mmWave Radar-Based Human Understanding: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept Relations: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
A Graph-Based Reinforcement Learning Framework for Structured Drift Diagnosis and Recovery in Autonomous LLM Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
P2Skill: Privacy Preserving Skill Distillation for Cloud-Local LLM Inference Systems: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
Regime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety Classifiers: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.