Make agents use tools and reusable skills more reliably, Improve model reasoning, planning, and verification
Today tracks: make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 497 candidate papers from the 2026-09-24 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure🔗
- 2From Interests to Semantic IDs: Retrieval-Grounded Credit Assignment for Generative Recommendation🔗
- 3Multi-Agent Orchestration of 3GPP Channel Estimators🔗
- 4SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance🔗
- 5Return or Revise? Learning When Revision Helps Retrieval-Augmented QA🔗
- 6AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, safety, evaluation, benchmark to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: A central concern in AI safety is that agents may treat oversight as an obstacle when it conflicts with completing their goals
Keywordsagentsafetyevaluationbenchmark
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around retrieval, code, reasoning, search to frame the reasoning and planning task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Semantic IDs (SIDs) encode each catalog item as a short token sequence, enabling generative recommenders to predict the next item autoregressively
Keywordsretrievalcodereasoningsearch
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, deployment, latency, agents to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Pilot-aided channel estimation is a decisive block in orthogonal frequency-division multiplexing (OFDM) receivers for both 5G New Radio (5G-NR) and Long-Term Evolution (LTE)
Keywordsagentdeploymentlatencyagents
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around benchmark, code, reasoning, Reasoning and Planning to frame the reasoning and planning task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Long-horizon reasoning remains a central challenge for large language models (LLMs) under sparse-reward regimes
KeywordsbenchmarkcodereasoningReasoning and Planning
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, retrieval, evaluation, Retrieval and RAG to frame the retrieval and rag task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: We consider the decision of whether to return an existing draft answer or revise it using retrieved evidence, as in answer-revision systems
KeywordsragretrievalevaluationRetrieval and RAG
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around alignment, evaluation, code, multimodal to frame the training and post-training task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Recent years have witnessed major progress in joint audio-video generation
Keywordsalignmentevaluationcodemultimodal
Code/DataCheck the source paper
Other papers worth tracking
TTLab at AlexandriaX-2026: A Fine-Tuned Surface Tagger for Arabic Machine-Translation Error-Span Detection and Classification: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
SWE-Prometheus: Measuring Engineering Governance Improvements in Real-World Repositories: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Decoupled Early Exits for Task-Dependent Compute Allocation in Flow-Matching VLAs: Covers a concrete robotics and embodied ai signal; useful as a follow-up candidate.
UpDown-SC: Gravity-Canonicalized Dual-Envelope Scan Context for Indoor LiDAR Place Recognition: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
CrossSafe: Towards Cross-Embodiment Latent Safety Filters: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Learning Better Reasoning for Generative Recommendation with Semantic IDs: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
An Empirical Study of VLM Pipelines for Long-Document QA: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Anatomy-Aligned Surface Field Learning for Myocardial Reconstruction from Sparse Short-Axis Cine MRI: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
SEE Challenge 2026: Event-Guided Brightness Adjustment Across a Broad Illumination Range: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
ASIRF: An Agentic Framework for Context-Dependent Sensitive Information Redaction: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
DAWN: Noise-Robust Quadruped Parkour via Depth-Denoising World Models: Covers a concrete video generation signal; useful as a follow-up candidate.
Learning from Mixed-Quality Deployment Experience for Robot Manipulation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
MoVISA: Multi-Token Reasoning for Video Object Segmentation: Covers a concrete video generation signal; useful as a follow-up candidate.
Control the Harness, Control the Cost: Routing and Governing AI Coding Agents in the Enterprise: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Towards Practical Compression of 3D Gaussian Splatting: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
SemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis with Incomplete Data: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
A Living Benchmark for Information Retrieval from Electronic Health Records: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
AERIAL: Adversarial Evaluation of Robustness in Accuracy-Preserving Low-Precision EEG Decoders: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.