Make agents use tools and reusable skills more reliably, Improve model reasoning, planning, and verification
Today tracks: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 436 candidate papers from the 2026-09-14 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Safe Meta-Reinforcement Learning via Information Space Reachability🔗
- 2The Troy Moment of AI: Why SomeWill Cheat and SomeWill Follow?🔗
- 3SkillLift: Learning Dense Rubrics from Sparse Oracles for Efficient Skill Evolution🔗
- 4CWM: Controllable White-Box Meta-Prompting for Adaptive Retrieval-Augmented Generation and Reasoning Ability🔗
- 5Disentangling Representation Evolution in Transformers through Directional Decomposition🔗
- 6Personalizing Personal Health Interfaces: Co-Design with Generative AI🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, safety, benchmark to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Meta-reinforcement learning (meta-RL) enables agents to adapt to unseen tasks with limited experience
Keywordsagentragsafetybenchmark
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, serving, benchmark, code to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Recent investigations of the July 2026 OpenAI--Hugging Face incident motivate two questions about agent behavior under task failure: when an assigned task becomes impossible, does an agent stop or escalate, and can observing another agent's behavior change tha
Keywordsagentservingbenchmarkcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, evaluation, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: LLM-based agents increasingly rely on persistent skills, i.e., reusable procedural prompts, to adapt without weight updates
Keywordsagentevaluationbenchmarkcode
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around rag, retrieval, benchmark, code to frame the retrieval and rag task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Recently, Large Language Models (LLMs) have gained significant attention due to their strong language understanding and generation capabilities, demonstrating impressive reasoning abilities as well as effective utilization of external knowledge
Keywordsragretrievalbenchmarkcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, serving, compression, code to frame the training and post-training task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it
Keywordsragservingcompressioncode
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around serving, latency, safety, systems to frame the systems and deployment task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Personal health interfaces present wellbeing data through standardized dashboards that rarely fit how people interpret or act on it
Keywordsservinglatencysafetysystems
Code/DataCheck the source paper
Other papers worth tracking
OpenAI4S: Code as Action, Science as Sessions: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Route Me If You Can: A Benchmark for Query Reformulation Selection: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Look Before You Leap: Factual Decoding with Internal Attribution Signals: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Learning Multimodal One-step Flow Policy via Value-weighted Optimal Transport: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Can a Neural Encoding Model Replicate an fMRI Visualization Study?: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
SAM3D-Part: Interactive Part Selection and Generation from 3D Objects: Covers a concrete multimodal models signal; useful as a follow-up candidate.
VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Accelerating the Solving of Many Tiny General Linear Systems on GPUs: Application to Constitutive Laws: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
ActGuard: Pre-execution Action Auditing against Indirect Prompt Injection in LLM Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Online Language Adaptive Sampling for Better Distributed Cross-lingual Gains: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Inoculation Midtraining with Learned Neologisms: Covers a concrete training and post-training signal; useful as a follow-up candidate.
CiteGuard-RAG: A Validation-Centered AI System for Evidence-Grounded Question Answering: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Atria Dawn: The Dawn of Agentic Superintelligence: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
KnowBench: Effort Reduction as a Unified, Deployment-Grounded Benchmark for Clinical AI: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
EvoOntology: A Self-Evolving Ontology Layer for Data Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.