Make agents use tools and reusable skills more reliably, Make RAG retrieval and knowledge-base QA more reliable, Improve image generation, visual understanding, and controllable rendering
Today tracks: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable.
This issue fetched and deduplicated 398 candidate papers from the 2026-09-08 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Miles v0.1: Production-Level Post-Training🔗
- 2ExecCritic: Learn to Test, Test to Improve for Coding Agents🔗
- 3From Coordinates to Candidate Regions: Temporal Change Localization via Region Selection in Remote Sensing Multimodal LLMs🔗
- 4Supervised Cross-Modal Feature Alignment for Zero-Wearable Freezing of Gait Detection in Parkinsonism🔗
- 5What Eviction Destroys: A Restore-Counterfactual Audit of Forgetting in Agent Memory🔗
- 6SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, deployment, alignment, fine-tuning to frame the training and post-training task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: We present Miles v0.1, a full-stack, production-ready system for frontier post-training
Keywordsagentdeploymentalignmentfine-tuning
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, evaluation, code, post-training to frame the code intelligence task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture the behavior requested by the issue
Keywordsagentevaluationcodepost-training
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, evaluation, code, multimodal to frame the multimodal models task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Remote sensing multimodal large language models (RS-MLLMs) have advanced scene understanding and visual question answering over satellite imagery, yet localizing specific objects or changed regions remains challenging
Keywordsragevaluationcodemultimodal
Code/DataCheck the source paper
improve image generation, visual understanding, and controllable rendering
Signalthis paper targets the concrete research problem behind improve image generation, visual understanding, and controllable rendering. It uses the title, abstract, and public signals around inference, deployment, alignment, evaluation to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve image generation, visual understanding, and controllable rendering. The main claim is the title, abstract, and public signals indicate: Objective assessment of Freezing of Gait (FoG) in Parkinson's disease (PD) relies predominantly on wearable Inertial Measurement Units (IMUs)
Keywordsinferencedeploymentalignmentevaluation
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, retrieval, benchmark, memory to frame the retrieval and rag task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Agent memory systems must discard stored information when their history exceeds a fixed token budget
Keywordsagentretrievalbenchmarkmemory
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, alignment, evaluation, code to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure safe alignment
Keywordsagentalignmentevaluationcode
Code/DataCheck the source paper
Other papers worth tracking
Evaluation Principles for MRI-MRA Registration in Trigeminal Neuralgia: An ROI-Centered Neurovascular Benchmark: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Structural Jailbreaks Generalize but Do Not Compound: A cross-provider and multilingual study of Involuntary In-Context Learning: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
A Measurement Study of LLM Inference Trade-offs Across Edge Continuum Hardware: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
Human-Centric Image Captioning with Subject-Centered Spatial Understanding: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Beyond Coherence: Benchmarking Professional Editing-Technique Execution in Multi-Shot Audio-Video Generation: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
SchemeArena: Factorized Stress Testing of Scheming in LLM Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
TASTE2: Text-Aligned Speech Modeling and Deployment toward Full-Duplex Voice Interaction: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
CoSA: Correlation-Guided Change A ttention with Learnable Residual Gating for Remote Sensing Change Detection: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
SynthRCT: Scalable Conditional Deformation Synthesis for Synthetic Repeat CT Generation: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Do New Attention Mechanisms Actually Fix Attention Sinks at Million-Token Context?: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Beyond Agent Harnesses: Cross-Substrate Authority for Multi-Agent Systems: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Safe Task Planning with Long-Term Graph Memory for Embodied Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
To Adapt or Not to Adapt? Selective Adaptation for Vision-Language Models: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
Rank Without an Oracle: Deviation-Aware Interaction-Rank Selection from Offline Multi-Agent Logs: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
HoneyRoute: Honeypot-Model Routing for Adversarial LLM Serving: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
Qiushi Engine on AstaBench E2E-Bench-Hard: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Snugi-AI-v2 @ eRisk 2026 Task 2: Early Depression Detection via a Learned Stopping Policy with Sustained Confidence Gate: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.