Make agents use tools and reusable skills more reliably, Improve code generation, execution feedback, and automated repair
Today tracks: make agents use tools and reusable skills more reliably, improve code generation, execution feedback, and automated repair, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 468 candidate papers from the 2026-09-17 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents🔗
- 2Beyond the Stability--Plasticity Frontier in Streaming Target Speaker Extraction🔗
- 3Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape🔗
- 4A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems🔗
- 5CoreSense: Traceable Failure Recall and Conflict-Aware Belief Gating for Auditable Robot Decisions🔗
- 6Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, improve code generation, execution feedback, and automated repair, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, safety, evaluation to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: To keep coding agents from going off the rails, production systems now review each proposed action with a blocking monitor that can reject it before it runs (Auto Mode in Claude Code, Guardian in OpenAI's Codex)
Keywordsagentragsafetyevaluation
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around code, memory, table, multimodal to frame the multimodal models task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Streaming target speaker extraction must maintain a representation of whom to extract while the target may fall silent, be masked by interference, or drift acoustically away from enrollment
Keywordscodememorytablemultimodal
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, inference, code to frame the code intelligence task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Frontier AI models are rapidly gaining the ability to exploit vulnerabilities in complex pieces of software
Keywordsagentraginferencecode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, deployment, safety, evaluation to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Benchmark scores alone provide an incomplete basis for assessing the trustworthiness of modern artificial intelligence systems
Keywordsagentdeploymentsafetyevaluation
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, deployment, safety, evaluation to frame the robotics and embodied ai task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Robots can recall prior failures without knowing whether recalled evidence remains valid, conflicts with current observations, or is sufficient to guide a decision
Keywordsragdeploymentsafetyevaluation
Code/DataCheck the source paper
improve image generation, visual understanding, and controllable rendering
Signalthis paper targets the concrete research problem behind improve image generation, visual understanding, and controllable rendering. It uses the title, abstract, and public signals around serving, alignment, memory, video to frame the video generation task, data, or evaluation flow to improve improve image generation, visual understanding, and controllable rendering. The main claim is the title, abstract, and public signals indicate: Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck
Keywordsservingalignmentmemoryvideo
Code/DataCheck the source paper
Other papers worth tracking
Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
CoRef-GS: Cooperative Referring Gaussian Splatting for Multi-Agent Scene Understanding: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
SenseFuse: Label-Free Fusion of Image and Shape Encoders for Open-Vocabulary 3D Instance Segmentation: Covers a concrete robotics and embodied ai signal; useful as a follow-up candidate.
A Multi-Modal Generative Model for Tomato Disease Leaves Understanding: Covers a concrete multimodal models signal; useful as a follow-up candidate.
Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
MILER: Semantic Mid-Level Representation for Sim-to-Real Reinforcement Learning in Unstructured Autonomous Driving: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Towards Scaling Marine Perception with Synthetic Data: Covers a concrete data engineering signal; useful as a follow-up candidate.
INSPECT: Learning Robot View Selection from Assistant Use: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images: Covers a concrete data engineering signal; useful as a follow-up candidate.
Reproducibility is not construct validity: LLM measurement of institutionally situated communication: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Towards High-DoF Dexterous Manipulation through VLA Post-Training: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
PosteriorBench: From Point Estimates to Posterior Matching in Evaluating Generative Inverse Solvers: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Should This Case Be Adapted? Prediction Fragmentation Controls Test-Time Adaptation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
RISC-V and machine learning: a survey: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
UniPolicy: Unified Objective-Specific Policies for Generative Search Advertising: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.