Improve code generation, execution feedback, and automated repair, Make agents use tools and reusable skills more reliably
Today tracks: improve code generation, execution feedback, and automated repair, make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 432 candidate papers from the 2026-09-21 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Toward a foundation model for forest point clouds🔗
- 2RRSI: Regularized Recursive Self-Improvement of Agent Harnesses🔗
- 3DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security🔗
- 4MedRSI: Recursive Self-Improvement for Medical Agents via Clinically Aligned Self-Evolution🔗
- 50.5\%>100\%: Bidirectional Reciprocal Learning for Referring Image Segmentation🔗
- 6High-Dimensional Online Change Point Detection with Adaptive Thresholding and Interpretability🔗
What is worth tracking today
Today’s high-signal papers point to: improve code generation, execution feedback, and automated repair, make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around benchmark, code, eval, evaluation to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Forest inventories increasingly rely on artificial intelligence (AI) models to derive forest attributes from large-scale 3D point clouds
Keywordsbenchmarkcodeevalevaluation
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, benchmark, code to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model
Keywordsagentragbenchmarkcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, deployment, evaluation to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: LLM-based agents increasingly operate in environments where they interact with users, tools, and external systems
Keywordsagentragdeploymentevaluation
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, deployment, safety, benchmark to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Medical agents increasingly combine general reasoning models with specialized clinical tools, yet their capabilities remain largely fixed by what clinicians and engineers design before deployment
Keywordsagentdeploymentsafetybenchmark
Code/DataCheck the source paper
strengthen multimodal understanding of charts, documents, and visual evidence
Signalthis paper targets the concrete research problem behind strengthen multimodal understanding of charts, documents, and visual evidence. It uses the title, abstract, and public signals around alignment, benchmark, code, vision-language to frame the training and post-training task, data, or evaluation flow to improve strengthen multimodal understanding of charts, documents, and visual evidence. The main claim is the title, abstract, and public signals indicate: Recent advances in vision foundation models (VFMs) have shown remarkable capabilities across diverse unimodal visual tasks
Keywordsalignmentbenchmarkcodevision-language
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, deployment, code, interpretability to frame the interpretability task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Change point detection (CPD) identifies abrupt and significant changes in sequential data, with applications in human activity recognition, financial markets, cybersecurity, manufacturing, and autonomous systems
Keywordsragdeploymentcodeinterpretability
Code/DataCheck the source paper
Other papers worth tracking
Incentive Noise and Structural Prior Infusion for Multi-modal Object Re-Identification: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
When Evidence Conflicts: Reliability-aware Meta-review Generation: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Who Pays for the KV Cache? Attributing Shared AI Inference Spend Across Kubernetes and LLM Provider Bills: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Rare Event Estimation via Iterative Unalignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
PrismGPT: Proxy-Guided Learning for Region-Aware Photo Editing with Self-Synthesized Reasoning: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
D-JEPA: A Decision-Aligned Latent World Model: Covers a concrete robotics and embodied ai signal; useful as a follow-up candidate.
AlignMorph: Tuning-Free Diffusion Image Morphing via Explicit Semantic Transport: Covers a concrete video generation signal; useful as a follow-up candidate.
A$^2$Safe: Counterfactual Evidence-Aligned Adaptive Agent Collaboration for Safe and Effective Visual Question Answering: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
ACLArena: Agent Continue Learning in Multi-stage Post-training: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
LEAP-NBV: Lightweight Edge Active-Perception for Foundation-Model Next-Best-View Planning: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
SPHQuant: Efficient extreme low bit weight quantization for Vision-Language Models: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
When Quantization Preserves Accuracy but Not Evidence: Explanation-Aware Post-Training Quantization for Medical LLMs: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
GraphSVR: q-Space--Aware Graph-Based Slice-to-Volume Registration for Diffusion MRI: Covers a concrete video generation signal; useful as a follow-up candidate.
ReSTI: A Source-Grounded Audit and Repair of STI-Bench: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Augmented Hypothesis Testing with Persona-Based LLM Simulations: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
FedMust: Semi-supervised Multi-task Student-Teacher Federated Learning for Multi-organ CT Segmentation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Learning tactile perception from high-bandwidth single-point sensing: Covers a concrete robotics and embodied ai signal; useful as a follow-up candidate.
Taking a Second Look: Correcting Sea Ice Forecasts with Sparse Observations: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
HyperCLIP++: Fine-tuning CLIP forOpen-vocabulary Semantic Segmentation in Hyperbolic Space: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.