Make agents use tools and reusable skills more reliably, Improve code generation, execution feedback, and automated repair
Today tracks: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, improve code generation, execution feedback, and automated repair.
This issue fetched and deduplicated 444 candidate papers from the 2026-09-25 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1WeaveAgent: A Two-Stage Tool-Routing Agent for Ultra-High-Resolution Remote Sensing Imagery🔗
- 2MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens🔗
- 3Region-Level Black-Box Defense Against Stealthy Embedding-Space Backdoors in CLIP🔗
- 4Muslim: A Deployed Arabic Voice AI Platform for Grounded Islamic Knowledge🔗
- 5Evolutionary Safety of Recursive Self-Improving AI: Taxonomy, Risk Discovery, and Evaluation🔗
- 6G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, improve code generation, execution feedback, and automated repair. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, compression, alignment, evaluation to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Problem
Keywordsagentcompressionalignmentevaluation
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, safety, evaluation, benchmark to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Most work on improving large language models treats accuracy as the sole objective
Keywordsagentsafetyevaluationbenchmark
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around serving, deployment, alignment, code to frame the systems and deployment task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Contrastive Language--Image Pretraining (CLIP) has emerged as a dominant vision backbone due to its strong transferability and zero-shot capabilities
Keywordsservingdeploymentalignmentcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, retrieval, serving, latency to frame the retrieval and rag task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: We present Muslim, a production Arabic voice AI platform serving grounded, sourced Islamic knowledge to real users
Keywordsagentretrievalservinglatency
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, safety, evaluation, risk to frame the safety and alignment task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Artificial intelligence is advancing rapidly, with increasingly capable systems taking larger roles in reasoning, decision-making, scientific discovery, and autonomous development
Keywordsagentsafetyevaluationrisk
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around alignment, code, post-training, memory to frame the training and post-training task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining
Keywordsalignmentcodepost-trainingmemory
Code/DataCheck the source paper
Other papers worth tracking
AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
ALF: An Active Learning Framework for Scientific Discovery: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
TRACKGRAPH: Online Open-Vocabulary 3D Scene Graphs via Image-Space Tracking: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
CCRV-Bench: Constraint-Based Evaluation of Causal Reasoning in Vision-Language Models: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Financial Fragility in Societies of LLM Agents: Coordination Failures and Stabilizing Mechanisms: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Learning Chance-Constrained MDPs with Bellman Distributional Certificates: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Learning Provable Neural Network Observer for Uncertain Dynamical Systems: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Evaluation Is All You Need for Multi-Modal Autonomous Driving: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Understanding the Role of Prompt Template in Knowledge Distillation for Safety Alignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Werracle: Sub-Cent Intra-Block AI Reflex Oracles and Flash-Loan Circuit Breakers for EVM Smart Contracts: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
CRC-Router: Risk-Constrained Routing for Medical Agentic AI Systems: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
How Far Can INRs Go? Cross-Domain Parameter-efficient INR-Based Semantic Segmentation for Brain MRI: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Stale-Document Poisoning: When Outdated Retrieval Overrides Correct Model Answers: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
When the Model Retires: An Empirical Study of LLM Migration in Open-Source Applications: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
MetaPermit: Scalable and Auditable Access Control for AI Agents via LLM-Inferred Meta-Attributes: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
FLIP: Final Layer Inference-Time Probing for Vision-Language Models: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
The KV Cache Is the New Memory Wall: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.