Make agents use tools and reusable skills more reliably, Improve code generation, execution feedback, and automated repair
Today tracks: make agents use tools and reusable skills more reliably, improve code generation, execution feedback, and automated repair, improve code generation, execution feedback, and automated repair.
This issue fetched and deduplicated 486 candidate papers from the 2026-09-01 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Compile, Don't Memorize: A Context Compilation Architecture (CCA) for In-Context Learning🔗
- 2Benchmarking Spatial, Spectral, and Self-Supervised Cues for Face Forgery Detection under Realistic Degradation🔗
- 3RadMatch: Auditable Radiology Report Evaluation via Finding-Level Matching🔗
- 4Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds🔗
- 5On the Design Fundamentals of Pixel Text Representation Learning🔗
- 6P-PatchDiff: Progressive Patch Diffusion Models for Low-light Image Enhancement🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, improve code generation, execution feedback, and automated repair, improve code generation, execution feedback, and automated repair. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, retrieval, benchmark, code to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Large language models (LLMs) increasingly handle in-context learning (ICL) tasks where a long, novel context defines the rules, knowledge, and output schema for a series of questions
Keywordsagentretrievalbenchmarkcode
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around compression, evaluation, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Face forgery detectors often achieve strong results on controlled benchmarks, but their reliability under realistic image degradations remains limited
Keywordscompressionevaluationbenchmarkcode
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around deployment, safety, evaluation, benchmark to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: As AI systems are increasingly used to draft radiology reports, reliably evaluating their clinical quality remains a critical challenge
Keywordsdeploymentsafetyevaluationbenchmark
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around alignment, evaluation, code, robot to frame the robotics and embodied ai task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Dexterous manipulation policies learned by imitation are typically evaluated for robustness to variation in scenes, objects, or instructions, but their performance across task execution speeds is less often examined
Keywordsalignmentevaluationcoderobot
Code/DataCheck the source paper
strengthen multimodal understanding of charts, documents, and visual evidence
Signalthis paper targets the concrete research problem behind strengthen multimodal understanding of charts, documents, and visual evidence. It uses the title, abstract, and public signals around compression, alignment, evaluation, code to frame the multimodal models task, data, or evaluation flow to improve strengthen multimodal understanding of charts, documents, and visual evidence. The main claim is the title, abstract, and public signals indicate: Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet existing pixel-text encoders struggle with fixed resolution pretraining, visual shortcut learning, weak visual grounding, and multilingual visual
Keywordscompressionalignmentevaluationcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, alignment, code, memory to frame the vision and image generation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Recent advancements in low-light image enhancement have leveraged diffusion models for their strong ability to generate perceptually realistic, detailed images
Keywordsragalignmentcodememory
Code/DataCheck the source paper
Other papers worth tracking
One Prompt Is Enough: Watermark Laundering Through Foundation Image Models: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Dotting the Eye: An Intent-Driven Image Retouching Agent for Visual Focus Enhancement: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Post-hoc Alignment of LLM-judges to Human Judgment Distribution: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
On-the-Fly3R: Towards Robust Online 3D Reconstruction with Feed-Forward 3R Models for Large-Scale UAV Scenarios: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
A multicenter benchmark and clinically structured metric for coronary CTA report generation: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Escaping Redundant Reasoning: Structure-Aware Search for Inference-Time LLMs: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
SCoNE: Selective Context-aware Neuron Editing for Robust Retrieval-Augmented Generation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
DGNet: Dual-knowledge Guided Network for Infrared Small Target Detection: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
REVISE: Validity-Guided Recovery for Online Revisions in Agent Workflows: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
EM^2Mem: Event-Centric Multimodal Memory for Large Language Models: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Closing Cost-Quality Gap in Document VLMs: Difficulty-Aware Data Curation and Quality-Adjusted Deployment Economics: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
Gradient-Update Mismatch: Rethinking Conflict-Free Training of Physics-Informed Neural Networks: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
Bandits in Prod: Hyperparameter Optimization at Inference Time: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
ReFlowSET: Representation-Aligned Latent Flow Matching for SAR-to-EO Image Translation: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Candidate-Expanding Routing with Permutation-Stabilized Experts for Mixed-Format Medical VQA: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Differentially Private Paired Table-Image Multimodal Synthesis: Covers a concrete multimodal models signal; useful as a follow-up candidate.
Trust Your Guide Only When Certain: Uncertainty-Aware Sparse Alignment at Inference Time: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.