Improve code generation, execution feedback, and automated repair, Make agents use tools and reusable skills more reliably
Today tracks: improve code generation, execution feedback, and automated repair, improve code generation, execution feedback, and automated repair, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 369 candidate papers from the 2026-07-06 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1InFlux++: Real and Synthetic Data for Estimating Dynamic Camera Intrinsics🔗
- 2ReCal3R: Reliability-Calibrated Learning Rates for Streaming 3D Reconstruction🔗
- 3Toward Trustworthy Large Language Model Agents in Healthcare🔗
- 4TGRIP: A Text-Guided Approach to Vehicle Instance Prediction in Autonomous Driving🔗
- 5From Open Loop to Closed Loop: A Test-Time Iterative Optimization Framework for Reference-Consistent Image Generation🔗
- 6QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding🔗
What is worth tracking today
Today’s high-signal papers point to: improve code generation, execution feedback, and automated repair, improve code generation, execution feedback, and automated repair, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around benchmark, code, synthetic data, eval to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Camera intrinsics are vital for recovering 3D structure from 2D video
Keywordsbenchmarkcodesynthetic dataeval
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around serving, alignment, code, memory to frame the training and post-training task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Streaming 3D reconstruction relies on a compact recurrent scene state to process long image streams in linear time and bounded memory
Keywordsservingalignmentcodememory
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, workflow, rag, retrieval to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Healthcare appointment scheduling remains a persistent operational bottleneck, driven by manual coordination, fragmented legacy systems, and high administrative overhead
Keywordsagentworkflowragretrieval
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, code, vision-language, temporal to frame the video generation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Bird's-Eye View (BEV) end-to-end instance prediction has emerged as a robust paradigm for autonomous driving perception, effectively mitigating the error propagation inherent in traditional modular pipelines
Keywordsagentcodevision-languagetemporal
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around serving, alignment, evaluation, code to frame the vision and image generation task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: While controllable image generation has made significant strides by incorporating visual reference conditions, existing methods predominantly operate as open-loop systems
Keywordsservingalignmentevaluationcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around retrieval, alignment, benchmark, code to frame the multimodal models task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: The performance of vision-language models (VLMs) in video understanding declines with increasing video duration, as video moments unrelated to the query confuse their language components
Keywordsretrievalalignmentbenchmarkcode
Code/DataCheck the source paper
Other papers worth tracking
G2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal Disentanglement: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Displacement Preserving Relational Distillation for Robust Medical Segmentation: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Multiplayer Interactive World Models with Representation Autoencoders: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Dynamic Airspace Management for UAVs in Evolving Urban Environments: Collaborative Coordination and Human Safety: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing: Covers a concrete training and post-training signal; useful as a follow-up candidate.
SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Selective Disclosure Watermarking for Large Language Models: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
Deep Learning for Semen Analysis in Male Infertility: Computer Vision, Multimodal Fusion, and Clinical Translation: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Evaluating and Understanding Model Editing for Medical Vision Language Models: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Biologically Informed Deep Neural Networks for Multi-Omic Integration, Pathway Activity Inference and Risk Stratification in Cancer: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
Routing Anonymity and Identifiability of Noisy Quantum Hardware: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
SteelBench: Evaluating Vision-Language Models in Real-World Industrial Environments: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Streaming Neural Speech Codecs through Time-Invariant Representations: Covers a concrete interpretability signal; useful as a follow-up candidate.
Unified Audio Intelligence Without Regressing on Text Intelligence: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
VLM-CASE: Vision-Language Model Enabled Context-Adaptive Safety Envelopes for Anticipatory Safe Autonomous Driving: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
On the risk of coding before testing: An empirical study on LLM-based test generation workflow: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
LangLoc: "Tell Me What You See": Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.