Improve code generation, execution feedback, and automated repair, Make agents use tools and reusable skills more reliably
Today tracks: improve code generation, execution feedback, and automated repair, make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 328 candidate papers from the 2026-08-12 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1HCGRec: Hint-Conditioned Generative Recommendation with Semantic IDs🔗
- 2Certifying What Helps Customer-Return Timing: A Screen-and-Confirm Test for Conditioning Signals, and Why Decay Is Nearly Enough🔗
- 3Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams🔗
- 4Rethinking Agent Security as a Networking Problem🔗
- 5Localizing to Debias: A Patch-Level Benchmark and Baseline for Weakly Supervised Spatial Anomaly Detection🔗
- 6Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder🔗
What is worth tracking today
Today’s high-signal papers point to: improve code generation, execution feedback, and automated repair, make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around alignment, benchmark, code, fine-tuning to frame the training and post-training task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Semantic-ID generative recommenders represent each item as a short sequence of discrete semantic tokens and predict the next item by autoregressively generating this token sequence
Keywordsalignmentbenchmarkcodefine-tuning
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around evaluation, benchmark, memory, eval to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Practitioners enrich customer-return models with ever more signals (lifetime value, category, recency/frequency, calendar, geography), and the temporal-point-process (TPP) literature follows suit with covariate- and external-covariate-conditioned intensities
Keywordsevaluationbenchmarkmemoryeval
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, evaluation, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration
Keywordsagentevaluationbenchmarkcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, serving, safety, agents to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: AI agents are rapidly becoming more capable and widely deployed, promising substantial gains in productivity and enabling new classes of applications
Keywordsagentservingsafetyagents
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around serving, evaluation, benchmark, vision-language to frame the video generation task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Despite growing interest in weakly supervised video anomaly detection (WSVAD), current methods struggle to bridge the gap between coarse temporal supervision and fine-grained spatial reasoning
Keywordsservingevaluationbenchmarkvision-language
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, evaluation, benchmark to frame the speech and audio task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity
Keywordsraginferenceevaluationbenchmark
Code/DataCheck the source paper
Other papers worth tracking
LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Curvature-Aware Zeroth-Order Optimization for Memory-Efficient Test-Time Adaptation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
VICBench: A Multi-Language Benchmark for Code Vulnerability Detection: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
ScaleVid: Geometry-Aware Video Object Scaling with Mesh-Free Inference: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
ScreenShot: A Foundation Model for Few-Shot Combination Drug Screening: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge: Covers a concrete training and post-training signal; useful as a follow-up candidate.
NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Context Blindness in DPO: Mitigating Object Hallucination in MLLMs via Context-Calibrated Preference Optimization: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Confidence Calibration of Deep Learning Systems: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Look What the Probes Dragged In! Real-World Chest X-ray Shortcuts in MedCLIP: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Faithful, Sufficient and Understandable: Rethinking Graph Counterfactual Explanations via Discrete Diffusion Inversion: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Do Not Forget the Obvious - RISC: A Risk-Informed Slice-Coverage Protocol for Safe Autonomous Driving: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.