Improve code generation, execution feedback, and automated repair, Improve model reasoning, planning, and verification, Make agents use tools and reusable skills more reliably
Today tracks: improve code generation, execution feedback, and automated repair, improve model reasoning, planning, and verification, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 479 candidate papers from the 2026-08-05 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1VQ-VAD: Vector-quantized Motion Representation Learning for Human-centric Video Anomaly Detection🔗
- 2Unified Planning-Learning Framework for Robust UUV Navigation Under Partial Observability🔗
- 3Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent🔗
- 4Toward Integrating Adaptive Experience Replay and Online Uncertainty Estimation in Safe Actor-Critic Optimal Control🔗
- 5RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists🔗
- 6Dense Metric Depth Completion from Sparse Direct Time-of-Flight Sensors🔗
What is worth tracking today
Today’s high-signal papers point to: improve code generation, execution feedback, and automated repair, improve model reasoning, planning, and verification, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around evaluation, benchmark, code, eval to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Video Anomaly Detection (VAD) is inherently challenging due to the scarcity of anomalies and the large visual variability in surveillance footage, including changes in lighting, viewpoint, and human appearance
Keywordsevaluationbenchmarkcodeeval
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around safety, evaluation, benchmark, training to frame the training and post-training task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: This paper presents an observation-only autonomy framework for Unmanned Underwater Vehicles (UUVs) navigation in dynamic underwater environments that integrates persistent occupancy mapping, global clearance-aware planning, and risk-aware local control
Keywordssafetyevaluationbenchmarktraining
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, inference, serving, safety to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Scaling supervision for multi-turn medical agents is difficult because expert dialogue annotation is costly and clinical conversations are privacy-restricted
Keywordsagentinferenceservingsafety
Code/DataCheck the source paper
identify and reduce safety, jailbreak, and alignment risks
Signalthis paper targets the concrete research problem behind identify and reduce safety, jailbreak, and alignment risks. It uses the title, abstract, and public signals around safety, evaluation, benchmark, post-training to frame the benchmarks and evaluation task, data, or evaluation flow to improve identify and reduce safety, jailbreak, and alignment risks. The main claim is the title, abstract, and public signals indicate: Safe actor-critic control often treats barrier filtering, uncertainty estimation, and experience replay as separate modules, even though each changes the data used for learning and control
Keywordssafetyevaluationbenchmarkpost-training
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around alignment, evaluation, benchmark, code to frame the code intelligence task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: The integration of Large Language Models (LLMs) into software engineering has shifted the focus from function-level generation to repository-scale assistance
Keywordsalignmentevaluationbenchmarkcode
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around code, robotics, synthetic data, open-source to frame the data engineering task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Direct Time-of-Flight (dToF) sensors provide highly accurate metric depth and are more robust than indirect ToF systems in challenging real-world conditions
Keywordscoderoboticssynthetic dataopen-source
Code/DataCheck the source paper
Other papers worth tracking
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
A GitOps-Driven Annotation Catalog for Fully Automatic Railway Operations: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
CSGen: A Multi-Domain Curvilinear Structure Generation Model via Hierarchical Multimodal Diffusion: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
DAC-Pose: Dual-Agent Collaborative Framework for Pose-Guided Human Generation: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles: Covers a concrete multimodal models signal; useful as a follow-up candidate.
MCHA: A Memory-Centric Hierarchical Architecture for Parallel-Sequential Computing: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
A Foundational EDM2-Based Generative Model for High-Resolution Synthetic Fetal Ultrasound Imaging from Open Datasets: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Robust Context-Aware Detection of Malicious Instructions in Text: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
RepairFormer: Automated Repair of Structured Inputs Using Transformers: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
A 6G Integrated Sensing and Communication Framework for Railway Intrusion Detection and Collision Prediction: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
GUARD: Grounding Uncertainty and Ablation-Based Risk Detection for Diffusion-Based VLAs: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Architectural Implications of Agentic AI Workflows: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
OmniRouting: A Semantic-Coupled Multimodal Benchmark for Constraint-Aware Spatial Reasoning in PCB Routing: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Predict, Then Retrieve: Cross-Instance Future-State Retrieval from Video Prefixes: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.