Make agents use tools and reusable skills more reliably, Improve model reasoning, planning, and verification, Improve code generation, execution feedback, and automated repair
Today tracks: make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification, improve code generation, execution feedback, and automated repair.
This issue fetched and deduplicated 326 candidate papers from the 2026-07-16 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space🔗
- 2Self-Evolving Human-Centered Framework for Explainable Depression Symptom Annotation🔗
- 3MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection🔗
- 4On-Policy Delta Distillation🔗
- 5DriftWorld: Fast World Modeling through Drifting🔗
- 6JADE-GS: Joint Alternating Deblurring Guided by Events in 3D Gaussian Splatting🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification, improve code generation, execution feedback, and automated repair. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, safety, benchmark, agents to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world
Keywordsagentsafetybenchmarkagents
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around alignment, evaluation, memory, eval to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Annotation quality is a major bottleneck in building reliable and explainable artificial intelligence (XAI) systems for mental health research
Keywordsalignmentevaluationmemoryeval
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around safety, benchmark, open-source, data to frame the data engineering task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Most medical AI benchmarks measure whether a model knows the correct answer
Keywordssafetybenchmarkopen-sourcedata
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around benchmark, code, post-training, training to frame the training and post-training task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constraints imposed by reward models by providing token-level supervision from a teacher model
Keywordsbenchmarkcodepost-trainingtraining
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around rag, inference, evaluation, benchmark to frame the robotics and embodied ai task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Predictive world models enable robots to plan by imagining the outcomes of their actions, but their value for control hinges on generating many rollouts quickly
Keywordsraginferenceevaluationbenchmark
Code/DataCheck the source paper
improve image generation, visual understanding, and controllable rendering
Signalthis paper targets the concrete research problem behind improve image generation, visual understanding, and controllable rendering. It uses the title, abstract, and public signals around serving, benchmark, rendering, vision-generation to frame the vision and image generation task, data, or evaluation flow to improve improve image generation, visual understanding, and controllable rendering. The main claim is the title, abstract, and public signals indicate: When a camera moves fast during exposure, blur destroys the intra-exposure motion a 3D model needs to recover the sharp scene, while event cameras capture exactly this signal at microsecond resolution
Keywordsservingbenchmarkrenderingvision-generation
Code/DataCheck the source paper
Other papers worth tracking
CoSimRec: Measuring Coordinated-Content Penetration in Recommender Feedback Loops: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Demographically-Conditioned Synthetic Medical Images for Bias Mitigation and Bias Detection in Disease Classifiers: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Physics-Informed Diffusion for Biomechanically Plausible 3D Sign Language Generation: Covers a concrete video generation signal; useful as a follow-up candidate.
Evaluating Epistemic Uncertainty: Beyond OOD Detection and Active Learning: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Transcoders for Investigating Deception in Language Models: Covers a concrete interpretability signal; useful as a follow-up candidate.
Rare Concept Generation via Counterfactual Inference in Diffusion Models: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
AE-UAV: An Air-to-Air Event-Based UAV Tracking Benchmark and a Real-Time Frequency-Domain Tracker: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Causal-Adversarial Probing of Clinical Covariates for Prostate MRI Grading: Covers a concrete interpretability signal; useful as a follow-up candidate.
Reinforcement Learning for the Full Strawberry Harvesting Process: Obstacle Separation, Detachment, and Placement: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Harnessing LLMs for Reliable Academic Supervision: A Comparative Study: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Dendrite: A Real-Time Python Application for Online Brain-Computer Interface Research and Development: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
NavCMPO: Critic-Guided MeanFlow Policy Optimization for Adaptive Navigation: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
ExaGEMM: Exploration Framework for CPU-Driven ML Inference via Associative In-Register Computing for Low-Bit GEMM: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Auditing Fairness-Privacy Trade-offs: Subpopulation-Level Effects of Fairness-Enhancing Algorithms: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Breaking the Model Forgetting Cycle in Long-Incremental 3D Object Detection: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Multi-Scale ViT Inference with Habitat-Fit Priors and kNN Retrieval for Multi-Species Plant Identification: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Hierarchical Denoising For Multi-Step Visual Reasoning: Covers a concrete video generation signal; useful as a follow-up candidate.
Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.