Make agents use tools and reusable skills more reliably, Strengthen multimodal understanding of charts, documents, and visual evidence, Improve code generation, execution feedback, and automated repair
Today tracks: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, strengthen multimodal understanding of charts, documents, and visual evidence.
This issue fetched and deduplicated 321 candidate papers from the 2026-07-23 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation🔗
- 2Future Rendering $\neq$ Future Surface: A Benchmark and Dataset for Dynamic Surface Reconstruction Beyond the Observed Window🔗
- 3Multimodal Pretraining for Generalizable EEG Representation Learning🔗
- 4Inference-Time Scaling of Diffusion Models via Progressive Seed Pruning🔗
- 5When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation🔗
- 6Declarative Problem Solving in UAM Strategic Deconfliction🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, strengthen multimodal understanding of charts, documents, and visual evidence. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, workflow, serving, safety to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction
Keywordsagentworkflowservingsafety
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around deployment, evaluation, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Dynamic-scene reconstruction is almost always evaluated inside the observed time window, yet deployment settings such as AR overlays, robot interaction, and anticipatory planning need the future surface: the geometry at times beyond those captured
Keywordsdeploymentevaluationbenchmarkcode
Code/DataCheck the source paper
strengthen multimodal understanding of charts, documents, and visual evidence
Signalthis paper targets the concrete research problem behind strengthen multimodal understanding of charts, documents, and visual evidence. It uses the title, abstract, and public signals around alignment, evaluation, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve strengthen multimodal understanding of charts, documents, and visual evidence. The main claim is the title, abstract, and public signals indicate: Electroencephalography (EEG) models used for epilepsy are often limited to specific datasets and tasks
Keywordsalignmentevaluationbenchmarkcode
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around inference, alignment, evaluation, code to frame the vision and image generation task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Diffusion and flow-matching models dominate conditional image generation, yet inference-time scaling for these models is far less developed than for autoregressive language models
Keywordsinferencealignmentevaluationcode
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around safety, benchmark, code, multimodal to frame the multimodal models task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: A vision-language AI assistant returns its answer as a stream of generated tokens
Keywordssafetybenchmarkcodemultimodal
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around benchmark, memory, program, execution to frame the code intelligence task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: The growing demand for Urban Air Mobility (UAM) introduces significant challenges in airspace management, particularly within densely populated metropolitan regions
Keywordsbenchmarkmemoryprogramexecution
Code/DataCheck the source paper
Other papers worth tracking
Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Sidewalk Moments: Are Richer Representations Always More Human-Aligned? Evidence from City-Walk Videos: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
HGeo-TopoMap: Boosting Topological Mapping with Hierarchical Geometric Priors: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Boosting Robustness for All-Weather Self-Supervised Depth Estimation in Autonomous Driving: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Toward Continuous Assurance for the Democratization of AI Agent Creation in Industry: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Euclid-MCP: A Model Context Protocol Server for Deterministic Logical Reasoning via Prolog: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
SlerpFlow: Spherical Trajectory Correction for Rectified Flow Inversion: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Adaptive Depth Sparse Framework: Similarity-Driven Resource Allocation for Pre-Trained LLMs: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Detectors Learn the Wrong Thing: Shortcut-Resistant Adversarial Training Against Physically Realizable Attacks: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Safety-oriented sidewalk and road segmentation for smartphone-based assistive navigation: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
C-PTQ: Fisher-weighted Channel-wise Sensitivity for Post-training Quantization of MLLMs: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Spectral Transformation for Layer-wise Global Rank Discovery in Federated LoRA for Vision Transformers: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
From Evaluation to Optimisation: Hierarchy-Aware Training Signals for CWE Prediction in Python: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.