Strengthen multimodal understanding of charts, documents, and visual evidence, Improve model reasoning, planning, and verification, Make agents use tools and reusable skills more reliably
Today tracks: strengthen multimodal understanding of charts, documents, and visual evidence, improve model reasoning, planning, and verification, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 337 candidate papers from the 2026-06-24 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Action ControlNet: A Lightweight Delay-Aware Adapter for Smooth Asynchronous Control in Vision-Language-Action Models🔗
- 2Dual Agreement Consistency Learning for Semi-Supervised Fetal Ultrasound Segmentation🔗
- 3Leaking Circuit Secrets: Gradient Leakage Attacks on Graph Neural Networks🔗
- 4From Sparse and Imperfect 2D Anchors to Consistent 3D Gaussian Street Scenes: Support-Aware Appearance🔗
- 5Dual Distribution Estimation for Zero-shot Noisy Test-Time Adaptation with VLMs🔗
- 6Efficient Real-World Dehazing via Physics-Inspired Global-Local Decoupling🔗
What is worth tracking today
Today’s high-signal papers point to: strengthen multimodal understanding of charts, documents, and visual evidence, improve model reasoning, planning, and verification, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
strengthen multimodal understanding of charts, documents, and visual evidence
Signalthis paper targets the concrete research problem behind strengthen multimodal understanding of charts, documents, and visual evidence. It uses the title, abstract, and public signals around inference, latency, vision-language, robot to frame the robotics and embodied ai task, data, or evaluation flow to improve strengthen multimodal understanding of charts, documents, and visual evidence. The main claim is the title, abstract, and public signals indicate: Vision-language-action (VLA) models have shown strong potential for general-purpose robot manipulation, but their inference latency remains a major obstacle to stable high-frequency control
Keywordsinferencelatencyvision-languagerobot
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around rag, deployment, alignment, code to frame the data engineering task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Maternal-fetal US is the primary imaging modality for monitoring fetal development, yet accurate automated segmentation remains challenging due to the scarcity of pixel-level annotations
Keywordsragdeploymentalignmentcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around serving, compression, evaluation, benchmark to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: As graph neural networks (GNNs) become standard tools for critical tasks in circuit design and analysis, their security and privacy risks require careful attention
Keywordsservingcompressionevaluationbenchmark
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around inference, deployment, alignment, evaluation to frame the training and post-training task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Image priors can synthesize target conditions for 3D Gaussian street scenes, but independently edited views do not define a coherent 3D target
Keywordsinferencedeploymentalignmentevaluation
Code/DataCheck the source paper
strengthen multimodal understanding of charts, documents, and visual evidence
Signalthis paper targets the concrete research problem behind strengthen multimodal understanding of charts, documents, and visual evidence. It uses the title, abstract, and public signals around inference, benchmark, code, vision-language to frame the multimodal models task, data, or evaluation flow to improve strengthen multimodal understanding of charts, documents, and visual evidence. The main claim is the title, abstract, and public signals indicate: While test-time adaptation (TTA) empowers vision-language models to adapt without costly retraining, it remains highly vulnerable to out-of-distribution (OOD) outliers prevalent in real-world applications
Keywordsinferencebenchmarkcodevision-language
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, deployment, latency to frame the systems and deployment task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Real-world single image dehazing is highly ill-posed due to spatially and spectrally varying scattering, while practical deployment demands lightweight and low-latency models
Keywordsraginferencedeploymentlatency
Code/DataCheck the source paper
Other papers worth tracking
OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
MedGuards: Multi-Agent System for Reliable Medical Error Detection and Correction: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
TACO: Towards Task-Consistent Open-Vocabulary Adaptation in Video Recognition: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Beyond Visual Forensics: Auditing Multimodal Robustness for Synthetic Medical Image Detection: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
OrthoTrack: Continuous 6-DoF UAV Trajectory Estimation Anchored in Public Orthophotos: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Autodata: An agentic data scientist to create high quality synthetic data: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
How Robust is OCR-Reasoning? Evaluating OCR-Reasoning Robustness of Vision-Language Models under Visual Perturbations: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Detect, Unlearn, Restore: Defending Text Summarization Models Against Data Poisoning: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Privacy Vulnerabilities of Attention Layers in Tabular Foundation Models and Protection of High-Risk Queries: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Models: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Pulmonary Embolism Risk Stratification from CTPA and Medical Records: Vascular Graphs Are Not All You Need: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
DSP-SLAM++: A Unified Framework for Multi-Class, High-Fidelity Object SLAM in the Wild: Covers a concrete data engineering signal; useful as a follow-up candidate.
USS: Unified Spatial-Semantic Prompts for Embodied Visual Tracking with Latent Dynamics Learning: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.