Strengthen multimodal understanding of charts, documents, and visual evidence, Make agents use tools and reusable skills more reliably, Improve model reasoning, planning, and verification
Today tracks: strengthen multimodal understanding of charts, documents, and visual evidence, make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification.
This issue fetched and deduplicated 374 candidate papers from the 2026-07-13 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1LaGuadia: Language-Guided Adaptive Distillation from Pathology Foundation Models🔗
- 2MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents🔗
- 3RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM🔗
- 4MAGIC: Transition-Aware Generation of Navigable Multi-Scene Game Worlds with Large Language Models🔗
- 5TreeThink: A Modular Tree Search Library for Mathematical Reasoning with LLMs🔗
- 6Input-Aware Dynamic Backdoor Attack Against Quantum Neural Networks🔗
What is worth tracking today
Today’s high-signal papers point to: strengthen multimodal understanding of charts, documents, and visual evidence, make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
strengthen multimodal understanding of charts, documents, and visual evidence
Signalthis paper targets the concrete research problem behind strengthen multimodal understanding of charts, documents, and visual evidence. It uses the title, abstract, and public signals around alignment, code, vision-language, image to frame the multimodal models task, data, or evaluation flow to improve strengthen multimodal understanding of charts, documents, and visual evidence. The main claim is the title, abstract, and public signals indicate: Pathology Foundation Models (PFMs) offer powerful Whole Slide Image (WSI) representations but suffer from massive computational costs
Keywordsalignmentcodevision-languageimage
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, workflow, evaluation, benchmark to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents
Keywordsagentworkflowevaluationbenchmark
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around rag, retrieval, code, open-source to frame the retrieval and rag task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Graph retrieval-augmented generation (GraphRAG) enhances large language models with structured knowledge, yet existing systems construct knowledge graphs in a single extraction pass, producing noisy entities and brittle retrieval
Keywordsragretrievalcodeopen-source
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, evaluation, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Multi-scene navigation (clearing an objective in one bounded space and then crossing a portal into the next) is a defining feature of contemporary 3D games, but authoring it is laborious: every portal must have consistent endpoints on both sides, each interior
Keywordsagentevaluationbenchmarkcode
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around inference, evaluation, code, open-source to frame the reasoning and planning task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Tree search algorithms enable systematic exploration of the proof space in neural theorem proving
Keywordsinferenceevaluationcodeopen-source
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, deployment, fine-tuning, visual to frame the vision and image generation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Quantum Neural Networks (QNNs) are a promising framework for quantum machine learning on near-term quantum devices, but their security risks remain insufficiently understood
Keywordsragdeploymentfine-tuningvisual
Code/DataCheck the source paper
Other papers worth tracking
Technical Report on the CVPR 2026@AdvML Workshop Challenge: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
ToFu: A White-Box, Token-Efficient Agent Harness for Researchers: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
HASTE: A Platform for Rapid Post-Disaster Building Damage Assessment: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
HierCAD: Hierarchical Text-to-CAD Design via Structure Alignment and Parameter Grounding: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
SISA-Rec: A Semantically Integrated Sequential Recommender with Contrastive Alignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
PREF-Gate: Provenance-Constrained Relational Evidence Fusion with Validation-Gated Selection for Graph Fraud Detection: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Revisiting Matching Response and Swept Feature Volumes for Wide-baseline Omnidirectional Stereo: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
BackendForge: Benchmarking Agentic End-to-End Code Generation with Backend Services: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
TreeSoc: Tree-Structured Dynamic Reasoning and Tool Synergy for Soccer Video Understanding: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Evidence-Backed Video Question Answering: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
MicroCharNet: Less is More for License Plate Character Detection: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Higher-Order Cell Tracking Transformer: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
GFR-SAM: Training-Free Referring Camouflaged Object Segmentation via Cross-Image Prompting: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
FoundationGeo: Learning Spatial Pixel-Wise Fields for Monocular Metric Geometry: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.