Make agents use tools and reusable skills more reliably, Make RAG retrieval and knowledge-base QA more reliable, Improve model reasoning, planning, and verification
Today tracks: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable.
This issue fetched and deduplicated 388 candidate papers from the 2026-08-06 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models🔗
- 2DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model🔗
- 3DistMedVL: Distributional Vision-Language Alignment for Uncertainty-Aware Medical Image Segmentation🔗
- 4ChronoVision: Temporal Reasoning via Latent State Reconstruction🔗
- 5SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation🔗
- 6TruthLens: Object Hallucination Detection via Self-Evaluating Truthfulness Scores in LVLMs🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, alignment, code, fine-tuning to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: World models enable agents to perform forward rollout and planning without real-world interaction
Keywordsagentalignmentcodefine-tuning
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, latency, safety to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible consequences on external states, user data, and downstream services
Keywordsagentraglatencysafety
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, alignment, benchmark, code to frame the multimodal models task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Cross-modal alignment of visual and textual representations is fundamental to multimodal medical image understanding, yet remains hindered by uncertainty in both modalities under real-world clinical conditions
Keywordsragalignmentbenchmarkcode
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around alignment, benchmark, multimodal, fine-tuning to frame the training and post-training task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning
Keywordsalignmentbenchmarkmultimodalfine-tuning
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, deployment, alignment to frame the retrieval and rag task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Although agent skills equip LLMs with reusable procedural knowledge, manual maintenance suffers from high costs, unscalability, and misalignment
Keywordsagentragdeploymentalignment
Code/DataCheck the source paper
strengthen multimodal understanding of charts, documents, and visual evidence
Signalthis paper targets the concrete research problem behind strengthen multimodal understanding of charts, documents, and visual evidence. It uses the title, abstract, and public signals around inference, deployment, evaluation, benchmark to frame the benchmarks and evaluation task, data, or evaluation flow to improve strengthen multimodal understanding of charts, documents, and visual evidence. The main claim is the title, abstract, and public signals indicate: Despite the remarkable progress of large vision language models (LVLMs), object hallucination remains a fundamental challenge that hinders their trustworthy deployment
Keywordsinferencedeploymentevaluationbenchmark
Code/DataCheck the source paper
Other papers worth tracking
FOCUS: Decoupling Expert Personas in LLMs to Enhance Domain Expert Capabilities: Covers a concrete training and post-training signal; useful as a follow-up candidate.
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
From Sports to Safety: Benchmarking Proactive Risk Inference in MLLMs: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Scalable estimation of VARMA models: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Minimax Optimal Early-Stopped Gradient Descent for Gaussian Mixture Classification: Covers a concrete training and post-training signal; useful as a follow-up candidate.
DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
Reversible Unlearnable Examples: Towards the Copyright Protection in Deep Learning Era: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Beyond Marginal Validity: Finite-Sample Guarantees for Localized Conformal Prediction: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
CFGPNet: Cross-Attention-Based Fused Gradient Programmed Network Framework for Multispectral Object Detection: Covers a concrete multimodal models signal; useful as a follow-up candidate.
EvReflection: Event-Driven Micro-Dynamics for Reflection Removal: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
LLM Inference Under Bursty Workload Distribution: Modifying the WAIT Algorithm: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Is Self-Pretraining really useful to improve diagnosis in medical Time Series?: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
BioKD: Selective Physiology-to-Video Knowledge Distillation via Reliability Gate for Emotion Recognition: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Iterate or Widen? When Test-Time Refinement Helps LiDAR Scene Completion: A Controlled Study of Evidence Geometry, Training Coverage, and Compute: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Wan-Animate-2: Pushing the Application Boundaries of Character Animation: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
Operating Multi-Node Full Fine-Tuning on NVIDIA B300: A Field Report on Telemetry-Based Triage, Negative Results, and Operational Hardening: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Floating Radiance Networks: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
GSBF: Gaussian Splatting for Environment-Aware Beamforming: Covers a concrete data engineering signal; useful as a follow-up candidate.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.