Make agents use tools and reusable skills more reliably, Improve model reasoning, planning, and verification, Make RAG retrieval and knowledge-base QA more reliable
Today tracks: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification.
This issue fetched and deduplicated 352 candidate papers from the 2026-09-04 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Qlippy: A Retrieval-Augmented GenAI Assistant for Reproducible Quantum Workflows and Experiment Tracking🔗
- 2BIT.UA at BioASQ 14B: Modular Retrieval with pg_textsearch and Qdrant, and Agent-Based Answer Generation🔗
- 3A Tree-based RAG Framework for Evidence-Intensive QA via Adaptive Planning and Topology-Aware Evidence Gathering🔗
- 4SAM-D2Q: Aligning Multimodal Doc2Query with Search Demand and Conversion for E-commerce🔗
- 5MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression🔗
- 6One Diffusion Model, Two Roles: Guided Trajectory Planning and Safety-Critical Scenario Generation in Closed-Loop Simulation🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around workflow, retrieval, serving, deployment to frame the retrieval and rag task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Quantum software development is iterative and error-prone
Keywordsworkflowretrievalservingdeployment
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, retrieval, code to frame the retrieval and rag task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: This paper describes the participation of the BIT.UA team from the University of Aveiro in the 14th edition of the BioASQ Task B challenge on biomedical question answering
Keywordsagentragretrievalcode
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around rag, retrieval, benchmark, code to frame the reasoning and planning task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Recent structured RAG methods leverage tree- or graph-based reasoning structures to improve multi-hop QA
Keywordsragretrievalbenchmarkcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, retrieval, alignment, multimodal to frame the multimodal models task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: E-commerce search often suffers from vocabulary mismatch between user queries and merchant-authored product titles, since short titles cannot fully cover diverse user expressions or visual product attributes
Keywordsragretrievalalignmentmultimodal
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around inference, serving, compression, alignment to frame the systems and deployment task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Recently, multimodal large-scale reasoning models have demonstrated remarkable capabilities in solving complex tasks through long Chains-of-Thought (M-CoT)
Keywordsinferenceservingcompressionalignment
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, inference, serving to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Diffusion probabilistic models can capture the multi-modal, interaction-rich distribution of joint future trajectories in driving scenes
Keywordsagentraginferenceserving
Code/DataCheck the source paper
Other papers worth tracking
Artificial Intelligence in Equity and Crypto Markets: Progress, Profitability Evidence, and the Limits of Automated Investing: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Adaptation Interfaces for In-Context Tabular Foundation Models in Time-to-Event Prediction: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
AutoLR: Automating the Path from Research to Launch Review in Industrial Recommender Systems: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
CoLMIN: LLM-based Multi-Decision Path Negotiation for Cooperative Autonomous Driving: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Intrinsic Temporal Adaptation of CLIP for Partially Relevant Video Retrieval: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Diffusion Language Models for Mobile Edge Agentic AI: Foundations, Applications, and Challenges: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Shadow Queries for Private Retrieval in Vector Databases: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
AngelFingerprint: A Traceable, Explainable, and White-Box Stealthy Watermark for Text-Guided Image Editing: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
How Developers Discuss Generative AI: A Longitudinal Study of the Visual Studio Code Community: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Latent-Aligned Reasoning for Multimodal Recommendation: Covers a concrete multimodal models signal; useful as a follow-up candidate.
CrossDepth: Geometry-Constrained Attention for Generalizable Multi-View Surround Depth Estimation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Think-Verify-Revise: Neuro-Symbolic Visual Reasoning with Vision-Language Models and Dynamic Logic Tensor Networks: Covers a concrete multimodal models signal; useful as a follow-up candidate.
Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation: Covers a concrete video generation signal; useful as a follow-up candidate.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.