Make RAG retrieval and knowledge-base QA more reliable, Make agents use tools and reusable skills more reliably
Today tracks: make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 383 candidate papers from the 2026-06-10 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Performance Analysis of YOLOv11 and YOLOv8 for Mixed Traffic Object Detection under Adverse Weather Conditions in Developing Countries🔗
- 2ParseFixer: An Agentic Framework for Document Parsing via Selective Multimodal Correction🔗
- 3AutoMine Solution for AV2 2026 Scenario Mining Challenge🔗
- 4PROJECTMEM: A Local-First, Event-Sourced Memory and Judgment Layer for AI Coding Agents🔗
- 5Corpus Augmentation for Sign Language Translation via LLM-Guided Video Stitching🔗
- 6From Uniform to Learned Graph Priors: Diffusion for Structure Discovery🔗
What is worth tracking today
Today’s high-signal papers point to: make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, deployment, safety to frame the benchmarks and evaluation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: In modern vehicular systems, robust performance under harsh conditions has become a critical problem of autonomous driving
Keywordsraginferencedeploymentsafety
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, serving, code, multimodal to frame the multimodal models task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: In this report, we present our third-place solution for the DataMFM Challenge Track 1: Document Parsing
Keywordsagentservingcodemultimodal
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around serving, safety, evaluation, code to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: With the development of autonomous driving systems, mining high-value, safety-critical, and planning-relevant scenarios from large-scale driving logs has become essential for data-driven evaluation
Keywordsservingsafetyevaluationcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, code, open-source to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: AI coding assistants now support a growing share of software work, from quick scripts to production applications
Keywordsagentragcodeopen-source
Code/DataCheck the source paper
strengthen multimodal understanding of charts, documents, and visual evidence
Signalthis paper targets the concrete research problem behind strengthen multimodal understanding of charts, documents, and visual evidence. It uses the title, abstract, and public signals around alignment, code, vision-language, synthetic data to frame the data engineering task, data, or evaluation flow to improve strengthen multimodal understanding of charts, documents, and visual evidence. The main claim is the title, abstract, and public signals indicate: Sign language translation (SLT) converts sign language video into spoken language text and holds significant promise for improving accessibility and enabling communication between signing and non-signing communities
Keywordsalignmentcodevision-languagesynthetic data
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around inference, alignment, benchmark, code to frame the training and post-training task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Neural relational inference (NRI) methods discover interaction graphs from trajectories through variational reasoning on discrete potential edges
Keywordsinferencealignmentbenchmarkcode
Code/DataCheck the source paper
Other papers worth tracking
FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
A Lightweight Multi-Agent Framework for Automated Concrete Barrier Design: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Human-Enhanced Loop Modeling (HELM): Agent-Based Finite Element Modeling of Concrete Bridge Barriers: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Tail-Aware Adaptive-k: Query-Adaptive Context Selection for Retrieval-Augmented Generation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
MoCA-Agent: A Market-of-Claims Code Agent for Financial and Numerical Reasoning: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
CCKS: Consensus-based Communication and Knowledge Sharing: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
AGE-MIL: Anchor-Guided Evidence Learning for Patient-Level Prediction: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
GraspLLM: Towards Zero-Shot Generalization on Text-Attributed Graphs with LLMs: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Critic Architecture Matters: Dual vs. Unified Critics for Humanoid Loco-Manipulation: Covers a concrete training and post-training signal; useful as a follow-up candidate.
When Do Data-Driven Systems Exhibit the Capability to Infer?: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
SAFER-Nav: Enhancing Safety for Visual Robot Navigation via Segmentation-Aware Fine-Tuning: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence Vectors: Covers a concrete data engineering signal; useful as a follow-up candidate.
FreqKD: Frequency-Decoupled Cross-Modal Knowledge Distillation for Infrared Object Detection: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Context-Driven Incremental Compression for Multi-Turn Dialogue Generation: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Redesign Mixture-of-Experts Routers with Manifold Power Iteration: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
System Report for CCL25-Eval Task 5: New Dataset and LoRA-Fine-Tuned Qwen2.5: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Latent World Recovery for Multimodal Learning with Missing Modalities: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.