Make RAG retrieval and knowledge-base QA more reliable, Make agents use tools and reusable skills more reliably
Today tracks: make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable.
This issue fetched and deduplicated 481 candidate papers from the 2026-06-17 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Toten: Knowledge-Based Ontological Tokenization Of Physical Quantities And Technical Notation In Brazilian Portuguese🔗
- 2StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns🔗
- 3LEAP: Layer-skipping Efficiency via Adaptive Progression for Vision Transformer Distillation🔗
- 4A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2🔗
- 5A Technical Taxonomy of LLM Agent Communication Protocols🔗
- 6REVES: REvision and VErification--Augmented Training for Test-Time Scaling🔗
What is worth tracking today
Today’s high-signal papers point to: make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, compression, evaluation, benchmark to frame the retrieval and rag task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Byte-Pair Encoding tokenization is statistically efficient for vocabulary compression, but semantically blind to structured technical entities, fragmenting physical quantities, numbers, units, and symbolic expressions into lexically arbitrary subwords
Keywordsragcompressionevaluationbenchmark
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, benchmark, code, open-source to frame the code intelligence task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: We introduce StaminaBench, a benchmark that measures the stamina of coding agents: how many consecutive interaction turns (change requests) they can handle before failing
Keywordsagentbenchmarkcodeopen-source
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around retrieval, inference, deployment, code to frame the training and post-training task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Vision Foundation Models (VFMs) with Vision Transformer (ViT) backbones, such as DINOv2, have become essential for downstream tasks like object recognition and semantic segmentation
Keywordsretrievalinferencedeploymentcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, compression, evaluation, benchmark to frame the multimodal models task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Text-rich images often contain privacy-sensitive, transactional, or decision-relevant information
Keywordsragcompressionevaluationbenchmark
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, open-source, agents to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: As large language models (LLMs) advance and multi-agent systems aim to overcome the limits of standalone agents, robust communication protocols are becoming essential infrastructure for distributed agent networks
Keywordsagentragopen-sourceagents
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around inference, alignment, code, post-training to frame the reasoning and planning task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Test-time scaling via sequential revision has emerged as a powerful paradigm for enhancing Large Language Model (LLM) reasoning
Keywordsinferencealignmentcodepost-training
Code/DataCheck the source paper
Other papers worth tracking
Language Models as Interfaces, Not Oracles: A Hybrid LLM-ML System for Pediatric Appendicitis: Covers a concrete multimodal models signal; useful as a follow-up candidate.
FloatDoor: Platform-Triggered Backdoors in LLMs: Covers a concrete code intelligence signal; useful as a follow-up candidate.
OpenAnt: LLM-Powered Vulnerability Discovery Through Code Decomposition, Adversarial Verification, and Dynamic Testing: Covers a concrete code intelligence signal; useful as a follow-up candidate.
Adaptive Speech-to-Spike Encoding for Spiking Neural Networks: Covers a concrete code intelligence signal; useful as a follow-up candidate.
Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
DeXposure-Claw: An Agentic System for DeFi Risk Supervision: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
TurboServe: Serving Streaming Video Generation Efficiently and Economically: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Scoring Backends Matter More Than Pooling: A Systematic Study of Training-Free Anomalous Sound Detection under Domain Shift: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Physics-IQ Verified: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Scaling Learning-based AEB with Massive Unlabeled Data: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Rethinking Air-Ground Collaboration: A Progressive Cross-Task Benchmark and Socialized Learning Framework: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Skill-MAS: Evolving Meta-Skill for Automatic Multi-Agent Systems: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Formal Verification of Learned Multi-Agent Communication Policies via Decision Tree Distillation: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Predicting Mergeability of Parameter-Efficient Fine-Tuning Updates: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Measuring Curriculum Alignment across Topical Coverage, Competency, and Cognitive Depth: A Longitudinal Framework Applied to CS2013 and CS2023: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Deontic Policies for Runtime Governance of Agentic AI Systems: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Confidence is Not Reliability: Rethinking MC Dropout in Brain Tumour Segmentation: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
DreamReasoner-8B: Block-Size Curriculum Learning for Diffusion Reasoning Models: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.