Make RAG retrieval and knowledge-base QA more reliable, Make agents use tools and reusable skills more reliably
Today tracks: make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable.
This issue fetched and deduplicated 332 candidate papers from the 2026-07-07 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Industry Classification of GitHub Repositories Using the North American Industry Classification System (NAICS)🔗
- 2RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications🔗
- 3Energy-Efficient GPU DVFS for Fine-Tuning of SLMs on Resource-constrained Embedded Devices🔗
- 4KOAL: Knowledge-Driven Prostate Cancer Grading with Ordinal-Aware Learning🔗
- 5Breaking Spurious Correlations via Generative Randomization and Cross-Variant Self-Supervised Learning🔗
- 6SparseCtrl-HOI: Sparse Temporal Control for Human-Object Interaction Video Generation🔗
What is worth tracking today
Today’s high-signal papers point to: make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around retrieval, benchmark, code, open-source to frame the benchmarks and evaluation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: GitHub hosts hundreds of millions of public repositories, but the platform exposes no native mapping from repositories to standardized industry sectors
Keywordsretrievalbenchmarkcodeopen-source
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, benchmark, code, open-source to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Developers increasingly delegate real maintenance work to product-grade coding agents, and many state tasks in their native language, in the style of a customer request rather than a curated English issue
Keywordsagentbenchmarkcodeopen-source
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, benchmark, code to frame the code intelligence task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Dynamic Voltage Frequency Scaling (DVFS) on resource-constrained embedded GPU platforms is essential for energy-efficient small language model (SLM) fine-tuning, as privacy- and personalization-driven adaptation increasingly requires local execution and involv
Keywordsraginferencebenchmarkcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, alignment, code to frame the training and post-training task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Non-invasive prediction of Gleason Grade Group (GGG) in prostate cancer using multiparametric MRI (mpMRI) is clinically vital for reducing unnecessary biopsies
Keywordsraginferencealignmentcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, serving, benchmark, code to frame the vision and image generation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Deep neural networks trained with Empirical Risk Minimization (ERM) often fail under distribution shifts because they exploit spurious correlations between object labels and background context
Keywordsragservingbenchmarkcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, serving, evaluation, code to frame the video generation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Human-Object Interaction (HOI) video generation aims to synthesize realistic videos of humans manipulating diverse objects, serving as a promising avenue for AI-driven live streaming e-commerce
Keywordsragservingevaluationcode
Code/DataCheck the source paper
Other papers worth tracking
Verification of Dynamic Holographic Behavior in Identity Documents: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
From Application-Layer Simulation to Native Meta-Architecture: Structural Tension as an Endogenous Driver for Heterogeneous AI Evolution: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Tuning-Free Latent Diffusion Models for Ultrahigh-Resolution Image Editing: Covers a concrete vision and image generation signal; useful as a follow-up candidate.
Decoupled Single-Mask Annotation Noise Detection via Cross-Sectional Patch Self-Consistency: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Think Before You Grid-Search: Floor-First Triage for LLM Serving: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
AbICL: In-Context Learning for Antigen-Specific Antibody Affinity Ranking: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
The Large Cancer Assistant (LCA): A Model-Agnostic Orchestration Framework for Scalable Clinical Decision Support in Oncology: Covers a concrete multimodal models signal; useful as a follow-up candidate.
FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Prompt-Adapter Context Routing for Parameter-Efficient Multi-Shot Long Video Extrapolation: Covers a concrete video generation signal; useful as a follow-up candidate.
From Voting to Agent Collaboration: Answer-Type-Aware LLM Pipelines for BioASQ 14b: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
TILDE: TILt-based Distributional Erasure for Concept Unlearning: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Generalized Synthetic Image Detection with Enhanced RGB-Noise Representation Learning: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Bridging Diffusion Pruning and Step Distillation with Teacher-Aligned Repair: Covers a concrete training and post-training signal; useful as a follow-up candidate.
DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Synthetic-to-Real Translation for Class-Agnostic Motion Prediction: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Spider 2.0-AIFunc: Extending Real-World Text-to-SQL to AI-Native SQL Workflows: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.