Make RAG retrieval and knowledge-base QA more reliable, Improve model reasoning, planning, and verification, Make agents use tools and reusable skills more reliably
Today tracks: make RAG retrieval and knowledge-base QA more reliable, make RAG retrieval and knowledge-base QA more reliable, improve model reasoning, planning, and verification.
This issue fetched and deduplicated 382 candidate papers from the 2026-08-11 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1ConfTriage: A Calibration-Aware LLM Triage Framework for Pulmonary Nodule Malignancy with Selective Specialist Deferral🔗
- 2FedCGR: Federated Cross-Domain Generative Recommendation🔗
- 3Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian Splatting🔗
- 4Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition Agents🔗
- 5From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop🔗
- 6TACTICL: Task-Aware Compression of Tabular ICL Models🔗
What is worth tracking today
Today’s high-signal papers point to: make RAG retrieval and knowledge-base QA more reliable, make RAG retrieval and knowledge-base QA more reliable, improve model reasoning, planning, and verification. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, safety, code to frame the safety and alignment task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Pulmonary nodule malignancy prediction typically depends on image-trained specialist deep learning (DL) models that require substantial annotated imaging data and task-specific training
Keywordsraginferencesafetycode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, deployment, alignment, evaluation to frame the training and post-training task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Cross-domain recommendation (CDR) transfers preference knowledge across related domains, but federated deployment makes cross-domain alignment difficult because the behavioral anchors that align item spaces, such as overlapping users and shared interaction sig
Keywordsragdeploymentalignmentevaluation
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around evaluation, multimodal, vision-language, robot to frame the robotics and embodied ai task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Embodied mobile manipulation requires language, visual observations, three-dimensional scene structure, and action feasibility to be aligned before execution
Keywordsevaluationmultimodalvision-languagerobot
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, code, data to frame the data engineering task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expensive measurement -- and then must decide when the acquired signal is worth using
Keywordsagentragcodedata
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around rag, alignment, safety, search to frame the retrieval and rag task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static mode
Keywordsragalignmentsafetysearch
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around inference, compression, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: The strong performance of foundation models for tabular tasks comes at substantial inference costs
Keywordsinferencecompressionbenchmarkcode
Code/DataCheck the source paper
Other papers worth tracking
PEAK: Precise and Persistent Concept Erasure via k-Sparse Autoencoders: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
GitSkills: A Dataset of Agent Skills on GitHub: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
SQuaT: Self-Supervised Knowledge Distillation via Student-Aware Quantized Teacher Features: Covers a concrete training and post-training signal; useful as a follow-up candidate.
MEGA: Self-Evolving Agent Optimization Infrastructure via Wisdom Graph: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Are We Really Making Progress in Group Recommendation? Unmasking the Tie-Breaking Illusion: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
How to Verify Consistency of Probabilistic Claims: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Agentic Configuration Management (ACM): A Reference Configuration Model for Governed Agentic Systems: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
PRMU: A Corpus-Free Benchmark for Person-Centric Knowledge Unlearning in Multimodal Large Language Models: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
The Illusion of Cross-Lingual Safety in Low-Resource Languages: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Static in Frames, Dynamic in Events: Rethinking Features in Event Cameras as Motion Cues: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
CapProbe: Evaluating Detailed Image Captions via Full-Scene Dense Question Answering: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Mapping and Measuring the Behavioral Evolution of Large Language Models: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Data Attribution of Emergent Misalignment with Persona Features: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
TimeRoute: Time-Aware Modality Routing and Diffusion for Multi-Modal Recommendation: Covers a concrete vision and image generation signal; useful as a follow-up candidate.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.