Make RAG retrieval and knowledge-base QA more reliable, Make agents use tools and reusable skills more reliably, Improve code generation, execution feedback, and automated repair
Today tracks: make RAG retrieval and knowledge-base QA more reliable, make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 376 candidate papers from the 2026-08-24 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Better Retrieval, Worse Robustness:How Multi-hop RAG Amplifies Upstream ASR Errors🔗
- 2RAD: Rule-Augmented Relational Anomaly Detection🔗
- 3Dynamic Topic Modeling for Cross-Corpus Temporal Analysis🔗
- 4TEE-X: TEE-aware Acceleration Framework for Large Vision Models at the Edge🔗
- 5MetaCaster: Meta-Harness-Optimized Agent for End-to-End Few-Shot Learning of Lightweight Time Series Forecasters🔗
- 6OptiSight: Bridging Semantic Reasoning and Geometric Control for Embodied Navigation🔗
What is worth tracking today
Today’s high-signal papers point to: make RAG retrieval and knowledge-base QA more reliable, make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, retrieval, benchmark, code to frame the speech and audio task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Speech-based applications pass spoken queries through automatic speech recognition (ASR) before any retrieval module, so ASR errors enter the pipeline as a fixed upstream constraint
Keywordsragretrievalbenchmarkcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, serving, benchmark, code to frame the retrieval and rag task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Anomaly detection is often applied to data stored in relational databases, yet most existing methods require flattening multiple tables into a single feature matrix
Keywordsragservingbenchmarkcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around retrieval, serving, alignment, fine-tuning to frame the retrieval and rag task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Dynamic Embedded Topic Models (D-ETM) provide an interpretable framework for modeling temporal semantic evolution, but cross-corpus comparison remains difficult because topics are often learned independently and aligned only after training, a process that does
Keywordsretrievalservingalignmentfine-tuning
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around inference, latency, safety, memory to frame the systems and deployment task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Despite their remarkable success, machine learning models, particularly in vision applications, are alarmingly vulnerable to a range of security threats
Keywordsinferencelatencysafetymemory
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, deployment, multimodal, agents to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Time series forecasting (TSF) is evolving toward multimodal and agentic settings, yet using foundation models remains uneconomical in resource-constrained scenarios, where compact, specialized forecasters are more desirable
Keywordsagentdeploymentmultimodalagents
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around code, multimodal, vision-language, vlm to frame the multimodal models task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Autonomous indoor navigation requires both semantic understanding and precise geometric control
Keywordscodemultimodalvision-languagevlm
Code/DataCheck the source paper
Other papers worth tracking
Do Spoken Language Models Hear Speech as They Read Text? Bridging Structural Gaps Between Speech and Text: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Risk-Aware Reranking for Agentic Tool Retrieval: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Prime Agent: A Self-Improving RLM Harness: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
The Disconnect Between Better Descriptive Reasoning Trace Quality and Recommendation Effectiveness: Covers a concrete training and post-training signal; useful as a follow-up candidate.
AraDetox: A Multi-Dialect Arabic Detoxification Dataset: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
VeCAS: Vessel-Focused Contrast-Free Angiogram Synthesis for Vascular Interventions: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
ByteAction: Byte-space Action Recognition Foundation Model: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
On the Threat Model of Weird Generalization and Emergent Misalignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Traceable Spectral Inference via Influence Functions: Efficient Data Attribution and Error Proxies for the Ariel Mission: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Towards Actionable Surgical Team Dynamics: from Teamwork to Counterfactual Annotations: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
EviSafe: Evidence-Grounded Safety Evaluation for Vision-Language Models: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Grounding Free-Form Instructions for Fashion Complementary Image Generation: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
What Remains Normal? Clean Images Miss Useful Near-Defect Normal Patches for Anomaly Detection: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Spotter: Efficient Urban Visual Localization via Geo-Referenced Facade Landmarks in GPS-Degraded Environments: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.