Make RAG retrieval and knowledge-base QA more reliable, Make agents use tools and reusable skills more reliably
Today tracks: make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 386 candidate papers from the 2026-06-25 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1TGHE: Template-based Graph Homomorphic Encryption for Privacy-Preserving GNN Inference in Edge-Cloud Systems🔗
- 2LAMP: Lane-Aligned Motion Primitives for Feasible Trajectory Prediction🔗
- 3Autoformalization of Agent Instructions into Policy-as-Code🔗
- 4LayersReg: A Layer-by-Layer Progressive Regressor for Reliable Intraoperative 3D/2D Registration🔗
- 5SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference🔗
- 6IDEA: Insensitive to Dynamics Mismatch via Effect Alignment for Sim-to-Real Transfer in Multi-Agent Control🔗
What is worth tracking today
Today’s high-signal papers point to: make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, serving, evaluation to frame the benchmarks and evaluation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Existing homomorphic encryption (HE)-based GNN systems adopt a graph-centric paradigm that couples per-query cost to global graph size, limiting evaluations to at most ~20k nodes and making them incompatible with dynamic, large-scale financial graphs
Keywordsraginferenceservingevaluation
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around serving, safety, code, multimodal to frame the video generation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Motion forecasting is essential for autonomous driving systems to enable safe decision-making and planning in complex driving scenarios
Keywordsservingsafetycodemultimodal
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, safety, benchmark, code to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Agent safety in high-stakes domains requires formal policy enforcement, but most existing approaches either rely on probabilistic guardrails (fine-tuned classifiers, prompt-based steering) that offer no formal guarantees, or on hand-coded symbolic enforcement
Keywordsagentsafetybenchmarkcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around retrieval, alignment, multimodal, memory to frame the retrieval and rag task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: 3D/2D registration serves as a cornerstone technique in surgical navigation
Keywordsretrievalalignmentmultimodalmemory
Code/DataCheck the source paper
strengthen multimodal understanding of charts, documents, and visual evidence
Signalthis paper targets the concrete research problem behind strengthen multimodal understanding of charts, documents, and visual evidence. It uses the title, abstract, and public signals around inference, serving, latency, compression to frame the systems and deployment task, data, or evaluation flow to improve strengthen multimodal understanding of charts, documents, and visual evidence. The main claim is the title, abstract, and public signals indicate: Low-bit floating-point formats and semi-structured sparsity are increasingly supported by modern accelerators, yet combining them for LLM activation compression remains challenging: activations contain input-dependent outliers that dominate block scales in FP4
Keywordsinferenceservinglatencycompression
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, deployment, alignment to frame the training and post-training task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Complex multi-agent control tasks remain challenging for traditional rule-based and model-based approaches, motivating the adoption of learning-based methods
Keywordsagentragdeploymentalignment
Code/DataCheck the source paper
Other papers worth tracking
Agents That Know Too Much: A Data-Centric Survey of Privacy in LLM Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Can Large Language Models Reliably Code Qualitative Humanitarian Data? A Benchmark Study Against Human Expert Adjudication: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
NeuraDock Visual Cognitive Load Agent Tutorial: A Quality-Gated Open-Source EEG Workflow for Alpha Dynamics and Real-Time Applications: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Speaking Numbers to LLMs: Multi-Wavelet Number Embeddings for Time Series Forecasting: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Don't Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
NOVA: A Verification-Aware Agent Harness for Architecture Evolution in Industrial Recommender Systems: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline): Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Mostly Automatic Translation of Language Interpreters from C to Safe Rust: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
TMP: Tree-structured Mixed-policy Pruning for Large-scale Image Generation and Editing: Covers a concrete vision and image generation signal; useful as a follow-up candidate.
ShareLock: A Stealthy Multi-Tool Threshold Poisoning Attack Against MCP: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
MinGram: A Minimalist Unigram Tokenizer with High Compression and Competitive Morphological Alignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Adaptive Utility driven Resource Orchestration for Resilient AI (AURORA-AI): Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs: Covers a concrete code intelligence signal; useful as a follow-up candidate.
In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics: Covers a concrete video generation signal; useful as a follow-up candidate.
How Much Static Structure Do Code Agents Need? A Study of Deterministic Anchoring: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
RedVox: Safety and Fairness Gaps in Speech Models Across Languages: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Learning to Recover Task Experts from a Multi-Task Merged Model: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.