Make RAG retrieval and knowledge-base QA more reliable, Improve model reasoning, planning, and verification, Make agents use tools and reusable skills more reliably
Today tracks: make RAG retrieval and knowledge-base QA more reliable, improve model reasoning, planning, and verification, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 374 candidate papers from the 2026-08-26 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings🔗
- 2OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora🔗
- 3Beyond Pairwise Feedback: Listwise Vision-Language Supervision for Preference-Based Reward Learning🔗
- 4When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs🔗
- 5PRISM: Projection-Integrated Sampling-Based MPC with Bayesian Cost Tuning for Bimanual Manipulation🔗
- 6A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks🔗
What is worth tracking today
Today’s high-signal papers point to: make RAG retrieval and knowledge-base QA more reliable, improve model reasoning, planning, and verification, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around workflow, rag, retrieval, deployment to frame the retrieval and rag task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Addressing critical global challenges, from food security and disaster risk to disease outbreaks and socio-economic vulnerability, demands high-fidelity geospatial modeling
Keywordsworkflowragretrievaldeployment
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around evaluation, benchmark, code, multimodal to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Multimodal Large Language Models (MLLMs) have demonstrated strong abilities in solving diverse visual and textual reasoning tasks
Keywordsevaluationbenchmarkcodemultimodal
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, vision-language, training to frame the training and post-training task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Vision-language models (VLMs) have emerged as a powerful source of supervision for reinforcement learning, enabling agents to leverage rich semantic knowledge during training
Keywordsagentragvision-languagetraining
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, serving, compression, code to frame the interpretability task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Sparse autoencoders (SAEs) are widely used to interpret the internal representations of large language models (LLMs), yet their reliability under post-hoc model compression remains poorly understood
Keywordsragservingcompressioncode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around evaluation, code, video, motion to frame the video generation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Bimanual manipulation in cluttered, contact-rich environments remains challenging because it requires coordinated motion generation, interaction-aware planning, and reliable execution under tight kinematic constraints
Keywordsevaluationcodevideomotion
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, serving, deployment, safety to frame the safety and alignment task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Large language models (LLMs) remain vulnerable to jailbreak attacks that exploit techniques such as role-playing, obfuscation, code transformation, and multi-step indirection to elicit harmful outputs
Keywordsagentservingdeploymentsafety
Code/DataCheck the source paper
Other papers worth tracking
Pointing the Way, Hiding the Destination: Practical Private Dense Retrieval at Scale: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
ClueWeaver: Reward-Guided Dual-Agent Evidence Reasoning for Compact LLMs on Literary Long Narratives: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Distance Is Not Enough: Forget-Retain Alignment Gap Predicts LLM Relearning Robustness: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
RSFusionDet: Underwater RGB-Sonar Multimodal Object Detection: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
ProgRouter: Online Progress-Guided Orchestration for Multi-Agent LLM Workflows under Quality-Cost Tradeoffs: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
MMJailBench: A Factorized Benchmark for Disentangling Multimodal Jailbreak Vulnerabilities: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
A Storage-Retrieval Gap in Parametric Knowledge Graph Memory: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models: Covers a concrete multimodal models signal; useful as a follow-up candidate.
Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
When Personality Meets Quantization: A Layer-wise MBTI Analysis of Quantized LLMs: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
SciMIF: Understanding Multimodal Instruction Following in Scientific Domains: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
PANDA - Prototype-Anchored Alignment for Partially Unpaired Multimodal Learning, with Applications to Alzheimers MRI and TCGA Pathology: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
TAU-Agent: An Agentic Retrieval-Augmented Framework for Traffic Anomaly Understanding: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Query-Side Attacks on GNN-Based KGQA: Tracing Failures from Entity Linking to Answer Generation: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
VISTA: Visually Inferred Spatial ConTact Attention for Contact-Rich Manipulation: Covers a concrete vision and image generation signal; useful as a follow-up candidate.
DEFUSE: Generalizable Backdoor Defense for Self-Supervised Encoders with Generative Priors: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Learning Continuous Regional Temperature Fields with Lead-Time and Resolution Queries: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Cooperative Multi-Agent Reinforcement Learning for Adaptive Aggregation in Semi-Supervised Federated Learning with non-IID Data: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Why Does Graph Learning Fail to Fully Benefit from a Text Teacher?: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.