Make RAG retrieval and knowledge-base QA more reliable, Make agents use tools and reusable skills more reliably, Improve model reasoning, planning, and verification
Today tracks: make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification.
This issue fetched and deduplicated 469 candidate papers from the 2026-07-30 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1ReToken: One Token to Improve Vision-Language Models for Visual Retrieval🔗
- 2OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models🔗
- 3Inducing language models to assert their own consciousness restores human beliefs and values🔗
- 4SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute🔗
- 5ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow🔗
- 6HARGO: Heterogeneity-Aware Reward-Guided Optimization for RL Post-Training of LLMs on HPC Tasks🔗
What is worth tracking today
Today’s high-signal papers point to: make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around retrieval, inference, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints
Keywordsretrievalinferencebenchmarkcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, alignment, evaluation, benchmark to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Computer-using agents (CUAs) are advancing rapidly across the digital world
Keywordsagentalignmentevaluationbenchmark
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around alignment, safety, fine-tuning, harmful to frame the safety and alignment task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values
Keywordsalignmentsafetyfine-tuningharmful
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around rag, inference, benchmark, reasoning to frame the reasoning and planning task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Scaling test-time computation can improve language-model reasoning, but uniform budgets waste computation on easy inputs, while verifier-guided refinement relies on external feedback
Keywordsraginferencebenchmarkreasoning
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, code, fine-tuning, video to frame the video generation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models
Keywordsragcodefine-tuningvideo
Code/DataCheck the source paper
identify and reduce safety, jailbreak, and alignment risks
Signalthis paper targets the concrete research problem behind identify and reduce safety, jailbreak, and alignment risks. It uses the title, abstract, and public signals around alignment, benchmark, fine-tuning, post-training to frame the training and post-training task, data, or evaluation flow to improve identify and reduce safety, jailbreak, and alignment risks. The main claim is the title, abstract, and public signals indicate: Supervised fine-tuning (SFT) can equip large language models (LLMs) with domain knowledge for high-performance computing (HPC) tasks such as data race detection and benchmark question answering
Keywordsalignmentbenchmarkfine-tuningpost-training
Code/DataCheck the source paper
Other papers worth tracking
Oracle-Budgeted Molecular Optimization with Short-Term Graph Memory: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Structural Validation of LLM-Generated Microservice Decompositions Using Source-Code Dependencies: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
CDAE: Enhancing Perturbation Robustness in Pretrained Language Models with Contrastive Denoising: Covers a concrete code intelligence signal; useful as a follow-up candidate.
EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Scaling Vision-Language Models Is Not Enough to Mitigate Bias: Covers a concrete multimodal models signal; useful as a follow-up candidate.
Old Tricks, New Models: How Simple Image Transformations Break Modern AI-based Content Moderation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
LM-GRASP: Instance-Specific Language Models for Combinatorial Construction via Online Imitation Learning: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
ConMem: Contribution-Aware Memory for Long-Horizon Manufacturing Inspection Logs: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
VIG-RL: Learning to Search and Insert for Verified Image Grounding: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
RaDiVe: Robust 4D Radar Odometry with Distance-Bounded NDT and Velocity-Discrepancy Point Uncertainty: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.