Make RAG retrieval and knowledge-base QA more reliable, Make agents use tools and reusable skills more reliably, Improve model reasoning, planning, and verification
Today tracks: make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification.
This issue fetched and deduplicated 499 candidate papers from the 2026-06-16 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams🔗
- 2All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code🔗
- 3Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs🔗
- 4RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills🔗
- 5Do Generative Recommenders Deepen the Information Cocoon? A Closed-Loop Simulation with LLM-powered User Simulators🔗
- 6Temporal Preference Optimization for Unsupervised Retrieval🔗
What is worth tracking today
Today’s high-signal papers point to: make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around retrieval, inference, alignment, evaluation to frame the benchmarks and evaluation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Despite the remarkable progress of Video Large Language Models (Video-LLMs), current online architectures still struggle to simultaneously process continuous video streams, decide autonomously when to respond, and preserve long-horizon contextual memory
Keywordsretrievalinferencealignmentevaluation
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, code, coding, software to frame the code intelligence task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Software practitioners increasingly use AI coding agents that generate test code alongside production code in open source pull requests (PRs)
Keywordsagentcodecodingsoftware
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around rag, alignment, benchmark, multimodal to frame the video generation task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Reinforcement learning has improved the reasoning ability of large language models, but applying outcome-only rewards to video multimodal large language models (Video-MLLMs) provides limited guidance on which visual evidence should support the answer
Keywordsragalignmentbenchmarkmultimodal
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, deployment, alignment, evaluation to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: The LLM-empowered personal health agents with user health (sensor) metrics have offered a promising pathway to alleviate global disparities in healthcare access
Keywordsagentdeploymentalignmentevaluation
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, serving, code, agents to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Recommender systems alleviate information overload, yet repeated feedback between recommendations and user interactions can reinforce existing preferences and narrow users' exposure, forming information cocoons
Keywordsagentservingcodeagents
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, retrieval, alignment, code to frame the training and post-training task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Unsupervised dense retrievers offer scalability by learning semantic similarity from unlabeled documents via contrastive learning, but they struggle to capture the temporal relevance, retrieving semantically related but temporally misaligned documents-an impor
Keywordsragretrievalalignmentcode
Code/DataCheck the source paper
Other papers worth tracking
Spatio-Temporal Fusion Model for Standard View Classification of Echocardiographic Videos: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Monotonic Kolmogorov-Arnold Networks: A Theoretical and Empirical Study of Monotonicity as an Inductive Bias: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Flux-Guard: Facial Identity Protection using diffusion models: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
RSRank: Learning Relevance from Representational Shifts: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
CheckMIABench: Firm Foundations For Membership Inference Attacks on Language Models: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
TS-Fault: Benchmarking Time Series Forecasters Against Structural Faults: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Learning-Based Decision Making for Combustion Phasing Control in Multi-Fuel CI Engines with Latent Fuel Reactivity Estimation: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
EventDrive: Event Cameras for Vision-Language Driving Intelligence: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
The Stanford EDGAR Filings Dataset: Reconstructing U.S. Corporate and Financial Disclosures into Layout-Faithful and Token-Efficient Pretraining Data: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Knowledge Reutilization in Meta-Reinforcement Learning: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
MoonSplat: Monocular Online Gaussian Splatting with Sim(3) Global Optimization: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
When Multiple Scripts Matter: Evaluating ASR in Clinical Settings: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
MGUP: A Momentum-Gradient Alignment Update Policy for Stochastic Optimization: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
When Robots Sleep: Offline Skill Consolidation for Shared-Policy Robot Learning: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
StereoFactory: A Unified Merging Framework for Robust Stereo Matching: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Memory-Efficient Meta-Reinforcement Learning for Adaptive Safety-Critical Control in Adversarial Spacecraft Proximity Operations: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Do as the Romans Do: Learning Universal Behaviors from Heterogeneous Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.