Make agents use tools and reusable skills more reliably, Make RAG retrieval and knowledge-base QA more reliable, Improve model reasoning, planning, and verification
Today tracks: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable.
This issue fetched and deduplicated 402 candidate papers from the 2026-08-17 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1GEO-Flag: Detecting and Measuring GEO-Optimized Web Content🔗
- 2Orbit-Planner: Towards Latent World Models for On-Orbit Obstacle Avoidance of Satellite Agents🔗
- 3SIGMA-Lane: Scale-pyramId Gated MAmba for Temporally Consistent Video Lane Detection🔗
- 4Seeing Before Answering: Training-Free Visual Layer Profiling for Vision-Language Models🔗
- 5PixRestore: Unified Image Restoration via Pixel Diffusion Transformer🔗
- 6Think Inside the Chunk: RegulaRAG for Regulation-Compliant Scenario Generation using LLMs: A Case Study of UN Regulation No. 152🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, retrieval, evaluation, benchmark to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Generative Engine Optimization (GEO) modifies web content to increase its likelihood of being selected and cited by generative search engines
Keywordsagentretrievalevaluationbenchmark
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, code, agents, Agents and Tool Use to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Satellite agents for on-orbit navigation tasks need to predict collision risks using limited onboard observations
KeywordsagentcodeagentsAgents and Tool Use
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around retrieval, alignment, memory, video to frame the video generation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Video lane detection requires predictions that remain stable across frames, yet severe vehicle occlusions can break temporal cues
Keywordsretrievalalignmentmemoryvideo
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around rag, inference, alignment, benchmark to frame the multimodal models task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: LLaVA-style Vision-Language Models (VLMs) pass visual tokens from a fixed late layer of the vision backbone, typically the penultimate one, to the language model
Keywordsraginferencealignmentbenchmark
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, serving, benchmark to frame the training and post-training task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Unified image restoration (UIR) aims to recover high-quality (HQ) content from low-quality (LQ) images with different degradations using a single model
Keywordsraginferenceservingbenchmark
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, retrieval, safety, search to frame the retrieval and rag task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Generating regulation-compliant test scenarios is essential for validating safety-critical automotive systems, yet Large Language Models (LLMs) struggle to ground outputs in long, hierarchical standards
Keywordsragretrievalsafetysearch
Code/DataCheck the source paper
Other papers worth tracking
Toward Better Assessment of LLMs' Performance in Clinical Error Detection: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
CACSurv: Concordance-Aligned Comparative Learning with Large Language Models for Cancer Survival Prediction: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization: Covers a concrete training and post-training signal; useful as a follow-up candidate.
JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Beyond Similarity Matching: Structured Reasoning for Open-Vocabulary Referring Segmentation in 3DGS: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Bridging the Gap between Labeled and Unlabeled Data via Unified Flow with Feature Memory Bank: Covers a concrete vision and image generation signal; useful as a follow-up candidate.
HarmTrace: Anchor-Calibrated Decoupled Optimization for Fine-Grained Target Identification in Harmful Memes: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
MLLM-Guided Semantic Correction for Text-to-Video Generation: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Towards Risk-free AI Agent Deployment: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
What Does Context Compression Cost an Agent? Interaction Costs Unrevealed by Task-Completion Metrics: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Unlocking Motion in Expressions: Temporal Calibration for Referring Video Object Segmentation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Multi-Granularity Sentiment Integration for LLM-Based Multimodal Sentiment Analysis: Covers a concrete video generation signal; useful as a follow-up candidate.
Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
KeyID: Decoupled Drafting and Keyframe Editing for Identity-Preserving Video Generation: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
Spatial Temporal Synergy: Balancing Change and Invariance in Text Driven 3D Human Motion Editing: Covers a concrete video generation signal; useful as a follow-up candidate.
Non-Crossing Deep Quantile Regression for Distributional Survival Prediction: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
UniDot: A Unified Network for Sequence Modeling and Feature Interaction in Large-scale Recommendation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Revisiting Classifier-Free Guidance Methods in Latent Diffusion Models: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.