Make RAG retrieval and knowledge-base QA more reliable, Make agents use tools and reusable skills more reliably, Test temporal consistency and motion realism in video generation
Today tracks: make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably, test temporal consistency and motion realism in video generation.
This issue fetched and deduplicated 363 candidate papers from the 2026-08-13 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ🔗
- 2Heterogeneity-Aware Belief Synchronization for Semantic Communication in AI-Native 6G Networks🔗
- 3Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation🔗
- 4AmalthAI: An Open-Source Computer Vision Platform for Cultural Heritage🔗
- 5Predictive Relative-Velocity Steering for Safe Robotic Manipulator Teleoperation in Dynamic Environments🔗
- 6Chance-constrained selection of sequential intervention strategies from counterfactual estimates🔗
What is worth tracking today
Today’s high-signal papers point to: make RAG retrieval and knowledge-base QA more reliable, make agents use tools and reusable skills more reliably, test temporal consistency and motion realism in video generation. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, evaluation, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Although multimodal large language models (MLLMs) have shown substantial potential in visual understanding and graphic code generation, editing scientific figures through code presents a greater challenge: a model must jointly recover visual structure, ground
Keywordsragevaluationbenchmarkcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, serving, alignment, code to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: 6G networks will not be serving as communication infrastructures only; rather, they are expected to evolve into intelligent systems, where thousands of autonomous artificial intelligence (AI) agents are interconnected
Keywordsagentservingalignmentcode
Code/DataCheck the source paper
test temporal consistency and motion realism in video generation
Signalthis paper targets the concrete research problem behind test temporal consistency and motion realism in video generation. It uses the title, abstract, and public signals around inference, serving, latency, alignment to frame the systems and deployment task, data, or evaluation flow to improve test temporal consistency and motion realism in video generation. The main claim is the title, abstract, and public signals indicate: Interactive autoregressive video generation demands both low-latency rollouts and precise online control
Keywordsinferenceservinglatencyalignment
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around inference, deployment, code, vision-language to frame the multimodal models task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Computer vision (CV) and machine learning (ML) offer new tools for cultural heritage (CH) artifact analysis, but the CV/ML pipeline remains largely inaccessible to CH domain experts, who lack the background to configure, train, or assess models
Keywordsinferencedeploymentcodevision-language
Code/DataCheck the source paper
test temporal consistency and motion realism in video generation
Signalthis paper targets the concrete research problem behind test temporal consistency and motion realism in video generation. It uses the title, abstract, and public signals around serving, latency, safety, motion to frame the video generation task, data, or evaluation flow to improve test temporal consistency and motion realism in video generation. The main claim is the title, abstract, and public signals indicate: Recent advances in teleoperation have enabled robotic manipulators to perform dexterous, human-arm-like motions
Keywordsservinglatencysafetymotion
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around safety, code, training, Training and Post-training to frame the training and post-training task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Many operational decisions are sequences of interventions under a cumulative resource limit, such as a maintenance schedule within a crew-hour budget
KeywordssafetycodetrainingTraining and Post-training
Code/DataCheck the source paper
Other papers worth tracking
Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Into the ORBIT for Time Series: Training Regimes for Foundation Models: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Incremental Evaluation and Training in Relational Deep Learning: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
RAGSieve: Self-Referenced Local Contrast for Knowledge-Poison Detection in Retrieval-Augmented Generation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reconcile Once, Write Anytime: A Trust-Tiered Librarian and a Multi-Agent Writer for Drift-Free, Point-in-Time Research: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
DiCoR: Decoupled Referent Disambiguation and Contour Recalibration for Efficient Referring Remote Sensing Image Segmentation: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
Bias Mitigation in Face Recognition via Demographic-based Supervised Contrastive Learning: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Understanding Backdoor Vulnerabilities in Vertical Federated Learning: The Gap Between Research and Practice: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
EGRL: Edge generation-guided relation-aware learning for RNA-protein interaction prediction: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Adaptive $k$ Nearest Neighbors Classifier via Granular Ball Computing: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Labels Are Not Endpoints: Treatment Leakage and Construct Validity in MCP Agent Security Evaluation: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
The Embedder's Dilemma: LLMs Are Better, but at What Cost?: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Semantic Steering for Controllable Generation: Tuning-Free Concept Erasure in Multimodal Diffusion Transformers: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
PIPES: Securing Agent Perception with Provenance and Priors: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.