Test temporal consistency and motion realism in video generation, Make agents use tools and reusable skills more reliably, Improve code generation, execution feedback, and automated repair
Today tracks: test temporal consistency and motion realism in video generation, make agents use tools and reusable skills more reliably, improve code generation, execution feedback, and automated repair.
This issue fetched and deduplicated 330 candidate papers from the 2026-08-28 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia🔗
- 2AGENT-O: A Semantic Agent Card Framework for Interoperable and Governed Healthcare AI Agents🔗
- 3ZipMVS: Multi-View Stereo with Compressed Cost Volumes🔗
- 4FedEHR-Agents: Federated Agentic Optimization for Automated EHR Modeling🔗
- 5MaCoPlanner: LLM-Assisted Manual-Compiled Task Planning with Proactive Safety Verification for Robotic Industrial Panel Operation🔗
- 6Synth-JDoc: Synthesizing a Japanese Document Image Dataset for OCR with Diverse Layouts and Embedded Images🔗
What is worth tracking today
Today’s high-signal papers point to: test temporal consistency and motion realism in video generation, make agents use tools and reusable skills more reliably, improve code generation, execution feedback, and automated repair. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
test temporal consistency and motion realism in video generation
Signalthis paper targets the concrete research problem behind test temporal consistency and motion realism in video generation. It uses the title, abstract, and public signals around safety, evaluation, benchmark, fine-tuning to frame the benchmarks and evaluation task, data, or evaluation flow to improve test temporal consistency and motion realism in video generation. The main claim is the title, abstract, and public signals indicate: Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common use case: users seeking practical help over multiple turns in culturally grounded scenarios
Keywordssafetyevaluationbenchmarkfine-tuning
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, workflow, deployment, alignment to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: AGENT-O is a modular ontology framework that defines a semantic Agent Card for representing health-oriented AI agent systems and supports assessment of reporting completeness in scientific publications
Keywordsagentworkflowdeploymentalignment
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around serving, deployment, compression, code to frame the systems and deployment task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Multi-view stereo (MVS) methods typically deliver highly accurate 3D reconstructions from multiple registered RGB images, thanks to the highly informative, geometric constraints between them
Keywordsservingdeploymentcompressioncode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, workflow, serving, evaluation to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Recent advances in large language models are enabling autonomous clinical agents to perform increasingly complex electronic health record (EHR) modeling workflows
Keywordsagentworkflowservingevaluation
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around deployment, safety, evaluation, execution to frame the code intelligence task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Robotic industrial panel operation requires not only accurate control localization but also compliance with operating procedures, safety rules, and device-state constraints distributed across heterogeneous manuals
Keywordsdeploymentsafetyevaluationexecution
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, evaluation, code, synthetic data to frame the multimodal models task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: The ability of Large Vision Language Models (LVLMs) to read text within document images is crucial, as it enables various applications such as Document Visual Question Answering
Keywordsragevaluationcodesynthetic data
Code/DataCheck the source paper
Other papers worth tracking
Dual-Stream Semantic Guidance with Prototype Anchor Calibration for Source-Fully-Free Adaptation of Vision-Language Models: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Attribute Token Arithmetic: Disentangled and Continuous Semantic Control for Visual Autoregressive Models: Covers a concrete vision and image generation signal; useful as a follow-up candidate.
QUORUM: QUality-Optimized Routing Using Multiple annotators: Covers a concrete data engineering signal; useful as a follow-up candidate.
Information-Guided Selective Modality-Interest Alignment for Multimodal Recommendation: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Cross-Session Decomposition Attacks: Scaling Risk and Intent-Aligned Retrieval Defense: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
A Deep Learning-Based Stacking Ensemble Framework for Turbofan Engine Remaining Useful Life Prediction: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
ITER: Interaction-Aware Retrieval for Agentic Search: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Phoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment Evaluation: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
COVER: Identifiable Evaluation of Coalition Routing: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
ARC-CT: Anatomy-Routed Contrastive Vision-Language Learning for 3D Chest CT: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
GeoFF3D: Coordinate-Anchored Feed-Forward Reconstruction for Large-Scale UAV Mapping: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Beyond Global Scalars: Synergizing Token-Level Statistics and Deep Semantics for Adversarial AIGC Text Detection: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Anchored Scenario Coverage for Failure-Aware First-Hit Batch Inverse Design: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation: Covers a concrete data engineering signal; useful as a follow-up candidate.
LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Prompt-Guided Interactive Segmentation of Interstitial Lung Disease in Thoracic CT: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
LongPIBench: A Long-Context Benchmark for Prompt Injection: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.