Improve model reasoning, planning, and verification, Make agents use tools and reusable skills more reliably
Today tracks: improve model reasoning, planning, and verification, improve model reasoning, planning, and verification, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 385 candidate papers from the 2026-08-25 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos🔗
- 2Bridging Adversarial and Collaborative Learning for AI-Generated Image Quality Assessment🔗
- 3Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses🔗
- 4TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency Distillation🔗
- 5SeisMamba: Low-Latency Single-Station Seismic Magnitude Estimation for Spatially Distributed Earthquake Early Warning🔗
- 6SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction🔗
What is worth tracking today
Today’s high-signal papers point to: improve model reasoning, planning, and verification, improve model reasoning, planning, and verification, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around rag, inference, serving, alignment to frame the multimodal models task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Video multimodal large language models support language guided video segmentation, but they often show spatio temporal inconsistencies, e.g., jitter, drift, and identity switches
Keywordsraginferenceservingalignment
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around inference, alignment, benchmark, code to frame the multimodal models task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: AI-generated image quality assessment (AIGIQA) requires jointly reasoning about perceptual fidelity and prompt alignment, two quality dimensions that are often treated as independent in existing AIGIQA models
Keywordsinferencealignmentbenchmarkcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, benchmark, code, memory to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state and misalign skill invocation
Keywordsagentbenchmarkcodememory
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around inference, serving, deployment, latency to frame the systems and deployment task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Joint text-to-video-audio generation produces synchronized visual and acoustic content, but the long sampling trajectories and heterogeneous multimodal computation of large models make inference prohibitively expensive
Keywordsinferenceservingdeploymentlatency
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, deployment, latency, benchmark to frame the code intelligence task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Rapid earthquake magnitude estimation is central to earthquake early warning, yet many operational systems depend on dense regional seismic networks and region-specific calibration
Keywordsragdeploymentlatencybenchmark
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, alignment, evaluation to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: LLM agents can generate paper reproduction code, yet often produce scientifically unfaithful implementations
Keywordsagentragalignmentevaluation
Code/DataCheck the source paper
Other papers worth tracking
RecurSE: Bounded Recursive Self-Evaluation for LLM Rubric Judges: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Curved Inference II: Sleeper Agent Geometry - Extending Interpretability Beyond Probes: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
GlanceWAM: Sparse Test-Time Imagination for World-Action Models: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Meta$^n$: Recursive Self-Improvement through Emergent Depth: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Parameter-Efficient Self-Supervised Adaptation for EEG-FM under Fixed Computational Budgets: Covers a concrete training and post-training signal; useful as a follow-up candidate.
'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Rethinking Pre-Training and Augmentation for Zero-Shot Cross-City Object Detection: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Rules Before Oracles: Auditable, User-Configurable Argument Selection for Deliberative Polling: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Bellman Calibration for Marginalized Importance Weighting in Offline Reinforcement Learning: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Strictly Causal Streaming Video Anomaly Detection with a Theoretically-Grounded State-Space Core: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Expectation, Backlash, Recovery, and Excitement: How Model Releases Shape Reddit Perceptions of Conversational AI Systems: Covers a concrete code intelligence signal; useful as a follow-up candidate.
Hierarchical Prototype-Memory Adaptation of SAM for Surgical Instrument Segmentation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Dataset Scarcity Limits Robust Evaluation of Multilingual Embedding Models: A Case Study of Slavic Languages: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
SQLite is Enough. Lexical, Semantic, and Hybrid Search with scrydb: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Safety-aware Model Predictive Path Integral Control with Signal Temporal Logic: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
RefineRank: Joint Box Refinement and Ranking for Surgical Spatio-Temporal Grounding: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Do Robotic World Models Really Follow Actions? Diagnosing and Aligning Action-Conditioned Generation for Policy Learning: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.