Make agents use tools and reusable skills more reliably, Test temporal consistency and motion realism in video generation, Improve code generation, execution feedback, and automated repair
Today tracks: make agents use tools and reusable skills more reliably, test temporal consistency and motion realism in video generation, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 399 candidate papers from the 2026-08-27 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1BekchiAI: Measuring, Observing, and Controlling LLM Agents in One Click🔗
- 2Robust Neural Stimulation Response Modeling Through Meta-Learning and Pretraining🔗
- 3INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment🔗
- 4TempJail: Temporal Jailbreak Attacks against Image-to-Video Generation Models🔗
- 5Instruction Quality Matters: Refining Instructions for Effective Preference Learning🔗
- 6When Privacy Hurts Mergeability: Geometry-Aware Model Merging under Differential Privacy🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, test temporal consistency and motion realism in video generation, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, serving, latency, evaluation to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Large language model agents reason, call tools, and act autonomously over many steps, but their agentic skills-correctly sequencing tools, planning under dependencies, judging untrusted inputs, and grounding generated arguments-are hard to measure with accurac
Keywordsagentservinglatencyevaluation
Code/DataCheck the source paper
test temporal consistency and motion realism in video generation
Signalthis paper targets the concrete research problem behind test temporal consistency and motion realism in video generation. It uses the title, abstract, and public signals around deployment, eval, dataset, test to frame the benchmarks and evaluation task, data, or evaluation flow to improve test temporal consistency and motion realism in video generation. The main claim is the title, abstract, and public signals indicate: Objective: Model-based closed-loop neural stimulation holds promise for therapeutic applications ranging from Parkinson's disease to sensory restoration, but deployment has been limited by two obstacles: 1) forecasting models for predicting the consequences of
Keywordsdeploymentevaldatasettest
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, alignment, safety, code to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: As large language models (LLMs) are deployed as autonomous agents, safety failures increasingly involve consequential actions
Keywordsagentalignmentsafetycode
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around inference, serving, safety, evaluation to frame the video generation task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: In recent years, image-to-video (I2V) generation models have made remarkable progress in subject consistency and temporal coherence, enabling high quality video synthesis
Keywordsinferenceservingsafetyevaluation
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around alignment, benchmark, code, data to frame the data engineering task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Preference learning optimizes models using response pairs, yet the informativeness of these pairs is fundamentally shaped by the instructions from which they are generated
Keywordsalignmentbenchmarkcodedata
Code/DataCheck the source paper
test temporal consistency and motion realism in video generation
Signalthis paper targets the concrete research problem behind test temporal consistency and motion realism in video generation. It uses the title, abstract, and public signals around serving, alignment, fine-tuning, rl to frame the training and post-training task, data, or evaluation flow to improve test temporal consistency and motion realism in video generation. The main claim is the title, abstract, and public signals indicate: Model merging promises to construct a single multi-task model from independently fine-tuned task models without accessing the original task data
Keywordsservingalignmentfine-tuningrl
Code/DataCheck the source paper
Other papers worth tracking
RuleWeaver: Benchmarking Rule-Centered Scenario Reasoning for Large Language Models: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Your Voice Cloning System is Secretly a Voice Anonymizer: Covers a concrete speech and audio signal; useful as a follow-up candidate.
MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Multi-Image Visual Token Pruning in Large Visual Language Models: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
A Unified Descriptive-Complexity Framework for Model Selection under Correlated Designs: Covers a concrete code intelligence signal; useful as a follow-up candidate.
SPEAR: Distilling Domain-Adaptive Reasoning Skeletons via Sequential Symbolic Alignment in Reinforcement Learning: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
HUG-VIS: A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reconstructing Humans and Objects in Interaction using Large Reconstruction Models: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
BTS-AgentBench: A Deterministic, Replayable Pipeline from Read-Only Telemetry Logs to Agent Benchmarks: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Verify Smarter, Evolve Further: Efficient Harness Evolution through Behavior-Aware Verification: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
A Point-of-Prescription Safety-Check System for Adverse Drug Reactions in Rural Bangladeshi Hospitals: A Feasibility Study: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
Unsupervised Adaptation of 3D CT Foundation Models for 3D CBCT Segmentation: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
SSMB: Self-Supervised Local Feature Detection under Motion Blur: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Parameter-Efficient pretrained-CT-to-MRI Transfer for Rectal Cancer Segmentation: Performance-Calibration Trade-offs: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Ancient-Bench: A Comprehensive Multi-millennial, Multi-medium, and Multi-script Benchmark for Ancient Chinese Artifact Text Recognition: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
ANTShapes Benchmarking Datasets for Event-Based Neuromorphic Object Classification: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.