Make agents use tools and reusable skills more reliably, Identify and reduce safety, jailbreak, and alignment risks, Improve code generation, execution feedback, and automated repair
Today tracks: make agents use tools and reusable skills more reliably, identify and reduce safety, jailbreak, and alignment risks, improve code generation, execution feedback, and automated repair.
This issue fetched and deduplicated 503 candidate papers from the 2026-08-04 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Accelerating Dynamic Graph Clustering on GPU Architectures with cuGraph🔗
- 2How Closely Do LLM Reviews Align with Human Peer Review?🔗
- 3JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion🔗
- 4PRISM: Powerful Time Series to Image (TS2I) Representations for Multivariate Anomaly Detection🔗
- 5SeqLLM: Augmenting LLMs with Behavioral-Sequence Modeling for High-Stakes Decisions at WeChat Pay🔗
- 6ATLAS: Learning to Recommend Across Unseen Domains🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, identify and reduce safety, jailbreak, and alignment risks, improve code generation, execution feedback, and automated repair. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around serving, code, synthetic data, open-source to frame the data engineering task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: This work addresses community detection in temporal networks through GPU-accelerated extensions of spectral clustering and modularity-based algorithms originally designed for static graphs
Keywordsservingcodesynthetic dataopen-source
Code/DataCheck the source paper
identify and reduce safety, jailbreak, and alignment risks
Signalthis paper targets the concrete research problem behind identify and reduce safety, jailbreak, and alignment risks. It uses the title, abstract, and public signals around alignment, evaluation, eval, Benchmarks and Evaluation to frame the benchmarks and evaluation task, data, or evaluation flow to improve identify and reduce safety, jailbreak, and alignment risks. The main claim is the title, abstract, and public signals indicate: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting
KeywordsalignmentevaluationevalBenchmarks and Evaluation
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around inference, serving, latency, evaluation to frame the video generation task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency
Keywordsinferenceservinglatencyevaluation
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around workflow, rag, evaluation, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Time series anomaly detection (TSAD) underpins applications in predictive maintenance, finance, and cloud computing, however performance remains sensitive to representation choices, especially in multivariate settings
Keywordsworkflowragevaluationcode
Code/DataCheck the source paper
test temporal consistency and motion realism in video generation
Signalthis paper targets the concrete research problem behind test temporal consistency and motion realism in video generation. It uses the title, abstract, and public signals around serving, alignment, benchmark, fine-tuning to frame the safety and alignment task, data, or evaluation flow to improve test temporal consistency and motion realism in video generation. The main claim is the title, abstract, and public signals indicate: Merchant risk control at large payment platforms screens tens of millions of merchants daily, where false positives harm legitimate merchants and false negatives leave harmful activity undetected
Keywordsservingalignmentbenchmarkfine-tuning
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, alignment, code, knowledge to frame the retrieval and rag task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Recommender systems remain domain-bound: a model trained on one interaction environment typically requires retraining or target-domain adaptation before it can operate on a new catalogue
Keywordsragalignmentcodeknowledge
Code/DataCheck the source paper
Other papers worth tracking
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Binding: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Latent Reward Registers for Diffusion Preference Alignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Risky Business: Measuring The Faithfulness-Safety Tension: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Shielding for Higher-Order Safety: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Policy Fragmentation or Institutional Alignment? Institutional Governance of AI in Universities and Business Schools: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Solver-Aware Decompositions for Programming-by-Example: When Dividing Requires Knowing how to Conquer: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
Test-Time Scaling for Safe Text-Guided Image Generation via Intermediate Clean Estimates: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
CIGTSurv: Clinical Information Guided Tri-modal Survival Prediction with Local Prototype Association and Global Feature Alignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Channel-wise Dynamic Knowledge Distillation via Adaptive Sample Generation for Action Recognition: Covers a concrete training and post-training signal; useful as a follow-up candidate.
CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generation: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Clinically-Grounded Hierarchical Classification for Consistent Chest X-ray Interpretation: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems via Deterministic Side Channels: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs: Covers a concrete training and post-training signal; useful as a follow-up candidate.
TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.