Make agents use tools and reusable skills more reliably, Improve model reasoning, planning, and verification
Today tracks: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification.
This issue fetched and deduplicated 339 candidate papers from the 2026-07-29 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1FARI: Robust One-Step Inversion for Watermarking in Diffusion Models🔗
- 2MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis🔗
- 3Constitutional Midtraining: Content Presence Drives Alignment Gains🔗
- 4Conformal Changepoint Localization and Root Cause Analysis with Corrupted Observations🔗
- 5SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response🔗
- 6Semantic-Aware Temporal Adaptation for UAV Anti-UAV Tracking🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around inference, evaluation, code, fine-tuning to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Inversion-based watermarking is a promising approach to authenticate diffusion-generated images, yet practical use is bottlenecked by inversion that is both slow and error-prone
Keywordsinferenceevaluationcodefine-tuning
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, benchmark, code to frame the code intelligence task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation
Keywordsagentragbenchmarkcode
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around rag, alignment, benchmark, code to frame the training and post-training task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Post-training alignment is often shallow, eroding under fine-tuning
Keywordsragalignmentbenchmarkcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, deployment, safety to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Detecting when the statistical behavior of an engineered system changes, and identifying which component is responsible, are core problems in the monitoring of telecommunication networks, robotic platforms, security infrastructure, and multi-agent systems
Keywordsagentragdeploymentsafety
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, workflow, benchmark, code to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities
Keywordsagentworkflowbenchmarkcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, alignment, benchmark to frame the video generation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: UAV Anti-UAV tracking is an emerging low-altitude security task for localizing an adversarial UAV using the onboard camera of a moving observer UAV
Keywordsraginferencealignmentbenchmark
Code/DataCheck the source paper
Other papers worth tracking
Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
PowerAtlas: Towards Electricity-Computing Co-Scheduling for Power Systems: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
NMKFR: A Robust Framework for Time-Aware Cold-Start Recommendation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Can AI agents conduct open-ended AI research? Early evidence from two case studies: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Minimal Markovization via Stable Quotients in Holonomy-Cover Decision Processes: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
When Knowledge Changes: Metamorphic Testing of RAG Systems with Mutations: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
FedTopo: Relation-Level Topology Sharing for Model-Heterogeneous Federated Learning: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Safety-Gated Autoscaling: A Multi-Layered Defense Architecture for Kubernetes Vertical Resource Optimization: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
HumanCLAW: Can Vision-Language Models Act Through a Body?: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
KAMR: Grounding Generation via Knowledge-Aligned Multi-hop Retrieval: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.