Make agents use tools and reusable skills more reliably, Improve model reasoning, planning, and verification
Today tracks: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification.
This issue fetched and deduplicated 999 candidate papers from the 2026-09-28 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1When Does Structured Knowledge Help Neural Theorem Proving?🔗
- 2eval-unlearn: Benchmarking unlearning in Text-to-Image Diffusion Models🔗
- 3Imprint Reader: From Weight-Update Readout to Behavioral Intervention🔗
- 4Improving Large Language Models for Code through Runtime Program-State Reasoning🔗
- 5Denoising Multi-Robot Trajectories🔗
- 6Rethinking Visual Token Compression for Video Large Language Models: A Simple Yet Strong Baseline🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, retrieval, code, fine-tuning to frame the retrieval and rag task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Does structured mathematical knowledge help LLMs prove theorems in Lean 4?
Keywordsagentretrievalcodefine-tuning
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around rag, inference, evaluation, benchmark to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: The rising number of concept unlearning techniques for text-to-image (T2I) diffusion models has produced a fragmented evaluation landscape
Keywordsraginferenceevaluationbenchmark
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around rag, safety, code, reasoning to frame the reasoning and planning task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: As language models take a growing role in AI development, a natural aspiration is for them to reflect on their own learning process, as humans do, and use that reflection to improve themselves
Keywordsragsafetycodereasoning
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, code, fine-tuning, post-training to frame the code intelligence task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Large language models receive limited explicit training in reasoning about runtime program states
Keywordsagentcodefine-tuningpost-training
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around rag, deployment, evaluation, code to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Multi-robot trajectory planning is a fundamental problem in multi-robot coordination but remains computationally challenging due to its nonconvex, multimodal, and high-dimensional nature
Keywordsragdeploymentevaluationcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, serving, compression to frame the systems and deployment task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Video Large Language Models (Video LLMs) have achieved remarkable progress in video understanding, but their inference efficiency is constrained by the large number of visual tokens produced by long videos
Keywordsraginferenceservingcompression
Code/DataCheck the source paper
Other papers worth tracking
d-OPD: Future-Aware On-Policy Distillation for Block Diffusion Language Models: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
SCBO: Semantically Coherent Batching and Ordering for LLM-Based Social Surveys: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
VEX-Bench: Benchmarking Verification Complexity of LLM-Generated Misinformation: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
TermJudge: A Document-Level Metric Judging, Not Counting, Terminology in Machine Translation Evaluation: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
ActionUNet: Improving Robustness of VLA Models with Efficient Multi-scale Fine-tuning: Covers a concrete video generation signal; useful as a follow-up candidate.
See it, Say it, Sorted: Mechanistic Diagnosis and Parameter-Space Mitigation of Emergent Misalignment in LLMs: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
SEAD: A State-Based Perspective on Attack and Defense in Tool-Using Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
HPMD: A Historical Persian Manuscript Dataset for Word Spotting with Line-Level Annotation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
ABC-Align: Prediction-Powered Alignment with Adaptive Bias Control: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Predictive Semantic Safety: From Visual Physical Reasoning to Safety-Critical Control: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors: Covers a concrete training and post-training signal; useful as a follow-up candidate.
AdaGuard: An Adaptive Guard Model with User-defined Policies: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Behavior-Grounded Semantic Enrichment for Financial Fraud Modeling and Reasoning: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Functional Autoencoders for Amplitude-Phase Representation Learning: Covers a concrete video generation signal; useful as a follow-up candidate.
What Does a Stream Model Buy You in Flow Matching?: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
GPUPhysBench: Benchmarking Coding Agents for Correct and Efficient GPU Physics Simulation: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Source-preserving alignment for robust evidence localization in scientific PDFS: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
EdgeVLN: Runtime-Aware Deployment Ready Quantized Vision Language Navigation Model: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.