Make agents use tools and reusable skills more reliably, Improve model reasoning, planning, and verification
Today tracks: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification.
This issue fetched and deduplicated 453 candidate papers from the 2026-08-10 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1DREAM Technical Report🔗
- 2Stealing Reasoning Traces from Proprietary LLM APIs🔗
- 3MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation🔗
- 4ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents🔗
- 5Entropy-based Code Adversarial Translation for Real-world Repository Migration🔗
- 6Multimodal Federated Learning under Dual-Axis Modality Missingness🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, retrieval, serving to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines
Keywordsagentragretrievalserving
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, code, reasoning to frame the reasoning and planning task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage
Keywordsagentragcodereasoning
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around benchmark, code, vision-language, fine-tuning to frame the data engineering task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding
Keywordsbenchmarkcodevision-languagefine-tuning
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, safety, benchmark, open-source to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API
Keywordsagentsafetybenchmarkopen-source
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, alignment, benchmark, code to frame the code intelligence task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: LLMs have demonstrated strong capabilities in code generation and automated program repair, but migrating an entire repository rarely produces a runnable application because long-horizon translation challenges LLM-based agents' ability to maintain repository-l
Keywordsagentalignmentbenchmarkcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, deployment, code, multimodal to frame the multimodal models task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Multimodal federated learning (FL) supports collaborative modeling in privacy-sensitive health-sensing and medical settings, but realistic deployments often exhibit dual-axis modality missingness: clients have different modality sets, and individual samples ma
Keywordsragdeploymentcodemultimodal
Code/DataCheck the source paper
Other papers worth tracking
GRASP: Granularity-Aware Region Alignment and Semantic Prototype Learning for Fine-Grained Cross-Modal Understanding in Drone Views: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Governing the KV Cache: Preventing Timing Side-Channel Leakage in Multi-Tenant LLM Inference: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Multi-Submap Implicit Neural SLAM with Local-to-Global Loop Closure for Large-Scale Scene Reconstruction: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
PolicyKG: An Agentic LLM Pipeline for Translating Institutional Policies into SHACL Knowledge Graphs: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
TeaMatch: Teachable Cross-Modal Representation Learning for 2D-3D Matching: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework: Covers a concrete multimodal models signal; useful as a follow-up candidate.
Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
SAFE-CHEM: Uncertainty-Aware Policy Switching for Robust Robotic Chemistry: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Bootstrapping Vision-Language Model for Hysteroscopic Surgical Scene Segmentation: Covers a concrete multimodal models signal; useful as a follow-up candidate.
Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Multimodal Model Diffing for Feature Discovery and Control: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Beyond Hazard Resemblance: Contrastive Event Adjudication for Training-Free Video Anomaly Detection: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
DistMoE: Private-data Rehearsal-free Routing in Mixture-of-Experts for Distributed Instruction Tuning: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
SHE: Trajectory-driven Safety Harness Evolution for LLM Agents: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Mismatch Matters: On-Policy Distillation Beyond Token Agreement: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Multi-Agent AI Safety as an Institutional Design Problem: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.