Make agents use tools and reusable skills more reliably, Make RAG retrieval and knowledge-base QA more reliable, Improve model reasoning, planning, and verification
Today tracks: make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable, improve model reasoning, planning, and verification.
This issue fetched and deduplicated 354 candidate papers from the 2026-09-09 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents🔗
- 2Cross-Species Animal Re-Identification with Semantic Consistency Learning🔗
- 3MOONWALK: Mediating Operations with Intent-Evidence-Action Alignment Across Junior-Supervisor Review Workflows in Animation/VFX Pre-Production🔗
- 4When Fusion Fails: Corruption-Aware Rebalanced Fusion for Multi-Modal Medical Image Segmentation🔗
- 5What Makes Adversarial Examples Transfer Across Deepfake Detectors?🔗
- 6Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable, improve model reasoning, planning, and verification. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, alignment, evaluation, benchmark to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Existing evaluation frameworks mostly assess only one part of AI agents, such as task completion (AgentBench) or security robustness (AgentDojo, ASB), rather than the complete pipeline of planning, tool selection, tool execution, memory and reasoning
Keywordsagentalignmentevaluationbenchmark
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, serving, evaluation, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Generalizable animal Re-Identification (ReID) aims to recognize individual animals across species with diverse morphologies and ecological contexts
Keywordsragservingevaluationcode
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around workflow, alignment, evaluation, code to frame the agents and tool use task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Animation and VFX pre-production review requires teams to translate loosely specified creative intent--briefs, evolving specifications, heterogeneous references, and verbal decisions--into revisions that junior artists can execute without repeated clarificatio
Keywordsworkflowalignmentevaluationcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, alignment, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Multi-modal medical image segmentation leverages complementary diagnostic information, yet fusion can underperform single-modality baselines when spatially aligned inputs differ in quality
Keywordsraginferencealignmentcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, evaluation, code, training to frame the training and post-training task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Deepfake detectors remain vulnerable to transfer-based black-box attacks, in which adversarial examples are generated on a source surrogate model and transferred to a target model, unknown to the attacker
Keywordsragevaluationcodetraining
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around alignment, safety, fine-tuning, jailbreak to frame the safety and alignment task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Large language model safety and security research is preoccupied with, among other things, detecting and preventing jailbreak attacks: alignment bypasses that allow an adversarial user to elicit unwanted or harmful outputs from models
Keywordsalignmentsafetyfine-tuningjailbreak
Code/DataCheck the source paper
Other papers worth tracking
3rd Place Solution to Human Motion Challenges in Real-World and Clinical Settings (MoCha) @ECCV2026: Language-Aligned Motion Representations for Domain-Generalizable UPDRS-Gait Severity Estimation: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
EEGBind: Detecting Source-Level Interictal Epileptiform Discharges via EEG-Centric Multimodal Binding: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Watermarks Without Verification: AI Text Watermarking After the EU AI Act: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
ProbPlug: A Plugin Uncertainty Network for Reliable Confidence in LLM Binary Classification: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
When Does Low-Bit Quantization Preserve the Decisions of Vector Search?: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Subgroup Membership Inference Audits of Differentially Private Synthetic Text: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Distilling Image Prototypes for Guided Test-Time Adaptation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Settling: Equilibrium Inference for Non-Convex Validity Sets: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
TEFM: Token-Efficient Faithful Modeling for Structured Data: Covers a concrete data engineering signal; useful as a follow-up candidate.
IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
ConvMem: Convolutional Memory for Long-Context Reasoning: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
Glyph: A Multi-Strategy Agentic System for Column Description and Sensitivity-Ontology Tagging of Enterprise Data Catalogs: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
PACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Cyber-Financial Contagion: Modeling the Propagation of an AI Vendor Compromise Through the Banking System: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Learning to Adapt and Calibrate: Score Distribution Alignment for Few-Shot Uncertainty Prediction in Medical VLMs: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.