Make agents use tools and reusable skills more reliably, Identify and reduce safety, jailbreak, and alignment risks, Make RAG retrieval and knowledge-base QA more reliable
Today tracks: make agents use tools and reusable skills more reliably, identify and reduce safety, jailbreak, and alignment risks, make RAG retrieval and knowledge-base QA more reliable.
This issue fetched and deduplicated 414 candidate papers from the 2026-09-22 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving🔗
- 2Calibration as a First-Class Criterion in LLM Evaluation🔗
- 3IndustrialVLA-Bench: A Traceable Multi-Axis Evaluation of Open Robot Policy Models🔗
- 4TriWorldBench: A Tri-View Consistency Perspective on Embodied World Models🔗
- 5Governed AI-Agent Coordination for Dementia Care: Architecture, Safety Contracts, and Evidence-Derived Workflow Verification🔗
- 6A Data-Interventional Framework for Auditing Privacy and Fairness in Generative Medical Imaging🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, identify and reduce safety, jailbreak, and alignment risks, make RAG retrieval and knowledge-base QA more reliable. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, inference, serving to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: We introduce SWE-Serve, a benchmark for evaluating agents on production inference engineering tasks
Keywordsagentraginferenceserving
Code/DataCheck the source paper
identify and reduce safety, jailbreak, and alignment risks
Signalthis paper targets the concrete research problem behind identify and reduce safety, jailbreak, and alignment risks. It uses the title, abstract, and public signals around deployment, alignment, evaluation, benchmark to frame the benchmarks and evaluation task, data, or evaluation flow to improve identify and reduce safety, jailbreak, and alignment risks. The main claim is the title, abstract, and public signals indicate: Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP
Keywordsdeploymentalignmentevaluationbenchmark
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, deployment, latency to frame the robotics and embodied ai task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Open robot policies increasingly follow two paradigms: vision-language-action models (VLAs) directly map observations and instructions to actions, whereas world-action models (WAMs) incorporate learned video or world dynamics into policy learning or action gen
Keywordsraginferencedeploymentlatency
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around alignment, evaluation, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Embodied world models predict the outcomes of robot actions to support learning and planning
Keywordsalignmentevaluationbenchmarkcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, workflow, serving, safety to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Dementia care increasingly involves connected sensors, medication devices, electronic records, and assistive technologies
Keywordsagentworkflowservingsafety
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around retrieval, code, synthetic data, frame to frame the video generation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Diffusion-based synthetic data generation offers a promising route for sharing medical imaging data without releasing sensitive patient records
Keywordsretrievalcodesynthetic dataframe
Code/DataCheck the source paper
Other papers worth tracking
Testing-Driven Reliability Audit of Trajectory-Based Early Outcome Prediction for LLM Agents: Target-Specific Calibration Transfer Persists Within a Single Benchmark: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
EquivSVA: A Formally Verified Dataset of Behavioral Assertions Across Equivalent RTL Implementations: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Beyond End-Task Success: How to Audit Visual Experience Retrieval in Robotics: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
SafeLoop: Risk-Aware Rollback for Vision-Language-Action Manipulation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Certified Against Which Oracle? Execution Labels Set the Reported Risk of Conformal Abstention for Text-to-SQL: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Shallow to Deep: Aligning Token Pruning with Stage-wise Roles in LVLMs: Covers a concrete multimodal models signal; useful as a follow-up candidate.
HABILIS Brain 0: Geometry-Change Supervision for Vision-Language-Action and Residual Flow Recovery: Covers a concrete robotics and embodied ai signal; useful as a follow-up candidate.
The Sirens' Song: When Proximal Background Context Overshadows Distant Evidence: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Optimal Sequential Annotations for Off-Policy Evaluation: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
GTR: Gated Token Recurrence for Efficient Dense Prediction: Covers a concrete vision and image generation signal; useful as a follow-up candidate.
Combining Hierarchical Cognitive Process with Process Supervision for Interpretable Scene Safety Understanding: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
KwaiMind Technical Report: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
PACT: From Credit Assignment to Critic Alignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Towards Systematic Qualification of Vision-Language Models for Automotive Perception Systems: Covers a concrete multimodal models signal; useful as a follow-up candidate.
When Does Execution Provenance Help Agent Memory Retrieval?: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Risk-Aware Online Conformal State Probing: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
LoRango: It Takes Two LoRAs to Unlock Hidden Behaviors in Diffusion Models: Covers a concrete vision and image generation signal; useful as a follow-up candidate.
AgenticSizing: A Large Language Model-based Multi-Agent Framework for Analog Circuit Sizing: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.