Make agents use tools and reusable skills more reliably, Improve model reasoning, planning, and verification, Make RAG retrieval and knowledge-base QA more reliable
Today tracks: make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification, make RAG retrieval and knowledge-base QA more reliable.
This issue fetched and deduplicated 468 candidate papers from the 2026-09-17 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Chronicle: Cut-Point Replay for Regression Testing of LLM Agents🔗
- 2Reasoning Quality Matters: Combating Reasoning Collapse in LLM-based Embedding Learning🔗
- 3SCGFM-ART: Amortized Relational Transport for Structure-Centric Graph Foundation Models🔗
- 4Stress-testing Alignment Midtraining🔗
- 5QoS-Aware Federated Learning for Multimodal In-Cabin Interaction in Smart Vehicles🔗
- 6Marginal utility, matrix factorization, and the Key-Value (KV) cache: a unified information-economic framework for sovereign geo-mining inference🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification, make RAG retrieval and knowledge-base QA more reliable. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, inference, benchmark, code to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Large language model responses are non-deterministic, so failures in LLM agents are hard to reproduce: a failure depends on inference that is not bitwise reproducible, on tools that read changing state, and on a multi-step trajectory that a re-run rarely repea
Keywordsagentinferencebenchmarkcode
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around rag, retrieval, serving, alignment to frame the retrieval and rag task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Large Language Models (LLMs) have recently shown strong potential for producing context-rich text embeddings for retrieval
Keywordsragretrievalservingalignment
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, alignment, benchmark to frame the benchmarks and evaluation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Graph foundation models (GFMs) aim to learn transferable representations across severely heterogeneous graph domains
Keywordsraginferencealignmentbenchmark
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, deployment, alignment, post-training to frame the benchmarks and evaluation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: When aligning frontier models through post-training techniques, it is not possible to directly demonstrate all of the behaviours we want a model to exhibit in all possible deployment environments; our model must generalise outside of the post-training distribu
Keywordsragdeploymentalignmentpost-training
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, deployment, latency, safety to frame the training and post-training task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Modern smart vehicles leverage multimodal sensors, ranging from high-bandwidth vision systems to low-rate physiological monitors, to provide personalized in-cabin services
Keywordsragdeploymentlatencysafety
Code/DataCheck the source paper
strengthen multimodal understanding of charts, documents, and visual evidence
Signalthis paper targets the concrete research problem behind strengthen multimodal understanding of charts, documents, and visual evidence. It uses the title, abstract, and public signals around inference, latency, compression, benchmark to frame the systems and deployment task, data, or evaluation flow to improve strengthen multimodal understanding of charts, documents, and visual evidence. The main claim is the title, abstract, and public signals indicate: This paper builds a theoretical bridge between the economic notion of marginal utility and two machine-learning constructs, matrix factorization and the Key--Value cache of transformer language models
Keywordsinferencelatencycompressionbenchmark
Code/DataCheck the source paper
Other papers worth tracking
PointEvent: Rethinking Event-based Tiny Object Detection via Serialized Motion Evidence Accumulation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
GRF-Recon: Global Ray-Field Optimization for Long-Sequence Feed-forward Reconstruction: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Benchmarking LLM Compliance with China AI Generated Content Regulations: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
From "Who Is This User?" to "What Does This Purchase Mean?": A Deployed Pipeline for Semantic User Profiling at Bank Scale: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
ClashBench: Conflicts Leading Agents to Seize and Harm: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Beyond Flattened Tokens: Structure-Preserving EEG Decoding with Reusable TriDim Blocks: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Absence is Presence: Understanding Visual Scene Negative Events Under Safety Cognitive Constraint: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Learn Before You Judge: Progressive Knowledge-to-Decision Alignment for Explainable Hateful Meme Detection: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Benchmarking MLLMs via Cognitive Expected Scene Graph for Safety-Critical Visual Negation Understanding: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Federated Learning Framework for Privacy-Preserving Kidney Stone Detection: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Opinion Dynamics-based Coalition Formation for Federated Learning in Heterogeneous IoT Systems: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
TacSushi: Tactile-Grounded World-Action Modeling for Dexterous Sushi Manipulation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
FootprintRAG: Visual Analytics for Evidence Context Refinement in RAG-based Scientific Literature Exploration: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Form Over Content In Gradient-Based Data Attribution Methods: Covers a concrete training and post-training signal; useful as a follow-up candidate.
AMB3R-SLAM: Kilometer-scale SLAM with Hierarchical Backend: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
NS3Learn: Transferring 5G NR Mode-2 Reception Realism from ns-3 to the Veins/SUMO Stack for Connected-Vehicle Safety Assessment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Accelerating Visual Policy Learning with Sampling-Based Model Predictive Control: Covers a concrete robotics and embodied ai signal; useful as a follow-up candidate.
Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.