Improve model reasoning, planning, and verification, Make agents use tools and reusable skills more reliably, Make RAG retrieval and knowledge-base QA more reliable
Today tracks: improve model reasoning, planning, and verification, make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable.
This issue fetched and deduplicated 497 candidate papers from the 2026-09-24 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback🔗
- 2Low-Cost Assays for Measuring Model Behavior Across Vendors and Releases🔗
- 3Diverse Geometries, Frozen Weights: Robust Heterogeneous Treatment-Effect Estimation via Causal Expert Ensembles🔗
- 4Beyond Spatial Benchmarks: From Spatial Reasoning to Navigation🔗
- 5BeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular Video🔗
- 6OREO: Fidelity Alignment in 3D Generation via On-the-fly Rendering-Editing Optimization🔗
What is worth tracking today
Today’s high-signal papers point to: improve model reasoning, planning, and verification, make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around workflow, benchmark, code, coding to frame the code intelligence task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Improving the scientific coding capabilities of large language models (LLMs) requires high-quality training data
Keywordsworkflowbenchmarkcodecoding
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, code, open-source, coding to frame the code intelligence task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Language models advise people, keep them company, and write software while they sleep
Keywordsagentcodeopen-sourcecoding
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, evaluation, benchmark, eval to frame the benchmarks and evaluation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Estimating heterogeneous treatment effects from observational data is difficult because the most appropriate inductive bias varies with overlap, treatment imbalance, prognostic structure, and sample size
Keywordsragevaluationbenchmarkeval
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around inference, benchmark, code, dataset to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Does progress on spatial reasoning benchmarks translate into better navigation?
Keywordsinferencebenchmarkcodedataset
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, latency, video, temporal to frame the video generation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Learning executable motions from human videos offers a scalable solution for humanoid robots to acquire demonstration motions
Keywordsraglatencyvideotemporal
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, serving, alignment, diffusion to frame the vision and image generation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Despite recent advancements in 3D generation, models often struggle to produce assets with high visual fidelity
Keywordsragservingalignmentdiffusion
Code/DataCheck the source paper
Other papers worth tracking
Industrial Anomaly Detection via Defect-Grounded Reasoning in Visual Latent Space: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Decoupled Learning and Selection in Slate Recommendation for Privacy and Stability Under Noisy Scores: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
SPADE-DFL: Communication-Efficient Decentralized Federated Learning via Derivative-Free Linearized ADMM: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
When Misalignment Becomes Supervision: Structured Label Noise in Supervised Synthetic CT Generation: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
WST-Graph: Topology-Preserving Wavelet Scattering Front-End for Speech Deepfake Detection: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
ArGuard Shared Task: Harmful Content Detection in Arabic Memes and LLM Prompts: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Off-manifold robustness in synthesizer inversion with joint distribution flow matching: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Baszta: Data-Centric Fine-Tuning of a Polish Multi-Label Safety Classifier: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
FounRef: Robust, Structure-Preserving, and Fast Metric Refinement of Frozen Monocular Foundation Priors with Sparse Anchors: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Through Human Eyes and Machine Eyes: Understanding View Mismatch in Video See-Through Extended Reality: Covers a concrete video generation signal; useful as a follow-up candidate.
Edge AI on Constrained Devices for Binary Sleep-Wake Classification in Dynamic Environments: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
A Particle-Swarm-Assisted Gradient Meta-Learning Algorithm for Joint Transmit Precoding and STAR-RIS Coefficient Optimization: Covers a concrete code intelligence signal; useful as a follow-up candidate.
Tag-Aware Structured Text Translation: Towards a Systematic Understanding: Covers a concrete training and post-training signal; useful as a follow-up candidate.
SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
FluidRain: Incompressible Rain Flow as an Attention Bias for Loop-in-Loop Video Deraining: Covers a concrete video generation signal; useful as a follow-up candidate.
Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
From Static Personal Values to Contextualized Personalization: Bayesian Personalized Value Alignment for LLMs: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
HelloWorld: Towards Practical Applications of Generative Driving World Models: Covers a concrete video generation signal; useful as a follow-up candidate.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.