Make RAG retrieval and knowledge-base QA more reliable, Improve model reasoning, planning, and verification, Make agents use tools and reusable skills more reliably
Today tracks: make RAG retrieval and knowledge-base QA more reliable, improve model reasoning, planning, and verification, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 399 candidate papers from the 2026-08-27 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Systematic Literature Review of Machine Learning Models and Applications for Text Recognition🔗
- 2CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes🔗
- 3SWE-Prime: Fewer Trajectories, Better Performance🔗
- 4Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners🔗
- 5Sidecar: Training-Free Semantic Reuse for Character-Consistent Free-form Visual Storytelling🔗
- 6Dose-PlanNet: Physics Based Radiotherapy Dose Prediction with Deep Learning🔗
What is worth tracking today
Today’s high-signal papers point to: make RAG retrieval and knowledge-base QA more reliable, improve model reasoning, planning, and verification, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, evaluation, multimodal, search to frame the retrieval and rag task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Optical Character Recognition (OCR) for text recognition using machine vision has significantly improved, particularly when handling heterogeneous textual data
Keywordsragevaluationmultimodalsearch
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around rag, inference, code, reasoning to frame the reasoning and planning task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Recent advances in inference-time scaling have significantly improved the reasoning performance of large language models (LLMs)
Keywordsraginferencecodereasoning
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, fine-tuning, training to frame the training and post-training task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: To improve large language models' ability to resolve real-world software issues, prior work has focused on constructing large-scale agent trajectory datasets and performing supervised fine-tuning (SFT) on successful trajectories
Keywordsagentragfine-tuningtraining
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around rag, evaluation, benchmark, eval to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Static scanners are increasingly used to identify executable or otherwise unsafe content in machine- learning artifacts, yet conventional evaluation metrics characterize only cases where a scanner yields a usable security judgment
Keywordsragevaluationbenchmarkeval
Code/DataCheck the source paper
improve image generation, visual understanding, and controllable rendering
Signalthis paper targets the concrete research problem behind improve image generation, visual understanding, and controllable rendering. It uses the title, abstract, and public signals around serving, alignment, diffusion, visual to frame the vision and image generation task, data, or evaluation flow to improve improve image generation, visual understanding, and controllable rendering. The main claim is the title, abstract, and public signals indicate: Visual storytelling requires generating images that follow a narrative while preserving consistent character identities across frames
Keywordsservingalignmentdiffusionvisual
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around workflow, rag, deployment, planning to frame the agents and tool use task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Automating prostate radiotherapy treatment planning is dosimetrically complex, particularly for extreme hypofractionated regimens
Keywordsworkflowragdeploymentplanning
Code/DataCheck the source paper
Other papers worth tracking
PAWBench: How Far Are We from Probabilistically Aligned World Modeling?: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection: Covers a concrete vision and image generation signal; useful as a follow-up candidate.
KinyaEmbed: Contrastive Sentence Embeddings for Kinyarwanda via Multi-Stage Curriculum Training: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
CGS-SLAM: Collaborative Gaussian Splatting based SLAM for Multi-Agent Reconstruction: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reinforcement Learning-Based Control of CAV Platoon Joining Maneuvers in Mixed Traffic: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
A Geometry-Driven, Framework-Agnostic Optimization for Object Pose Estimation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Cross-Architecture Knowledge Distillation from a Vision Foundation Model to a Lightweight Visual State Space Model for Tea Leaf Disease Classification: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
DEEPCHART: How Far are LLMs from Faithful Data-Science Chart Generation?: Covers a concrete multimodal models signal; useful as a follow-up candidate.
Glass Surface Detection Grounded in 3D Visual Geometry: Covers a concrete vision and image generation signal; useful as a follow-up candidate.
Discovering Relationships in Data Lakes Using Large Language Models: An Industrial Case: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
When Does Supervised Fine-Tuning Reduce Instruction Sensitivity?: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Meta-Learning Where to Allocate Experts: Task-Conditioned Layer-Wise Compression for MoEs: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Real-time Unsupervised Object Discovery from Asynchronous Event Streams: Covers a concrete video generation signal; useful as a follow-up candidate.
Information-Guided Frontier Decoding: Contextual Utility-Driven Commitment in dMLLMs: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Who Remains, What Changes: Identity Anchored Composed Gait Retrieval: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
The Thousand-Graph Hypothesis: A Testable Hypothesis of Task-Conditioned Relation Materialization in Repository-Level Code Reasoning: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Not Just Reason, Not Just Scan: Reinforcement Learning for Proactive Scientific Error Verification over Academic Paper: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.