Improve model reasoning, planning, and verification, Make agents use tools and reusable skills more reliably, Make RAG retrieval and knowledge-base QA more reliable
Today tracks: improve model reasoning, planning, and verification, make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 386 candidate papers from the 2026-06-25 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP🔗
- 2Empirical Software Engineering TerraProbe: A Layered-Oracle Framework for Detecting Deceptive Fixes in LLM-Assisted Terraform🔗
- 3Retrieval-Warmed Energy-Based Reasoning: A Five-Arm Ablation Methodology for Diffusion-as-Inference on Structured Reasoning Tasks🔗
- 4Prompt Injection in Automated Résumé Screening with Large Language Models: Single and Multi-Injection Settings🔗
- 5Do Safety Guardrails Need to Reason? LeanGuard: A Fast and Light Approach for Robust Moderation🔗
- 6\textsc{DiARC}: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models🔗
What is worth tracking today
Today’s high-signal papers point to: improve model reasoning, planning, and verification, make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around rag, alignment, evaluation, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Current VLM evaluations often conflate language priors with genuine spatial reasoning
Keywordsragalignmentevaluationcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, deployment, evaluation, code to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Security misconfigurations in Terraform Infrastructure-as-Code are a growing risk in cloud deployments, and large language models are increasingly used as automated repair agents
Keywordsagentdeploymentevaluationcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around rag, retrieval, inference, alignment to frame the retrieval and rag task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Warm-started diffusion samplers accelerate iterative inference, but it is rarely clear which part of the pipeline carries the gain
Keywordsragretrievalinferencealignment
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, evaluation, code, api to frame the code intelligence task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Large language models (LLMs) are increasingly used to screen and rank job applicants, creating incentives for candidates to strategically manipulate algorithmic hiring systems
Keywordsragevaluationcodeapi
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around rag, inference, safety, benchmark to frame the reasoning and planning task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: In order to screen a prompt or a response, the recent guardrail methods generate a chain-of-thought (CoT) before they issue a verdict
Keywordsraginferencesafetybenchmark
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around alignment, benchmark, code, fine-tuning to frame the training and post-training task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: The Abstraction and Reasoning Corpus (ARC;~\citealp{chollet2019measure}) contains tasks that require summarizing patterns from limited grid samples and predicting output grids
Keywordsalignmentbenchmarkcodefine-tuning
Code/DataCheck the source paper
Other papers worth tracking
FlameVQA: A Physically-Grounded UAV Wildfire VQA Benchmark with Radiometric Thermal Supervision: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and Generation: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Geometric Gradient Rectification for Safe Open-Set Semi-Supervised Learning: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
LCAi: Life Cycle Assessment with big data fusion and retrieval-augmented generation-assisted interpretation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Adversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation for Text, Vision, and Vision-Language Models: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
OctoSense: Self-Supervised Learning for Multimodal Robot Perception: Covers a concrete video generation signal; useful as a follow-up candidate.
Automating Potential-based Reward Shaping with Vision Language Model Guidance: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Cascaded Multi-Granularity Pruning for On-Device LLM Inference in Industrial IoT: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
The Capability Frontier: Benchmarks Miss 82% of Model Performance: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Dual-Prior Guided Null-Space Learning with Mixture-of-Splines for Arbitrary Medical Slice Super-Resolution: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
MLFFM-SegDiff: A Multi-Level Feature Fusion Diffusion Model for Skin Lesion Segmentation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Autoregressive Boltzmann Generators: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
See & Sniff: Learning Visuo-Olfactory Representations: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
How Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple Mitigation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
SatSplatDiff: Geometry-preserving generative refinement for high-fidelity satellite Gaussian Splatting: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Paved with True Intents: Intent-Aware Training Improves LLM Safety Classification Across Training Regimes: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.