Improve model reasoning, planning, and verification, Make RAG retrieval and knowledge-base QA more reliable
Today tracks: improve model reasoning, planning, and verification, improve model reasoning, planning, and verification, improve model reasoning, planning, and verification.
This issue fetched and deduplicated 445 candidate papers from the 2026-07-02 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Fast Multi-dimensional Refusal Subspaces via RFM-AGOP🔗
- 2DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing🔗
- 3AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models🔗
- 4ArcAD: Anomaly-Rectified Calibration for Cold-Start Supervised Anomaly Detection🔗
- 5RadiomicNet: A Hybrid Radiomics-Guided Lightweight Architecture for Interpretable Medical Image Segmentation🔗
- 6AdaCount: Training-Free Similarity-Guided Spatial and Feature Adaptation for Zero-Shot Object Counting🔗
What is worth tracking today
Today’s high-signal papers point to: improve model reasoning, planning, and verification, improve model reasoning, planning, and verification, improve model reasoning, planning, and verification. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around safety, code, harmful, Safety and Alignment to frame the safety and alignment task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Steering and monitoring activations in Large Language Models (LLMs) are increasingly used for both safety and interpretability
KeywordssafetycodeharmfulSafety and Alignment
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around evaluation, benchmark, code, open-source to frame the data engineering task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Recent image generation and editing models can produce visually appealing natural images, yet they remain unreliable when the target image is a knowledge-intensive diagram whose correctness depends on disciplinary concepts, symbolic structure, and precise spat
Keywordsevaluationbenchmarkcodeopen-source
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around evaluation, benchmark, vision-language, eval to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG)
Keywordsevaluationbenchmarkvision-languageeval
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, deployment, code, rl to frame the training and post-training task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: The deployment of Industrial Anomaly Detection (IAD) in real-world manufacturing frequently encounters a challenging cold-start bottleneck, in which limited normal samples fail to represent the full normal distribution and only a few anomalies are available
Keywordsragdeploymentcoderl
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around rag, alignment, code, knowledge to frame the retrieval and rag task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Deep learning has achieved remarkable performance in medical image segmentation, yet it suffers from critical limitations: mathematical intractability, substantial parameter requirements, and lack of clinical interpretability
Keywordsragalignmentcodeknowledge
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, serving, benchmark, code to frame the multimodal models task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Zero-shot object counting (ZOC) aims to count instances of arbitrary object categories specified only through textual prompts
Keywordsragservingbenchmarkcode
Code/DataCheck the source paper
Other papers worth tracking
Privacy-Preserving and Verifiable Approximate Distributed Coded Computing: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
ContextNest: Verifiable Context Governance for Autonomous AI Agent: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Hidden Forgetting in Continual Multimodal Learning: When Accuracy Survives but Grounding Fails: Covers a concrete multimodal models signal; useful as a follow-up candidate.
Liquid Latent State Dynamics for Interpretable Turbofan Degradation Modeling: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Towards a Phonology-Informed Evaluation of Multilingual TTS: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
NAVER LABS Europe Submission to the Instruction-following 2026 Short Track: Covers a concrete data engineering signal; useful as a follow-up candidate.
Personalized 4D Whole-Heart Mesh Reconstruction from Cine MRI via Multi-Scale Temporal Modeling and Differentiable Contour Rendering: Covers a concrete video generation signal; useful as a follow-up candidate.
AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K-12 Educational Explanations: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Rethinking Post-Hoc Calibration in Semantic Segmentation: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Gaming Consensus: Coordinated Manipulation in Crowdsourced Fact-Checking: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
On the Limits of Steering Vectors for Preference-Aligned Generation: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Verifiable Knowledge Expansion through Retrieval-Grounded Formal Concept Analysis: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Denser $\neq$ Better: Limits of On-Policy Self-Distillation for Continual Post-Training: Covers a concrete training and post-training signal; useful as a follow-up candidate.
InterCMDM: Block-Causal Diffusion for Autoregressive Human Interaction Generation: Covers a concrete video generation signal; useful as a follow-up candidate.
Pmeta-TLA: Backdoor Attacks for Speech Classification Models via Meta-Learning with Timbre Leakage Attack: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Epistemic Goggles: A Pretrained Module that Induces an Epistemic Frame via Gradient Editing: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Model Merging as Probabilistic Inference in Fine-Tuning Parameter Space: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.