Make agents use tools and reusable skills more reliably, Make RAG retrieval and knowledge-base QA more reliable, Improve model reasoning, planning, and verification
Today tracks: make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable, improve model reasoning, planning, and verification.
This issue fetched and deduplicated 307 candidate papers from the 2026-07-14 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction🔗
- 2Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference🔗
- 3Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?🔗
- 4Label-Decoupled Style Augmentation for Domain Generalization in Multi-Label Remote Sensing Scene Classification🔗
- 5Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution🔗
- 6LARAD: Layout-Aware Road Anomaly Detection via Spatial-Logic Reasoning🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable, improve model reasoning, planning, and verification. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, alignment, evaluation, benchmark to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Rubric-based evaluation is a promising approach for assessing open-ended outputs from LLM-based research agents, particularly in paper reproduction, where direct paper-to-repository comparison is prone to hallucination
Keywordsagentalignmentevaluationbenchmark
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, deployment, latency to frame the systems and deployment task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Vision-Language-Action (VLA) models have achieved impressive performance on diverse embodied tasks
Keywordsraginferencedeploymentlatency
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around inference, deployment, alignment, benchmark to frame the video generation task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Recent advances in multimodal large language models (MLLMs) have significantly improved the performance of multimodal emotion recognition (MER) and enabled interpretable description generation by jointly modeling video, audio, and language, etc
Keywordsinferencedeploymentalignmentbenchmark
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Multi-label classification assigns several co-occurring labels to each aerial scene, yet deployed models often encounter data distributions different from their training
Keywordsraginferencebenchmarkcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, workflow, retrieval, benchmark to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they rarely ask how much effort a task actually requires
Keywordsagentworkflowretrievalbenchmark
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around inference, latency, training, rl to frame the training and post-training task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Accurate open-world obstacle detection is critical for autonomous driving
Keywordsinferencelatencytrainingrl
Code/DataCheck the source paper
Other papers worth tracking
Rethinking the Evaluation of Harness Evolution for Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Decouple and Reason: Anatomically Guided Two-Stage Voxel-Level Grounding of Free-Text Findings in 3D Chest CT: Covers a concrete vision and image generation signal; useful as a follow-up candidate.
Lightweight Multi-Scale Anomaly Detection for Resource-Constrained Edge Devices: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
StratMamba: Strategic and Reactive Stream Partitioning for Path-Efficient LiDAR-Based Obstacle Avoidance: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
PalmClaw: A Native On-Device Agent Framework for Mobile Phones: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
DermDepth: Toward Monocular Metric Scale 3D Reconstruction Models for Dermatology: Covers a concrete training and post-training signal; useful as a follow-up candidate.
ViHoRec: A Quality-Controlled Vietnamese Hotel Recommendation Dataset and Cold-Start Benchmark: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Open-KNEAD: Knowledge-grounded Nutrition Estimation via Agentic Decomposition: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
UniMedSeg: Unified In-Context Learning for Multi-Paradigm 2D/3D Medical Image Segmentation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Inhibited Self-Attention: Sharpening Focus in Vision Transformers: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Verifier-Based Reinforcement Fine-Tuning of Reasoning Models for Thermal Energy Storage Control: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Breaking Déjà Vu: Independent Auditing of Visual Place Recognition through Vision-Language Reasoning: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models: Covers a concrete multimodal models signal; useful as a follow-up candidate.
Directional Constraints for Efficient Exploration in Safe Reinforcement Learning: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Constraint-Aware Aggregation for Federated Reinforcement Learning in Microgrid Energy Coordination: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
RFMSR: Residual Flow Matching for Image Super-Resolution: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Epistemic Stance Flexibility Probing: Measuring Prompt-Conditioned Register Shift in Large Language Models: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.