Make agents use tools and reusable skills more reliably, Improve model reasoning, planning, and verification, Make RAG retrieval and knowledge-base QA more reliable
Today tracks: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification.
This issue fetched and deduplicated 558 candidate papers from the 2026-06-09 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1MPC-Patch-Bench: Security-Aware LLM Code Patch for Multi-Party Computation🔗
- 2The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes🔗
- 3Globally Localizing Lunar Rover in Pixels via Graph Alignment🔗
- 4i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models🔗
- 5ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity🔗
- 6SceneMiner: Identity-Preserving Multi-Task Fine-Tuning for Unified BEV Scene Mining🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, improve model reasoning, planning, and verification. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, serving, evaluation, benchmark to frame the code intelligence task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Repository-level benchmarks for evaluating Large Language Model (LLM) code repair on Secure Multi-Party Computation (MPC) software do not yet exist, and directly transplanting general-purpose benchmarks such as SWE-bench fails on three structural fronts: (i) M
Keywordsagentservingevaluationbenchmark
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, retrieval, inference, evaluation to frame the reasoning and planning task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Large Language Models (LLMs) have achieved strong performance across natural language processing tasks, yet reliable reasoning remains an open challenge
Keywordsagentretrievalinferenceevaluation
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around rag, alignment, code, data to frame the data engineering task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Precise rover localization is a prerequisite for autonomous lunar exploration, yet the absence of Global Navigation Satellite System (GNSS) signals and the cumulative drift of local localization methods severely constrain long-range missions
Keywordsragalignmentcodedata
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Diffusion models have consistently driven progress in text-to-image generation
Keywordsraginferencebenchmarkcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, benchmark, code to frame the retrieval and rag task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Large language models (LLMs) are rapidly acquiring capabilities relevant to biological research, from literature synthesis to interpretation of experimental data
Keywordsagentragbenchmarkcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around retrieval, serving, safety, code to frame the training and post-training task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Mining hard, safety-critical scenes from driving logs is bottlenecked by the absence of difficulty labels, and no single proxy, collision risk, trajectory ambiguity, or semantic rarity suffices to find such scenes on its own
Keywordsretrievalservingsafetycode
Code/DataCheck the source paper
Other papers worth tracking
Energy-Efficient On-Device RAG on a Mobile NPU: System Design and Benchmark on Snapdragon X Elite: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Agent Skill Evaluation and Evolution: Frameworks and Benchmarks: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
GLACIER: A Multimodal Student-Teacher Foundation Model for Molecular Property Prediction: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Schützen: Evaluating LLM Safety in Bulgarian and German Contexts: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
When to Align, When to Predict: A Phase Diagram for Multimodal Learning: Covers a concrete multimodal models signal; useful as a follow-up candidate.
ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations: Covers a concrete multimodal models signal; useful as a follow-up candidate.
FADA: Accessible fetal ultrasound interpretation and annotation with a selectively distilled unified vision-language model: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
LIBERO-Occ: Evaluating and Improving Vision-Language-Action Models under Scene-Induced Occlusion via Viewpoint Imagination: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Dep-LLM: Training-Free Depression Diagnosis via Evidence-Guided Structured Multi-factor with Reliable LLM Reasoning: Covers a concrete training and post-training signal; useful as a follow-up candidate.
The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
GUI-AC: Enhancing Continual Learning in GUI Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
KCSAT-ML: Probing Reasoning Models with Nationwide-Cohort Human Difficulty: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Dissect and Prune: Enhancing Robustness in AI-Generated Image Detection: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
When Metrics Disagree: A Meta-Analysis of Knowledge-Graph-Completion Model Benchmarking: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
From Simulation to Real-World: An In-Field 6D Pose Dataset and Baseline for Robotic Strawberry Harvesting: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Overcoming Rank Collapse in Feedback Alignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Do Transformers Actually Help Intrusion Detection? A Temporal Sequence Evaluation on CIC-IDS2017: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
RedAct: Redacting Agent Capability Traces for Procedural Skill Protection: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
ChartLens: A Dual-Branch Framework for Chart Data Correction and Factual Summary Refinement: Covers a concrete multimodal models signal; useful as a follow-up candidate.
STORM: Stepwise Token Optimization with Reward-Guided Beam Search: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.