Make agents use tools and reusable skills more reliably, Strengthen multimodal understanding of charts, documents, and visual evidence, Make RAG retrieval and knowledge-base QA more reliable
Today tracks: make agents use tools and reusable skills more reliably, strengthen multimodal understanding of charts, documents, and visual evidence, strengthen multimodal understanding of charts, documents, and visual evidence.
This issue fetched and deduplicated 419 candidate papers from the 2026-06-18 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning🔗
- 2FFinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming🔗
- 3StylisticBias: A Few Human Visual Cues Drive Most Social Biases in MLLMs🔗
- 4SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm🔗
- 5FlowBender: Feedback-Aware Training for Self-Correcting Conditional Flows🔗
- 6Towards Modality-imbalanced Federated Graph Learning: A Data Synthesis-based Approach🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, strengthen multimodal understanding of charts, documents, and visual evidence, strengthen multimodal understanding of charts, documents, and visual evidence. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, evaluation, agents, Agents and Tool Use to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: This work presents a general framework for training large language models (LLMs) to "Connect the Dots" (CoD), a meta-capability required by long-lifecycle agents: as an LLM-based AI agent gets deployed in an environment, it solves a long sequence of tasks whil
KeywordsagentevaluationagentsAgents and Tool Use
Code/DataCheck the source paper
strengthen multimodal understanding of charts, documents, and visual evidence
Signalthis paper targets the concrete research problem behind strengthen multimodal understanding of charts, documents, and visual evidence. It uses the title, abstract, and public signals around safety, evaluation, benchmark, eval to frame the benchmarks and evaluation task, data, or evaluation flow to improve strengthen multimodal understanding of charts, documents, and visual evidence. The main claim is the title, abstract, and public signals indicate: Existing safety benchmarks target general adversarial scenarios but miss finance-specific risks
Keywordssafetyevaluationbenchmarkeval
Code/DataCheck the source paper
strengthen multimodal understanding of charts, documents, and visual evidence
Signalthis paper targets the concrete research problem behind strengthen multimodal understanding of charts, documents, and visual evidence. It uses the title, abstract, and public signals around evaluation, benchmark, code, multimodal to frame the benchmarks and evaluation task, data, or evaluation flow to improve strengthen multimodal understanding of charts, documents, and visual evidence. The main claim is the title, abstract, and public signals indicate: Multimodal large language models (MLLMs) are increasingly deployed in personally and societally consequential settings, yet the visual cues that shape how these models judge people remain poorly understood
Keywordsevaluationbenchmarkcodemultimodal
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around retrieval, alignment, evaluation, benchmark to frame the benchmarks and evaluation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Multimodal foundation models have advanced rapidly thanks to large optical benchmarks, but comparable resources for synthetic aperture radar (SAR) remain limited
Keywordsretrievalalignmentevaluationbenchmark
Code/DataCheck the source paper
improve image generation, visual understanding, and controllable rendering
Signalthis paper targets the concrete research problem behind improve image generation, visual understanding, and controllable rendering. It uses the title, abstract, and public signals around inference, compression, alignment, systems to frame the systems and deployment task, data, or evaluation flow to improve improve image generation, visual understanding, and controllable rendering. The main claim is the title, abstract, and public signals indicate: Conditional diffusion and flow models routinely fail to satisfy the very constraints that define their task
Keywordsinferencecompressionalignmentsystems
Code/DataCheck the source paper
strengthen multimodal understanding of charts, documents, and visual evidence
Signalthis paper targets the concrete research problem behind strengthen multimodal understanding of charts, documents, and visual evidence. It uses the title, abstract, and public signals around deployment, alignment, code, multimodal to frame the vision and image generation task, data, or evaluation flow to improve strengthen multimodal understanding of charts, documents, and visual evidence. The main claim is the title, abstract, and public signals indicate: MultiModal Federated Graph Learning (MM-FGL) offers a natural collaborative training paradigm, but its practical deployment is challenged by two granularities of modality imbalance
Keywordsdeploymentalignmentcodemultimodal
Code/DataCheck the source paper
Other papers worth tracking
Sovereign Execution Brokers: Enforcing Certificate-Bound Authority in Agentic Control Planes: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
CRAX: Fast Safe Reinforcement Learning Benchmarking: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
The Register Gap: A Meaning Intelligence Framework for Nigerian Public Discourse: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think: Covers a concrete training and post-training signal; useful as a follow-up candidate.
ScholarQuest: A Taxonomy-Guided Benchmark for Agentic Academic Paper Search in Open Literature Environments: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
DeepForestVisionV2: Ecology-Driven Taxonomy Expansion for Camera-Trap Monitoring in African Tropical Forests: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
ReNikud: Audio-Supervised Hebrew Grapheme-to-Phoneme Conversion: Covers a concrete speech and audio signal; useful as a follow-up candidate.
MedRLM: Recursive Multimodal Health Intelligence for Long-Context Clinical Reasoning, Sensor-Guided Screening, Evidence-Grounded Decision Support, and Community-to-Tertiary Referral Optimization: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Modularity-Free Conflict-Averse Training for Generalized PINNs: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Personalized Keyword Spotting for User-Defined Keywords Leveraging Text-Independent Speaker Verification: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Repository-Level Solidity Code Generation with Large Language Models: From Prompting to Fine-Tuning: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Advancing DialNav through Automatic Embodied Dialog Augmentation: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Timage: A Generative Text-in-Image Paradigm for Fine-Tuning Vision-Language Models: Covers a concrete multimodal models signal; useful as a follow-up candidate.
PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation: Covers a concrete data engineering signal; useful as a follow-up candidate.
Speeding up the annotation process in semantic segmentation industrial applications: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Co-policy: Responsive Human-Robot Co-Creation for Musical Performances: Covers a concrete robotics and embodied ai signal; useful as a follow-up candidate.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.