增强多模态模型理解图表和文档的能力、让 Agent 更可靠地调用工具和复用技能、提升 RAG 检索和知识库问答可靠性
今天主要跟进:增强多模态模型理解图表和文档的能力、让 Agent 更可靠地调用工具和复用技能、提升 RAG 检索和知识库问答可靠性。
本期从 2026-06-02 论文源抓取并去重 264 篇候选论文,筛选 5 篇重点论文与 15 篇补充关注。
本期重点
- 1KODA:Contrastive Representation Comparison 与 Alignment 面向 Vision-Language Foundation Models🔗
- 2Impact of Configuring Agentic AI Coding Tools on Build-vs-Buy Decisions:A Study Protocol🔗
- 3Automating In面向mation 抽取 与 Retrieval 面向 Industrial Spare Parts Pooling🔗
- 4Stationarity-Aware Retrieval-Augmented Time Series 面向ecasting🔗
- 5Entropy Gate:Entropy Quenching 面向 Near-Lossless Token Compression in LLM Pipelines🔗
今天最值得跟进的方向
今天的高分论文主要指向:增强多模态模型理解图表和文档的能力、让 Agent 更可靠地调用工具和复用技能、提升 RAG 检索和知识库问答可靠性。下面按核心问题、方法线索、主要论点和关键词整理。
重点论文:题目、看点与核验线索
增强多模态模型理解图表和文档的能力
增强多模态模型理解图表和文档的能力。核心线索:Vision-language foundation models such as CLIP and SigLIP provide widely used representations for multimodal learning systems. 代码/数据可用性需查看原文确认。
让 Agent 更可靠地调用工具和复用技能
让 Agent 更可靠地调用工具和复用技能。核心线索:Agentic AI coding tools write code with increasing autonomy and in doing so decide when to import a library and when to implement functionality from scratch. 代码/数据可用性需查看原文确认。
提升 RAG 检索和知识库问答可靠性
提升 RAG 检索和知识库问答可靠性。核心线索:Maintenance organizations in manufacturing try to avoid downtime and unnecessary purchasing by reusing existing assets, but the main obstacle is not a lack of parts but a lack of actionable visibility across sites and partners. 代码/数据可用性需查看原文确认。
提升 RAG 检索和知识库问答可靠性
提升 RAG 检索和知识库问答可靠性。核心线索:Time series forecasting relies on historical patterns, but real-world series often exhibit non-stationarity and regime shifts that challenge fully parametric forecasters. 代码/数据可用性需查看原文确认。
让 Agent 更可靠地调用工具和复用技能
让 Agent 更可靠地调用工具和复用技能。核心线索:LLM pipelines waste substantial token budgets on low-information content: repeated context, verbose responses, and redundant boilerplate. 代码/数据可用性需查看原文确认。
其他值得关注
VLESA: Vision-Language Embodied Safety Agent for Human Activity Monitoring:关注模型安全、护栏路由、风险分类或治理评测,适合跟进安全评测与治理工具链。
MM-BizRAG: Rethinking Multimodal Retrieval-Augmented Generation for General Purpose Enterprise Q&A:关注检索、知识库问答与证据可靠性,适合跟进 RAG 评测和企业知识系统。
MimeLens: Position-Agnostic Content-Type Detection for Binary Fragments:关注检索、知识库问答与证据可靠性,适合跟进 RAG 评测和企业知识系统。
When Autoregressive Consistency Hurts Safety Alignment:关注模型安全、护栏路由、风险分类或治理评测,适合跟进安全评测与治理工具链。
End-to-End Text Line Detection and Ordering:关注检索、知识库问答与证据可靠性,适合跟进 RAG 评测和企业知识系统。
Expert-Aware Refusal Steering:关注模型安全、护栏路由、风险分类或治理评测,适合跟进安全评测与治理工具链。
HighTide: An Agent-Curated Open-Source VLSI Benchmark Suite:关注工具调用、执行反馈和可复用能力,适合跟进 Agent 工作流和工程可靠性。
UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing:关注推理成本、延迟、吞吐和部署约束,适合跟进系统优化。
MAOAM: Unified Object and Material Selection with Vision-Language Models:关注检索、知识库问答与证据可靠性,适合跟进 RAG 评测和企业知识系统。
Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill:关注工具调用、执行反馈和可复用能力,适合跟进 Agent 工作流和工程可靠性。
Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning:关注工具调用、执行反馈和可复用能力,适合跟进 Agent 工作流和工程可靠性。
AgenticRL: Self-Refining Agentic Reinforcement Learning for Vision-Conditioned UAV Navigation:关注工具调用、执行反馈和可复用能力,适合跟进 Agent 工作流和工程可靠性。
SEAOTTER: Sensor Embedded Autoencoding with One-Time Transcode for Efficient Reconstruction:关注多模态模型中的新任务、数据或系统线索,适合快速判断是否值得阅读全文。
Visual Instruction Tuning Aligns Modalities through Abstraction:关注训练与后训练中的新任务、数据或系统线索,适合快速判断是否值得阅读全文。
Leveraging BART to Assess CS1 C++ Programming Assignments using Rubric-based Criteria:关注检索、知识库问答与证据可靠性,适合跟进 RAG 评测和企业知识系统。
阅读边界
- 自动排序会偏向有社区信号、代码信号和工程关键词的论文。
- 简报默认基于标题、摘要和公开元数据,不替代全文精读。
- 外部 API 限流或不可用时,相关信号会降级为空并在内部记录中保留说明。