提升代码生成、执行反馈和自动修复能力、提升 RAG 检索和知识库问答可靠性、让 Agent 更可靠地调用工具和复用技能
今天主要跟进:提升代码生成、执行反馈和自动修复能力、提升 RAG 检索和知识库问答可靠性、让 Agent 更可靠地调用工具和复用技能。
本期从 2026-06-11 论文源抓取并去重 395 篇候选论文,筛选 6 篇重点论文与 20 篇补充关注。
本期重点
- 1CRAFTIIF:Cross-Resolution Analytic Four-Type Interpretable Isolation 面向est 面向 Multivariate Time Series Anomaly Detection🔗
- 2MOSAIC:Modality-Specific Adaptation 面向 Incremental Continual Learning in Parkinson's Disease Gait 评估🔗
- 3X-MADAM-RAG:Diagnosing 与 H与ling Chinese-English Evidence Conflict in Retrieval-Augmented Generation🔗
- 4Topical Phase Transitions in Artificial Intelligence Research:Large-Scale Evidence 与 an Early-Warning Signature 面向 Emerging Topics🔗
- 5Containment Gap:How Deployed Agentic AI 框架s Fail Public-Facing Safety Requirements🔗
- 6Budget-Constrained Step-Level Diffusion Caching🔗
今天最值得跟进的方向
今天的高分论文主要指向:提升代码生成、执行反馈和自动修复能力、提升 RAG 检索和知识库问答可靠性、让 Agent 更可靠地调用工具和复用技能。下面按核心问题、方法线索、主要论点和关键词整理,便于快速判断后续跟进价值。
重点论文:核心问题、方法线索与关键词
提升代码生成、执行反馈和自动修复能力
核心:这篇论文主要解决提升代码生成、执行反馈和自动修复能力这一方向中的具体研究问题;方法上通过题名、摘要和公开信号中的 benchmark、code、interpretability、attribution 等线索来组织可解释性任务、数据或评测流程实现提升代码生成、执行反馈和自动修复能力;主要论点是标题、摘要和公开信号显示:Anomaly detection in multivariate time series is challenged by four structurally distinct anomaly types -- point (isolated spikes), distributional (level shifts), temporal (rhythm changes), and collective (inter-sensor correlation breakdowns) -- each requiring。关键词:benchmark、code、interpretability、attribution。代码/数据可用性需查看原文确认。
提升 RAG 检索和知识库问答可靠性
核心:这篇论文主要解决提升 RAG 检索和知识库问答可靠性这一方向中的具体研究问题;方法上通过题名、摘要和公开信号中的 rag、serving、deployment、code 等线索来组织检索与 RAG任务、数据或评测流程实现提升 RAG 检索和知识库问答可靠性;主要论点是标题、摘要和公开信号显示:Gait-based Parkinson's disease assessment increasingly relies on heterogeneous sensors, but clinical systems rarely collect all modalities simultaneously。关键词:rag、serving、deployment、code。代码/数据可用性需查看原文确认。
让 Agent 更可靠地调用工具和复用技能
核心:这篇论文主要解决让 Agent 更可靠地调用工具和复用技能这一方向中的具体研究问题;方法上通过题名、摘要和公开信号中的 rag、retrieval、serving、benchmark 等线索来组织基准与评测任务、数据或评测流程实现让 Agent 更可靠地调用工具和复用技能;主要论点是标题、摘要和公开信号显示:Retrieval-augmented generation (RAG) systems may receive evidence that is not merely noisy but mutually contradictory。关键词:rag、retrieval、serving、benchmark。代码/数据可用性需查看原文确认。
让 Agent 更可靠地调用工具和复用技能
核心:这篇论文主要解决让 Agent 更可靠地调用工具和复用技能这一方向中的具体研究问题;方法上通过题名、摘要和公开信号中的 agent、retrieval、code、multimodal 等线索来组织多模态模型任务、数据或评测流程实现让 Agent 更可靠地调用工具和复用技能;主要论点是标题、摘要和公开信号显示:Do research topics in artificial intelligence grow gradually, or do they advance through abrupt, detectable jumps?。关键词:agent、retrieval、code、multimodal。代码/数据可用性需查看原文确认。
让 Agent 更可靠地调用工具和复用技能
核心:这篇论文主要解决让 Agent 更可靠地调用工具和复用技能这一方向中的具体研究问题;方法上通过题名、摘要和公开信号中的 agent、deployment、safety、memory 等线索来组织Agent 与工具调用任务、数据或评测流程实现让 Agent 更可靠地调用工具和复用技能;主要论点是标题、摘要和公开信号显示:Agentic large language model systems that autonomously invoke tools, maintain persistent memory, and execute multi-step plans are increasingly deployed in public-facing domains, including government services, healthcare triage, and financial advising。关键词:agent、deployment、safety、memory。代码/数据可用性需查看原文确认。
让 Agent 更可靠地调用工具和复用技能
核心:这篇论文主要解决让 Agent 更可靠地调用工具和复用技能这一方向中的具体研究问题;方法上通过题名、摘要和公开信号中的 inference、deployment、latency、alignment 等线索来组织系统与部署任务、数据或评测流程实现让 Agent 更可靠地调用工具和复用技能;主要论点是标题、摘要和公开信号显示:Step-level caching accelerates diffusion models by exploiting temporal redundancy across denoising steps。关键词:inference、deployment、latency、alignment。代码/数据可用性需查看原文确认。
其他值得关注
Existence Precedes Value: Joint Modeling of Observational Existence and Evolving States in Time Series Forecasting:涉及任务设置、指标和失效案例,可补充模型评测与回归测试。
PolyFlow: Safe and Efficient Polytope-Constrained Flow Matching with Constraint Embedding and Projection-free Update:涉及推理成本、延迟、吞吐和部署约束,可补充系统优化方向。
Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents:涉及工具调用、执行反馈和可复用能力,可作为 Agent 工作流可靠性的补充线索。
Nous: An Attempt to Extract and Inject the Cognition Behind Prediction-Market Behavior:涉及工具调用、执行反馈和可复用能力,可作为 Agent 工作流可靠性的补充线索。
YOLO-AMC: An Improved YOLO Architecture with Attention Mechanisms for Building Crack Detection:涉及任务设置、指标和失效案例,可补充模型评测与回归测试。
Zero-source LLM Hallucination Detection with Human-like Criteria Probing:涉及模型安全、护栏路由、风险分类或治理评测,可作为安全评测与治理工具链的补充线索。
MLUBench: A Benchmark for Lifelong Unlearning Evaluation in MLLMs:涉及任务设置、指标和失效案例,可补充模型评测与回归测试。
Recursive Agent Harnesses:涉及工具调用、执行反馈和可复用能力,可作为 Agent 工作流可靠性的补充线索。
One Polluted Page Is Enough: Evaluating Web Content Pollution in Generative Recommenders:涉及检索、知识库问答与证据可靠性,可作为 RAG 评测和企业知识系统的补充线索。
EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis:涉及工具调用、执行反馈和可复用能力,可作为 Agent 工作流可靠性的补充线索。
EvTexture++: Event-Driven Texture Enhancement for Video Super-Resolution:涉及工具调用、执行反馈和可复用能力,可作为 Agent 工作流可靠性的补充线索。
AgentRivet: an automated system for producing Rivet routines from journal publications:涉及工具调用、执行反馈和可复用能力,可作为 Agent 工作流可靠性的补充线索。
Mining Architectural Quality Under Agentic AI Adoption: A Causal Study of Java Repositories:涉及工具调用、执行反馈和可复用能力,可作为 Agent 工作流可靠性的补充线索。
Multi-Field Hybrid Retrieval-Augmented Generation for Maritime Accident Root Cause Analysis:涉及检索、知识库问答与证据可靠性,可作为 RAG 评测和企业知识系统的补充线索。
DIG: Oracle-Guided Directed Input Generation for One-Day Vulnerabilities:涉及工具调用、执行反馈和可复用能力,可作为 Agent 工作流可靠性的补充线索。
A Multi-Modal Framework with Cross-Subject Pseudo-Labeling and Semantic Alignment for Micro-Gesture Recognition:涉及模型安全、护栏路由、风险分类或治理评测,可作为安全评测与治理工具链的补充线索。
Structured Testbench Generation for LLM-Driven HDL Design and Verification-Oriented Data Curation:涉及工具调用、执行反馈和可复用能力,可作为 Agent 工作流可靠性的补充线索。
MARS: Margin-Adversarial Risk-controlled Stopping for Parallel LLM Test-time Scaling:涉及推理与规划中的新任务、数据或系统线索,可作为后续跟进清单的一部分。
Order Is Not Control:涉及检索、知识库问答与证据可靠性,可作为 RAG 评测和企业知识系统的补充线索。
Language-Guided Abstraction for Visual Reasoning:涉及检索、知识库问答与证据可靠性,可作为 RAG 评测和企业知识系统的补充线索。
阅读边界
- 自动排序会偏向有社区信号、代码信号和工程关键词的论文。
- 简报默认基于标题、摘要和公开元数据,不替代全文精读。
- 外部 API 限流或不可用时,相关信号会降级为空并在内部记录中保留说明。