提升代码生成、执行反馈和自动修复能力、提升 RAG 检索和知识库问答可靠性、增强多模态模型理解图表和文档的能力
今天主要跟进:提升代码生成、执行反馈和自动修复能力、提升 RAG 检索和知识库问答可靠性、提升 RAG 检索和知识库问答可靠性。
本期从 2026-09-04 论文源抓取并去重 352 篇候选论文,筛选 6 篇重点论文与 20 篇补充关注。
本期重点
- 1MePo++:Unifying Representation Refinement 与 Reconciliation 面向 General Continual Learning🔗
- 2GLASS:Graph-Language Alignment with Spherical Scoring 面向 Transferable Graph-Level Anomaly Detection🔗
- 3MURAL:Multimodal Uncertainty-aware Recommendation 通过 Adaptive edge Learning🔗
- 4What Matters,When? Diagnosing 与 Improving Conditional Visual 落地 in Visuomotor Imitation Policies🔗
- 5TruthInsightBench:An Evidence-Grounded 基准 面向 Automated 评测 of Open-Ended Scientific Discovery Agents🔗
- 6CoSkill:Joint Rein面向cement Learning of Reasoning 与 Meta-Skill Agents 面向 Hierarchical Skill Evolution🔗
今天最值得跟进的方向
今天的高分论文主要指向:提升代码生成、执行反馈和自动修复能力、提升 RAG 检索和知识库问答可靠性、提升 RAG 检索和知识库问答可靠性。下面按核心问题、方法线索、主要论点和关键词整理,便于快速判断后续跟进价值。
重点论文:核心问题、方法线索与关键词
提升代码生成、执行反馈和自动修复能力
中文标题MePo++:Unifying Representation Refinement 与 Reconciliation 面向 General Continual Learning
信号显示该研究围绕题名与摘要所揭示的问题展开,具体信号需结合原文进一步核验:General continual learning (GCL) aims to learn from evolving data streams without task identities, explicit boundaries, or repeated access to previous data, making it a realistic yet challenging setting for continual intelligence
关键词servingalignmentcodepost-training
代码/数据需查看原文确认
提升 RAG 检索和知识库问答可靠性
中文标题GLASS:Graph-Language Alignment with Spherical Scoring 面向 Transferable Graph-Level Anomaly Detection
信号显示该研究围绕题名与摘要所揭示的问题展开,具体信号需结合原文进一步核验:We introduce GLASS, a framework for graph-level anomaly detection (GLAD) that achieves robust cross-domain transferability through graph-language alignment on the unit hypersphere
关键词ragalignmentbenchmarkcode
代码/数据需查看原文确认
提升 RAG 检索和知识库问答可靠性
中文标题MURAL:Multimodal Uncertainty-aware Recommendation 通过 Adaptive edge Learning
信号显示该研究围绕题名与摘要所揭示的问题展开,具体信号需结合原文进一步核验:Multimodal Graph Neural Networks have become standard for recommendation by augmenting sparse interaction data with content features
关键词ragretrievalalignmentbenchmark
代码/数据需查看原文确认
增强多模态模型理解图表和文档的能力
中文标题What Matters,When? Diagnosing 与 Improving Conditional Visual 落地 in Visuomotor Imitation Policies
信号显示该研究围绕题名与摘要所揭示的问题展开,具体信号需结合原文进一步核验:Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail when visually similar objects or receptacles are introduced
关键词servingvision-languagemanipulationpolicy
代码/数据需查看原文确认
让 Agent 更可靠地调用工具和复用技能
中文标题TruthInsightBench:An Evidence-Grounded 基准 面向 Automated 评测 of Open-Ended Scientific Discovery Agents
信号显示该研究围绕题名与摘要所揭示的问题展开,具体信号需结合原文进一步核验:Autonomous coding agents are increasingly proposed as AI-scientist systems that conduct analyses and write research reports, but executing a prescribed analysis is not the same as making a discovery
关键词agentevaluationbenchmarkcode
代码/数据需查看原文确认
让 Agent 更可靠地调用工具和复用技能
中文标题CoSkill:Joint Reinforcement Learning of Reasoning 与 Meta-Skill Agents 面向 Hierarchical Skill Evolution
信号显示Skill libraries improve the sample efficiency of agentic 强化学习 (RL) by enabling large language model (LLM) agents to reuse procedural knowledge
关键词agentworkflowcodeagents
代码/数据需查看原文确认
其他值得关注
HiSfM: Disambiguating Structure-from-Motion via Scaffold-Anchored Hierarchical Reconstruction
中文标题HiSfM:Disambiguating Structure-来自-Motion via Scaffold-Anchored Hierarchical Reconstruction
关注理由涉及工具调用、执行反馈和可复用能力,可作为 Agent 工作流可靠性的补充线索。
中文标题RegionFed:Federated Learning 面向 Personalized Query Understanding in Heterogeneous Retail Environments
关注理由涉及推理与规划中的新任务、数据或系统线索,可作为后续跟进清单的一部分。
CABAL: Multi-Agent Simulacra for Tracing the Effects of Collusive Bidding in Peer Review
中文标题CABAL:Multi-Agent Simulacra 面向 Tracing the Effects of Collusive Bidding in Peer Review
关注理由涉及工具调用、执行反馈和可复用能力,可作为 Agent 工作流可靠性的补充线索。
SwanWeave:One-Stage Multi-Task Instruction-Guided 3D Spatial Audio Editing
中文标题SwanWeave:One-Stage Multi-Task Instruction-Guided 3D Spatial Audio Editing
关注理由涉及语音与音频中的新任务、数据或系统线索,可作为后续跟进清单的一部分。
VICAL: Vicinal Consistency Alignment for Long-Tailed Visual Recognition
中文标题VICAL:Vicinal Consistency Alignment 面向 Long-Tailed Visual Recognition
关注理由涉及模型安全、护栏路由、风险分类或治理评测,可作为安全评测与治理工具链的补充线索。
Cost-Aware Hierarchical Multi-Agent Ransomware Detection and Family Attribution
中文标题Cost-Aware Hierarchical Multi-Agent Ransomware Detection 与 Family Attribution
关注理由涉及工具调用、执行反馈和可复用能力,可作为 Agent 工作流可靠性的补充线索。
中文标题Beneath the Surface of Chains-of-Thought:A Mechanistic Interpretation of Reasoning Operations in LLMs
关注理由涉及可解释性中的新任务、数据或系统线索,可作为后续跟进清单的一部分。
Leveraging Imperfect Restoration for Data Availability Attack
中文标题Leveraging Imperfect Restoration 面向 Data Availability Attack
关注理由涉及检索、知识库问答与证据可靠性,可作为 RAG 评测和企业知识系统的补充线索。
中文标题Same Trajectory,Contradictory Rewards (ROBORMBENCH):Paraphrase Fragility in Vision Language Reward Models
关注理由涉及检索、知识库问答与证据可靠性,可作为 RAG 评测和企业知识系统的补充线索。
When LLM Decompilers Recompile More and Preserve Less
中文标题When LLM Decompilers Recompile More 与 Preserve Less
关注理由涉及工具调用、执行反馈和可复用能力,可作为 Agent 工作流可靠性的补充线索。
中文标题Distill Globally,Adapt Locally:Reasoning Distillation 与 Product-Type Test-Time Training 面向 可扩展 Trade-Up Recommendation
关注理由涉及任务设置、指标和失效案例,可补充模型评测与回归测试。
MEOX: Compact Multimodal Mixture-of-Experts for Earth Observation
中文标题MEOX:Compact Multimodal Mixture-of-Experts 面向 Earth Observation
关注理由涉及检索、知识库问答与证据可靠性,可作为 RAG 评测和企业知识系统的补充线索。
RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
中文标题RoboSPA:Can VLA Models Go Beyond Simple Scenes 与 Short-Horizon Tasks?
关注理由涉及工具调用、执行反馈和可复用能力,可作为 Agent 工作流可靠性的补充线索。
中文标题Large Language Models 面向 HVAC Operations in Building Energy Systems:A Critical Review of Methods,Applications,与 Deployment Readiness
关注理由涉及工具调用、执行反馈和可复用能力,可作为 Agent 工作流可靠性的补充线索。
中文标题Beyond Aggregate Scores:Behavioral Correctness Assumptions 面向 Assessing Reference-Based Automatic 评测 Methods
关注理由涉及任务设置、指标和失效案例,可补充模型评测与回归测试。
Cross-Domain Tracker Adaptation Without Target-Domain Labels via Vision-Language Agents
中文标题Cross-Domain Tracker Adaptation Without Target-Domain Labels 通过 Vision-Language Agents
关注理由涉及工具调用、执行反馈和可复用能力,可作为 Agent 工作流可靠性的补充线索。
Substrate-Aware AI Agents: Execution Context as a First-Class Input
中文标题Substrate-Aware AI Agents:Execution Context as a First-Class Input
关注理由涉及工具调用、执行反馈和可复用能力,可作为 Agent 工作流可靠性的补充线索。
First Things First: Teaching LLM-Based Agents to Prioritize Must-Haves before Nice-to-Haves
中文标题First Things First:Teaching 基于 LLM 的 Agents to Prioritize Must-Haves before Nice-to-Haves
关注理由涉及工具调用、执行反馈和可复用能力,可作为 Agent 工作流可靠性的补充线索。
A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment
中文标题一种Human-in-the-Loop 框架 面向 AI-Assisted Scoring in Large-Scale Writing 评估
关注理由涉及任务设置、指标和失效案例,可补充模型评测与回归测试。
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
中文标题SciDocBench:A Workflow-Centered 基准 与 Data Pipeline 面向 Scientific Document Understanding
关注理由涉及工具调用、执行反馈和可复用能力,可作为 Agent 工作流可靠性的补充线索。
阅读边界
- 自动排序会偏向有社区信号、代码信号和工程关键词的论文。
- 简报默认基于标题、摘要和公开元数据,不替代全文精读。
- 外部 API 限流或不可用时,相关信号会降级为空并在内部记录中保留说明。