Make agents use tools and reusable skills more reliably, Strengthen multimodal understanding of charts, documents, and visual evidence, Make RAG retrieval and knowledge-base QA more reliable
Today tracks: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, strengthen multimodal understanding of charts, documents, and visual evidence.
This issue fetched and deduplicated 363 candidate papers from the 2026-08-13 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1ProME: Prototype-Margin Environments with Repair-Aware Selection for Group-Robust Learning🔗
- 2TopoIntent: Compiling Security Intent into Executable, Compliance-Checked Network Topologies🔗
- 3Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding🔗
- 4Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings🔗
- 5Learning Unified Video and Image Representation for Video Face Forgery Detection🔗
- 6Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, strengthen multimodal understanding of charts, documents, and visual evidence. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around rag, deployment, alignment, evaluation to frame the safety and alignment task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Group-robust learning is crucial for maintaining accuracy on rare subpopulations when training-group labels are unavailable
Keywordsragdeploymentalignmentevaluation
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around rag, retrieval, serving, alignment to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Enterprise security topology design requires translating business intent, regulatory requirements, and risk assumptions into zones, boundary devices, inter-zone paths, and access-control policies
Keywordsragretrievalservingalignment
Code/DataCheck the source paper
strengthen multimodal understanding of charts, documents, and visual evidence
Signalthis paper targets the concrete research problem behind strengthen multimodal understanding of charts, documents, and visual evidence. It uses the title, abstract, and public signals around inference, alignment, benchmark, vision-language to frame the benchmarks and evaluation task, data, or evaluation flow to improve strengthen multimodal understanding of charts, documents, and visual evidence. The main claim is the title, abstract, and public signals indicate: Referring Expression Comprehension (REC) is commonly studied under dataset-specific fine-tuning, resulting in specialist models with limited cross-dataset generalization
Keywordsinferencealignmentbenchmarkvision-language
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, retrieval, evaluation, search to frame the retrieval and rag task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Parliamentary proceedings are a primary record of democratic deliberation, yet their volume and fragmentation make multi-perspective access difficult for citizens, journalists, and researchers
Keywordsragretrievalevaluationsearch
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around serving, alignment, benchmark, code to frame the video generation task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Face forgery detection is crucial for preserving the security and integrity of facial data given the rapid developments in face manipulation techniques and deep generative models
Keywordsservingalignmentbenchmarkcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, retrieval, safety, benchmark to frame the code intelligence task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Self-improving LLM agents convert successful trajectories into persistent cross-task state
Keywordsagentretrievalsafetybenchmark
Code/DataCheck the source paper
Other papers worth tracking
How Good are Foundation Models in Longitudinal MRI Disease Progression Reasoning?: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
VALG: An Agentic System for ML Theory Research: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
P2Fusion: Prompt-based Progressive Infrared-Visible Image Fusion via Dual-Prior Distillation: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
PatchGen: Learning Soft Intra-Image Predictive Subsets for Visual Generalization: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Rules or Character? Scaling Laws for AI Safety Design: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Operationalizing Cyber Threat Intelligence with GraphRAG: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
STAR: Structured Tokenization and Target-Aware Interest Representation for PCVR Prediction: Covers a concrete data engineering signal; useful as a follow-up candidate.
Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Fast A/B/n Testing: Exact Multi-Policy Comparison via Tree-Coupled Feedback Sharing: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Query Translation vs. Cross-Lingual Embeddings for Sinhala-Tamil E-Government Information Retrieval: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Vero: Can AI Agents Build Formally Verified Software Repositories?: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Mind the Context: Continual Learning of Socially Appropriate Robot Actions via Environmental-Social Disentanglement: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.