Strengthen multimodal understanding of charts, documents, and visual evidence, Make agents use tools and reusable skills more reliably, Improve model reasoning, planning, and verification
Today tracks: strengthen multimodal understanding of charts, documents, and visual evidence, make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 477 candidate papers from the 2026-08-31 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Modality Disentangled Learning for Incomplete Multimodal Emotion Recognition: A Primitive Memory Distillation Perspective🔗
- 2Beyond the Payload: How User Invocation Shapes Coding Agent Vulnerability to Repository Poisoning🔗
- 3SIR: Self-improving Red-teaming for Compute Use Agents🔗
- 4OCR-MetaReasoning Benchmark: Evaluating the Meta-Reasoning Ability of MLLMs in Text-Rich Image Understanding🔗
- 5Reactivating Test-Time Scaling for Plane Geometry Problem Solving🔗
- 6Preference Shapes Relevance: Cross-component Hierarchical Semantic Alignment for Personalized Generative Retrieval🔗
What is worth tracking today
Today’s high-signal papers point to: strengthen multimodal understanding of charts, documents, and visual evidence, make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
strengthen multimodal understanding of charts, documents, and visual evidence
Signalthis paper targets the concrete research problem behind strengthen multimodal understanding of charts, documents, and visual evidence. It uses the title, abstract, and public signals around inference, code, multimodal, memory to frame the training and post-training task, data, or evaluation flow to improve strengthen multimodal understanding of charts, documents, and visual evidence. The main claim is the title, abstract, and public signals indicate: Multimodal Emotion Recognition (MER) systems often suffer from missing modalities in real-world scenarios
Keywordsinferencecodemultimodalmemory
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, evaluation, benchmark, eval to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Coding agents are increasingly used for software engineering tasks, including bootstrapping projects from third-party repositories whose integrity cannot be assumed
Keywordsagentevaluationbenchmarkeval
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, safety, benchmark, vision-language to frame the safety and alignment task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Computer use agents (CUAs) are vision-language models that perceive a screen and act on a real operating system through mouse, keyboard, and terminal, and they are increasingly deployed to automate everyday digital tasks
Keywordsagentsafetybenchmarkvision-language
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around inference, evaluation, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Text-rich image understanding requires multimodal large language models (MLLMs) to organize OCR (Optical Character Recognition)-grounded evidence across words, layout, fields, charts, and visual correspondences
Keywordsinferenceevaluationbenchmarkcode
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around inference, benchmark, code, multimodal to frame the reasoning and planning task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Plane geometry problem (PGP) solving has become a critical benchmark for multimodal reasoning because it requires accurate visual perception and precise multi-step symbolic deduction
Keywordsinferencebenchmarkcodemultimodal
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around retrieval, inference, latency, alignment to frame the systems and deployment task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Generative Retrieval (GR) has emerged as a promising paradigm by mapping queries directly to Semantic IDs (SIDs) with powerful representation capabilities for candidate items
Keywordsretrievalinferencelatencyalignment
Code/DataCheck the source paper
Other papers worth tracking
Beyond Token-Level Guidance: Inference-Time Alignment of Specialized LLMs via Cross-Family Representation Steering: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Pretrained, Curriculum-Tuned, and Ensembled: A Tracer-Aware Interactive Segmentation Pipeline for AutoPET V: Covers a concrete training and post-training signal; useful as a follow-up candidate.
UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Knowing Beyond the Known: Reinforced Knowledge Specification for Multi-Label Class-Incremental Learning: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
DICS: Exploring Data Intrinsic Consistency for Visual Instruction Selection: Covers a concrete training and post-training signal; useful as a follow-up candidate.
A Simple Transformer Pipeline for Full-Key Side-Channel Attacks on Uncropped Datasets: Covers a concrete data engineering signal; useful as a follow-up candidate.
Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Multimodal Shared Latent Representation of Narration, Microscope and iOCT Images for Phase Recognition in Vitreoretinal Surgery: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Responsible Integration of AI in Cancer Genomics: Barriers, Risks, and Pathways to Trustworthy Clinical Translation: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Low-Resource Preference Adaptation of LLMs via Activation-Based Label Propagation: Covers a concrete data engineering signal; useful as a follow-up candidate.
You Shouldn't Have Asked: A Pragmatics-Inspired Taxonomy for Evaluating LLM Refusals: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
The Fragility of Jailbreak Robustness Across Operational States: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
SingProbe Technical Report: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Failure or Drift? Evaluating Monocular SLAM under Synthetic and Real-World Corruptions: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
AdaPath: Query-Adaptive Path-Finding via Path-Bank for Multi-Hop Implicit Biomedical KGQA: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
ALTSTEER: Selective Safety Steering for Moving Beyond Hard Refusals to Constructive Alternatives: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
VeriCam: A Verification Baseline for the Classification of Unknown Data: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.