Improve code generation, execution feedback, and automated repair, Strengthen multimodal understanding of charts, documents, and visual evidence, Make agents use tools and reusable skills more reliably
Today tracks: improve code generation, execution feedback, and automated repair, strengthen multimodal understanding of charts, documents, and visual evidence, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 466 candidate papers from the 2026-06-29 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Curvature-Guided Sheaf Diffusion for Unsupervised Community Detection on Heterophilic Graphs🔗
- 2RiverONE: Generating Knowledge-Intensive VLM by Simulated Quantum Machines🔗
- 3Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?🔗
- 4On the Vulnerability of Parameter-Level Defenses to Model Merging🔗
- 5Internal-State Probes Read the Situation, Not the Action: Three Negative Results for Pre-Action Misalignment Monitoring🔗
- 6Emergence of a Shared Canonical Object Frame from In-the-Wild Videos🔗
What is worth tracking today
Today’s high-signal papers point to: improve code generation, execution feedback, and automated repair, strengthen multimodal understanding of charts, documents, and visual evidence, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around evaluation, benchmark, code, open-source to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Detecting communities in heterophilic graphs -- where connected nodes often belong to different classes -- is hard for unsupervised methods: classical modularity and spectral methods are feature agnostic, while deep graph-clustering methods rely on contrastive
Keywordsevaluationbenchmarkcodeopen-source
Code/DataCheck the source paper
strengthen multimodal understanding of charts, documents, and visual evidence
Signalthis paper targets the concrete research problem behind strengthen multimodal understanding of charts, documents, and visual evidence. It uses the title, abstract, and public signals around inference, compression, code, vision-language to frame the multimodal models task, data, or evaluation flow to improve strengthen multimodal understanding of charts, documents, and visual evidence. The main claim is the title, abstract, and public signals indicate: Quantum computing provides a powerful paradigm for representing and transforming high-dimensional information through superposition, entanglement, and measurement-induced nonlinear features
Keywordsinferencecompressioncodevision-language
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, evaluation, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Rubric-based scoring has become a widely used paradigm in model evaluation, typically with LLM-as-a-Judge (LaaJ) for rubric scoring
Keywordsagentevaluationbenchmarkcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, evaluation, code, fine-tuning to frame the training and post-training task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: The training-free integration of expert models via model merging has exposed significant security risks, enabling free-riders to combine specialized models without authorization
Keywordsragevaluationcodefine-tuning
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, alignment, code, tool to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Probes on model internals could help monitor agentic systems if they identify harmful text or tool actions before those actions are generated
Keywordsagentalignmentcodetool
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around alignment, benchmark, code, video to frame the video generation task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Comparing object orientations and positions across different instances requires their poses to be expressed in a shared canonical frame
Keywordsalignmentbenchmarkcodevideo
Code/DataCheck the source paper
Other papers worth tracking
BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
LETT-NeXt: A Lightweight RECIST-Guided Model for 3D CT Lesion Segmentation: Covers a concrete multimodal models signal; useful as a follow-up candidate.
IHDec: Divergence-Steered Contrastive Decoding for Securing Multi-Turn Instruction Hierarchies: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Redefining Maritime Anomaly Detection via Equation-Grounded Synthetic Anomalies: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Toward Secure and Reliable PDDL Formalization of Large Language Models with Planner-in-the-Loop Feedback: Covers a concrete training and post-training signal; useful as a follow-up candidate.
SA-Homo: Scale Adaptive Homography Estimation for Scale Variation Scenarios: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Diagnosing and Mitigating Retrieval Bottlenecks in LLM-Based Cold-Start Recommendation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Timesteps of Mamba Align with Human Reading Times: Covers a concrete data engineering signal; useful as a follow-up candidate.
Experience Graphs: The Data Foundation for Self-Improving Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
UrbanCDNet: Appearance-Robust and Boundary-Aware Bitemporal Change Detection for Korean Urban Building Monitoring: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
AerialMetric: Benchmarking and Adapting UAV Monocular Metric Depth Estimation in the Real World: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
C$^{2}$R: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse Autoencoders: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
MESA: Prioritizing Vulnerable Communication Channels for Securing Multi-Agent Systems: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
The Human Creativity Benchmark: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
TraceLab: Characterizing Coding Agent Workloads for LLM Serving: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.