Improve code generation, execution feedback, and automated repair, Make RAG retrieval and knowledge-base QA more reliable
Today tracks: improve code generation, execution feedback, and automated repair, improve code generation, execution feedback, and automated repair, improve code generation, execution feedback, and automated repair.
This issue fetched and deduplicated 355 candidate papers from the 2026-07-28 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1OmniPhys: Knowledge-Graph-Driven Benchmarking and Collective Optimization for Physical Commonsense in Text-to-Image Generation🔗
- 2Face De-Identification: A Domain-Centric Survey from Capture to Processing🔗
- 3A Unified Benchmark and Modality-Adaptive Network for Day-and-Night Drone-View Geo-Localization🔗
- 4ReLATE: Reliability-Guided Evidence Fusion for Robust UAV--Satellite cross-view Geo-Localization🔗
- 5Group Equivariant Diffusion for Anomaly Detection in Computational Cytology🔗
- 6DensFiLM: Density-Conditioned Video Saliency for Crowd Scenes🔗
What is worth tracking today
Today’s high-signal papers point to: improve code generation, execution feedback, and automated repair, improve code generation, execution feedback, and automated repair, improve code generation, execution feedback, and automated repair. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around alignment, evaluation, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: While text-to-image models exhibit remarkable visual fidelity, they frequently violate fundamental physical commonsense
Keywordsalignmentevaluationbenchmarkcode
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around serving, evaluation, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Face de-identification (De-ID) aims to remove or conceal personally identifiable facial features in images or videos to prevent identity recognition while preserving utility for downstream tasks
Keywordsservingevaluationbenchmarkcode
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around alignment, evaluation, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Most existing drone-view geo-localization (DVGL) benchmarks contain drone imagery captured under a single illumination condition and lack geographically aligned visible drone images, infrared drone images, and satellite images from the same locations
Keywordsalignmentevaluationbenchmarkcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, retrieval, compression, benchmark to frame the benchmarks and evaluation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Unmanned aerial vehicle (UAV)-satellite cross-view geo-localization matches UAV images against satellite imagery and has achieved impressive accuracy on clean (non-degraded) image benchmarks
Keywordsragretrievalcompressionbenchmark
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, alignment, code to frame the multimodal models task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Computational cytology on whole-slide images is challenging because malignant cells are rare, heterogeneous, and annotated slides are scarce
Keywordsraginferencealignmentcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, evaluation, code, video to frame the video generation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Video saliency models typically apply a single fixation strategy across crowd scenes, despite systematic changes in attention with crowd density
Keywordsragevaluationcodevideo
Code/DataCheck the source paper
Other papers worth tracking
CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model: Covers a concrete multimodal models signal; useful as a follow-up candidate.
Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
ContractHIL-HLS: Contract-Aligned Multi-Agent Workflow with Hardware-in-the-Loop Feedback for HLS Design: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
ObliCity: A Benchmark and Baseline for Roof-to-Ground Projection Displacement Correction: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities: Covers a concrete multimodal models signal; useful as a follow-up candidate.
SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Less is More: Modality-Decoupling for General AIGC Audio-Video Detection: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
A Control System, a Dataset, and a Recipe for Making Frozen LLM Agents Learn a Domain: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
VetClaw: An Edge-Cloud Multimodal Agentic System for Veterinary Disease Screening: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation Models: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
MemLens: A Value-Aware Memory Management System with Interactive Analytics for LLM-based Agents: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Shieldstral: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Hypothesis-Driven Shelf Generation for Personalised Recommendation: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Fine-Grained Food Image Understanding via Target-Aware Data Alignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
FLASH: Efficient Impact Fall Detection with Unified Hypergraph State-Space Model: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Cooperative Multi-UAV Navigation in Complex Environments via Systematic Multi-Agent Deep Reinforcement Learning: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.