Improve code generation, execution feedback, and automated repair, Make agents use tools and reusable skills more reliably
Today tracks: improve code generation, execution feedback, and automated repair, improve code generation, execution feedback, and automated repair, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 393 candidate papers from the 2026-07-01 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1BaseRT: Best-in-Class LLM Inference on Apple Silicon via Native Metal🔗
- 2The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models🔗
- 3Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?🔗
- 4From Prediction Uncertainty to Conformalized Distance Fields for Safe Motion Planning🔗
- 5Rise From The Ashes: LLM-based Static Analysis for Deep Learning Framework Bugs🔗
- 6EquiSteer: Cross-Attention Steering Towards a Fairer Text-Guided Image Generation🔗
What is worth tracking today
Today’s high-signal papers point to: improve code generation, execution feedback, and automated repair, improve code generation, execution feedback, and automated repair, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around inference, deployment, latency, code to frame the systems and deployment task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: We present BaseRT, a native Metal inference runtime for large language models (LLMs) on Apple Silicon, and report the highest inference throughput on this hardware to date
Keywordsinferencedeploymentlatencycode
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around serving, alignment, safety, evaluation to frame the safety and alignment task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Safety alignment of text-to-image (T2I) diffusion models aims to suppress harmful generations while preserving utility on benign prompts
Keywordsservingalignmentsafetyevaluation
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents by applying patches to real repositories and comparing runtime against unoptimized baselines and official reference patches
Keywordsagentragbenchmarkcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around rag, safety, benchmark, planning to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Safe motion planning in dynamic environments requires reasoning about the uncertainty in predicted obstacle motion without sacrificing real-time performance
Keywordsragsafetybenchmarkplanning
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, workflow, rag, retrieval to frame the code intelligence task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Deep learning (DL) frameworks are critical AI infrastructures that often hide bugs with serious security implications
Keywordsagentworkflowragretrieval
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, inference, alignment, code to frame the vision and image generation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Text-to-image diffusion models power everyday creative tasks, but they still reproduce the demographic biases in their training data
Keywordsraginferencealignmentcode
Code/DataCheck the source paper
Other papers worth tracking
Seahorse: A Unified Benchmarking Framework for Spatiotemporal Event Modeling: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Human-Machine Collaboration on Generative Meta-Learning: Model and Algorithm: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
ClinRAG-GRAPH: Clinical-prior Retrieval-Augmented Graph Model with Domain Adversarial Learning for Breast pCR Prediction: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
From Structural Equation Modelling to Double Machine Learning: Robustness Analysis for Survey-Based Research: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
PAPA: Online Personalized Active Preference Alignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Beyond the Prompt: Jailbreaking Function-Calling LLMs via Simulated Moderation Traces: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Learning Gait-Aware Quadruped Locomotion with Temporal Logic Specifications: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
LIST3R: Long-sequence Instance-aware 3D Reconstruction: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
SoK: Attack and Defense Landscape of Mobile On-device AI Systems: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Managed Autonomy at Runtime: Gear-Based Safety and Governance for Single- and Multi-Agent Cyber-Physical Systems: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Autonomous Scientific Discovery via Iterative Meta-Reflection: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
MemSyco-Bench: Benchmarking Sycophancy in Agent Memory: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Svarna: An Open Corpus Workbench for Modern Greek: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Meta-Transfer Learning for mmWave Beam Alignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Active Spatial Guidance: Eliminating Injected Positional Mechanisms in Vision Transformers: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
MedCAGD: Context-Aware Gated Decoder for Efficient Medical Image Segmentation: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.