Improve code generation, execution feedback, and automated repair, Improve model reasoning, planning, and verification, Make agents use tools and reusable skills more reliably
Today tracks: improve code generation, execution feedback, and automated repair, improve model reasoning, planning, and verification, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 355 candidate papers from the 2026-09-02 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1From Multi-Fisheye Sensing to Panoramic Perception: A Parallax-Aware Onboard Platform for Ultra-Low-Altitude UAVs🔗
- 2H3DNAS: Hardware-Aware ONNX-Native 3D Point Cloud Model Compression🔗
- 3Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds🔗
- 4Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems🔗
- 5Removing Speech, Keeping Activities: A Privacy Firewall for Acoustic Sensing in Assisted Living🔗
- 6YesTrack: Referring Multi-Object Tracking via MLLM-based Yes/No Verification🔗
What is worth tracking today
Today’s high-signal papers point to: improve code generation, execution feedback, and automated repair, improve model reasoning, planning, and verification, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around alignment, evaluation, open-source, eval to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Ultra-low-altitude unmanned aerial vehicles (UAVs) require surround vision near buildings, vegetation, and other obstacles
Keywordsalignmentevaluationopen-sourceeval
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around inference, compression, code, memory to frame the reasoning and planning task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Deploying 3D point cloud models on edge hardware such as the NVIDIA Jetson Orin Nano is severely constrained by compute and memory budgets
Keywordsinferencecompressioncodememory
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, inference, deployment, alignment to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluation can support
Keywordsagentinferencedeploymentalignment
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, evaluation, code, memory to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through textual reflection
Keywordsagentevaluationcodememory
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around serving, deployment, evaluation, code to frame the speech and audio task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Acoustic sensing offers a promising non-intrusive approach for monitoring daily activities of older adults, yet speech privacy concerns remain a critical barrier to real-world deployment
Keywordsservingdeploymentevaluationcode
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, latency, alignment, code to frame the multimodal models task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Referring multi-object tracking (RMOT) aims to track every instance in a video that matches a given language expression
Keywordsraglatencyalignmentcode
Code/DataCheck the source paper
Other papers worth tracking
If It Moves, Radar Knows: A Physics-Aware Radar Transformer for Class-Agnostic Moving-Object Detection: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
TAME: Temporal-Aware Mixture-of-Experts for Text-Video Retrieval: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Beyond Modality Harmony: Orthogonal Purification and Topology-Guided MoE for Conflict-Aware Multimodal Recommendation: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Transfer Safety Awareness for Cross-Modal Safety Drift in Multimodal Large Language Models: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Benchmarking RAW and RGB Restoration in Image Signal Processors: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
SPADE: SPaT Attack Detection from the Connected Vehicle's Perspective: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Rethinking the Teacher-Student Framework for Test-Time Adaptation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Signal or Noise? Auditing Rotation-Induced Saliency Drift in Medical and Aerial Imaging: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
DynG-Diff: A State-Aware Dynamic Guidance Diffusion Framework for Probabilistic Time Series Forecasting: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Train What You Deploy: Closing the MLP Reachability Gap in Low-Rank Clone Distillation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
GRADSOLVE: fast exact gradients for ODE ensembles on GPUs: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Video-Based Palm-Vein Authentication under Challenging Conditions: Covers a concrete video generation signal; useful as a follow-up candidate.
CodePoisonRAG: Knowledge Poisoning Attacks on Retrieval-Augmented Code Generation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
RVSD: Retrieval Vision Sparse Decoding for Mitigating Visual Hallucinations in Large Vision-Language Models: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
CORAL: An LLM-Native Harness for Production Recommender Systems: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Query Rewriting for Complex Object Segmentation in 4D Gaussian Representations: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.