Improve code generation, execution feedback, and automated repair, Make agents use tools and reusable skills more reliably, Make RAG retrieval and knowledge-base QA more reliable
Today tracks: improve code generation, execution feedback, and automated repair, make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable.
This issue fetched and deduplicated 328 candidate papers from the 2026-08-18 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1ArguLens: An Open-Source System for Automated Essay Scoring and Label-Aware Feedback Generation🔗
- 2aDSL: Agentic 3D Creation via Joint Agent-Program Design🔗
- 3Learnware for CSI Feedback: Scene-specific Small Models Can Do Big🔗
- 4Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges🔗
- 5The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges🔗
- 6Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering🔗
What is worth tracking today
Today’s high-signal papers point to: improve code generation, execution feedback, and automated repair, make agents use tools and reusable skills more reliably, make RAG retrieval and knowledge-base QA more reliable. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around inference, evaluation, open-source, eval to frame the benchmarks and evaluation task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Most automated essay scoring (AES) systems output a single holistic score without interpretable evidence and rely on closed APIs that introduce data privacy and cost barriers
Keywordsinferenceevaluationopen-sourceeval
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, workflow, rag, serving to frame the code intelligence task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Programmatic representations provide a compelling paradigm for 3D content creation, enabling fine-grained edits, interpretability, and explicit structural control
Keywordsagentworkflowragserving
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around retrieval, deployment, latency, code to frame the training and post-training task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Intelligent channel state information (CSI) feedback is essential for realizing the high capacity and spectral efficiency goals of future 6G systems, yet existing deep learning solutions face a trade-off between model generalization and scenario-specific perfo
Keywordsretrievaldeploymentlatencycode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, alignment, evaluation, benchmark to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction
Keywordsagentalignmentevaluationbenchmark
Code/DataCheck the source paper
strengthen multimodal understanding of charts, documents, and visual evidence
Signalthis paper targets the concrete research problem behind strengthen multimodal understanding of charts, documents, and visual evidence. It uses the title, abstract, and public signals around deployment, latency, code, table to frame the multimodal models task, data, or evaluation flow to improve strengthen multimodal understanding of charts, documents, and visual evidence. The main claim is the title, abstract, and public signals indicate: LLMs increasingly rely on external contexts, such as pre-defined system prompts or retrieved documents, to improve generation quality
Keywordsdeploymentlatencycodetable
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, retrieval, inference, code to frame the retrieval and rag task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Answering questions accurately and efficiently in embodied scenarios presents significant challenges due to limited computational and memory resources for Vision Language Model (VLM) inference
Keywordsragretrievalinferencecode
Code/DataCheck the source paper
Other papers worth tracking
DMT-Dens: Density-preserving manifold visualization for biological data: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Debate Training Reduces Reward Hacking in RLAIF: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Benchmarking Automated Security Patch Backporting: How Far Are We?: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models: Covers a concrete training and post-training signal; useful as a follow-up candidate.
Leveraging generative hallucination and biophysics-informed modeling for unified biomolecular sequence-structure co-design: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
TokEval: A Tokenizer Evaluation Suite: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
PerFact: Perception-Derived Fact Prompting for 3D Brain MRI Report Generation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Training with synthetic data for drone detection in thermal imagery: Covers a concrete data engineering signal; useful as a follow-up candidate.
TraceSQL: Traceable Answerability Estimation for Reference-Free Text-to-SQL Verification: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Vision-Language Models for Analog Gauge Reading: An Empirical Study of Specialization, Transfer and Reliability: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Cross-View Correspondence Is a Measurement Intervention: Two-Sided Validation for Agent Evaluation and Credit Assignment: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
GADR: Gathering Architecture Decision Records from Meeting Transcriptions: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android Apps: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Denoised Variance-Based Pruning with Optimal Brain Bias Compensation: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
DEPT: Document Embedding Preservation Tuning for Unified Query Expansion and Retrieval: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
OOD Detection for EEG-based Machine Learning in High-Risk Environments: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.