Make agents use tools and reusable skills more reliably, Make RAG retrieval and knowledge-base QA more reliable
Today tracks: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 304 candidate papers from the 2026-08-21 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Large Language Models at the Intersection of Software Engineering and Software Security:An Evidence-Centered Structured Survey and Research Agenda🔗
- 2Specification Portability Across LLM Development Agents: Cross-Agent Compatibility in Specification-Driven Software Migration🔗
- 3ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models🔗
- 4When Adaptation Hurts: Connecting Representational Drift to OOD Failures in MedSAM Fine-Tuning🔗
- 5Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems🔗
- 6PromptResponse: Optimizing Prompts for LLM Coding Tasks🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, workflow, evaluation, benchmark to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Large Language Models (LLMs) are moving from code completion toward repository-scale agents that retrieve context, edit files, execute tools, and participate in security-sensitive workflows
Keywordsagentworkflowevaluationbenchmark
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, workflow, retrieval, compression to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: This paper investigates cross-agent specification portability using Oracle-to-PostgreSQL migration as a controlled software transformation task
Keywordsagentworkflowretrievalcompression
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, serving, alignment, safety to frame the safety and alignment task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: While multimodal large language models (MLLMs) extend model capabilities beyond text, they also make safety alignment increasingly challenging
Keywordsagentservingalignmentsafety
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around serving, alignment, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Foundation models for medical image segmentation, like prompt-based MedSAM, generalize well across domains and modalities, often in zero or few-shot setups
Keywordsservingalignmentbenchmarkcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, rag, retrieval, inference to frame the retrieval and rag task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Retrieval-Augmented Generation (RAG) grounds Large Language Model (LLM) outputs in external knowledge, but RAG systems usually trust whatever they retrieve, creating a Security-Reliability Gap: high semantic relevance does not guarantee factual truth
Keywordsagentragretrievalinference
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around workflow, alignment, evaluation, code to frame the code intelligence task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Large language models (LLMs) are increasingly used in research workflows and software development pipelines, yet their output remains sensitive to input prompt variations
Keywordsworkflowalignmentevaluationcode
Code/DataCheck the source paper
Other papers worth tracking
AudioWorldSim: Realistic Binaural Audio Datasets For World Models: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
RDANet: Relative Degradation Aware Network for Infrared Small Target Detection: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and Aggregation: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
CoST: Semantic-Aware Urban Understanding via Spatial-Temporal Alignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Trojaning the Alignment: Stealthy Backdoor Attacks against Graph Foundation Models: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Breaking High Confidence: Practical Face Impersonation under High-Security Thresholds: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
BC-Bench: Evaluating Agentic Engineering in a Domain-Specific Language for ERP: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Routing Before Looking: Query-Adaptive Evidence Acquisition for Long-form Video Understanding: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Generating Multi-view Adversarial Examples for Visual Geometry Grounded Transformer: Covers a concrete multimodal models signal; useful as a follow-up candidate.
AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Equivalence Verification: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
TopoSurfel: Closing the Loop between Gaussian Surfels and Meshes for Surface Reconstruction: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
RiskTraf: Risk-Extrapolated Residual Learning for Multi-Variate Traffic Flow Prediction: Covers a concrete code intelligence signal; useful as a follow-up candidate.
Meta-clustering of milk mid-infrared spectra identifies dairy cow groups associated with negative energy balance in early lactation: Covers a concrete code intelligence signal; useful as a follow-up candidate.
Rethinking Expressivity and Efficiency in Test-Time Training: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
VT-MUSE: Multimodal Unified Sequential Visuotactile Representation Learning for Manipulation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
EnSI-RAG: Entity-Structure-Indexed Retrieval-Augmented Generation for Long-Document Question Answering: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Benchmarking Patent Drafting from Inventor-Style Disclosures: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Personalized Privacy Control in LLMs via Attention Head Intervention: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.