Make RAG retrieval and knowledge-base QA more reliable, Improve code generation, execution feedback, and automated repair, Make agents use tools and reusable skills more reliably
Today tracks: make RAG retrieval and knowledge-base QA more reliable, improve code generation, execution feedback, and automated repair, make agents use tools and reusable skills more reliably.
This issue fetched and deduplicated 497 candidate papers from the 2026-09-24 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1Beyond Average Safety: Chance-Constrained LLM Fine-tuning🔗
- 2ENDOPROMPT: Victim-Side Pseudo-References for Utility Degradation🔗
- 3PUBG Ally: A Conversational Embodied Agent as an AI Teammate🔗
- 4Hallucination Neurons and Where to Find Them: An Investigation into the existence of Hallucination Neurons🔗
- 5TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening🔗
- 6C3M: Cross-Session Multimodal Memory Maintenance for Long-Horizon Tasks🔗
What is worth tracking today
Today’s high-signal papers point to: make RAG retrieval and knowledge-base QA more reliable, improve code generation, execution feedback, and automated repair, make agents use tools and reusable skills more reliably. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around rag, serving, safety, fine-tuning to frame the safety and alignment task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Fine-tuning large language models on new objectives can improve helpfulness, instruction following, or domain-specific performance, but it can also induce regressions on safety-critical prompts
Keywordsragservingsafetyfine-tuning
Code/DataCheck the source paper
improve code generation, execution feedback, and automated repair
Signalthis paper targets the concrete research problem behind improve code generation, execution feedback, and automated repair. It uses the title, abstract, and public signals around deployment, benchmark, code, search to frame the reasoning and planning task, data, or evaluation flow to improve improve code generation, execution feedback, and automated repair. The main claim is the title, abstract, and public signals indicate: Prompt injection can degrade benign task performance without eliciting harmful content
Keywordsdeploymentbenchmarkcodesearch
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, tool use, latency, compression to frame the training and post-training task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: We introduce PUBG Ally, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate
Keywordsagenttool uselatencycompression
Code/DataCheck the source paper
identify and reduce safety, jailbreak, and alignment risks
Signalthis paper targets the concrete research problem behind identify and reduce safety, jailbreak, and alignment risks. It uses the title, abstract, and public signals around alignment, safety, evaluation, open-source to frame the interpretability task, data, or evaluation flow to improve identify and reduce safety, jailbreak, and alignment risks. The main claim is the title, abstract, and public signals indicate: Interpretable machine learning for Large Language Models (LLMs) increasingly relies on sparse probing methods that identify small sets of neurons claimed to detect and causally influence behaviors such as factuality recall, safety alignment, and hallucination
Keywordsalignmentsafetyevaluationopen-source
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around rag, evaluation, benchmark, code to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Ligand-based virtual screening (LBVS) is a practical first-pass tool in early-stage drug discovery, but existing benchmarks can overestimate performance through random negatives, easy decoys, limited target coverage, and non-standardized evaluation protocols
Keywordsragevaluationbenchmarkcode
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around serving, compression, code, multimodal to frame the multimodal models task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Long-horizon tasks require preserving and later recovering cross-session evidence under a bounded, query-blind memory budget
Keywordsservingcompressioncodemultimodal
Code/DataCheck the source paper
Other papers worth tracking
Dense Coverage, Sparse Refinement: Byte-Constrained Cooperative Perception: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
A Systematic Multi-Domain Evaluation of Document Retrievers: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Demystifying Agent Skills for Smart Contract Auditing: Design, Effectiveness, Behavioral Impact: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
An auditable conditional-strategy framework for open-ended decision-making in complex lung cancer: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Learning a Flow to Self-Supervised Representations: Covers a concrete training and post-training signal; useful as a follow-up candidate.
SkinAgent AI: A Safety-Grounded Multimodal Agentic Framework for Non-Diagnostic Skincare Support: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
AEGIS: Audio Endogenous Guarding via Internal Signals Against Large Audio-Language Model Jailbreaks: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
From Text Decisions to Pixels: An Study of Jev-Style Visual Choice Model: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
ComplexSync: High-Fidelity and Real-Time Lip Sync in Complex Scenarios: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Med-AR: Autoregressive Vision-Language Pretraining for Long-Tailed Chest X-Ray Classification and Uncertainty-Aware Evaluation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Scope Before You Persist: Preventing Cross-Family Interference in Agent Memory: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
FoCal: Frequency-Oriented Cross-Modal Interaction and Spectral Calibration for Aerial Visible-Infrared Object Detection: Covers a concrete reasoning and planning signal; useful as a follow-up candidate.
TRACE: Interactive Bi-Directional Tracing of Monochrome Cables Amid Clutter: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
ELF-REG: Scaling Continuous Diffusion Language Models to Reasoning Tasks: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Calibrated Decision Models for Autonomous Penetration-Testing Harnesses: JEV and Laya as System One Decision Layers for LLM-Driven Pentest Agents: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
LLM Agents Can Easily Tamper With Their Own Traces: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Smartphone-Based Method for Automated Speed Enforcement: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.