Make agents use tools and reusable skills more reliably, Test temporal consistency and motion realism in video generation, Make RAG retrieval and knowledge-base QA more reliable
Today tracks: make agents use tools and reusable skills more reliably, test temporal consistency and motion realism in video generation, make RAG retrieval and knowledge-base QA more reliable.
This issue fetched and deduplicated 292 candidate papers from the 2026-07-22 source date, then selected 6 featured papers and 20 additional mentions.
Featured
- 1PerfAgent: Profiler-Guided Iterative Refinement for Repository-Level Code Optimization🔗
- 2LKValues: Aligning Large Language Models with Sri Lankan Societal Values🔗
- 3Not All Patches are Equal: Sampling Matters for Visible-Infrared Pre-Training🔗
- 4Small, Free, and Effective: Orchestrating Open-Weight Small Language Models to Outperform Single LLM for Malware Analysis🔗
- 5Development of an automated, reliable, and clinically meaningful artificial intelligence (AI) tool for diagnosing cardiac disease from conventional cardiovascular magnetic resonance (CMR) images🔗
- 6PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference🔗
What is worth tracking today
Today’s high-signal papers point to: make agents use tools and reusable skills more reliably, test temporal consistency and motion realism in video generation, make RAG retrieval and knowledge-base QA more reliable. The notes below focus on the core problem, method signal, main claim, and keywords for each featured paper.
Featured papers: core problem, method signal, and keywords
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, workflow, serving, benchmark to frame the agents and tool use task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Large language model (LLM) agents now perform well on correctness-oriented repository-level tasks, including SWE-Bench issue resolution and feature implementation in real codebases
Keywordsagentworkflowservingbenchmark
Code/DataCheck the source paper
test temporal consistency and motion realism in video generation
Signalthis paper targets the concrete research problem behind test temporal consistency and motion realism in video generation. It uses the title, abstract, and public signals around alignment, evaluation, benchmark, fine-tuning to frame the benchmarks and evaluation task, data, or evaluation flow to improve test temporal consistency and motion realism in video generation. The main claim is the title, abstract, and public signals indicate: Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms
Keywordsalignmentevaluationbenchmarkfine-tuning
Code/DataCheck the source paper
make RAG retrieval and knowledge-base QA more reliable
Signalthis paper targets the concrete research problem behind make RAG retrieval and knowledge-base QA more reliable. It uses the title, abstract, and public signals around retrieval, alignment, benchmark, code to frame the vision and image generation task, data, or evaluation flow to improve make RAG retrieval and knowledge-base QA more reliable. The main claim is the title, abstract, and public signals indicate: Visible-infrared (VIS-IR) alignment is a key pre-training task for robust multi-sensor perception
Keywordsretrievalalignmentbenchmarkcode
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around agent, deployment, benchmark, eval to frame the benchmarks and evaluation task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Malware analysis demands rapid interpretation of complex detonation reports spanning filesystem, network, and process behaviours
Keywordsagentdeploymentbenchmarkeval
Code/DataCheck the source paper
make agents use tools and reusable skills more reliably
Signalthis paper targets the concrete research problem behind make agents use tools and reusable skills more reliably. It uses the title, abstract, and public signals around inference, code, multimodal, open-source to frame the data engineering task, data, or evaluation flow to improve make agents use tools and reusable skills more reliably. The main claim is the title, abstract, and public signals indicate: Aims: Cardiovascular magnetic resonance (CMR) imaging enables non-invasive assessment of myocardial structure, function, and pathology, but requires substantial experience in interpretation of CMR images that could be supported by artificial intelligence (AI)-
Keywordsinferencecodemultimodalopen-source
Code/DataCheck the source paper
improve model reasoning, planning, and verification
Signalthis paper targets the concrete research problem behind improve model reasoning, planning, and verification. It uses the title, abstract, and public signals around rag, inference, serving, alignment to frame the reasoning and planning task, data, or evaluation flow to improve improve model reasoning, planning, and verification. The main claim is the title, abstract, and public signals indicate: Large language models (LLMs) provide strong reasoning capabilities but are expensive to serve at scale, whereas small language models (SLMs) are cheaper but less reliable on difficult problems
Keywordsraginferenceservingalignment
Code/DataCheck the source paper
Other papers worth tracking
D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Lean-SAM2: Target-Anchored Memory and Encoder Acceleration for SAM2: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Evolving Cache Schedules for Fast Diffusion Policy Inference: Covers inference cost, latency, throughput, and deployment constraints; useful for systems optimization.
Interpretable Fuzzy Rule-Based Regression Extension for Ex-Fuzzy Library: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Adaptive Bayesian Online Learning via Expert Aggregation: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
RIM: A Retrieval-In-Matching Framework for Cross-Domain Global Visual Localization of UAVs: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Cumsum-Composable Phase Transport for Low-Cost Streaming Keyword Spotting: Covers a concrete video generation signal; useful as a follow-up candidate.
A Framework of User Experience Principles for Human-AI Agent Interaction in the Workplace: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Harnessing Disagreement: Detecting Correlated Agreement Blindness in Multi-Agent Triage: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Efficient Clustering with Provable Guardrails for LLM Inference at Scale: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Edge Intelligence in Civil Aviation: Paradigms, Techniques, and Applications: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
FMRP-LEAN: A HIPAA-Compliant AI-Augmented LIMS Architecture for End-to-End Clinical Assay Workflow Optimization: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
Closing the Lab-to-Store Gap: A Data-Efficient Post-Training and Experience-Driven Learning VLA Framework for Retail Humanoids: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills: Covers model safety, guardrail routing, risk classification, or governance evaluation; useful as a safety workflow lead.
Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
Unified Prediction and Planning via Conflict-Aware Disjoint Parameter Training: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
When Shippers Become Algorithms: Candidate Exposure, Information Design, and the Concentration of LLM-Mediated Freight Markets: Covers tool use, execution feedback, and reusable capabilities; useful as an agent reliability lead.
When Does Knowledge Distillation Hurt? Reliability-Aware Distillation for Low-Resource Language Summarization: Covers retrieval, knowledge-base QA, and evidence reliability; useful as a RAG evaluation lead.
G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection: Covers task design, metrics, and failure cases; useful for model evaluation and regression tests.
Reading boundaries
- Automated ranking favors papers with community, code, and applied-engineering signals.
- Briefs are based on titles, abstracts, and public metadata by default, not full-paper review.
- External API failures degrade optional signals and are reflected in internal records.