2026-07-17

Internal Generation Record

Internal generation metadata: 289 candidate papers.

Published 2026-07-17 Target source 2026-07-15 Actual source 2026-07-15 Candidates 289 Featured 6 Tracked 20

Generation Record

This page preserves selected papers, candidate scale, and source-date metadata for traceability. The page only changes presentation, not selected papers, ordering, or counts.

Internal generation record. Fetched at 2026-07-16T22:12:39.870800+00:00. Generated at 2026-07-16T22:13:50.865860+00:00. Machine-readable details stay under data/processed and data/reports.

Selected papers

RankTakeawayTopicarXiv
1make agents use tools and reusable skills more reliablyRobotics and Embodied AI2607.13653
2make agents use tools and reusable skills more reliablyAgents and Tool Use2607.13602
3test temporal consistency and motion realism in video generationBenchmarks and Evaluation2607.13457
4improve model reasoning, planning, and verificationTraining and Post-training2607.13394
7improve code generation, execution feedback, and automated repairCode Intelligence2607.13674
8make agents use tools and reusable skills more reliablyRetrieval and RAG2607.14044
5improve code generation, execution feedback, and automated repairTraining and Post-training2607.14070
6make agents use tools and reusable skills more reliablyBenchmarks and Evaluation2607.13940
9make agents use tools and reusable skills more reliablyBenchmarks and Evaluation2607.14004
10improve code generation, execution feedback, and automated repairBenchmarks and Evaluation2607.13941
11make RAG retrieval and knowledge-base QA more reliableSafety and Alignment2607.13596
12improve model reasoning, planning, and verificationBenchmarks and Evaluation2607.13586
13make agents use tools and reusable skills more reliablyBenchmarks and Evaluation2607.13551
14identify and reduce safety, jailbreak, and alignment risksData Engineering2607.13423
15make agents use tools and reusable skills more reliablyRetrieval and RAG2607.14046
16make RAG retrieval and knowledge-base QA more reliableRetrieval and RAG2607.14035
17make agents use tools and reusable skills more reliablyAgents and Tool Use2607.14006
18make agents use tools and reusable skills more reliablyBenchmarks and Evaluation2607.13998
19make RAG retrieval and knowledge-base QA more reliableData Engineering2607.13976
20make agents use tools and reusable skills more reliablyAgents and Tool Use2607.13938
21improve model reasoning, planning, and verificationMultimodal Models2607.13931
22make RAG retrieval and knowledge-base QA more reliableMultimodal Models2607.13892
23improve model reasoning, planning, and verificationBenchmarks and Evaluation2607.13860
24improve model reasoning, planning, and verificationTraining and Post-training2607.13753
25make agents use tools and reusable skills more reliablyAgents and Tool Use2607.13716
26use benchmarks and evaluations to expose model weaknessesTraining and Post-training2607.13689