← Discover MCPs and Agents
D
AgentAI & MLGitHub

Deep-Research-Survey

A Systematic Survey of Deep Research

Links

README

From the repo.

🌐 A Systematic Survey of Deep Research

Project PDF Web Status Preprints arXiv

We will continuously update this repo.

If you like our project, please give us a star ⭐ on GitHub for the latest update.

Overview

This repository contains a systematic collection of research papers on Deep Research (DR). We organize papers across key categories including four key components in DR, widely-used training paradigms in DR and relevant benchmark & resource.

Deep Research Overview

For more details, please check our survey! Our survey collection presents a comprehensive and systematic overview of deep research systems, including a clear roadmap, foundational components, practical implementation techniques, important challenges, and future directions. As the field of deep research continues to evolve rapidly, we are committed to continuously updating this survey to reflect the latest progress in this area

πŸ“£ Latest News

[2025.11.25] πŸŽ‰πŸŽ‰πŸŽ‰ We release our survey Deep Research: A systematic Survey. Thanks to my awesome co-authors🀩. Feel free to contact me if you are interested in this topic and want to discuss me.

🎬 Table of Content

πŸ“š Reading List

will be updated as soon as possible!

To get started with Deep Research, we recommend the representative and often seminal papers listed below. Reviewing this selection will provide a solid overview of the field.

Query Planning

VenueDatePaper TitleURL
ICLR 202321 May 2022Least-to-Most Prompting Enables Complex Reasoning in Large Language Modelshttps://arxiv.org/abs/2205.10625
NeurIPS 202317 May 2023Tree of Thoughts: Deliberate Problem Solving with Large Language Modelshttps://arxiv.org/abs/2305.10601
ACL 202421 Jun 2024Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answeringhttps://aclanthology.org/2024.acl-long.397/
Arxiv20 Sep 2023Chain-of-Verification Reduces Hallucination in Large Language Modelshttps://arxiv.org/abs/2309.11495
EMNLP 202323 May 2023Query Rewriting for Retrieval-Augmented Large Language Modelshttps://aclanthology.org/2023.emnlp-main.322/
COLM 202528 Feb 2025DeepRetrieval: Hacking Real Search Engines and Retrievers with LLMs via RLhttps://arxiv.org/abs/2503.00223
Arxiv11 Oct 2025CardRewriter: Leveraging Knowledge Cards for Long-Tail Query Rewriting on Short-Video Platformshttps://arxiv.org/abs/2510.10095
NeurIPS 202525 Jan 2025Improving Retrieval-Augmented Generation through Multi-Agent Reinforcement Learninghttps://arxiv.org/abs/2501.15228
NAACL 202414 Nov 2023LLatrieval: LLM-Verified Retrieval for Verifiable Generationhttps://aclanthology.org/2024.naacl-long.305/
ACL 2024β€”DRAGIN: Dynamic Retrieval Augmented Generation based on the Information Needs of LLMshttps://aclanthology.org/2024.acl-long.702/
WWW 202518 Jul 2024Retrieve, Summarize, Plan: Advancing Multi-hop QA with an Iterative Approachhttps://dl.acm.org/doi/10.1145/3701716.3716889
Arxiv10 Jun 2025RAISE: Enhancing Scientific Reasoning in LLMs via Step-by-Step Retrievalhttps://arxiv.org/abs/2506.08625
Arxiv20 May 2025s3: You Don’t Need That Much Data to Train a Search Agent via RLhttps://arxiv.org/abs/2505.14146
Arxiv28 Aug 2025AI-SearchPlanner: Modular Agentic Search via Pareto-Optimal Multi-Objective RLhttps://arxiv.org/abs/2508.20368
COLM 202512 Mar 2025Search-r1: Training LLMs to Reason and Leverage Search Engines with RLhttps://arxiv.org/abs/2503.09516
Arxiv7 Mar 2025R1-Searcher: Incentivizing the Search Capability in LLMs via RLhttps://arxiv.org/abs/2503.05592
Arxiv22 May 2025R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via RLhttps://arxiv.org/abs/2505.17005
NAACL 202517 Dec 2024RAG-Star: Enhancing Deliberative Reasoning with Retrieval-Augmented Verification and Refinementhttps://aclanthology.org/2025.naacl-long.361/
ACL 202521 Jan 2025Divide-Then-Aggregate: An Efficient Tool Learning Method via Parallel Tool Invocationhttps://aclanthology.org/2025.acl-long.1401/
Arxiv29 Jul 2025DeepSieve: Information Sieving via LLM-as-a-Knowledge-Routerhttps://arxiv.org/abs/2507.22050
Arxiv3 Feb 2025DeepRAG: Thinking to Retrieve Step by Step for Large Language Modelshttps://arxiv.org/abs/2502.01142
Arxiv1 Aug 2025MAO-ARAG: Multi-Agent Orchestration for Adaptive Retrieval-Augmented Generationhttps://arxiv.org/abs/2508.01005

Information Acquisition

Knowledge Boundary

VenueDatePaper TitleURL
ICML 201714 Jun 2017On Calibration of Modern Neural Networkshttps://arxiv.org/abs/1706.04599
EMNLP 202017 May 2020Calibration of Pre-trained Transformershttps://arxiv.org/pdf/2003.07892
TACL 20212 Dec 2020How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answeringhttps://arxiv.org/abs/2012.00955
Anthropic11 Jul 2022Language Models (Mostly) Know What They Knowhttps://arxiv.org/abs/2207.05221
ACL 20243 Jul 2023Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Modelshttps://arxiv.org/abs/2307.01379
ICLR 202319 June 2024Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generationhttps://arxiv.org/abs/2302.09664
EMNLP 202315 Mar 2023Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language modelshttps://arxiv.org/abs/2303.08896
EMNLP 20233 Nov 2023SAC3: Reliable Hallucination Detection in Black-Box Language Models via Semantic-aware Cross-check Consistencyhttps://arxiv.org/abs/2311.01740
EMNLP 202326 Apr 2023The Internal State of an LLM Knows When It's Lyinghttps://arxiv.org/pdf/2304.13734
ACL 202517 Feb 2025Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perceptionhttps://www.arxiv.org/abs/2502.11677
TMLR 202228 May 2022Teaching Models to Express Their Uncertainty in Wordshttps://arxiv.org/abs/2205.14334
EMNLP 202324 May 2023Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedbackhttps://arxiv.org/abs/2305.14975
ICLR 202422 Jun 2023Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMshttps://arxiv.org/abs/2306.13063
NAACL 20247 Jun 2024R-Tuning: Instructing Large Language Models to Say β€˜I Don’t Know’https://arxiv.org/pdf/2311.09677
NeurIPS 202312 Dec 2023Alignment for Honestyhttps://arxiv.org/abs/2312.07000
Arxiv20 Oct 2025Annotation-Efficient Universal Honesty Alignmenthttps://arxiv.org/abs/2510.17509

Retrieval Timing

VenueDatePaper TitleURL
EMNLP 202311 May 2023Active Retrieval Augmented Generationhttps://arxiv.org/abs/2305.06983
ACL 202418 Feb 2024When Do LLMs Need Retrieval Augmentation? Mitigating LLMs' Overconfidence Helps Retrieval Augmentationhttps://aclanthology.org/2024.findings-acl.675/
ACL 202412 Mar 2024DRAGIN: Dynamic Retrieval Augmented Generation based on the Information Needs of Large Language Modelshttps://arxiv.org/pdf/2403.10081
SIGIR-AP 202516 Feb 2024Retrieve Only When It Needs: Adaptive Retrieval Augmentation for Hallucination Mitigation in Large Language Modelshttps://arxiv.org/abs/2402.10612
ACL 202529 May 2024CtrlA: Adaptive Retrieval-Augmented Generation via Inherent Controlhttps://arxiv.org/abs/2405.18727
EMNLP 202418 Jun 2024Unified active retrieval for retrieval augmented generationhttps://arxiv.org/abs/2406.12534
ICLR 20236 Oct 2022ReAct: Synergizing Reasoning and Acting in Language Modelshttps://arxiv.org/abs/2210.03629
ACL 202320 Dec 2022Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questionshttps://arxiv.org/abs/2212.10509
ICLR 202417 Oct 2023Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflectionhttps://arxiv.org/abs/2310.11511
EMNLP 20259 Jan 2025Search-o1: Agentic Search-Enhanced Large Reasoning Modelshttps://arxiv.org/abs/2501.05366
COLM 202512 Mar 2025Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learninghttps://arxiv.org/abs/2503.09516
EMNLP 202522 May 2025Search Wisely: Mitigating Sub-optimal Agentic Searches By Reducing Uncertaintyhttps://arxiv.org/abs/2505.17281
EMNLP 202521 May 2025StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimizationhttps://arxiv.org/abs/2505.15107

Information Filtering

VenueDatePaper TitleURL
EMNLP 202419 Apr 2023Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agentshttps://arxiv.org/abs/2304.09542
NAACL 2024 Findings30 Jun 2023Large Language Models are Effective Text Rankers with Pairwise Ranking Promptinghttps://arxiv.org/abs/2306.17563
ICLR 202413 Jul 2023In-context Autoencoder for Context Compression in a Large Language Modelhttps://arxiv.org/abs/2307.06945
ICLR 202406 Oct 2023RECOMP: Improving Retrieval-Augmented LMs with Compression and Selective Augmentationhttps://arxiv.org/abs/2310.04408
ICLR 202417 Oct 2023Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflectionhttps://arxiv.org/abs/2310.11511
EMNLP 202414 Nov 2023Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Modelshttps://arxiv.org/abs/2311.09210
ACL 2024 Findings19 Feb 2024BIDER: Bridging Knowledge Inconsistency for Efficient Retrieval-Augmented LLMs via Key Supporting Evidencehttps://arxiv.org/abs/2402.12174
ACL 202424 Feb 2024ListT5: Listwise Reranking with Fusion-in-Decoder Improves Zero-shot Retrievalhttps://arxiv.org/abs/2402.15838
NeurIPS 202422 May 2024xRAG: Extreme Context Compression for Retrieval-augmented Generation with One Tokenhttps://arxiv.org/abs/2405.13792
ACL 202403 Jun 2024An Information Bottleneck Perspective for Effective Noise Filtering on Retrieval-Augmented Generationhttps://arxiv.org/abs/2406.01549
ACL 202404 Jun 2024Retaining Key Information under High Compression Ratios: Query-Guided Compressor for LLMshttps://arxiv.org/abs/2406.02376
WWW 202517 Jun 2024TourRank: Utilizing Large Language Models for Documents Ranking with a Tournament-Inspired Strategyhttps://arxiv.org/abs/2406.11678
ICLR 202519 Jun 2024InstructRAG: Instructing Retrieval-Augmented Generation via Self-Synthesized Rationaleshttps://arxiv.org/abs/2406.13629
WWW 202526 Jun 2024Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generationhttps://arxiv.org/abs/2406.18676
NeurIPS 202402 Jul 2024RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMshttps://arxiv.org/abs/2407.02485
WSDM 202512 Jul 2024Context Embeddings for Efficient Answer Generation in RAGhttps://arxiv.org/abs/2407.09252
NeurIPS 202407 Oct 2024TableRAG: Million-Token Table Understanding with Language Modelhttps://arxiv.org/abs/2410.04739
WWW 202405 Nov 2024HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systemshttps://arxiv.org/abs/2411.02959
ACL 202525 Feb 2025RankCoT: Refining Knowledge for Retrieval-Augmented Generation through Ranking Chain-of-Thoughtshttps://arxiv.org/abs/2502.17888
EMNLP 202508 Mar 2025Rank-R1: Enhancing Reasoning in LLM-based Document Rerankers via Reinforcement Learninghttps://arxiv.org/abs/2503.06034
EMNLP 2025 Findings24 Jul 2025Dynamic Context Compression for Efficient RAGhttps://arxiv.org/abs/2507.22931v2

Memory Management

VenueDatePaper TitleURL
ACL 202501 Jul 2025In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agentshttps://aclanthology.org/2025.acl-long.413/
arXiv preprint13 Aug 2025MemGuide: Intent-Driven Memory Selection for Goal-Oriented Multi-Session LLM Agentshttps://arxiv.org/abs/2505.20231
arXiv preprint10 Jul 2025MIRIX: Multi-Agent Memory System for LLM-Based Agentshttps://arxiv.org/abs/2507.07957
arXiv preprint06 Jun 2025PersonaAgent: When Large Language Model Agents Meet Personalization at Test Timehttps://arxiv.org/abs/2506.06254
arXiv preprint29 Apr 2025PaRT: Enhancing Proactive Social Chatbots with Personalized Real-Time Retrievalhttps://arxiv.org/abs/2504.20624
ACL 202501 Jul 2025Recursive Question Understanding for Complex Question Answering over Heterogeneous Personal Datahttps://arxiv.org/abs/2505.11900
arXiv preprint23 Jul 2025H-MEM: Hierarchical Memory for High-Efficiency Long-Term Reasoning in LLM Agentshttps://arxiv.org/abs/2507.22925
arXiv preprint28 Apr 2025MemO: Building Production-Ready AI Agents with Scalable Long-Term Memoryhttps://arxiv.org/abs/2504.19413
ACL 202501 Jul 2025Memory-augmented Query Reconstruction for LLM-based Knowledge Graph Reasoninghttps://aclanthology.org/2025.findings-acl.1234/
arXiv preprint27 Aug 2025Nemori: Self-Organizing Agent Memory Inspired by Cognitive Sciencehttps://arxiv.org/abs/2508.03341
arXiv preprint20 Jan 2025Zep: A Temporal Knowledge Graph Architecture for Agent Memoryhttps://arxiv.org/abs/2501.13956
NeurIPS 202508 Oct 2025Mem: Agentic Memory for LLM Agentshttps://arxiv.org/abs/2502.12110
arXiv preprint15 Nov 2023Think-in-Memory: Recalling and Post-thinking Enable LLMs with Long-Term Memoryhttps://arxiv.org/abs/2311.08719
arXiv preprint09 Oct 2025Multiple Memory Systems for Enhancing the Long-term Memory of Agentshttps://arxiv.org/abs/2508.15294
EMNLP 202501 Nov 2025Coarse-to-Fine Grounded Memory for LLM Agent Planninghttps://aclanthology.org/2025.emnlp-main.659/
AAAI 202612 Nov 2025ComoRAG: A Cognitive-Inspired Memory-Organized RAG for Stateful Long Narrative Reasoninghttps://arxiv.org/abs/2508.10419
arXiv preprint12 Feb 2024MemGPT: Towards LLMs as Operating Systemshttps://arxiv.org/abs/2310.08560
arXiv preprint03 Jul 2025MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agenthttps://arxiv.org/abs/2507.02259
arXiv preprint17 Jul 2025MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agentshttps://arxiv.org/abs/2506.15841
arXiv preprint09 Oct 2025Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memoryhttps://arxiv.org/abs/2508.09736
arXiv preprint25 Aug 2025Memento: Fine-tuning LLM Agents without Fine-tuning LLMshttps://arxiv.org/abs/2508.16153
arXiv preprint08 Oct 2025Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via RLhttps://arxiv.org/abs/2508.19828
arXiv preprint15 Aug 2025Learn to Memorize: Optimizing LLM-based Agents with Adaptive Memory Frameworkhttps://arxiv.org/abs/2508.16629
arXiv preprint23 Oct 2025MLP Memory: A Retriever-Pretrained Memory for Large Language Modelshttps://arxiv.org/abs/2508.01832v3
arXiv preprint23 Oct 2025Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Modelshttps://arxiv.org/abs/2508.09874

Answer Generation

VenueDatePaper TitleURL
TPAMI13 Mar 2022Towards Visual-Prompt Temporal Answer Grounding in Instructional Videohttps://ieeexplore.ieee.org/document/10552074
ACL 202306 Mar 2023LIDA: A Tool for Automatic Generation of Grammar-Agnostic Visualizations and Infographics Using LLMshttps://aclanthology.org/2023.acl-demo.11/
NeurIPS 202319 May 2023Any-to-Any Generation via Composable DiffusionNeurIPS paper Link
TVCG03 Nov 2023ChartGPT: Leveraging LLMs to Generate Charts from Abstract Natural Languagehttps://ieeexplore.ieee.org/document/10443572
ICLR 202513 Aug 2024LongWriter: Unleashing 10,000+ Word Generation from Long-Context LLMshttps://openreview.net/forum?id=kQ5s9Yh0WI
EMNLP 202507 Jan 2025PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slideshttps://aclanthology.org/2025.emnlp-main.728/
NeurIPS 202527 May 2025Paper2Poster: Towards Multimodal Poster Automation from Scientific Papershttps://arxiv.org/abs/2505.21497
arXiv04 Jun 2025SuperWriter: Reflection-Driven Long-Form Generation with LLMshttps://arxiv.org/abs/2506.04180
EMNLP 202505 Jul 2025PresentAgent: Multimodal Agent for Presentation Video Generationhttps://aclanthology.org/2025.emnlp-demos.58/
arXiv24 Aug 2025PosterGen: Aesthetic-Aware Paper-to-Poster Generation via Multi-Agent LLMshttps://arxiv.org/abs/2508.17188
arXiv06 Oct 2025Paper2Video: Automatic Video Generation from Scientific Papershttps://arxiv.org/abs/2510.05096

Training Paradigm

Supervised Fine-tuning

Most work below focuses on data synthesis, i.e., designing scalable approaches or frameworks to construct high-quality, large-scale training datasets to train LLM-based agents.

VenueDatePaper TitleURL
NeurIPS 20252025.05.28WebDancer: Towards Autonomous Information Seeking Agencyhttps://arxiv.org/abs/2505.22648
arXiv preprint2025.07.03WebSailor: Navigating Super-human Reasoning for Web Agenthttps://arxiv.org/abs/2507.02592
arXiv preprint2025.07.20WebShaper: Agentically Data Synthesizing via Information-Seeking Formalizationhttps://arxiv.org/abs/2507.15061
NeurIPS 20252025.04.30WebThinker: Empowering Large Reasoning Models with Deep Research Capabilityhttps://arxiv.org/abs/2504.21776
arXiv preprint2025.07.06WebSynthesis: World-Model-Guided MCTS for Efficient WebUI-Trajectory Synthesishttps://arxiv.org/abs/2507.04370
arXiv preprint2025.05.26MaskSearch: A Universal Pre-Training Framework to Enhance Agentic Search Capabilityhttps://arxiv.org/abs/2505.20285
arXiv preprint2025.08.06Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RLhttps://arxiv.org/abs/2508.13167
Findings of EMNLP 20252025.05.26WebCoT: Enhancing Web Agent Reasoning by Reconstructing Chain-of-Thought in Reflection, Branching, and Rollbackhttps://aclanthology.org/2025.findings-emnlp.276/
ACL 20252024.10.18Synthesizing Post-Training Data for LLMs through Multi-Agent Simulationhttps://aclanthology.org/2025.acl-long.1136/
arXiv preprint2024.06.28Scaling Synthetic Data Creation with 1,000,000,000 Personashttps://arxiv.org/abs/2406.20094
NeurIPS 20252025.05.26Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchershttps://arxiv.org/abs/2505.20128
EMNLP 20252025.05.28EvolveSearch: An Iterative Self-Evolving Search Agenthttps://arxiv.org/abs/2505.22501
NeurIPS 20252025.05.06Absolute Zero: Reinforced Self-Play Reasoning with Zero Datahttps://arxiv.org/abs/2505.03335

Agentic End-to-End Reinforcement Learning

VenueDatePaper TitleURL
Arxiv30 Sep 2025Planner-R1: Reward Shaping Enables Efficient Agentic RL with Smaller LLMshttps://arxiv.org/abs/2509.25779
Arxiv21 May 2025ConvSearch-R1: Enhancing Query Reformulation for Conversational Search with Reasoning via RLhttps://arxiv.org/abs/2505.15776
AAAI 202622 Aug 2025OPERA: A RL-Enhanced Orchestrated Planner-Executor Architecture for Multi-Hop Retrievalhttps://arxiv.org/abs/2508.16438
Arxiv1 Aug 2025MAO-ARAG: Multi-Agent Orchestration for Adaptive Retrieval-Augmented Generationhttps://arxiv.org/abs/2508.01005
Arxiv28 Aug 2025AI-SearchPlanner: Modular Agentic Search via Pareto-Optimal Multi-Objective RLhttps://arxiv.org/abs/2508.20368
Arxiv7 Mar 2025R1-Searcher: Incentivizing the Search Capability in LLMs via RLhttps://arxiv.org/abs/2503.05592
Arxiv22 May 2025R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via RLhttps://arxiv.org/abs/2505.17005
Arxiv4 Apr 2025DeepResearcher: Scaling Deep Research via RL in Real-world Environmentshttps://arxiv.org/abs/2504.03160
Arxiv25 Jun 2025MMSearch-R1: Incentivizing LMMs to Searchhttps://arxiv.org/abs/2506.20670
COLM 202512 Mar 2025Search-r1: Training LLMs to Reason and Leverage Search Engines with RLhttps://arxiv.org/abs/2503.09516
Arxiv21 May 2025An Empirical Study on RL for Reasoning-Search Interleaved LLM Agentshttps://arxiv.org/abs/2505.15117
Arxiv4 Jun 2025R-Search: Empowering LLM Reasoning with Search via Multi-Reward RLhttps://arxiv.org/abs/2506.04185
Arxiv7 May 2025ZeroSearch: Incentivize the Search Capability of LLMs without Searchinghttps://arxiv.org/abs/2505.04588
Arxiv22 May 2025O2-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended QAhttps://arxiv.org/abs/2505.16582
Arxiv11 Aug 2025HierSearch: A Hierarchical Enterprise Deep Search Frameworkhttps://arxiv.org/abs/2508.08088
Arxiv29 Jul 2025Graph-R1: Towards Agentic GraphRAG Framework via End-to-end RLhttps://arxiv.org/abs/2507.21892
Arxiv23 Jul 2025DynaSearcher: Dynamic Knowledge Graph Augmented Search Agent via Multi-Reward RLhttps://arxiv.org/abs/2507.17365
Arxiv28 May 2025WebDancer: Towards Autonomous Information Seeking Agencyhttps://arxiv.org/abs/2505.22648
Arxiv16 Sep 2025WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data & RLhttps://arxiv.org/abs/2509.13305
Arxiv28 Jul 2025Kimi k2: Open Agentic Intelligencehttps://arxiv.org/abs/2507.20534
Arxiv11 Aug 2025Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Async RLhttps://arxiv.org/abs/2508.07976
Arxiv30 May 2025Pangu DeepDiver: Adaptive Search Intensity Scaling via Open-Web RLhttps://arxiv.org/html/2505.24332v1
Arxiv6 Aug 2025Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation & RLhttps://arxiv.org/abs/2508.13167
Arxiv22 May 2025Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via RLhttps://arxiv.org/abs/2505.16410
Arxiv26 Jul 2025Agentic Reinforced Policy Optimizationhttps://arxiv.org/abs/2507.19849
Arxiv16 Oct 2025Agentic Entropy-Balanced Policy Optimizationhttps://arxiv.org/abs/2510.14545

Datasets & Benchmarks

VenueDatePaper TitlePaper URLDataset/Code/Leaderboard URL
NAACL 202519 Sep 2024Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented GenerationPaperHuggingface
arXiv21 May 2025InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented GenerationPaperHuggingface
EMNLP 202422 Jul 2024AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?PaperHuggingface
NeurIPS 202309 Jun 2023Mind2Web: Towards a Generalist Agent for the WebPaperHuggingface
NeurIPS 202526 Jun 2025Mind2Web 2: Evaluating Agentic Search with Agent-as-a-JudgePaperHuggingface
arXiv06 May 2025Deep Research Bench: Evaluating AI Web Research AgentsPaperWebsite
arXiv25 May 2025DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep ResearchPaperWebsite
ICLR 202425 Jul 2023WebArena: A Realistic Web Environment for Building Autonomous AgentsPaperGitHub
arXiv13 Jan 2025WebWalker: Benchmarking LLMs in Web TraversalPaperHuggingface
arXiv11 Aug 2025WideSearch: Benchmarking Agentic Broad Info-SeekingPaperHuggingface
ACL 2025 Findings15 Apr 2024MMInA: Benchmarking Multihop Multimodal Internet AgentsPaperHuggingface
NeurIPS 202410 Jun 2024AutoSurvey: Large Language Models Can Automatically Write SurveysPaperGitHub
arXiv14 Aug 2025ReportBench: Evaluating Deep Research Agents via Academic Survey TasksPaperHuggingface
EMNLP 202525 Aug 2025SurveyGen: Quality-Aware Scientific Survey Generation with Large Language ModelsPaperHuggingface
arXiv07 Jul 2025Deep Research Comparator: A Platform for Fine-Grained Human Annotations of Deep Research AgentsPaperGitHub
arXiv22 Jul 2025ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific InquiryPaperGitHub
arXiv29 Sep 2025Towards Personalized Deep Research: Benchmarks and EvaluationsPaperHuggingface
arXiv06 Aug 2025Characterizing Deep Research: A Benchmark and Formal DefinitionPaperGitHub
ACL 202426 Jan 2024ProxyQA: An Alternative Framework for Evaluating Long-Form Text Generation with LLMsPaperGitHub
arXiv21 Nov 2024OpenScholar: Synthesizing Scientific Literature with Retrieval-Augmented Language ModelsPaperGitHub
arXiv27 May 2025Paper2Poster: Towards Multimodal Poster Automation from Scientific PapersPaperHuggingface
arXiv24 Aug 2025PosterGen: Aesthetic-Aware Paper-to-Poster Generation via Multi-Agent LLMsPaperGitHub
arXiv21 May 2025P2P: Automated Paper-to-Poster Generation and Fine-Grained BenchmarkPaperGithub
AAAI 202228 Jan 2021DOC2PPT: Automatic Presentation Slides Generation from Scientific DocumentsPaperWebsite
CVPR 202501 Jan 2025AutoPresent: Designing Structured Visuals from ScratchPaperGitHub
arXiv07 Jan 2025PPTAgent: Generating and Evaluating Presentations Beyond Text-to-SlidesPaperHuggingface
arXiv16 May 2025Talk to Your Slides: Language-Driven Agents for Efficient Slide EditingPaperGitHub
arXiv19 Apr 2025AI Idea Bench 2025: AI Research Idea Generation BenchmarkPaperGitHub
arXiv24 May 2025AI-Researcher: Autonomous Scientific InnovationPaperGitHub
arXiv02 Apr 2025PaperBench: Evaluating AI’s Ability to Replicate AI ResearchPaperHuggingface
JAIR 202230 Jan 2021Can We Automate Scientific Reviewing?PaperWebsite
ICLR 202511 Mar 2025DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking ProcessPaperGitHub
ICLR 202410 Oct 2023SWE-bench: Can Language Models Resolve Real-World GitHub Issues?PaperHuggingface
ACL 202410 Jul 2024Can Language Models Serve as Text-Based World Simulators?PaperGitHub
EMNLP 202214 Mar 2022ScienceWorld: Is your Agent Smarter than a 5th Grader?PaperGitHub
NeurIPS 202410 Jun 2024DiscoveryWorld: A Virtual Environment for Scientific Discovery AgentsPaperGitHub
arXiv17 Sep 2024CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent BenchmarkPaperGitHub
ICLR 202409 Oct 2024MLE-bench: Evaluating Machine Learning Agents on Machine Learning EngineeringPaperGitHub
ICML 202522 Nov 2024RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human expertsPaperGitHub
ICLR 202511 Sep 2024DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?PaperGitHub
NeurIPS 202415 Jul 2024Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?PaperHuggingface
ACL 202527 Feb 2024Benchmarking Data Science AgentsPaperGitHub
arXiv16 Apr 2025UnivEARTH: Towards LLM Agents for Earth ObservationPaperWebsite
arXiv02 Dec 2024Commit0: Library Generation from ScratchPaperGitHub

❀️ Acknowledgement

This project benefits from deepresearch, Tongyi-DeepResearch, Search Agent, and Knowledge-Boundary Thanks for their wonderful works and collective efforts.

πŸ“ž Contact

Feel free to contact us if there are any problems: zhengliang.shii@gmail.com; shizhl@mail.sdu.edu.cn

πŸ₯³ Citation

If you find this work useful, please cite:

@article{shi2025deep,
  title={Deep Research: A Systematic Survey},
  author={Shi, Zhengliang and Chen, Yiqun and Li, Haitao and Sun, Weiwei and Ni, Shiyu and Lyu, Yougang and Fan, Run-Ze and Jin, Bowen and Weng, Yixuan and Zhu, Minjun and others},
  journal={arXiv preprint arXiv:2512.02038},
  year={2025}
}

Collected info

  • β˜… 319 stars
  • βŽ‡ 15 forks
  • Source updated: 6/24/2026