Survey on AI Search with Large Language Models 1Tencent YouTu Lab, 2RUC, 3NJU
Curated collection of papers and resources on AI Search: Methods, Benchmarks, Software, and Products.
🗂️ Table of Contents
We will actively maintain this repository and incorporate new research as it emerges. If you have any questions, please contact me.
- RAG "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks". Lewis P, Perez E, Piktus A, et al.. NeurIPS 2020. [Paper]
- REALM "REALM: Retrieval-Augmented Language Model Pre-Training". Guu K, Lee K, Tung Z, et al.. ICML 2020. [Paper]
- RETRO "Improving Language Models by Retrieving from Trillions of Tokens". Borgeaud S, Mensch A, Hoffmann J, et al.. ICML 2022. [Paper]
- Query Rewriting "Query Rewriting for Retrieval-Augmented Large Language Models". Ma X, Gong Y, He P, et al.. EMNLP 2023. [Paper]
- In-context RAG "In-Context Retrieval-Augmented Generation for Large Language Models". Ram O, Levine Y, Dalmedigos I, et al.. TACL 2023. [Paper]
- LLMLingua "LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models". Jiang H, Wu Q, Lin C Y, et al.. EMNLP 2023. [Paper]
- RECOMP "RECOMP: Improving Retrieval-Augmented LMs with Context Compression and Selective Augmentation". Xu F, Shi W, Choi E.. ICLR 2024. [Paper]
- Hierarchical Document Refinement "Enhancing Retrieval-Augmented Generation with Hierarchical Document Refinement". Anonymous.. arXiv 2024. [Paper]
- Long-context LLMs Meet RAG "Long-context LLMs Meet RAG". Jin B, Yoon J, Han J, et al.. ICLR 2025. [Paper]
- Inference Scaling for Long-context RAG "Inference Scaling for Long-context RAG". Yue Z, Zhuang H, Bai A, et al.. ICLR 2025. [Paper]
- ChatQA 2 "ChatQA 2: Bridging the Gap between Open-source and Proprietary LLMs in Long-context and RAG". Xu P, Ping W, Wu X, et al.. ICLR 2025. [Paper]
- RetrievalAttention "RetrievalAttention: Accelerating Long-context LLM Inference via Vector Database Retrieval". Liu D, Chen M, Lu B, et al.. NeurIPS 2025. [Paper]
- Tree of Clarifications "Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models". Kim G, Kim S, Lee B, et al.. EMNLP 2023. [Paper]
- GenRead "Generate rather than Retrieve: Large Language Models are Strong Context Generators". Yu W, Iter D, Wang S, et al.. ICLR 2023. [Paper]
- REPLUG "REPLUG: Retrieval-Augmented Black-Box Language Models". Shi W, Min S, Yasunaga M, et al.. NAACL 2024. [Paper]
- BlendFilter "BlendFilter: Advancing Retrieval-Augmented Large Language Models via Query Generation and Knowledge Filtering". Wang H, Li J, Wu H, et al.. EMNLP 2024. [Paper]
- Self-Knowledge Guided "Self-Knowledge Guided Retrieval Augmentation for Large Language Models". Wang Y, Li P, Sun M, et al.. EMNLP 2023. [Paper]
- Self-DC "Self-DC: When to Retrieve and When to Generate? Self Divide-and-Conquer for Compositional Unknown Questions". Li J, Wang H, Wu H, et al.. arXiv 2024. [Paper]
- Adaptive RAG "Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity". Jeong S, Baek J, Cho S, et al.. NAACL 2024. [Paper]
- Slim Proxy Models "Small Models, Big Insights: Leveraging Slim Proxy Models to Decide When and What to Retrieve for LLMs". Tan J, Dou Z, Wen J R.. ACL 2024. [Paper]
- ReAct "ReAct: Synergizing Reasoning and Acting in Language Models". Yao S, Zhao J, Yu D, et al.. ICLR 2023. [Paper]
- Iterative RAG "Demonstrate-Search-Predict: Composing Retrieval and Language Models for Knowledge-Intensive NLP". Khattab O, Santhanam K, Li X L, et al.. NeurIPS 2023. [Paper]
- IRCoT "Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions". Trivedi H, Balasubramanian N, Khot T, et al.. ACL 2023. [Paper]
- Active RAG "Active Retrieval Augmented Generation". Jiang Z, Xu F, Gao L, et al.. EMNLP 2023. [Paper]
- FLARE "Active Retrieval Augmented Generation". Jiang Z, Xu F, Gao L, et al.. EMNLP 2023. [Paper]
- Self-RAG "Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection". Asai A, Wu Z, Wang Y, et al.. ICLR 2024. [Paper]
- Robust RAG "Benchmarking Large Language Models in Retrieval-Augmented Generation". Chen J, Lin H, Han X, et al.. AAAI 2024. [Paper]
- RARE "RARE: Retrieval-Augmented Reasoning Enhancement". Tran H, Yao Z, Wang J, et al.. ACL 2025. [Paper]
- Astute RAG "Astute RAG: Robust Retrieval-Augmented Generation with Imperfect Retrieval and Knowledge Conflicts". Wang F, Wan X, Sun R, et al.. ACL 2025. [Paper]
- MemoRAG "MemoRAG: Enhancing Retrieval with Global Memory for Long-context Processing". Qian H, Liu Z, Zhang P, et al.. WWW 2025. [Paper]
- Bee-RAG "Bee-RAG: Balanced Entropy Engineering for Retrieval-Augmented Generation". Wang Y, Ren R, Wang Y, et al.. AAAI 2026. [Paper]
- LevelRAG "LevelRAG: Multi-hop Logical Planning Enhanced RAG". Zhang Z, Feng Y, Zhang M.. arXiv 2025. [Paper]
- HippoRAG "HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models". Gutiérrez B J, Shu Y, Gu Y.. NeurIPS 2024. [Paper]
- GNN-RAG "GNN-RAG: Graph Neural Network Retrieval-Augmented Generation for LLM Reasoning". Mavromatis C, Karypis G.. arXiv 2024. [Paper]
- GFM-RAG "GFM-RAG: Graph Foundation Model as Structured Knowledge Index for RAG". Luo L, Zhao Z, Haffari R, et al.. NeurIPS 2025. [Paper]
- Chameleon "Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models". Lu P, Peng B, Cheng H, et al.. NeurIPS 2023. [Paper]
- Search-o1 "Search-o1: Agentic Search-Enhanced Large Reasoning Models". Li X, Zhang Y, Wang Z, et al.. arXiv 2025. [Paper]
- WebThinker "WebThinker: Empowering Large Reasoning Models to Search the Web". Anonymous.. arXiv 2025. [Paper]
- WebDancer "WebDancer: Task-Driven Web Agent with Self-Adaptive Online Learning". Wu Z, Ma J, Zhang Y, et al.. arXiv 2025. [Paper]
- ManuSearch "ManuSearch: A Multi-Agent Collaborative Search Framework for Complex Knowledge Tasks". Anonymous.. arXiv 2025. [Paper]
- HiRA "HiRA: Hierarchical Retrieval Augmentation for Complex Reasoning". Anonymous.. arXiv 2025. [Paper]
- SearchAgent-X "SearchAgent-X: Efficient and Scalable Web Search Agent for Complex Reasoning". Anonymous.. arXiv 2025. [Paper] [Github]
- Search and Refine During Think "Search and Refine During Think: Identifying Knowledge Gaps and Refining via Retrieval". Shi Y, Li S, Wu C, et al.. NeurIPS 2025. [Paper]
- MaskSearch "MaskSearch: Learning to Search via Masked Token Prediction". Anonymous.. arXiv 2025. [Paper]
- CoRAG "CoRAG: Collaborative Retrieval-Augmented Generation for Multi-Hop Reasoning". Anonymous.. arXiv 2025. [Paper]
- ReaRAG "ReaRAG: Reasoning-Enhanced Retrieval-Augmented Generation". Anonymous.. arXiv 2025. [Paper]
- ExSearch "ExSearch: Exploratory Search with Large Language Models". Anonymous.. arXiv 2025. [Paper]
- SimpleDeepSearcher "SimpleDeepSearcher: A Minimalist Approach to Deep Search". Anonymous.. arXiv 2025. [Paper] [Github]
- WebCoT "WebCoT: Web-Augmented Chain-of-Thought Reasoning". Anonymous.. arXiv 2025. [Paper]
- DeepRAG "DeepRAG: Deep Retrieval-Augmented Generation with Iterative Self-Improvement". Anonymous.. arXiv 2025. [Paper]
- Search-R1 "Search-R1: Training LLMs to Reason and Search with Reinforcement Learning". Anonymous.. arXiv 2025. [Paper]
- R1-Searcher "R1-Searcher: Incentivizing the Search Capability of LLMs via Reinforcement Learning". Anonymous.. arXiv 2025. [Paper]
- ReSearch "ReSearch: Learning to Reason with Search via Iterative Self-Play". Anonymous.. arXiv 2025. [Paper]
- WebSailor "WebSailor: Navigating the Web with Large Language Models". Anonymous.. arXiv 2025. [Paper]
- DeepResearcher "DeepResearcher: Scaling Deep Search with Long-Context LLMs". Anonymous.. arXiv 2025. [Paper]
- Ô-searcher "Ô-Searcher: Optimizing Search-Augmented Reasoning with Outcome Supervision". Anonymous.. arXiv 2025. [Paper]
- StepSearch "StepSearch: Step-by-Step Search for Complex Reasoning". Anonymous.. arXiv 2025. [Paper]
- ZeroSearch "ZeroSearch: Zero-Shot Search-Augmented Reasoning". Anonymous.. arXiv 2025. [Paper]
- SEM "SEM: Search-Enhanced Memory for Large Language Models". Anonymous.. arXiv 2025. [Paper]
- β-GRPO "β-GRPO: Balancing Exploration and Exploitation in Search-Augmented RL". Anonymous.. arXiv 2025. [Paper]
- s3 "S3: Simple, Scalable Search-Augmented Reasoning". Anonymous.. arXiv 2025. [Paper]
- WebAgent-R1 "WebAgent-R1: Empowering Web Agents with Reinforcement Learning". Anonymous.. arXiv 2025. [Paper]
- LiteWebAgent "LiteWebAgent: A Lightweight Web Agent with Structured Reasoning". Anonymous.. arXiv 2025. [Paper]
- AgentOccam "AgentOccam: A Simple Yet Powerful Baseline for Web Agents". Anonymous.. arXiv 2025. [Paper]
- Pangu DeepDiver "Pangu DeepDiver: Deep Web Exploration with Large Language Models". Anonymous.. arXiv 2025. [Paper]
- AutoWebGLM "AutoWebGLM: A Large Language Model-based Web Navigating Agent". Lai H, Liu X, Wong I, et al.. KDD 2024. [Paper]
- WebVoyager "WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models". He H, Yao W, Ma K, et al.. ACL 2024. [Paper]
- WebVLN "WebVLN: Vision-and-Language Navigation on the Web". Chen Q, Pitawela D, Zhao C, et al.. AAAI 2024. [Paper]
- Dual-View Visual Contextualization "Dual-View Visual Contextualization for Web Navigation". Kil J, Song C H, Zheng B, et al.. CVPR 2024. [Paper]
- Falcon-UI "Falcon-UI: Understanding Graphical User Interfaces through Large Language Models". Anonymous.. arXiv 2025. [Paper]
- CAAP "CAAP: Context-Aware Action Prediction for Web Agents". Anonymous.. arXiv 2025. [Paper]
- ShowUI "ShowUI: Instruction-Following Agents for Screen Understanding". Anonymous.. arXiv 2025. [Paper]
- SeeClick "SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents". Cheng K, Sun Q, Chu Y, et al.. ACL 2024. [Paper]
- Aria-UI "Aria-UI: Visual Grounding for GUI Instructions". Yang Y, Wang Y, Li D, et al.. ACL 2025. [Paper]
- PrivWeb "PrivWeb: Unobtrusive and Content-Aware Privacy Protection for Web Agents". Zhang S, Jiang Y, Ma R, et al.. ACM 2026. [Paper]
- AgentDAM "AgentDAM: Evaluating Privacy Leakage Risks of Autonomous Web Agents". Zharmagambetov A, Guo C, Evtimov I.. NeurIPS 2025. [Paper]
- MMSearch "MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines". Dong J, Wang P, Liu Z, et al.. arXiv 2024. [Paper]
- MMSearch-R1 "MMSearch-R1: Reinforcing Multimodal Search with Reasoning". Anonymous.. arXiv 2025. [Paper]
- Video-RAG "Video-RAG: Visually-Aligned Retrieval-Augmented Generation for Long Video Understanding". Luo Y, Zheng X, Li G, et al.. NeurIPS 2025. [Paper]
- VideoRAG "VideoRAG: Retrieval-Augmented Generation for Ultra-Long Videos". Ren X, Xu L, Xia L, et al.. KDD 2026. [Paper]
- TV-RAG "TV-RAG: Temporal-Aware and Semantic Entropy Weighted Long Video Retrieval". Cao Z, He Y, Liu A, et al.. ACM 2025. [Paper]
- DrVideo "DrVideo: Document Retrieval Based Long Video Understanding". Ma Z, Gou C, Shi H, et al.. CVPR 2025. [Paper]
- MM-Embed "MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs". Lin S C, Lee C, Shoeybi M, et al.. ICLR 2025. [Paper]
- Omni-Embed-Nemotron "Omni-Embed-Nemotron: Unified Multimodal Retrieval Model". Xu M, Zhou W, Babakhin Y, et al.. arXiv 2025. [Paper]
- UniIR "UniIR: Training and Evaluating Universal Multimodal Information Retrievers". Wei C, Chen Y, Chen H, et al.. ECCV 2024. [Paper]
- DSE "DSE: Document Screenshot Embedding for Unified Visual and Text Retrieval". Ma X, Lin S C, Li M, et al.. EMNLP 2024. [Paper]
- M3DocRAG "M3DocRAG: Multi-modal Retrieval for Multi-page Document Understanding". Cho J, Mahata D, Irsoy O, et al.. arXiv 2024. [Paper]
- URaG "URaG: Unifying Retrieval and Generation in Multimodal LLMs". Shi Y, Wang J, Shan Z, et al.. AAAI 2026. [Paper]
- Generative Cross-modal Retrieval "Generative Cross-modal Retrieval: Memorizing Images in Multimodal Language Models". Li Y, Wang W, Qu L, et al.. ACL 2024. [Paper]
- Cross-modal Retrieval Survey "Cross-modal Retrieval: Methods and Future Directions". Wang T, Li F, Zhu L, et al.. IEEE 2025. [Paper]
- MuVR "MuVR: Multi-modal Untrimmed Video Retrieval Benchmark". Feng Y, Hu J, Lu Q, et al.. NeurIPS 2025. [Paper]
- MMDocIR "MMDocIR: Multi-modal Long Document Retrieval Benchmark". Dong K, Chang Y, Goh X D, et al.. arXiv 2025. [Paper]
- Explorer "Explorer: Scaling Web Exploration with Multimodal Agents". Anonymous.. arXiv 2025. [Paper]
- AdaptAgent "AdaptAgent: Adaptive Multimodal Web Agents with Dynamic Planning". Anonymous.. arXiv 2025. [Paper]
- OpenWebVoyager "OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration". Anonymous.. arXiv 2025. [Paper]
- InfoTech Assistant "InfoTech Assistant: A Multimodal Agent for Technical Web Navigation". Anonymous.. arXiv 2025. [Paper]
- GPT-4V Grounded Agent "A Grounded Agent for Web Navigation with GPT-4V". Anonymous.. arXiv 2024. [Paper]
- Natural Questions (NQ) "Natural Questions: A Benchmark for Question Answering Research". Kwiatkowski T, Palomaki J, Redfield O, et al.. TACL 2019. [Paper]
- TriviaQA "TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension". Joshi M, Choi E, Weld D, et al.. ACL 2017. [Paper]
- HotpotQA "HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering". Yang Z, Qi P, Zhang S, et al.. EMNLP 2018. [Paper]
- FEVER "FEVER: a Large-scale Dataset for Fact Extraction and VERification". Thorne J, Vlachos A, Christodoulopoulos C, et al.. NAACL 2018. [Paper]
- 2WikiMultiHopQA "Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps". Ho X, Nguyen A K D, Sugawara S, et al.. COLING 2020. [Paper]
- KILT "KILT: a Benchmark for Knowledge Intensive Language Tasks". Petroni F, Piktus A, Fan A, et al.. NAACL 2021. [Paper]
- TREC Health "TREC Health Misinformation Track". Clarke C L A, Maistro M, Smucker M D, et al.. TREC 2021. [Paper]
- MuSiQue "MuSiQue: Multihop Questions via Single-hop Question Composition". Trivedi H, Balasubramanian N, Khot T, et al.. TACL 2022. [Paper]
- PopQA "PopQA: A Large-scale Open-domain Question Answering Dataset". Mallen A, Asai A, Zhong V, et al.. ACL 2023. [Paper]
- TART "TART: Task-Aware Retrieval with Instructions". Asai A, Schick T, Lewis P, et al.. ACL 2023. [Paper]
- GAIA "GAIA: a Benchmark for General AI Assistants". Mialon G, Fourrier C, Swift C, et al.. ICLR 2024. [Paper]
- MultiHop-RAG "MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries". Tang Y, Yang Y.. arXiv 2024. [Paper]
- RAGBench "RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems". Friel R, Belyi M, Sanyal A.. arXiv 2024. [Paper]
- LitSearch "LitSearch: A Benchmark for Real-World Academic Literature Search". Ajith A, Xia M, Chevalier A, et al.. EMNLP 2024. [Paper]
- MAIR "MAIR: Massive Instruction-based Retrieval Benchmark". Sun W, Shi Z, Long W J, et al.. EMNLP 2024. [Paper]
- BrowseComp "BrowseComp: A Benchmark for Browsing-based Complex Question Answering". Anonymous.. arXiv 2025. [Paper]
- BrowseComp-ZH "BrowseComp-ZH: A Chinese Benchmark for Browsing-based Complex QA". Anonymous.. arXiv 2025. [Paper]
- Mind2Web 2 "Mind2Web 2: Towards Generalist Web Agents". Anonymous.. arXiv 2025. [Paper]
- BRIGHT "BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval". Su H, Yen H, Xia M, et al.. ICLR 2025. [Paper]
- LiveRAG Challenge "LiveRAG Challenge: Dynamic RAG Evaluation at SIGIR 2025". Carmel D, Filice S, Horowitz G, et al.. SIGIR 2025. [Paper]
- Ragnarök "Ragnarök: A Reusable Framework and Baselines for TREC 2024 RAG Track". Pradeep R, Thakur N, Sharifymoghaddam S.. ECIR 2025. [Paper]
-
Mind2Web "Mind2Web: Towards a Generalist Agent for the Web". Deng X, Gu Y, Zheng B, et al.. NeurIPS 2023. [Paper]
-
WebArena "WebArena: A Realistic Web Environment for Building Autonomous Agents". Zhou S, Xu F, Zhu H, et al.. ICLR 2024. [Paper]
-
AgentHarm "AgentHarm: A Comprehensive Benchmark for Measuring LLM Agent Harmfulness". Andriushchenko M, Souly A, Dziemian M.. arXiv 2024. [Paper]
-
Agent-SafetyBench "Agent-SafetyBench: Comprehensive Safety Evaluation for LLM Agents". Zhang Z, Cui S, Lu Y, et al.. arXiv 2024. [Paper]
-
R-Judge "R-Judge: Benchmarking Safety Risk Awareness of LLM Agents". Yuan T, He Z, Dong L, et al.. EMNLP 2024. [Paper]
-
WebChoreArena "WebChoreArena: Evaluating Web Agents on Daily Chores". Anonymous.. arXiv 2025. [Paper]
-
WebCanvas "WebCanvas: Benchmarking Web Agents for Online Task Completion". Anonymous.. arXiv 2025. [Paper]
-
TurkingBench "TurkingBench: A Benchmark for Web Agents on Crowdsourcing Platforms". Anonymous.. arXiv 2025. [Paper]
-
BearCubs "BearCubs: A Benchmark for Evaluating Web Agents on Complex Tasks". Anonymous.. arXiv 2025. [Paper]
-
REAL "REAL: A Representative Benchmark for Web Agent Evaluation". Anonymous.. arXiv 2025. [Paper]
-
DeepShop "DeepShop: A Benchmark for Shopping Web Agents". Anonymous.. arXiv 2025. [Paper]
-
SafeArena "SafeArena: Evaluating the Safety of Web Agents". Anonymous.. arXiv 2025. [Paper]
-
WASP "WASP: Web Agent Safety and Privacy Benchmark". Anonymous.. arXiv 2025. [Paper]
-
OS-Harm "OS-Harm: A Safety Evaluation Benchmark for Computer-Using Agents". Kuntz T, Duzan A, Zhao H, et al.. NeurIPS 2025. [Paper]
-
AgentAuditor "AgentAuditor: Human-Level Safety and Security Evaluation for Agents". Luo H, Dai S, Ni C, et al.. NeurIPS 2025. [Paper]
-
Tool Retrieval Benchmark "Tool Retrieval Benchmark: Evaluating LLM Tool Retrieval Capabilities". Shi Z, Wang Y, Yan L, et al.. ACL 2025. [Paper]
- MMSearch "MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines". Dong J, Wang P, Liu Z, et al.. arXiv 2024. [Paper]
- VisualWebArena "VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks". Koh J Y, Lo R, Jang L, et al.. NeurIPS 2024. [Paper]
- LIVEVQA "LIVEVQA: Live Visual Question Answering". Anonymous.. arXiv 2025. [Paper]
- MRAMG-Bench "MRAMG-Bench: A Multimodal Retrieval-Augmented Multimodal Generation Benchmark". Anonymous.. arXiv 2025. [Paper]
- ChatGPT Deep Research OpenAI. 2025. [Website]
- Perplexity Deep Research Perplexity AI. 2025. [Website]
- You.com You.com. 2024. [Website]
- Gemini Deep Research Google DeepMind. 2025. [Website]
- Doubao ByteDance. 2025. [Website]
- Yuanbao Tencent. 2025. [Website]
- Nano AI Nano. 2025. [Website]
- Kimi Moonshot AI. 2025. [Website]
- DeepSeek Search DeepSeek. 2025. [Website]
- Quark DeepSearch Alibaba. 2025. [Website]
- MediSearch MediSearch. 2025. [Website]
- Devv.ai Devv.ai. 2025. [Website]
- Consensus Consensus. 2025. [Website]
Last Updated: 2026-07-31
