Skip to content

Full CRAG Document Corpus #10

Description

@Rag-Ranger

Hi, could you please help me with a link to the entire CRAG document corpus?

For each of the three CRAG tasks, the dataset links provide JSONL files where each query includes a search_results field containing the top search hits (usually 5 or 50 pages of HTML text).

What I am looking for is:
A single downloadable corpus that contains all search result documents for all queries combined — essentially, the full set of retrieved pages used across the entire CRAG benchmark.I want to download this complete corpus so I can build a global index and evaluate my own RAG pipeline over it.Do you have such a consolidated corpus available, or a recommended source for obtaining it?

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions