A research-oriented collection of multimodal fake news / misinformation / rumor / disinformation / fact-checking / out-of-context / media manipulation / AI-generated news datasets.
This collection supports research on Multimodal Fake News Detection (MFND/MMD) and related topics. It brings together papers, official repositories, dataset access links, and structured metadataβincluding task, modality, language, domain, label type, and data originβto make dataset discovery and comparison faster.
- π Currently includes 33 datasets / benchmarks: 21 core datasets + 12 closely related datasets.
- π°οΈ Covers classic datasets such as Fakeddit, FakeNewsNet, Weibo, Weibo21, MuMiN, and FakeSV.
- π Continuously includes recent datasets and benchmarks such as ROM, MDSM, FineFake, AMG, MMFakeBench, MFND, MM-Health, VLDBench, DriftBench, DeceptionDecoded, ReMMDBench, and FakeVE.
- π Covers different settings including multi-domain, multilingual, social context, external evidence / knowledge, fine-grained labels, multi-image, audio-video, and generative AI.
- π§© Distinguishes between real-world data, curated real-world data, out-of-context construction, synthetic manipulation, mixed real + synthetic data, and GenAI-diversified data.
- ποΈ
catalog/datasets.yamlstores the structured metadata, whiledatasets/datasets.csvcan be directly used for statistical analysis.
| If you want to⦠| Start with⦠|
|---|---|
| Compare the main fake-news datasets | Core dataset catalog |
| Filter candidates by research capability | Dataset feature comparison |
| Explore adjacent fact-checking and manipulation benchmarks | Closely related benchmarks |
| Choose data for a specific research question | Dataset recommendations |
| Reuse the catalog programmatically | catalog/datasets.yaml and datasets/datasets.csv |
Reading tip: Use the core catalog for datasets whose primary task is misinformation detection. Use the related catalog when studying a neighboring task such as out-of-context detection, fact-checking, or media-manipulation localization.
- π Core Multimodal Fake News Datasets
- β Dataset Feature Comparison
- π Closely Related Benchmarks
- π Visual Summary
- π― How to Choose a Dataset
- π€ Modality Legend
- π§± Dataset Scope
β οΈ Data Origin- π Repository Structure
- π Dataset Link Policy
The following datasets directly support fake news, misinformation, rumor, disinformation, fine-grained attribution, or short-video fake-news detection.
Link legend:
πpaper Β·π»official repository or project page Β·π¦dataset access To keep the table readable on GitHub, access details and more complete metadata are stored indatasets/datasets.csvandcatalog/datasets.yaml.
| Dataset | Year | Language | Main Setting | Modalities | Labels / Task | Scale |
|---|---|---|---|---|---|---|
| DeceptionDecoded π π» π¦ |
2026 Β· ICLR | EN | Creator intent / deception intent | T+I+R |
Intent-centric multi-task / misleading intent detection | 12,000 |
| DriftBench π π» π¦ |
2026 Β· AAAI | EN | GenAI robustness | T+I+E |
Truth verification + 6 diversification categories | 16,000 |
| FakeVE π π» π¦ |
2026 Β· IP&M | EN | Video fake news | V+A+T |
Explainable fake-news video detection | 2,672 |
| FineFake π π» π¦ |
2026 Β· Information Fusion | EN | Fine-grained / multi-domain | T+I+S+M+K |
Binary + 6-way fine-grained | 16,909 |
| ReMMDBench π π» π¦ |
2026 | Multilingual | Multilingual / multi-image / evidence verification | T+MI+E |
5-way veracity + 8 distortion labels | 500 |
| VLDBench π π» π¦ |
2026 Β· Information Fusion | EN | Multi-category disinformation | T+I |
13 benchmark-specific categories | β62K |
| AMG π π» π¦ |
2025 Β· AAAI | ZH | Fine-grained fake-news attribution | T+I |
6-way classification / attribution | 4,922 |
| MFND π π» π¦ |
2025 Β· IJCAI | EN | Multimodal manipulation | T+I |
11 manipulation types + localization | β |
| MM-Health π π» π¦ |
2025 Β· EMNLP Findings | EN | Health misinformation / AI generation | T+I |
Reliability + originality + fine-grained labels | 34,746 |
| MMFakeBench π π» π¦ |
2025 Β· ICLR | EN | Mixed-source misinformation | T+I |
Binary + 3 coarse classes + 12 subtypes | 11,000 |
| FakeTT π π» π¦ |
2024 Β· ACM MM | EN | Short-video fake news | V+A+T |
Binary classification | 1,991 |
| MΒ³A π π» π¦ |
2024 Β· CVIU | Global | Multimedia authenticity | T+I+A+V |
Fine-grained / multi-task | β |
| FakeSV π π» π¦ |
2023 Β· AAAI | ZH | Short-video fake news | V+A+T+S+M |
Binary + debunking information | 5,538 |
| MRΒ² π π» π¦ |
2023 Β· SIGIR | EN / ZH | Retrieval-augmented rumor detection | T+I+S+M+E |
3-way classification | 14,700 |
| MuMiN π π» π¦ |
2022 Β· SIGIR | 41 languages | Multilingual misinformation | T+I+S+M+G |
Binary classification | 12,914 claims |
| CHECKED π π» π¦ |
2021 Β· SNAM | ZH | COVID-19 / health | T+I+S+M |
Binary classification | 2,104 |
| Weibo21 π π» π¦ |
2021 Β· CIKM | ZH | Multi-domain fake news | T+I+M |
Binary classification | 9,128 |
| Fakeddit π π» π¦ |
2020 Β· LREC | EN | Social media | T+I+S+M |
2 / 3 / 6-way classification | 1,063,106 |
| MM-COVID π π» π¦ |
2020 Β· IEEE BigData | 6 languages | COVID-19 / cross-lingual | T+S+M |
Binary classification | 11,173 |
| FakeNewsNet π π» π¦ |
2018 | EN | News + social propagation | T+I+S+M |
Binary classification | β |
| Weibo Multimodal Rumor Dataset π π» π¦ |
2017 Β· ACM MM | ZH | Weibo rumor detection | T+I+S+M |
Binary classification | 9,528 |
This table provides a quick overview of the major differences among the core datasets.
β
indicates that the feature is an important part of the dataset or an explicitly supported research setting, while β means it is not a primary feature.
| Dataset | Multi-domain | Multilingual | Social Context | External Evidence / Knowledge | Fine-grained Labels | Generated / Synthetic Data | Audio-Video / Multi-image |
|---|---|---|---|---|---|---|---|
| DeceptionDecoded | β | β | β | β | β | β | β |
| DriftBench | β | β | β | β | β | β | β |
| FakeVE | β | β | β | β | β | β | β |
| FineFake | β | β | β | β | β | β | β |
| ReMMDBench | β | β | β | β | β | β | β |
| VLDBench | β | β | β | β | β | β | β |
| AMG | β | β | β | β | β | β | β |
| MFND | β | β | β | β | β | β | β |
| MM-Health | β | β | β | β | β | β | β |
| MMFakeBench | β | β | β | β | β | β | β |
| FakeTT | β | β | β | β | β | β | β |
| MΒ³A | β | β | β | β | β | β | β |
| FakeSV | β | β | β | β | β | β | β |
| MRΒ² | β | β | β | β | β | β | β |
| MuMiN | β | β | β | β | β | β | β |
| CHECKED | β | β | β | β | β | β | β |
| Weibo21 | β | β | β | β | β | β | β |
| Fakeddit | β | β | β | β | β | β | β |
| MM-COVID | β | β | β | β | β | β | β |
| FakeNewsNet | β | β | β | β | β | β | β |
| Weibo Multimodal Rumor | β | β | β | β | β | β | β |
Note: This table is intended as a quick overview and does not replace the complete dataset definitions in the original papers. Some datasets support multiple tasks; more detailed metadata is available in
catalog/datasets.yaml.
The following datasets are highly relevant to multimodal fake-news research but mainly focus on adjacent tasks, such as out-of-context detection, fact-checking, media-manipulation localization, or AI-generated-content detection.
| Dataset | Year | Main Focus | Modalities | Labels / Task | Scale |
|---|---|---|---|---|---|
| ROM π π» π¦ |
2026 Β· ACL | Reasoning-enhanced multimodal manipulation | T+I+FR |
Detection + manipulation grounding + forensic reasoning | 704,456 |
| MDSM π π» π¦ |
2026 Β· CVPR | MLLM-driven semantic-aligned manipulation | T+I |
Detection + 5 manipulation types + image grounding | 441,423 |
| MiRAGeNews π π» π¦ |
2024 Β· EMNLP Findings | AI-generated news | T+I |
Real vs AI-generated | 15,000 |
| VERITE π π» π¦ |
2024 Β· IJMIR | Out-of-context misinformation | T+I |
3-way classification | 1,000 |
| COSMOS π π» π¦ |
2023 Β· AAAI | OOC detection | T+I |
Binary classification | 204,458 images |
| DGMβ΄ π π» π¦ |
2023 Β· CVPR | Multimodal media manipulation | T+I |
Detection + manipulation-type classification + image/text grounding | 230,000 |
| FACTIFY 2 π π» π¦ |
2023 | Multimodal fact-checking | T+I+E |
5-way classification | 50,000 |
| MOCHEG π π» π¦ |
2023 Β· SIGIR | Fact-checking + explanation | T+I+E |
Verification + explanation generation | β |
| FACTIFY π π» π¦ |
2022 | Multimodal fact-checking | T+I+E |
3-way classification | 50,000 |
| NewsCLIPpings π π» π¦ |
2021 Β· EMNLP | OOC image-text mismatch | T+I |
Binary classification | 988,283 |
| VisualNews π π» π¦ |
2021 Β· EMNLP | Real news image-caption corpus | T+I+M |
News image captioning / image-text source corpus | 1,080,595 images |
| NYTimes800k π π» π¦ |
2020 Β· CVPR | Real NYT image-caption corpus | T+I+M |
News image captioning / image-text source corpus | 792,971 images |
These figures summarize the full catalog and complement, rather than replace, the dataset-level metadata above.
Use this table as a first-pass shortlist, then verify licensing, access requirements, and task definitions on each dataset's official page.
| Research Direction | Recommended Datasets |
|---|---|
| πΌοΈ Classic image-text fake news detection | Fakeddit, Weibo, Weibo21 |
| π Social propagation / social context | FakeNewsNet, MuMiN, CHECKED, MRΒ² |
| π¬ Fine-grained fake type / attribution | FineFake, AMG, MMFakeBench, MFND |
| π§ Multi-domain generalization | FineFake, Weibo21, MΒ³A, VLDBench |
| π§© Out-of-context image-text misinformation | COSMOS, NewsCLIPpings, VERITE |
| π οΈ Image/text manipulation detection and grounding | ROM, MDSM, DGMβ΄, MFND |
| π° Real news image-text source corpora / news captioning | VisualNews, NYTimes800k |
| π External evidence / retrieval-augmented verification | MRΒ², MOCHEG, FACTIFY, FineFake, ReMMDBench |
| π€ Robustness in the GenAI era | MMFakeBench, MM-Health, VLDBench, DriftBench, DeceptionDecoded |
| β¨ AI-generated multimodal news | MiRAGeNews, MM-Health |
| π¬ Short-video fake news | FakeSV, FakeTT, FakeVE |
| π Multilingual / cross-lingual | MM-COVID, MuMiN, MRΒ², ReMMDBench |
| πΌοΈπΌοΈ Multi-image verification | ReMMDBench |
| Abbreviation | Meaning | Abbreviation | Meaning |
|---|---|---|---|
T |
Text | I |
Image |
MI |
Multiple Images | V |
Video |
A |
Audio | S |
Social Context |
M |
Metadata | E |
External Evidence |
K |
External Knowledge | G |
Knowledge Graph |
R |
Reference Article | FR |
Forensic Rationale |
Datasets directly designed for multimodal fake news / misinformation / rumor / disinformation detection, including:
- Image-text fake news detection
- Multi-domain / multilingual detection
- Fine-grained fake type and attribution
- Social-context modeling
- Short-video fake news detection
- Fake-news detection and robustness evaluation in GenAI settings
Datasets that are highly relevant to multimodal fake news research but mainly target adjacent tasks, such as:
- Out-of-Context (OOC) detection
- Multimodal fact-checking
- Media manipulation detection and localization
- AI-generated news detection
- Evidence retrieval and explanation generation
Separating these two groups makes it easier to distinguish their research objectives and evaluation settings.
The way "fake" or misleading content is created varies substantially across datasets. This distinction is important when comparing model performance.
| Type | Description |
|---|---|
real_world |
Naturally occurring misinformation from news websites, social media, or short-video platforms |
real_world_curated |
Real-world samples that are selected, curated, or manually verified for benchmark construction |
synthetic_pairing |
Real images, captions, or text are recombined to create out-of-context samples |
synthetic_manipulation |
Images or text are modified through controlled manipulation |
mixed_real_and_synthetic |
Combines real-world data with generated or manipulated samples |
genai_diversified |
Uses generative AI to rewrite, diversify, or regenerate news content |
When comparing results across datasets, consider data origin, task definition, label granularity, and modality settings rather than accuracy alone.
.
βββ README.md
βββ CONTRIBUTING.md
βββ CITATION.cff
βββ requirements.txt
βββ catalog/
β βββ datasets.yaml
βββ datasets/
β βββ datasets.csv
βββ analysis/
β βββ summary.json
β βββ datasets_by_year.png
β βββ modalities.png
β βββ scope.png
β βββ categories.png
βββ docs/
β βββ METADATA_SCHEMA.md
β βββ SOURCES.md
βββ scripts/
βββ build_assets.py
This repository is mainly used to collect and summarize dataset papers, official repositories, and dataset access links. Original dataset files are not re-uploaded or redistributed here.
For download procedures, application requirements, and usage restrictions, please refer to the official page of each dataset.
If this collection is useful for your research, a Star β is greatly appreciated.



