Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
13 changes: 13 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,19 @@ All notable changes to this project will be documented in this file.

The format is based on [Keep a Changelog](http://keepachangelog.com/en/1.0.0/).

## [v1.0.1] - 2025-08-22

### Added

- Added a script for syncing datasheets with our docs site
- Created simplified main "README" template for syncing
- Added a simplified language distribution plot. It only uses the "main" language.

### Changed

- Changed the document size distribution plots to use log bins instead of fixed width bins.
- Changed the dataset sizes plot to make sure all datasets are represented on the y-axis.

## [v1.0.0] - 2025-08-18

### Added
Expand Down
190 changes: 98 additions & 92 deletions README.md

Large diffs are not rendered by default.

4 changes: 2 additions & 2 deletions data/adl/adl.md
Original file line number Diff line number Diff line change
Expand Up @@ -62,7 +62,7 @@ An entry in the dataset consists of the following fields:

- `id` (`str`): An unique identifier for each document.
- `text`(`str`): The content of the document.
- `source` (`str`): The source of the document (see [Source Data](#source-data)).
- `source` (`str`): The source of the document.
- `added` (`str`): An date for when the document was added to this collection.
- `created` (`str`): An date range for when the document was originally created.
- `token_count` (`int`): The number of tokens in the sample computed using the Llama 8B tokenizer
Expand All @@ -74,7 +74,7 @@ An entry in the dataset consists of the following fields:

<!-- START-DATASET PLOTS -->
<p align="center">
<img src="./images/dist_document_length.png" width="600" style="margin-right: 10px;" />
<img src="./images/dist_document_length.svg" width="600" style="margin-right: 10px;" />
</p>
<!-- END-DATASET PLOTS -->

Expand Down
6 changes: 3 additions & 3 deletions data/adl/descriptive_stats.json
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
"number_of_tokens": 58493311,
"min_length_tokens": 53,
"max_length_tokens": 662143,
"number_of_characters": 161816257,
"min_length_characters": 136,
"max_length_characters": 1879004
"number_of_characters": 164742458,
"min_length_characters": 137,
"max_length_characters": 1928381
}
3,885 changes: 3,885 additions & 0 deletions data/adl/images/dist_document_length.html

Large diffs are not rendered by default.

Binary file removed data/adl/images/dist_document_length.png
Binary file not shown.
1 change: 1 addition & 0 deletions data/adl/images/dist_document_length.svg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
4 changes: 2 additions & 2 deletions data/ai-aktindsigt/ai-aktindsigt.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,7 +57,7 @@ An entry in the dataset consists of the following fields:

- `id` (`str`): An unique identifier for each document.
- `text`(`str`): The content of the document.
- `source` (`str`): The source of the document (see [Source Data](#source-data)).
- `source` (`str`): The source of the document.
- `added` (`str`): An date for when the document was added to this collection.
- `created` (`str`): An date range for when the document was originally created.
- `token_count` (`int`): The number of tokens in the sample computed using the Llama 8B tokenizer
Expand All @@ -68,7 +68,7 @@ An entry in the dataset consists of the following fields:

<!-- START-DATASET PLOTS -->
<p align="center">
<img src="./images/dist_document_length.png" width="600" style="margin-right: 10px;" />
<img src="./images/dist_document_length.svg" width="600" style="margin-right: 10px;" />
</p>
<!-- END-DATASET PLOTS -->

Expand Down
6 changes: 3 additions & 3 deletions data/ai-aktindsigt/descriptive_stats.json
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
"number_of_tokens": 139234696,
"min_length_tokens": 9,
"max_length_tokens": 152599,
"number_of_characters": 408005923,
"min_length_characters": 29,
"max_length_characters": 406832
"number_of_characters": 417940348,
"min_length_characters": 30,
"max_length_characters": 412486
}
3,885 changes: 3,885 additions & 0 deletions data/ai-aktindsigt/images/dist_document_length.html

Large diffs are not rendered by default.

Binary file removed data/ai-aktindsigt/images/dist_document_length.png
Binary file not shown.
1 change: 1 addition & 0 deletions data/ai-aktindsigt/images/dist_document_length.svg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
4 changes: 2 additions & 2 deletions data/arxiv_abstracts_filtered/arxiv_abstracts_filtered.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,7 @@ An entry in the dataset consists of the following fields:

- `id` (`str`): An unique identifier for each document.
- `text`(`str`): The content of the document.
- `source` (`str`): The source of the document (see [Source Data](#source-data)).
- `source` (`str`): The source of the document.
- `added` (`str`): An date for when the document was added to this collection.
- `created` (`str`): An date range for when the document was originally created.
- `token_count` (`int`): The number of tokens in the sample computed using the Llama 8B tokenizer
Expand All @@ -51,7 +51,7 @@ An entry in the dataset consists of the following fields:

<!-- START-DATASET PLOTS -->
<p align="center">
<img src="./images/dist_document_length.png" width="600" style="margin-right: 10px;" />
<img src="./images/dist_document_length.svg" width="600" style="margin-right: 10px;" />
</p>
<!-- END-DATASET PLOTS -->

Expand Down
2 changes: 1 addition & 1 deletion data/arxiv_abstracts_filtered/descriptive_stats.json
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
"number_of_tokens": 524449662,
"min_length_tokens": 4,
"max_length_tokens": 1308,
"number_of_characters": 2395960336,
"number_of_characters": 2395960436,
"min_length_characters": 6,
"max_length_characters": 6091
}
3,885 changes: 3,885 additions & 0 deletions data/arxiv_abstracts_filtered/images/dist_document_length.html

Large diffs are not rendered by default.

Binary file not shown.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
4 changes: 2 additions & 2 deletions data/arxiv_papers_filtered/arxiv_papers_filtered.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,7 @@ An entry in the dataset consists of the following fields:

- `id` (`str`): An unique identifier for each document.
- `text`(`str`): The content of the document.
- `source` (`str`): The source of the document (see [Source Data](#source-data)).
- `source` (`str`): The source of the document.
- `added` (`str`): An date for when the document was added to this collection.
- `created` (`str`): An date range for when the document was originally created.
- `token_count` (`int`): The number of tokens in the sample computed using the Llama 8B tokenizer
Expand All @@ -51,7 +51,7 @@ An entry in the dataset consists of the following fields:

<!-- START-DATASET PLOTS -->
<p align="center">
<img src="./images/dist_document_length.png" width="600" style="margin-right: 10px;" />
<img src="./images/dist_document_length.svg" width="600" style="margin-right: 10px;" />
</p>
<!-- END-DATASET PLOTS -->

Expand Down
4 changes: 2 additions & 2 deletions data/arxiv_papers_filtered/descriptive_stats.json
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
"number_of_tokens": 6110615238,
"min_length_tokens": 3,
"max_length_tokens": 994988,
"number_of_characters": 19425252893,
"number_of_characters": 2300402326,
"min_length_characters": 6,
"max_length_characters": 2478816
"max_length_characters": 2482793
}
3,885 changes: 3,885 additions & 0 deletions data/arxiv_papers_filtered/images/dist_document_length.html

Large diffs are not rendered by default.

Binary file not shown.
1 change: 1 addition & 0 deletions data/arxiv_papers_filtered/images/dist_document_length.svg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Original file line number Diff line number Diff line change
Expand Up @@ -40,7 +40,7 @@ An entry in the dataset consists of the following fields:

- `id` (`str`): An unique identifier for each document.
- `text`(`str`): The content of the document.
- `source` (`str`): The source of the document (see [Source Data](#source-data)).
- `source` (`str`): The source of the document.
- `added` (`str`): An date for when the document was added to this collection.
- `created` (`str`): An date range for when the document was originally created.
- `token_count` (`int`): The number of tokens in the sample computed using the Llama 8B tokenizer
Expand All @@ -53,7 +53,7 @@ An entry in the dataset consists of the following fields:

<!-- START-DATASET PLOTS -->
<p align="center">
<img src="./images/dist_document_length.png" width="600" style="margin-right: 10px;" />
<img src="./images/dist_document_length.svg" width="600" style="margin-right: 10px;" />
</p>
<!-- END-DATASET PLOTS -->

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
"number_of_tokens": 8615087224,
"min_length_tokens": 6,
"max_length_tokens": 16891,
"number_of_characters": 35371244933,
"number_of_characters": 1094752881,
"min_length_characters": 32,
"max_length_characters": 56738
"max_length_characters": 57806
}
3,885 changes: 3,885 additions & 0 deletions data/biodiversity_heritage_library_filtered/images/dist_document_length.html

Large diffs are not rendered by default.

Binary file not shown.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
4 changes: 2 additions & 2 deletions data/botxt/botxt.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,7 +59,7 @@ An entry in the dataset consists of the following fields:

- `id` (`str`): An unique identifier for each document.
- `text`(`str`): The content of the document.
- `source` (`str`): The source of the document (see [Source Data](#source-data)).
- `source` (`str`): The source of the document.
- `added` (`str`): An date for when the document was added to this collection.
- `created` (`str`): An date range for when the document was originally created.
- `token_count` (`int`): The number of tokens in the sample computed using the Llama 8B tokenizer
Expand All @@ -69,7 +69,7 @@ An entry in the dataset consists of the following fields:

<!-- START-DATASET PLOTS -->
<p align="center">
<img src="./images/dist_document_length.png" width="600" style="margin-right: 10px;" />
<img src="./images/dist_document_length.svg" width="600" style="margin-right: 10px;" />
</p>
<!-- END-DATASET PLOTS -->

Expand Down
6 changes: 3 additions & 3 deletions data/botxt/descriptive_stats.json
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
"number_of_tokens": 847973,
"min_length_tokens": 407,
"max_length_tokens": 83792,
"number_of_characters": 2011076,
"min_length_characters": 845,
"max_length_characters": 202015
"number_of_characters": 2159369,
"min_length_characters": 976,
"max_length_characters": 218596
}
3,885 changes: 3,885 additions & 0 deletions data/botxt/images/dist_document_length.html

Large diffs are not rendered by default.

Binary file removed data/botxt/images/dist_document_length.png
Binary file not shown.
Loading