Skip to content

Latest commit

 

History

History
486 lines (354 loc) · 28 KB

File metadata and controls

486 lines (354 loc) · 28 KB

Datasets

Check Links

Topics

QA

Persian (Farsi) Question Answering Dataset. with models: bert-base-fa-qa with 162M parameters fine-tuned on this dataset and xlm-roberta-large-fa-qa with 558M parameters fine-tuned on this dataset and SQuAD2.0 (English) dataset.

Medical Question Answering dataset consists of 15k dialogs in 70 specialities.

26k QA and related excerpt extracted from Persian wikipedia. Some of the questions can not be answered based on the given excerpt by design (like SQuAD2.0).

Persian Question Answering Dataset based on Machine Translation of SQuAD 2.0

Persian NLP team trained various mt5 and BERT models on their multiple-choice QA dataset.

It consists of 266k legal questions, answers and related tags.

Persian translation of 35k records of Stanford Alpaca Instruction dataset (52K records). There is also a version with different formatting.

This dataset contains 5900 Persian language question-answer pairs generated using the PersianAnswerGenerator class from answer.py. The answers are produced by an AI assistant leveraging the GPT-4o model through the Avala API service.

MauxiMix is a carefully curated dataset of 1,000 high-quality Persian conversations, translated from the SmolTalk dataset using advanced language models. This dataset is specifically designed for training and fine-tuning Large Language Models (LLMs) with Supervised Fine-Tuning (SFT) techniques, contributing to the development of open-source Persian language models.

This dataset consists of 30157 pairs of questions and answers.

This dataset includes over 570k questions and more than 1.9m answers, all in written form.

This dataset contains more than 211k questions and more than 700k answers, all produced in written form.

This dataset contains approximately 50k questions and around 60k answers, all produced in written form.

Datasets (classification)

This could be a nice tool for Persian writers or bloggers to automatically pick the suggested hashtag or even subject for their articles. We could even collect data from google trend for each hashtag or 'label' used in an article. Consists of 11k+ articles.

The file contains 3780 news articles published by BBC Persian. The articles mostly belong to the year 1399 and 1400, and are published before Aban 18th, 1400. Columns are: title, publish_name, link, related_topics, body, category.

Consists of 63k News articles with following columns: category, title, abstract, body, time.

Yearly collection of the Farsnews agency (1398). Contains 294k News article with following columns: title, abstract, paragraphs, cat, subcat, tags, link.

A total of 8,515 articles scraped from Digikala Online Magazine. This dataset includes seven different classes: Video Games, Shopping Guide, Health Beauty, Science Technology, General, Art Cinema and Books Literature.

Contains about 3K tweets, with each one of them labeled as either ironic or not.

4K of records of stance detection in headlines and bodies of News articles.

Consists of 5.5K pairs of tweets which the stance of the reply tweets have been marked as against, support or neither to the main tweet.

Consists of 3.8K tweets, in which the type of each claim in each tweet have been identified. But it does not show where is the claim located in the main tweet.

NER

Name Entity Recognition (NER) on the Persian Twitter dataset. Consists of 6 entity types: event, location, natinality, organization and pog (political organizations and historical dynasties). 12k Named Entities in 232k tokens.

Extends PEYMA corpus (300k tokens), with another 600k tokens. Consists of 16 entity types including: date, location, percent number, money, time, person and organization. 48k NEs in 884k tokens.

The dataset includes 250,015 tokens and 7,682 Persian sentences in total. Consists of 6 NE types including: facility, organization, location, event, person and proper noun. 37K NEs in 749k tokens.

Crowd-sourced NE dataset with 5 NE types. 2.2M NEs in 25M tokens.

These dataset is a mixed NER dataset collected from ARMAN, PEYMA, and WikiANN that covered ten types of entities including: Date, Event, Facility, Location, Money, Organization, Percent, Person, Product and Time. 140K NEs in 40k sentences.

It is a large Multilingual Dataset for Entity Linking containing data in 53 languages including Persian. DaMuEL consists of two components: a knowledge base that contains language-agnostic information about entities, including their claims from Wikidata and named entity types (PER, ORG, LOC, EVENT, BRAND, WORK_OF_ART, MANUFACTURED); and Wikipedia texts with entity mentions linked to the knowledge base, along with language-specific text from Wikidata such as labels, aliases, and descriptions, stored separately for each language. Paper. For this project UDPipe has been used.

XTREME is a benchmark for the evaluation of the cross-lingual generalization ability of pre-trained multilingual models that covers 40 typologically diverse languages and includes nine tasks. But for Persian it only consists of:

This repository contains a comprehensive Persian NER dataset with approximately 500,000 tokens. This dataset is a collection of all available Persian NER datasets, carefully cleaned and consolidated to ensure the highest quality for training, validating, and testing NER models in the Persian language.

Unlabled and Raw

Persian real SMS Dataset

Crawled more than 3k+ articles from tarjoman website.

27M tweets. Although these texts have been labeled or translated using various NLP toolkits, they have never been supervised.

Consists of 8M words with following columns: title, date, url and body.

219K abstracts collected from Ensani.ir papers.

Toxic text

We created a dataset of 33338 Persian tweets, of which 10% contained Abusive words and 90% were non-Abusive.

Persian Swear Dataset - you can use in your production to filter unwanted content. دیتاست کلمات نامناسب و بد فارسی برای فیلتر کردن متن ها

Stop word list

A collection of Persian stopwords. Consists of:

All combined

Different sources

Consists of about 2k stop words.

Encyclopedia and Word Set

Consists of following sets:

  • Words of Sareh Dictionary (Purified Persian Words)
  • Farhangestan chosen words for non-Persian equivalents.
  • Farhange Emlaee (A dictionary of Persian orthography and spelling)
  • A part of Ganjoor's website poetry repos.
  • Farhange Motaradef va Motazad (A dictionary of Persian synonyms and antonyms)
  • Farhange Teyfi (Persian Thesaurus)

Persian names dataset

A Python package for generating random Persian (Farsi) names.

A SQL database that includes a dictionary of 494,286 Persian words.

This repository is a Persian meaningful database with json

850k categorized Persian words.

pre-calculated list of similar Persian words ordered by rating and best match

List of ~240,000 Persian words

Useful Persian dictionary and more. Consists of:

  • Dehkhoda dictionary (36k)
  • Synonyms (20k)
  • Arabic to Persian dictionary (113k)
  • Persian to Arabic dictionary(32k)
  • Abjad Persian to Arabic dictionary (42k)
  • Arabic to Persian dictionary (8k)
  • Quran Mofradat (1.6k)
  • Arabic monolingual dictionary (4.6k)
  • Intermediate Arabic dictionary (41k)
  • Alamsal - Arabic proverbs dictionary (4.5k)

The "Iranian Job Title" dataset offers a comprehensive compilation of various job titles prevalent in Iran across diverse industries and sectors.

Moeen dictionary based Thesaurus for Persian.

It's an enhanced version of Flexicon word list with syllable, IPA procunciation and some refinements in word list itself.

Poetry and Literature

A simple Telegram bot implemented in Python.

Useful Persian dictionary and more. Consists of:

  • Persian poetry of Iranian poets:
    • Ahmad Shamlou
    • Baba-Taher
    • Parvin E'tesami
    • Hafez
    • Khayyam
    • Rahi-Moayeri
    • Roodaki
    • Sa'di
    • Sohrab Sepehri
    • Shahriar
    • Saeb Tabrizi
    • Onsori
    • Ferdowsi
    • Forugh Farrokhzad
    • Mehdi Akhavan Sales
    • Mowlavi
    • Nezami
    • Nima Yushij
  • Quran Database
    • Quran Surahs (114)
    • Quran Versus (6236)
    • Quran Versus Translation by Gomshe'i (6326)
    • Quran Translation Word by word (83668)
    • Reading voice of Famous Readers (48)

Collection of Persian Modernist Poetry from Iranian contemporary poets

Crawled Ganjoor for poems of 48 poets.

This model fine-tuned on ParsGPT2 with Chronological Persian poetry dataset and can generate poems by providing the name of the poet.

Dataset of poetry of 67 Persian poets of different times.

Audio

Persian spoken digit recognition

Simple Persian Questions aimed to use in a voice assistant in 4 Categories. Labeled NEs in command utterances (in text).

About 60 hours audio produced by various users reading sentences. All sentences with duplicates are 500h+.

This ~2.5-hour Single-Speaker Speech corpus.

A semi-natural db which contains emotional speech samples of Persian speakers. The database includes 3000 semi-natural utterances, equivalent to 3 h and 25 min of speech data extracted from online radio plays.

A Deep-Learning-Based Persian Speech Recognition System. Takes advantage of various ASR platforms to create models for ASR. Also it uses various datasets including Mozzila CommonVoice and their own dataset which consists of 300h+ audio and transcription.

Phoneme based speech dataset.

Open-source tool for speech recognition for various platforms and OSes, supprting 20 languages including Persian.

It is a wav2vec model fine-tuned on Mozzila CommonVoice Persian dataset. The model and the notebook to recreate the model with extra data are avaialble.

This dataset consists of over 385 hours of transcribed audio extracted from various YouTube videos in the Persian language (more than 400k rows).

This dataset consists of about 245 hours of transcribed audio extracted from various Filimo (an Iranian VOD) videos in the Persian language (more than 400k rows).

This datasets consists of a collection of 507 articles from the Tarjoman website until the end of 2023, each accompanied by corresponding audio recordings.

OCR

This is a dataset of handwritten cities in Iran in Arabic/Persian that has been used in my Master project. This dataset is collected for sorting postal packages.

Hand-written / typed names of different cities of Iran in image format.

50*50 Images of Persian letters (without dots) with 32 Different Fonts.

Consists of about 20k images of Persian subwords in different fonts and sizes to be used in ocr models.

Spam

persian sms spam word

Image Captioning

Coco 2017 translated to Persian language. 91k images with caption in Persian.

Dataset of Farsi License Plate Characters (83k).

The VQA dataset consists of almost 11k images and 28.5k question and answer pairs with short and long answers usable for both classification and generation VQA.

A dataset consists of 16M records of images and their corresponding texts. It also consists of a model traind on 400k of this dataset for searching images based on text and image.

Consists of about 26K records of images with th describing captions in Persian.

Translation

Persian language movies dataset from imvbox.com. 14k movies with storyline translated from Persian to English.

Quran ayat with translation in 21 languages.

A multilingual parallel corpus created from translations of the Bible. In 100 languages including Persian.

A set of corpora for 120 languages including Persian automatically collected from wikipedia and the web.

Persian NLP team trained various mt5 models on their translation dataset.

Summary

Consists of 63k News articles with following columns: category, title, abstract, body, time.

Yearly collection of the Farsnews agency (1398). Contains 294k News article with following columns: title, abstract, paragraphs, cat, subcat, tags, link.

95k documents with body and summery extracted from wikipedia Persian articles. There is also notebook to create and test models for summerization.

Statistical and Semantical Text Summarizer in Persian Language

A well-structured summarization dataset for the Persian language consists of 93,207 records. It is prepared for Abstractive/Extractive tasks (like cnn_dailymail for English). It can also be used in other scopes like Text Generation, Title Generation, and News Category Classification.

Consists of similar models fine-tuned on ParsBERT using three different datasets, these models can be utilized for various applications, including Text summarization.

MirasText has more than 2.8 million articles and over 1.4 billion content words. Consists of following columns: content, summary, keywords, title, url.

Paraphrase

Paraphrase data for Persian. It consists of 2.3M sentence pairs of which 1M of them are paraphrase and 1.3M are not parapharse of each other.

Persian NLP team trained various mt5 models on their query paraphrase dataset.

Consists of 800 pairs of Persian sentences wich are paraphrases of each other.

This dataset consists of about 1.5M of paraphrase sentences pairs.

WSD

SBU-WSD-Corpus: A Sense Annotated Corpus for Persian All-words Word Sense Disambiguation.

Sentiment Analysis

Awesome Persian Sentiment Analysis Resources - منابع مرتبط با تحلیل احساسات در زبان فارسی

  • Consists of following datasets:
    • Deep Neural Networks in Persian Sentiment Analysis
    • Sentiment Analysis Challenges
    • Sentiment Lexicon
    • Sentiment Tagged Corpus (dataset)
    • HesNegar: Persian Sentiment WordNet

Consists of data (3K) and code (notebook) to create a LSTM model for Sentiment Analysis.

Sentiment analysis using ML and DL models on Persian texts

A Sentiment Analysis Lexicon for Persian. Consists of 4k words

Persian book comment ratings dataset. Consists of about 70k comment about 11k books.

The Digikala (comments & products) dataset offers a comprehensive glimpse into the vast online marketplace of Digikala, comprising over 1.2 million products and more than 6 million comments.

3k comments with score and ratings.

93k digikala products comments with manual labeling.

20k tweets with emotion identification labels.

A Dataset of 30,000 emotion labeled Persian Tweets.

Consists of 5.56K tweets with labels (sadness, anger, happiness, hatred, wonder and fear) describing their emotions.

Consists of 7k docs with 6 emotion label types (sadness, anger, happiness, hatred, wonder, fear).

Snappfood (an online food delivery company) user comments containing 70,000 comments with two labels (i.e. polarity classification): Happy, Sad.

It is the Persian translation of NRC Emotion Lexicon which is a list of English words with their associate basic emotions in eigth categories( anger, fear, anticipation, trust, surprise, sadness, joy, and disgust).

Consists of 10k samples which each record focuses on one aspect (e.g. camera, screen resolution, etc of a comment about a cell phone) of a comment. Each comment may appear on more than one sample based on the number of aspects that exist in that comment.

Consists of 1500 words with their degrees of polarity.

Utilizes the SentiPers dataset, which consists of 7,400 sentences, and enhances it with various embeddings to develop both LSTM and CNN models. All the original and newly transformed data, along with the notebooks used to create the models, are available in this repository.

Dependency Parsing

The Persian Universal Dependency Treebank (Seraji) is based on Uppsala Persian Dependency Treebank (UPDT). The conversion of the UPDT to the Universal Dependencies was performed semi-automatically with extensive manual checks and corrections.

The Persian Universal Dependency Treebank (PerUDT) is the result of automatic coversion of Persian Dependency Treebank (PerDT) with extensive manual corrections. Consists of 29k sentences.

PARSEME is a verbal multiword expressions (VMWEs) corpus for Farsi. All the annotated data come from a subset of the Farsi section of the MULTEXT-East "1984" annotated corpus 4.0. More than colums of LEMMA UPOS, XPOS, FEATS, HEAD and DEPREL there is also PARSEME:MVE which is manually annotated.

Informal Persian Universal Dependency Treebank, consisting of 3000 sentences and 54,904 tokens, is an open source collection of colloquial informal texts from Persian blogs.