Our lab develops and releases datasets, models, and tools to support research in Natural Language Processing, Information Retrieval, and Speech Processing, with a particular focus on Persian and low-resource languages.
A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis
The largest Persian speech corpus designed specifically for TTS applications, featuring 1,804 hours of high-quality speech from 470+ speakers.
A Corpus for Persian Punctuation Restoration
A comprehensive dataset for Persian punctuation restoration, enabling models to accurately restore punctuation marks in unpunctuated Persian text.
Paper: Accepted at SilkRoad NLP Workshop @ EACL 2025
Dataset: Download
Year: 2025
A Story-Driven Cultural Evaluation of LLMs in Persian
Dataset designed to assess cultural sensitivity of LLMs toward Persian culture through story-based, multiple-choice questions.
U.S. Political Bluesky Dataset with User Stance Labels
First-of-its-kind dataset for stance detection on Bluesky, focused on the 2024 U.S. presidential election. Contains 16,044 user-target stance pairs with detailed metadata.
Evaluating Large Language Models on a Persian-English Bilingual Medical Question Answering Benchmark
Comprehensive evaluation dataset for assessing LLM performance in Persian medical domain, with parallel English translations.
Hate Speech and Offensive Language Detection
Dataset of 38K Persian tweets containing hate and offensive language, manually annotated with bias measurement metrics.
Paper: Link
Year: 2024
Pre-trained Model for Persian Abstractive Summarization
Transformer-based encoder-decoder model achieving state-of-the-art performance on Persian summarization tasks.
Many of our projects include open-source code repositories. Visit our GitHub organization for implementation details and tools.
If you use our resources in your research, please cite the relevant papers. See individual resource pages for specific citation information.