Kurdish Speech is the first research and knowledge enterprise to develop Automatic Speech Recognition (ASR), Speaker Recognition, and Speech Command software for Kurdish through Artificial Intelligence and Signal Processing.
Kurdish Speech is a research and business group working on natural language processing technologies for Kurdish language. Kurdish is a member of the Indo-Iranian branch of Indo-European languages spoken by over 40 million people mainly in Iraq, Turkey, Iran, Syria, Armenia, and Azerbaijan. Despite its diversity of dialects, Kurdish belongs to low-resourced languages in computational linguistics. Kurdish Speech conducts continuous research to develop computational resources, models, and real-world NLP applications.
Kurdish Speech is a knowledge-based enterprise founded and managed by experts in Artificial Intelligence, Computer Engineering, and Computational Linguistics.
A part from developing applications for Kurdish language, Kurdish Speech is working to provide language data and resources for Computer Speech and Language Processing of Kurdish language. Our activities to that end include, but are not limited to, providing text corpus, speech corpus, WordNet, lexicon and parallel corpora.
Kurdish language speech data and its related resources like tags are of most important language resources which are required for NLP research and applications such as automatic speech recognition, speaker recognition, etc. In this project, speech data for Kurdish language (Central Kurdish) was designed and collected so that it could be used in automatic speech recognition, speaker recognition, phonology researches, dialect analysis, etc. So far, approximately 30 hours of speech has been recorded and transcribed in order to produce this corpus.
Advanced natural language processing laboratory and editor tailored for Kurdish, Persian, and English text architectures
The NLP Lab is an interactive laboratory platform designed for comprehensive text analysis and transformation across Kurdish (Sorani & Kurmanji), Persian, and English. It incorporates advanced pipeline features including tokenization, stemming, lemmatization, NER, chunking, and keyword extraction.
Explore our state-of-the-art interactive lab platform designed for speech recognition, audio processing, and voice intelligence
The Kurdish Speech Laboratory integrates advanced machine learning models specifically trained on low-resource linguistic data. Researchers, linguists, and developers can interactively test automatic speech recognition (ASR), transformer-based text-to-speech (TTS), optical character recognition (OCR), and deep voice processing engines in real time.
Fast and accurate orthographic error detection and normalization tool for Kurdish text standardization
Precise and Fast for Editing Kurdish Texts .flags spelling errors and makes suggestions for correction ,flags punctuation errors and fixes them at ones,detects the various spellings of the words and suggests the standard form
Key artificial intelligence and NLP concepts researched by our engineering team
Enabling computers to interpret text and speech. Covers Tokenization, Lemmatization, POS Tagging, NER, and Chunking.
Clarifying the architectural differences and overlaps between Generative AI, Large Language Models, and Foundation Models.
Techniques including Zero-Shot, Few-Shot, Chain-of-Thought (CoT), Tree of Thoughts (ToT), and Retrieval Augmented Generation (RAG).
Parameter-Efficient Fine-Tuning methods (Prompt-tuning, P-tuning) to improve pretrained LLM performance with minimal compute cost.
Common questions regarding speech synthesis, recognition, and editing
Realistic Kurdish Text to Speech. TTS stands for Text-to-Speech. It is a technology that converts written text into spoken words. TTS systems take input text and use synthetic voices to produce spoken audio. These systems are widely used in accessibility features, automated customer service, navigation systems, and audiobook generation.
Automatic Speech Recognition for Central Kurdish. ASR converts spoken language into written text (STT). ASR systems analyze audio input and transcribe it into textual output, making it possible to convert spoken words into a digital format that can be processed, stored, or analyzed by computers.
Precise and Fast for Editing Kurdish Texts. Flags spelling errors and makes suggestions for correction, flags punctuation errors and fixes them at once, detects various spellings of words, and suggests the unified standard form.
Reach out to our research and engineering team
Iran, Tehran, Vanak
(Iran) +98 912 505 3609