Kurdish Speech is a research and business group working on natural language processing technologies for Kurdish language. Kurdish is a member of the Indo-Iranian branch of Indo-European languages spoken by over 40 million people mainly in Iraq, Turkey, Iran, Syria, Armenia, and Azerbaijan. Despite its diversity of dialects, Kurdish belongs to low-resourced languages in computational linguistics. Kurdish Speech conducts continuous research to develop computational resources, models, and real-world NLP applications.
Knowledge Enterprise
Kurdish Speech is a knowledge-based enterprise founded and managed by experts in Artificial Intelligence, Computer Engineering, and Computational Linguistics.
Language Data & Resources
A part from developing applications for Kurdish language, Kurdish Speech is working to provide language data and resources for Computer Speech and Language Processing of Kurdish language. Our activities to that end include, but are not limited to, providing text corpus, speech corpus, WordNet, lexicon and parallel corpora.
Speech Recognition & Tagging
Kurdish language speech data and its related resources like tags are of most important language resources which are required for NLP research and applications such as automatic speech recognition, speaker recognition, etc. In this project, speech data for Kurdish language (Central Kurdish) was designed and collected so that it could be used in automatic speech recognition, speaker recognition, phonology researches, dialect analysis, etc. So far, approximately 30 hours of speech has been recorded and transcribed in order to produce this corpus.