SADiLaR Language Resource Repository: Recent submissions
Now showing items 221-230 of 529
-
NCHLT isiXhosa Speech Corpus
(Meraka Institute, CSIR; North-West University, 2014-07-08) ~Resource Catalogue Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers. -
NCHLT Sesotho Speech Corpus
(Meraka Institute, CSIR; North-West University, 2014-07-08) ~Resource Catalogue Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers. -
NCHLT isiZulu Lemmatiser
(North-West University; Centre for Text Technology (CTexT), 2014-05-30) ~Resource Catalogue Lemmatiser developed during the NCHLT Text project. \n\n Available in the Readme.txt - Input format: Text data (encoding: UTF8 without BOM), one ... -
NCHLT Afrikaans Lemmatiser
(North-West University; Centre for Text Technology (CTexT), 2014-05-30) ~Resource Catalogue Lemmatiser developed during the NCHLT Text project. \n\n Available in the Readme.txt - Input format: Text data (encoding: UTF8 without BOM), one lowercase ... -
NCHLT Part of Speech Taggers
(North-West University; Centre for Text Technology (CTexT), 2014-05-30) ~Resource Catalogue Part of speech taggers developed during the NCHLT Text project. Available for the following languages: Afrikaans, English, isiNdebele, isiXhosa, isiZulu, ... -
Lara2
(North-West University; Centre for Text Technology (CTexT), 2014-05-30) ~Resource Catalogue Tool for annotating texts with lemma, part of speech and morphological analysis information -
CTexTools
(North-West University; Centre for Text Technology (CTexT), 2014-05-30) ~Resource Catalogue Corpus query and manipulation tool for performing tokenisation and sentencisation; extracting frequency list and word list; searching; and extracting ... -
NCHLT Tshivenda Text Corpora
(North-West University; Centre for Text Technology (CTexT), 2014-05-30) ~Resource Catalogue Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed ... -
NCHLT Xitsonga Text Corpora
(North-West University; Centre for Text Technology (CTexT), 2014-05-30) ~Resource Catalogue Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed ... -
NCHLT Setswana Text Corpora
(North-West University; Centre for Text Technology (CTexT), 2014-05-30) ~Resource Catalogue Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed ...