Browsing Resource Index by Title
Filter by:
Now showing items 253-272 of 412
-
NCHLT Sepedi Annotated Text Corpora
(North-West University; Centre for Text Technology (CTexT), 2014-05-30) ~Resource Catalogue Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Text project. -
NCHLT Sepedi Auxiliary Speech Corpus
(CSIR Meraka Institute; North-West University, 2019-06-01) ~Resource Catalogue The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven official languages. Transcriptions are provided in ... -
NCHLT Sepedi Lemmatiser
(North-West University; Centre for Text Technology (CTexT), 2014-05-30) ~Resource Catalogue Lemmatiser developed during the NCHLT Text project. \n\n Available in the Readme.txt - Input format: Text data (encoding: UTF8 without BOM), one ... -
NCHLT Sepedi Named Entity Annotated Corpus
(North-West University; Centre for Text Technology (CTexT), 2016-04-29) ~Resource Catalogue Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags. -
NCHLT Sepedi Phrase Chunk Annotated Corpus
(North-West University; Centre for Text Technology (CTexT), 2016-04-29) ~Resource Catalogue Phrase chunk annotated data for the NCHLT Text Resource Development: Phase II Project. The phrase chunk annotated data is a subset of the 50,000 tokens ... -
NCHLT Sepedi Speech Corpus
(Meraka Institute, CSIR; North-West University, 2014-07-08) ~Resource Catalogue Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers. -
NCHLT Sepedi Text Corpora
(North-West University; Centre for Text Technology (CTexT), 2014-05-30) ~Resource Catalogue Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed ... -
NCHLT Sesotho Morphological Decomposer
(North-West University; Centre for Text Technology (CTexT), 2014-05-30) ~Resource Catalogue Morphological decomposer developed during the NCHLT Text project. -
NCHLT Sesotho Annotated Text Corpora
(North-West University; Centre for Text Technology (CTexT), 2014-05-30) ~Resource Catalogue Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Text project. -
NCHLT Sesotho Auxiliary Speech Corpus
(CSIR Meraka Institute; North-West University, 2019-06-01) ~Resource Catalogue The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven official languages. Transcriptions are provided in ... -
NCHLT Sesotho Lemmatiser
(North-West University; Centre for Text Technology (CTexT), 2014-05-30) ~Resource Catalogue Lemmatiser developed during the NCHLT Text project. \n\n Available in the Readme.txt - Input format: Text data (encoding: UTF8 without BOM), one ... -
NCHLT Sesotho Named Entity Annotated Corpus
(North-West University; Centre for Text Technology (CTexT), 2016-04-29) ~Resource Catalogue Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags. -
NCHLT Sesotho Phrase Chunk Annotated Corpus
(North-West University; Centre for Text Technology (CTexT), 2016-04-29) ~Resource Catalogue Phrase chunk annotated data for the NCHLT Text Resource Development: Phase II Project. The phrase chunk annotated data is a subset of the 50,000 tokens ... -
NCHLT Sesotho Speech Corpus
(Meraka Institute, CSIR; North-West University, 2014-07-08) ~Resource Catalogue Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers. -
NCHLT Sesotho Text Corpora
(North-West University; Centre for Text Technology (CTexT), 2014-05-30) ~Resource Catalogue Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed ... -
NCHLT Setswana Annotated Text Corpora
(North-West University; Centre for Text Technology (CTexT), 2014-05-30) ~Resource Catalogue Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Text project. -
NCHLT Setswana Auxiliary Speech Corpus
(CSIR Meraka Institute; North-West University, 2019-06-01) ~Resource Catalogue The corpus contains orthographically transcribed broadband speech in each of South Africa's eleven official languages. Transcriptions are provided in ... -
NCHLT Setswana Lemmatiser
(North-West University; Centre for Text Technology (CTexT), 2014-05-30) ~Resource Catalogue Lemmatiser developed during the NCHLT Text project. \n\n Available in the Readme.txt - Input format: Text data (encoding: UTF8 without BOM), one ... -
NCHLT Setswana Morphological Decomposer
(North-West University; Centre for Text Technology (CTexT), 2014-05-30) ~Resource Catalogue Morphological decomposer developed during the NCHLT Text project. -
NCHLT Setswana Named Entity Annotated Corpus
(North-West University; Centre for Text Technology (CTexT), 2016-04-29) ~Resource Catalogue Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags.