Show simple item record

SPCS Speech Corpus
Broadband speech corpus of approximately 10 hours and the corresponding transcriptions. The development process of the corpus involved the recording and transcribing of radio broadcasts. The transcriptions were used to generate the Sepedi code-switched prompts to re-record speech from multiple speakers. The following sub-directories are found in this directory: Audio: Audio files for all the recorded code-switched speech Transcriptions: The corresponding orthographic transcriptions Metadata: Information about the speakers and the transcriptions Documentation: The directory structure and the Sepedi prompt list
Ulrike Janke
ulrike.must@gmail.com
Council for Scientific and Industrial Research; North-West University
Creative Commons Attribution 2.5 South Africa license: https://creativecommons.org/licenses/by/2.5/legalcode
English; Sepedi
Modipa, T. I.; Davel, M. H.; De Wet, F.
Sepedi; Sesotho sa Leboa; code-switching; orthographic transcription; English
T. I. Modipa, M. H. Davel, F. De Wet, "Implications of Sepedi/English code switching for ASR systems", Pattern Recognition Association of South Africa, pp. 112-117, 2015
https://hdl.handle.net/20.500.12185/530
Speech
Orthographic transcribed broadband speech corpus
Resource Catalogue
Resource Index
eng; nso
2020-04-21T10:15:12Z
2020-04-21T10:15:12Z
2015-11-25


Files in this item

Thumbnail

This item appears in the following Collection(s)

  • Resource Catalogue [239]
    A collection of language resources available for download from the RMA of SADiLaR. The collection mostly consists of resources developed with funding from the Department of Arts and Culture.
  • Resource Index [383]
    A collection of language resource metadata mostly collected during the NHN funded technology audit of 2009, as well as the SADiLaR technology audit of 2018. Not all resources in this collection are available for download.

Show simple item record