DSI LogoSADiLaR Logo
Clarin-ZA Logo
View Item 
  •   SADiLaR
  • Language Resource Management Agency
  • Resource Index
  • View Item
  •   SADiLaR
  • Language Resource Management Agency
  • Resource Index
  • View Item
    • Login
    JavaScript is disabled for your browser. Some features of this site may not work without it.

    Search form

    Browse

    All of SADiLaR

    Communities & CollectionsTitleProjectMedia type

    This Collection

    TitleProjectMedia type

    CKarma

    Thumbnail
    URI
    https://hdl.handle.net/20.500.12185/145
    Collections
    • Resource Index [409]
    Metadata
    Show full item record
    Description
    CKarma is a compound analyser for Afrikaans, to be used for the detection of word boundaries within compounds. It takes as input a string, and produces as output an analysed string, without any tags. For example, the string "hondehokdak" ('dog house roof') will be analysed as "hond _ e + hok + dak", where the plus sign indicates the beginning of an independent constituent, and the underscore the beginning of a dependent constituent (i.e. a valence morpheme). CKarma is a C5 classifier, trained on data consisting of circa 47,000 compound and 7,000 non-compounds. The resulting decision tree and cases can be converted to C code by means of a script written by MM van Zaanen. This C code can then be implemented in any other system.
    Contact person
    Martin Puttkammer
    Contact person's e-mail address
    Martin.Puttkammer@nwu.ac.za
    Publisher(s)
    North-West University
    Centre for Text Technology (CTexT)

    Copyright © 2018  SADiLaR. All Rights Reserved.
    Contact Us | Send Feedback
     

     


    Copyright © 2018  SADiLaR. All Rights Reserved.
    Contact Us | Send Feedback