• English
  • ÄŒeÅ¡tina
  • Deutsch
  • Español
  • Français
  • Gàidhlig
  • LatvieÅ¡u
  • Magyar
  • Nederlands
  • Português
  • Português do Brasil
  • Suomi
  • Svenska
  • Türkçe
  • Қазақ
  • বাংলা
  • हिंदी
  • Ελληνικά
  • Log In
  • Communities & Collections
  • Browse OpenUCT
  • English
  • ÄŒeÅ¡tina
  • Deutsch
  • Español
  • Français
  • Gàidhlig
  • LatvieÅ¡u
  • Magyar
  • Nederlands
  • Português
  • Português do Brasil
  • Suomi
  • Svenska
  • Türkçe
  • Қазақ
  • বাংলা
  • हिंदी
  • Ελληνικά
  • Log In
  1. Home
  2. Browse by Subject

Browsing by Subject "tokenization"

Now showing 1 - 1 of 1
Results Per Page
Sort Options
  • No Thumbnail Available
    Item
    Open Access
    Investigating subword tokenization strategies for the translation of English to low-resource South African languages
    (University of Cape Town, 2026) Tyobeka, Manala; Nyirenda, Juwa Chiza
    This thesis investigates how adaptations made to standard subword tokenization algorithms impact the translation performance of Transformer-based Neural Machine Translation (NMT) models that are trained for the translation of text from English into low-resource South African languages. While NMT models are data-hungry; limited research exists to determine how the vocabulary representation, architecture and training strategies of these models can be adapted to better suit South African languages, particularly in conditions with limited time and compute resources. This thesis sources parallel corpora from the Autshumato Project, using roughly 20 000 to 149 000 source-target pairs for model training, that contain translations from English into Northern Sesotho, isiNdebele and Xitsonga. NMT models are then trained on each of the parallel corpora to perform a comparative analysis of standard subword tokenization algorithms with subword regularization, and a task-specific tokenizer for target languages. The task-specific tokenizer narrows the gap between tokenization and translation by leveraging the loss function of a model, that has the same architecture as the NMT model used for translation, to constrain its parameter search. While standard subword tokenization algorithms generally outperform their corresponding subword regularization methods; subword regularization displays superior generalisation capabilities in the face noisy training examples. Relative to standard subword tokenization, the performance of the task-specific tokenizer is comparable; however, it has the capacity to improve the model's predictive accuracy with enough training examples.
UCT Libraries logo

Contact us

Jill Claassen

Manager: Scholarly Communication & Publishing

Email: openuct@uct.ac.za

+27 (0)21 650 1263

  • Open Access @ UCT

    • OpenUCT LibGuide
    • Open Access Policy
    • Open Scholarship at UCT
    • OpenUCT FAQs
  • UCT Publishing Platforms

    • UCT Open Access Journals
    • UCT Open Access Monographs
    • UCT Press Open Access Books
    • Zivahub - Open Data UCT
  • Site Usage

    • Cookie settings
    • Privacy policy
    • End User Agreement
    • Send Feedback

DSpace software copyright © 2002-2026 LYRASIS