Investigating subword tokenization strategies for the translation of English to low-resource South African languages

dc.contributor.advisorNyirenda, Juwa Chiza
dc.contributor.authorTyobeka, Manala
dc.date.accessioned2026-07-17T09:06:47Z
dc.date.available2026-07-17T09:06:47Z
dc.date.issued2026
dc.date.updated2026-07-17T09:05:00Z
dc.description.abstractThis thesis investigates how adaptations made to standard subword tokenization algorithms impact the translation performance of Transformer-based Neural Machine Translation (NMT) models that are trained for the translation of text from English into low-resource South African languages. While NMT models are data-hungry; limited research exists to determine how the vocabulary representation, architecture and training strategies of these models can be adapted to better suit South African languages, particularly in conditions with limited time and compute resources. This thesis sources parallel corpora from the Autshumato Project, using roughly 20 000 to 149 000 source-target pairs for model training, that contain translations from English into Northern Sesotho, isiNdebele and Xitsonga. NMT models are then trained on each of the parallel corpora to perform a comparative analysis of standard subword tokenization algorithms with subword regularization, and a task-specific tokenizer for target languages. The task-specific tokenizer narrows the gap between tokenization and translation by leveraging the loss function of a model, that has the same architecture as the NMT model used for translation, to constrain its parameter search. While standard subword tokenization algorithms generally outperform their corresponding subword regularization methods; subword regularization displays superior generalisation capabilities in the face noisy training examples. Relative to standard subword tokenization, the performance of the task-specific tokenizer is comparable; however, it has the capacity to improve the model's predictive accuracy with enough training examples.
dc.identifier.apacitationTyobeka, M. (2026). <i>Investigating subword tokenization strategies for the translation of English to low-resource South African languages</i>. (). University of Cape Town. Retrieved from http://hdl.handle.net/11427/43599en_ZA
dc.identifier.chicagocitationTyobeka, Manala. <i>"Investigating subword tokenization strategies for the translation of English to low-resource South African languages."</i> ., University of Cape Town, 2026. http://hdl.handle.net/11427/43599en_ZA
dc.identifier.citationTyobeka, M. 2026. Investigating subword tokenization strategies for the translation of English to low-resource South African languages. . University of Cape Town. http://hdl.handle.net/11427/43599en_ZA
dc.identifier.ris TY - Thesis / Dissertation AU - Tyobeka, Manala AB - This thesis investigates how adaptations made to standard subword tokenization algorithms impact the translation performance of Transformer-based Neural Machine Translation (NMT) models that are trained for the translation of text from English into low-resource South African languages. While NMT models are data-hungry; limited research exists to determine how the vocabulary representation, architecture and training strategies of these models can be adapted to better suit South African languages, particularly in conditions with limited time and compute resources. This thesis sources parallel corpora from the Autshumato Project, using roughly 20 000 to 149 000 source-target pairs for model training, that contain translations from English into Northern Sesotho, isiNdebele and Xitsonga. NMT models are then trained on each of the parallel corpora to perform a comparative analysis of standard subword tokenization algorithms with subword regularization, and a task-specific tokenizer for target languages. The task-specific tokenizer narrows the gap between tokenization and translation by leveraging the loss function of a model, that has the same architecture as the NMT model used for translation, to constrain its parameter search. While standard subword tokenization algorithms generally outperform their corresponding subword regularization methods; subword regularization displays superior generalisation capabilities in the face noisy training examples. Relative to standard subword tokenization, the performance of the task-specific tokenizer is comparable; however, it has the capacity to improve the model's predictive accuracy with enough training examples. DA - 2026 DB - OpenUCT DP - University of Cape Town KW - tokenization KW - translation LK - https://open.uct.ac.za PB - University of Cape Town PY - 2026 T1 - Investigating subword tokenization strategies for the translation of English to low-resource South African languages TI - Investigating subword tokenization strategies for the translation of English to low-resource South African languages UR - http://hdl.handle.net/11427/43599 ER - en_ZA
dc.identifier.urihttp://hdl.handle.net/11427/43599
dc.identifier.vancouvercitationTyobeka M. Investigating subword tokenization strategies for the translation of English to low-resource South African languages. []. University of Cape Town, 2026 [cited yyyy month dd]. Available from: http://hdl.handle.net/11427/43599en_ZA
dc.language.isoen
dc.language.rfc3066eng
dc.publisherUniversity of Cape Town
dc.publisher.departmentDepartment of Statistical Sciences
dc.publisher.facultyFaculty of Science
dc.publisher.institutionUniversity of Cape Town
dc.subjecttokenization
dc.subjecttranslation
dc.titleInvestigating subword tokenization strategies for the translation of English to low-resource South African languages
dc.typeThesis / Dissertation
dc.type.qualificationlevelMasters
dc.type.qualificationlevelMSc
Files
Original bundle
Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
thesis_sci_2026_tyobeka manala.pdf
Size:
2.53 MB
Format:
Adobe Portable Document Format
Description:
License bundle
Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.72 KB
Format:
Item-specific license agreed upon to submission
Description:
Collections