Investigating subword tokenization strategies for the translation of English to low-resource South African languages

Thesis / Dissertation

2026

Permanent link to this Item
Authors
Journal Title
Link to Journal
Journal ISSN
Volume Title
Publisher

University of Cape Town

Publisher

University of Cape Town

License
Series
Abstract
This thesis investigates how adaptations made to standard subword tokenization algorithms impact the translation performance of Transformer-based Neural Machine Translation (NMT) models that are trained for the translation of text from English into low-resource South African languages. While NMT models are data-hungry; limited research exists to determine how the vocabulary representation, architecture and training strategies of these models can be adapted to better suit South African languages, particularly in conditions with limited time and compute resources. This thesis sources parallel corpora from the Autshumato Project, using roughly 20 000 to 149 000 source-target pairs for model training, that contain translations from English into Northern Sesotho, isiNdebele and Xitsonga. NMT models are then trained on each of the parallel corpora to perform a comparative analysis of standard subword tokenization algorithms with subword regularization, and a task-specific tokenizer for target languages. The task-specific tokenizer narrows the gap between tokenization and translation by leveraging the loss function of a model, that has the same architecture as the NMT model used for translation, to constrain its parameter search. While standard subword tokenization algorithms generally outperform their corresponding subword regularization methods; subword regularization displays superior generalisation capabilities in the face noisy training examples. Relative to standard subword tokenization, the performance of the task-specific tokenizer is comparable; however, it has the capacity to improve the model's predictive accuracy with enough training examples.
Description

Reference:

Collections