MULTILINGUAL NEURAL MODELS FOR TRANSLATING CHAGATAI HISTORICAL TEXTS
DOI:
https://doi.org/10.37943/G0JE5782%20Keywords:
NLP, neural machine translation, multilingual models, low-resource languages, historical text translation, Chagatai languageAbstract
Recent advances in artificial intelligence have significantly improved automatic translation, yet historical and extremely low-resource languages remain challenging due to limited training data, complex morphology, and orthographic variation. This study investigates the effectiveness of modern neural translation approaches for translating Chagatai historical texts. The written heritage of Central Asia contains many historical documents composed in the Chagatai language, which served as a major literary and administrative language across the region from the Timurid period until the early twentieth century. Although many manuscripts have been digitized, the majority remain inaccessible to broader scholarly and public audiences because reliable translations into modern languages are scarce.
In this study we constructed an experimental parallel dataset by web scraping and custom extraction scripts. Dataset contains 10712 aligned entries in Chagatai paired with English and Kazakh translations. Preprocessing steps includes Unicode normalization, script standardization, and multilingual alignment. After that we use fine-tuning for two different translation architectures: a multilingual NLLB-200 neural machine translation model trained within the No Language Left Behind framework and a TranslateGemma generative large language model instruction-tuned for translation tasks. For the NLLB-200, the tokenizer was extended with a chg_arab token, and training used a Seq2Seq framework with AdamW optimization, mixed precision, and early stopping based on BLEU scores. For TranslateGemma, parameter-efficient fine-tuning with DoRA adapted the model’s key projection and embedding layers to Chagatai orthography, leveraging typological similarity with Uzbek language, while regularization and precision optimizations ensured stable convergence on the limited dataset.
The results show that the NLLB-200 multilingual model consistently produces more reliable translations, achieving higher scores in most evaluation metrics and demonstrating stronger semantic fidelity, particularly when translating into the related Turkic language Kazakh. In contrast, the generative model often produces fluent but less faithful translations and exhibits a higher rate of hallucinated content. For example, XCOMET showed that NLLB-200 consistently outperformed TranslateGemma, achieving 74.97 for English and 74.31 for Kazakh, compared to Gemma’s 73.80 and 69.75, respectively. These findings suggest that source-anchored multilingual translation models, such as NLLB-200, are better suited for translating extremely low-resource historical languages like Chagatai.
References
Abylkhozhin, Zh.B., Akulov, M.L., & Tsai, A.V. (Eds.). (2019). Zhivaya pamyat'. Stalinizm v Kazakhstane: Proshloe, Pamyat', Preodolenie. Daik-Press. ISBN: 978-601-290-110-8.
Kairanbayeva, N.N., & Shadkam, Z. (2019). Shag'ataj tilindegi sozdikterdi zertteulerge sholu. Journal of Oriental Studies, 91(4), 61-67. https://doi.org/10.26577/JOS-2019-4-o6
Tangsykbay, A. T. (2024). Qazaq tildi auditorijag’a shag’ataj tili leksikasyn oqytu zholdary («Babyrnama» mysalynda). Bulletin of LN Gumilyov Eurasian National University. PHILOLOGY Series, 149(4), 261-271. https://doi.org/10.32523/2616-678X-2024-149-4-261-271
Dieng, A. B., Wang, C., Gao, J., & Paisley, J. (2016). Topicrnn: A recurrent neural network with long-range semantic dependency. arXiv preprint arXiv:1611.01702. https://doi.org/10.48550/arXiv.1611.01702
Mirzakhalov, J., Babu, A., Ataman, D., Kariev, S., Tyers, F., Abduraufov, O., ... & Chellappan, S. (2021, November). A large-scale study of machine translation in Turkic languages. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (pp. 5876-5890). https://doi.org/10.18653/v1/2021.emnlp-main.475
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30. https://doi.org/10.48550/arXiv.1706.03762
Kornfilt, J. (2018). Turkish and the Turkic languages. In The world's major languages (pp. 536-561). Routledge. https://www.taylorfrancis.com/chapters/edit/10.4324/9780203216149-12/turkish-turkic-languages-jaklin-kornfilt
Kocmi, T., & Bojar, O. (2018, October). Trivial transfer learning for low-resource neural machine translation. In Proceedings of the third conference on machine translation: research papers (pp. 244-252). https://doi.org/10.18653/v1/W18-6325
Aliyu, Y., Sarlan, A., Danyaro, K. U., Rahman, A. S. B., & Abdullahi, M. (2024). Sentiment analysis in low-resource settings: a comprehensive review of approaches, languages, and data sources. IEEE Access, 12, 66883-66909. https://doi.org/10.1109/ACCESS.2024.3398635
Anfaresi, M. S. T. (2023). Developing multi translation chat application using Django frameworks and M2M100 model (Doctoral dissertation, Universitas Islam Indonesia).
Lee, E. S. A., Thillainathan, S., Nayak, S., Ranathunga, S., Adelani, D. I., Su, R., & McCarthy, A. D. (2022, May). Pre-trained multilingual sequence-to-sequence models: A hope for low-resource language translation?. In Findings of the Association for Computational Linguistics: ACL 2022 (pp. 58-67). https://doi.org/10.18653/v1/2022.findings-acl.6
Costa-Jussà, M. R., Cross, J., Çelebi, O., Elbayad, M., Heafield, K., Heffernan, K., ... & NLLB Team. (2022). No language left behind: Scaling human-centered machine translation. arXiv preprint arXiv:2207.04672. https://doi.org/10.48550/arXiv.2207.04672
Liang, X., Khaw, Y. M. J., Liew, S. Y., Tan, T. P., & Qin, D. (2025). Towards low-resource languages machine translation: A language-specific fine-tuning with LoRA for specialized large language models. IEEE Access. https://doi.org/10.1109/ACCESS.2025.3549795
Jiao, W., Wang, W., Huang, J. T., Wang, X., Shi, S., & Tu, Z. (2023). Is ChatGPT a good translator? Yes with GPT-4 as the engine. arXiv preprint arXiv:2301.08745. https://doi.org/10.48550/arXiv.2301.08745
Vilar, D., Freitag, M., Cherry, C., Luo, J., Ratnakar, V., & Foster, G. (2023, July). Prompting palm for translation: Assessing strategies and performance. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 15406-15427). https://doi.org/10.18653/v1/2023.acl-long.859
Nguyen, H. N., Do, D., Nguyen, B., Nguyen, L., Nguyen, D., & Tran, V. A. (2024, December). Advanced Language Model-Based Translator for English-Vietnamese Translation. In 2024 RIVF International Conference on Computing and Communication Technologies (RIVF) (pp. 398-402). IEEE. https://doi.org/10.1109/RIVF64335.2024.11009075
Chen, S., Shi, X., Li, P., Li, Y., & Liu, J. (2024). Refining translations with llms: A constraint-aware iterative prompting approach. arXiv preprint arXiv:2411.08348. https://doi.org/10.48550/arXiv.2411.08348
Hendy, A., Abdelrehim, M., Sharaf, A., Raunak, V., Gabr, M., Matsushita, H., ... & Awadalla, H. H. (2023). How good are gpt models at machine translation? a comprehensive evaluation. arXiv preprint arXiv:2302.09210. https://doi.org/10.48550/arXiv.2302.09210
Garcia, X., Bansal, Y., Cherry, C., Foster, G., Krikun, M., Johnson, M., & Firat, O. (2023, July). The unreasonable effectiveness of few-shot learning for machine translation. In International Conference on Machine Learning (pp. 10867-10878). PMLR. https://proceedings.mlr.press/v202/garcia23a.html
Finkelstein, M., Caswell, I., Domhan, T., Peter, J. T., Juraska, J., Riley, P., ... & Vilar, D. (2026). TranslateGemma Technical Report. arXiv preprint arXiv:2601.09012. https://doi.org/10.48550/arXiv.2601.09012
Buscaldi, D., & Rosso, P. (2023, November). How Good is NLLB for Low-resource Languages? A Study on the Genoese Language. In Proceedings of the 9th Italian Conference on Computational Linguistics (CLiC-it 2023) (pp. 490-493). https://aclanthology.org/2023.clicit-1.61/
Mesham, S., Hayward, L., Shapiro, J., & Buys, J. (2021). Low-resource language modelling of south african languages. arXiv preprint arXiv:2104.00772. https://doi.org/10.48550/arXiv.2104.00772
Ko, W., El-Kishky, A., Renduchintala, A., Chaudhary, V., Goyal, N., Guzmán, F., Fung, P., Koehn, P., & Diab, M. (2021). Adapting High-resource NMT Models to Translate Low-resource Related Languages without Parallel Data. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, 1, 802–812. https://doi.org/10.18653/v1/2021.acl-long.66
Schluessel, E. (2018). An introduction to Chaghatay: A graded textbook for reading Central Asian sources. Michigan Publishing. USA. https://doi.org/10.3998/MPUB.10110094
Mamyrbekova, G., & Seitbekova, A. (2013). Abilgazy bakһadur khannyn «Turki shezhiresinin» tezaurus sozdigi [Thesaurus dictionary of the “Turkic chronicle” by Abylgazy Bahadur Khan]. Almaty:«Firma Ornak»[in Kazakh].
Moran, S., & Cysouw, M. (2018). The Unicode Cookbook for Linguists: Managing writing systems using orthography profiles (Vol. 10). Language Science Press. https://doi.org/10.5281/zenodo.1296780
Asoka Chakravarthi, B. R. (2020). Leveraging orthographic information to improve machine translation of under-resourced languages (Doctoral dissertation, NUI Galway). https://doi.org/10.13025/17483
Fedus, W., Zoph, B., & Shazeer, N. (2022). Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120), 1-39. https://doi.org/10.48550/arXiv.2101.03961
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., ... & Wu, Y. (2023). Palm 2 technical report. arXiv preprint arXiv:2305.10403. https://doi.org/10.48550/arXiv.2305.10403
Post, M. (2018, October). A call for clarity in reporting BLEU scores. In Proceedings of the third conference on machine translation: Research papers (pp. 186-191). https://doi.org/10.18653/v1/W18-6319
Papineni, K., Roukos, S., Ward, T., & Zhu, W. J. (2002, July). Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics (pp. 311-318). https://doi.org/10.3115/1073083.1073135
Popović, M. (2017, September). chrF++: words helping character n-grams. In Proceedings of the second conference on machine translation (pp. 612-618). https://doi.org/10.18653/v1/W17-4770
Mozer, R., Miratrix, L., Kaufman, A. R., & Anastasopoulos, L. J. (2020). Matching with text data: An experimental evaluation of methods for matching documents and of measuring match quality. Political Analysis, 28(4), 445-468. https://doi.org/10.48550/arXiv.1801.00644
Alam, M. M. I., Kvapilíková, I., Anastasopoulos, A., Besacier, L., Dinu, G., Federico, M., ... & Nikoulina, V. (2021, November). Findings of the WMT shared task on machine translation using terminologies. In Proceedings of the Sixth Conference on Machine Translation (pp. 652-663). https://aclanthology.org/2021.wmt-1.69/
Guerreiro, N. M., Rei, R., Stigt, D. V., Coheur, L., Colombo, P., & Martins, A. F. (2024). xcomet: Transparent machine translation evaluation through fine-grained error detection. Transactions of the Association for Computational Linguistics, 12, 979-995. https://doi.org/10.1162/tacl_a_00683
Rei, R., Stewart, C., Farinha, A. C., & Lavie, A. (2020, November). COMET: A neural framework for MT evaluation. In Proceedings of the 2020 conference on empirical methods in natural language processing (emnlp) (pp. 2685-2702). https://doi.org/10.18653/v1/2020.emnlp-main.213
Pei, R., Liu, Y., Lin, P., Yvon, F., & Schütze, H. (2025, July). Understanding in-context machine translation for low-resource languages: A case study on Manchu. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 8767-8788). https://doi.org/10.18653/v1/2025.acl-long.429
Volk, M., Fischer, D. P., Fischer, L., Scheurer, P., & Ströbel, P. B. (2024, May). LLM-based machine translation and summarization for Latin. In Proceedings of the Third Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA)@ LREC-COLING-2024 (pp. 122-128). https://aclanthology.org/2024.lt4hala-1.15/
Zheng, F., Marrese-Taylor, E., & Matsuo, Y. (2024, August). Improving low-resource machine translation for formosan languages using bilingual lexical resources. In Findings of the Association for Computational Linguistics: ACL 2024 (pp. 11248-11259). https://doi.org/10.18653/v1/2024.findings-acl.670
Conia, S., Li, M., Navigli, R., & Potdar, S. (2025, July). SemEval-2025 task 2: Entity-aware machine translation. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025) (pp. 2535-2557). https://aclanthology.org/2025.semeval-1.326/
Waldendorf, J., Haddow, B., & Birch, A. (2024, March). Contrastive decoding reduces hallucinations in large multilingual machine translation models. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 2526-2539). https://doi.org/10.18653/v1/2024.eacl-long.155
Tang, Z., Chatterjee, R., & Garg, S. (2025, April). Mitigating hallucinated translations in large language models with hallucination-focused preference optimization. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) (pp. 3410-3433). https://doi.org/10.18653/v1/2025.naacl-long.175
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Articles are open access under the Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Authors who publish a manuscript in this journal agree to the following terms:
- The authors reserve the right to authorship of their work and transfer to the journal the right of first publication under the terms of the Creative Commons Attribution License, which allows others to freely distribute the published work with a mandatory link to the the original work and the first publication of the work in this journal.
- Authors have the right to conclude independent additional agreements that relate to the non-exclusive distribution of the work in the form in which it was published by this journal (for example, to post the work in the electronic repository of the institution or publish as part of a monograph), providing the link to the first publication of the work in this journal.
- Other terms stated in the Copyright Agreement.