REPRODUCIBLE EVALUATION OF ONTOLOGY-GUIDED HYBRID EXTRACTIVE QUESTION ANSWERING FOR KAZAKH HISTORY

Authors

DOI:

https://doi.org/10.37943/HRBD3124

Keywords:

domain-specific question answering, extractive question answering , ontology-guided filtering , hybrid neuro-symbolic systems , low-resource languages , Kazakh language , reproducible benchmarking , educational question answering

Abstract

Reproducibility and comparability remain important methodological concerns in domain-specific question answering, particularly in low-resource educational settings. This study examines the impact of ontology-guided context filtering on extractive question answering in the Kazakh language under controlled and reproducible conditions. A benchmark corpus on the History of Kazakhstan was constructed from 11 school textbooks for grades 5–10 and canonically indexed using deterministic section identifiers. Fixed train–test splits, standardized chunking, and unified preprocessing and evaluation scripts ensured consistency across experiments.

Three question answering systems were evaluated under identical constraints: a neural extractive baseline operating on unfiltered context, an ontology-only system relying on structured knowledge representations, and a hybrid system applying ontology-guided context filtering prior to neural answer extraction. Filtering was conducted under an inference setting that excludes access to gold document or section identifiers, preventing the use of privileged metadata.

The hybrid system demonstrates modest improvements over the neural-only baseline. Exact Match accuracy increases from 68.87% to 70.45%, while the token-level F1 score improves from 80.02% to 81.68%. Ranking-based retrieval metrics also show small gains, with Mean Reciprocal Rank increasing from 0.870 to 0.892 and Recall@10 from 95.1% to 96.9%. Inference latency is reduced from 20.01 ms to 6.83 ms per question. Paired statistical testing confirms that the improvement in F1 is statistically significant, whereas improvements in Exact Match are less conclusive.

Ablation experiments suggest that these differences are not attributable to context length reduction alone, but are associated with the interaction between ontology-guided context selection and neural reader training. Overall, ontology-guided filtering can provide incremental benefits for extractive question answering in a domain-specific, low-resource educational setting under transparent and reproducible experimental conditions.

Author Biography

Manas Yergesh, L.N. Gumilyov Eurasian National University

Манас Ергеш Жантуганулы — докторант компьютерных наук (код программы 8D06102) кафедры технологий искусственного интеллекта Евразийского национального университета имени Л.Н. Гумилева, Астана, Казахстан. Его исследования сосредоточены на искусственном интеллекте, обработке естественного языка, представлении знаний и интеллектуальных системах ответов на вопросы для казахского языка.

Ергеш Манас Жантуғанұлы — Астана қаласындағы Л.Н. Гумилев атындағы Евразия ұлттық университетінің Жасанды интеллектуальные технологии кафедрасында 8D06102 – Информатика білім беру бағдарламасы бойынша докторант. Ониң ғылыми зерттеу бағыты жасанды интеллект, табиғи тілді өңдеу, білімді ұсыну, сондай-ақ қазақ тілине арналған интеллектуалды сұрақ-жауап жүйелерімен байланысты.

References

Pineau, J., Vincent-Lamarre, P., Sinha, K., Larivière, V., Beygelzimer, A., d’Alché-Buc, F., Fox, E., & Larochelle, H. (2021). Improving reproducibility in machine learning research (a report from the NeurIPS 2019 reproducibility program). Journal of Machine Learning Research, 22(164), 1–20. http://jmlr.org/papers/v22/20-303.html

Gundersen, O. E. (2021). The fundamental principles of reproducibility. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 379(2197), 20200210. https://doi.org/10.1098/rsta.2020.0210

Polo, F. M., Izbicki, R., Lacerda, E. G., Ibieta-Jimenez, J. P., & Vicente, R. (2023). A unified framework for dataset shift diagnostics. Information Sciences, 649, Article 119612. https://doi.org/10.1016/j.ins.2023.119612

Chen, D., Fisch, A., Weston, J., & Bordes, A. (2017). Reading Wikipedia to answer open-domain questions. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (pp. 1870–1879). Association for Computational Linguistics. https://doi.org/10.18653/v1/P17-1171

Clark, C., & Gardner, M. (2018). Simple and effective multi-paragraph reading comprehension. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (pp. 845–855). Association for Computational Linguistics. https://doi.org/10.18653/v1/P18-1078

Lin, J., Nogueira, R., & Yates, A. (2021). Pretrained transformers for text ranking: BERT and beyond. Synthesis Lectures on Human Language Technologies, 14(4), 1–325. Morgan & Claypool. https://doi.org/10.2200/S01123ED1V01Y202108HLT053

Rajpurkar, P., Jia, R., & Liang, P. (2018). Know what you don’t know: Unanswerable questions for SQuAD. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (pp. 784–789). Association for Computational Linguistics. https://doi.org/10.18653/v1/P18-2124

Rogers, A., Kovaleva, O., & Rumshisky, A. (2020). A primer in BERTology: What we know about how BERT works. Transactions of the Association for Computational Linguistics, 8, 842–866. https://doi.org/10.1162/tacl_a_00349

Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423

Diefenbach, D., Lopez, V., Singh, K., & Maret, P. (2018). Core techniques of question answering systems over knowledge bases: A survey. Knowledge and Information Systems, 55(3), 529–569. https://doi.org/10.1007/s10115-017-1100-y

Ali, M., Jabeen, H., Hoyt, C. T., & Lehmann, J. (2019). The KEEN Universe: An ecosystem for knowledge graph embeddings with a focus on reproducibility and transferability. In The Semantic Web – ISWC 2019 (Lecture Notes in Computer Science, Vol. 11779, pp. 1–16). Springer. https://doi.org/10.1007/978-3-030-30796-7_1

Tleubayeva, A., & Shomanov, A. (2024). Comparative analysis of multilingual QA models and their adaptation to the Kazakh language. Scientific Journal of Astana IT University, 19, 89–97. https://doi.org/10.37943/19WHRK2878

Karpukhin, V., Oğuz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., & Yih, W. (2020). Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 6769–6781). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.emnlp-main.550

Mansurova, A., Tleubayeva, A., Nugumanova, A., Shomanov, A., & Seker, S. E. (2025). A systematic evaluation of large language models and retrieval-augmented generation for the task of Kazakh question answering. Information, 16(11), Article 943. https://doi.org/10.3390/info16110943

Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., & Zettlemoyer, L. (2020). BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 7871–7880). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.703

Hogan, A., Blomqvist, E., Cochez, M., d’Amato, C., de Melo, G., Gutiérrez, C., Kirrane, S., Labra Gayo, J. E., Navigli, R., Neumaier, S., Ngonga Ngomo, A.-C., Polleres, A., Rashid, S. M., Rula, A., Schmelzeisen, L., Sequeda, J., Staab, S., & Zimmermann, A. (2021). Knowledge graphs. ACM Computing Surveys, 54(4), Article 71. https://doi.org/10.1145/3447772

Song, Y., Zhang, H., Liu, Y., & Tang, J. (2023). Advancements in complex knowledge graph question answering: A survey. Electronics, 12(21), Article 4395. https://doi.org/10.3390/electronics12214395

Yu, D., Jiang, Y., Wang, Z., & Zhang, J. (2023). A survey on neural-symbolic learning systems. Neural Networks, 166, 364–382. https://doi.org/10.1016/j.neunet.2023.06.028

Bhuyan, B. P., Ramdane-Cherif, A., Tomar, R., & Singh, T. P. (2024). Neuro-symbolic artificial intelligence: A survey. Neural Computing and Applications, 36(21), 12809–12844. https://doi.org/10.1007/s00521-024-09960-z

Pires, T., Schlinger, E., & Garrette, D. (2019). How multilingual is multilingual BERT? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 4996–5001). Association for Computational Linguistics. https://doi.org/10.18653/v1/P19-1493

Downloads

Published

2026-06-30

How to Cite

Yergesh, M., Maxutova, K. ., Barlybayev, A. ., Tleubayeva, A. ., Orazayeva, A. ., & Abildina, A. . (2026). REPRODUCIBLE EVALUATION OF ONTOLOGY-GUIDED HYBRID EXTRACTIVE QUESTION ANSWERING FOR KAZAKH HISTORY. Scientific Journal of Astana IT University, 26(2), 151–166. https://doi.org/10.37943/HRBD3124

Issue

Section

Information Technologies