REPRODUCIBLE EVALUATION OF ONTOLOGY-GUIDED HYBRID EXTRACTIVE QUESTION ANSWERING FOR KAZAKH HISTORY
DOI:
https://doi.org/10.37943/HRBD3124Keywords:
domain-specific question answering, extractive question answering , ontology-guided filtering , hybrid neuro-symbolic systems , low-resource languages , Kazakh language , reproducible benchmarking , educational question answeringAbstract
Reproducibility and comparability remain important methodological concerns in domain-specific question answering, particularly in low-resource educational settings. This study examines the impact of ontology-guided context filtering on extractive question answering in the Kazakh language under controlled and reproducible conditions. A benchmark corpus on the History of Kazakhstan was constructed from 11 school textbooks for grades 5–10 and canonically indexed using deterministic section identifiers. Fixed train–test splits, standardized chunking, and unified preprocessing and evaluation scripts ensured consistency across experiments.
Three question answering systems were evaluated under identical constraints: a neural extractive baseline operating on unfiltered context, an ontology-only system relying on structured knowledge representations, and a hybrid system applying ontology-guided context filtering prior to neural answer extraction. Filtering was conducted under an inference setting that excludes access to gold document or section identifiers, preventing the use of privileged metadata.
The hybrid system demonstrates modest improvements over the neural-only baseline. Exact Match accuracy increases from 68.87% to 70.45%, while the token-level F1 score improves from 80.02% to 81.68%. Ranking-based retrieval metrics also show small gains, with Mean Reciprocal Rank increasing from 0.870 to 0.892 and Recall@10 from 95.1% to 96.9%. Inference latency is reduced from 20.01 ms to 6.83 ms per question. Paired statistical testing confirms that the improvement in F1 is statistically significant, whereas improvements in Exact Match are less conclusive.
Ablation experiments suggest that these differences are not attributable to context length reduction alone, but are associated with the interaction between ontology-guided context selection and neural reader training. Overall, ontology-guided filtering can provide incremental benefits for extractive question answering in a domain-specific, low-resource educational setting under transparent and reproducible experimental conditions.
References
Pineau, J., Vincent-Lamarre, P., Sinha, K., Larivière, V., Beygelzimer, A., d’Alché-Buc, F., Fox, E., & Larochelle, H. (2021). Improving reproducibility in machine learning research (a report from the NeurIPS 2019 reproducibility program). Journal of Machine Learning Research, 22(164), 1–20. http://jmlr.org/papers/v22/20-303.html
Gundersen, O. E. (2021). The fundamental principles of reproducibility. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 379(2197), 20200210. https://doi.org/10.1098/rsta.2020.0210
Polo, F. M., Izbicki, R., Lacerda, E. G., Ibieta-Jimenez, J. P., & Vicente, R. (2023). A unified framework for dataset shift diagnostics. Information Sciences, 649, Article 119612. https://doi.org/10.1016/j.ins.2023.119612
Chen, D., Fisch, A., Weston, J., & Bordes, A. (2017). Reading Wikipedia to answer open-domain questions. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (pp. 1870–1879). Association for Computational Linguistics. https://doi.org/10.18653/v1/P17-1171
Clark, C., & Gardner, M. (2018). Simple and effective multi-paragraph reading comprehension. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (pp. 845–855). Association for Computational Linguistics. https://doi.org/10.18653/v1/P18-1078
Lin, J., Nogueira, R., & Yates, A. (2021). Pretrained transformers for text ranking: BERT and beyond. Synthesis Lectures on Human Language Technologies, 14(4), 1–325. Morgan & Claypool. https://doi.org/10.2200/S01123ED1V01Y202108HLT053
Rajpurkar, P., Jia, R., & Liang, P. (2018). Know what you don’t know: Unanswerable questions for SQuAD. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (pp. 784–789). Association for Computational Linguistics. https://doi.org/10.18653/v1/P18-2124
Rogers, A., Kovaleva, O., & Rumshisky, A. (2020). A primer in BERTology: What we know about how BERT works. Transactions of the Association for Computational Linguistics, 8, 842–866. https://doi.org/10.1162/tacl_a_00349
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423
Diefenbach, D., Lopez, V., Singh, K., & Maret, P. (2018). Core techniques of question answering systems over knowledge bases: A survey. Knowledge and Information Systems, 55(3), 529–569. https://doi.org/10.1007/s10115-017-1100-y
Ali, M., Jabeen, H., Hoyt, C. T., & Lehmann, J. (2019). The KEEN Universe: An ecosystem for knowledge graph embeddings with a focus on reproducibility and transferability. In The Semantic Web – ISWC 2019 (Lecture Notes in Computer Science, Vol. 11779, pp. 1–16). Springer. https://doi.org/10.1007/978-3-030-30796-7_1
Tleubayeva, A., & Shomanov, A. (2024). Comparative analysis of multilingual QA models and their adaptation to the Kazakh language. Scientific Journal of Astana IT University, 19, 89–97. https://doi.org/10.37943/19WHRK2878
Karpukhin, V., Oğuz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., & Yih, W. (2020). Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 6769–6781). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.emnlp-main.550
Mansurova, A., Tleubayeva, A., Nugumanova, A., Shomanov, A., & Seker, S. E. (2025). A systematic evaluation of large language models and retrieval-augmented generation for the task of Kazakh question answering. Information, 16(11), Article 943. https://doi.org/10.3390/info16110943
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., & Zettlemoyer, L. (2020). BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 7871–7880). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.703
Hogan, A., Blomqvist, E., Cochez, M., d’Amato, C., de Melo, G., Gutiérrez, C., Kirrane, S., Labra Gayo, J. E., Navigli, R., Neumaier, S., Ngonga Ngomo, A.-C., Polleres, A., Rashid, S. M., Rula, A., Schmelzeisen, L., Sequeda, J., Staab, S., & Zimmermann, A. (2021). Knowledge graphs. ACM Computing Surveys, 54(4), Article 71. https://doi.org/10.1145/3447772
Song, Y., Zhang, H., Liu, Y., & Tang, J. (2023). Advancements in complex knowledge graph question answering: A survey. Electronics, 12(21), Article 4395. https://doi.org/10.3390/electronics12214395
Yu, D., Jiang, Y., Wang, Z., & Zhang, J. (2023). A survey on neural-symbolic learning systems. Neural Networks, 166, 364–382. https://doi.org/10.1016/j.neunet.2023.06.028
Bhuyan, B. P., Ramdane-Cherif, A., Tomar, R., & Singh, T. P. (2024). Neuro-symbolic artificial intelligence: A survey. Neural Computing and Applications, 36(21), 12809–12844. https://doi.org/10.1007/s00521-024-09960-z
Pires, T., Schlinger, E., & Garrette, D. (2019). How multilingual is multilingual BERT? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 4996–5001). Association for Computational Linguistics. https://doi.org/10.18653/v1/P19-1493
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Articles are open access under the Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Authors who publish a manuscript in this journal agree to the following terms:
- The authors reserve the right to authorship of their work and transfer to the journal the right of first publication under the terms of the Creative Commons Attribution License, which allows others to freely distribute the published work with a mandatory link to the the original work and the first publication of the work in this journal.
- Authors have the right to conclude independent additional agreements that relate to the non-exclusive distribution of the work in the form in which it was published by this journal (for example, to post the work in the electronic repository of the institution or publish as part of a monograph), providing the link to the first publication of the work in this journal.
- Other terms stated in the Copyright Agreement.