CRITERION-GUIDED AI ARCHITECTURE FOR LEARNING SUPPORT AND PRELIMINARY STUDENT ASSESSMENT

Authors

DOI:

https://doi.org/10.37943/LSGM5046

Keywords:

intelligent education systems; retrieval-augmented generation; rubric-based assessment; explainable artificial intelligence; human-in-the-loop; uncertainty estimation; higher education; decision support.

Abstract

This paper presents a pilot evaluation of a criterion-guided human-in-the-loop architecture for an intelligent educational system that supports learning and preliminary assessment of students’ written and oral responses. Instead of assigning a final score directly, the system combines relevant context, formalized criteria, interpretable response features, supporting evidence, uncertainty estimates, and instructor verification. The architecture separates probabilistic components, including retrieval, automatic speech recognition, natural language processing, and a language model, from deterministic assessment and governance mechanisms such as rubric weights, aggregation rules, thresholds, audit logging, and referral rules. Two modes were implemented: a retrieval-augmented generation assistant for formative learning support and an AI examiner for preliminary criterion-based assessment. In the assistant pilot, 336 of 347 queries were processed without recorded technical errors, giving a completion rate of 96.8%; source citations were included in 89.0% of responses, and median response time was 980 ms. Student perceptions of the examiner were evaluated using 52 questionnaires on a five-point Likert scale. Mean ratings were highest for question clarity at 4.81, usefulness of criterion-level feedback at 4.46, and sufficiency of response time at 4.42. Students also reported notable uncertainty about the grade, with a mean rating of 3.79, and concern about speech-recognition errors, with a mean rating of 3.31. Among 45 oral responses, 12 (26.7%) lacked usable word-level timestamps and were referred for review. A worked case demonstrated traceability from response features and evidence to criterion scores, uncertainty handling, aggregation, and instructor verification. A preliminary comparison of 23 paired AI and instructor-released scores was treated as diagnostic because of a severe ceiling effect. The results support technical feasibility and procedural acceptability, but not yet the validity or reliability of automated assessment.

Author Biographies

Syrym Zhakypbekov, International Information Technology University, Kazakhstan

Master’s Degree, Assistant Professor, Department of Computer Engineering

Artem Bykov, International Information Technology University, Kazakhstan

PhD, Associate Professor, Department of Information Systems

Vera Yermakova, International Information Technology University

Candidate of Philological Sciences, Associate Professor, Department of Languages

Regina Sharshova, International Information Technology University, Kazakhstan

Master of Humanities, Assistant Professor, Department of Languages

References

Zawacki-Richter, O., Marín, V. I., Bond, M., & Gouverneur, F. (2019). Systematic review of research on artificial intelligence applications in higher education—Where are the educators? International Journal of Educational Technology in Higher Education, 16, 39. https://doi.org/10.1186/s41239-019-0171-0

González-Calatayud, V., Prendes-Espinosa, P., & Roig-Vila, R. (2021). Artificial intelligence for student assessment: A systematic review. Applied Sciences, 11(12), 5467. https://doi.org/10.3390/app11125467

Manganello, F., Nico, A., & Boccuzzi, G. (2025). Theoretical foundations for governing AI-based learning outcome assessment in high-risk educational contexts. Information, 16(9), 814. https://doi.org/10.3390/info16090814

Fajardo-Ramos, D. C., Chiappe, A., & Mella-Norambuena, J. (2025). Human-in-the-loop assessment with AI: Implications for teacher education in Ibero-American universities. Frontiers in Education, 10, 1710992. https://doi.org/10.3389/feduc.2025.1710992

Khan, W., Topham, L., Jones, N., Atherton, P., Al-Shabandar, R., Kolivand, H., Khan, I., Alatrany, A., & Hussain, A. (2026). Auto-assessment of assessment: A human-in-the-loop AI framework addressing policy gaps in academic assessment. PLOS ONE, 21(4), e0346815. https://doi.org/10.1371/journal.pone.0346815

Mousavinasab, E., Zarifsanaiey, N., Niakan Kalhori, S. R., Rakhshan, M., Keikha, L., & Ghazi Saeedi, M. (2021). Intelligent tutoring systems: A systematic review of characteristics, applications, and evaluation methods. Interactive Learning Environments, 29(1), 142–163. https://doi.org/10.1080/10494820.2018.1558257

Karpukhin, V., Oğuz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., & Yih, W.-T. (2020). Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 6769–6781). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.emnlp-main.550

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-T., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474.

Stolfo, A. (2024). Groundedness in retrieval-augmented long-form generation: An empirical study. In Findings of the Association for Computational Linguistics: NAACL 2024 (pp. 1537–1552). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-naacl.100

Emirtekin, E. (2025). Large language model-powered automated assessment: A systematic review. Applied Sciences, 15(10), 5683. https://doi.org/10.3390/app15105683

Xie, W., Niu, J., Xue, C. J., & Guan, N. (2024). Grade like a human: Rethinking automated assessment with large language models. arXiv. https://doi.org/10.48550/arXiv.2405.19694

Cho, B., Jang, Y., & Yoon, J. (2023). Rubric-specific approach to automated essay scoring with augmentation training. arXiv. https://doi.org/10.48550/arXiv.2309.02740

Ercikan, K., & McCaffrey, D. F. (2022). Optimizing implementation of artificial-intelligence-based automated scoring: An evidence centered design approach for designing assessments for AI-based scoring. Journal of Educational Measurement, 59(3), 272–287. https://doi.org/10.1111/jedm.12332

Abdar, M., Pourpanah, F., Hussain, S., Rezazadegan, D., Liu, L., Ghavamzadeh, M., Fieguth, P., Cao, X., Khosravi, A., Acharya, U. R., Makarenkov, V., & Nahavandi, S. (2021). A review of uncertainty quantification in deep learning: Techniques, applications and challenges. Information Fusion, 76, 243–297. https://doi.org/10.1016/j.inffus.2021.05.008

Geifman, Y., & El-Yaniv, R. (2017). Selective classification for deep neural networks. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS '17) (pp. 4885–4894). Curran Associates Inc. https://dl.acm.org/doi/10.5555/3295222.3295241

Rožanec, J. M., Novalija, I., Zajec, P., Kenda, K., Tavakoli Ghinani, H., Suh, S., Bian, S., Veliou, E., Papamartzivanos, D., Giannetsos, T., Menesidou, S. A., Alonso, R., Cauli, N., Meloni, A., Reforgiato Recupero, D., Kyriazis, D., Sofianidis, G., Theodoropoulos, S., Fortuna, B., Mladenić, D., & Soldatos, J. (2023). Human-centric artificial intelligence architecture for Industry 5.0 applications. International Journal of Production Research, 61(20), 6847–6872. https://doi.org/10.1080/00207543.2022.2138611

Komarov, N., Mukhanov, S. B., Bazarbekov, I. M., Zhakypbekov, S. Zh., & Sibanbayeva, S. Y. (2025). Methods of processing and analyzing big data in machine learning tasks: Approaches and prospects. Herald of the Kazakh-British Technical University, 1, 25–35. https://doi.org/10.55452/1998-6688-2025-22-1-25-35

Kenzhebulatova, D., Grigoryeva, I., Yermakova, V., Bublikova, O., Tumanova, A., & Mukhamadiyev, K. (2025). The professional linguistic worldview of IT specialists as a component of linguistic culture. Forum for Linguistic Studies, 7(1), 98–111. https://doi.org/10.30564/fls.v7i1.7742

Mambetov, S., Joldasbayev, S., Bykov, A., Koishybay, S., & Dossanbek, K. (2026). Methods for automatic emotion recognition in hacker forum texts. In K. Kolesnikova & M. Ipalakova (Eds.), Proceedings of the Workshop on Cybersecurity, Infocommunication Systems and Networks 2025 (Vol. 4180, Paper 13). CEUR Workshop Proceedings. https://ceur-ws.org/Vol-4180/Paper13.pdf

Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In K. Inui, J. Jiang, V. Ng, & X. Wan (Eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 3982–3992). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1410

Cui, B., Li, Y., Zhang, Y., & Zhang, Z. (2017). Text coherence analysis based on deep neural network. In Proceedings of the 2017 ACM on Conference on Information and Knowledge Management (pp. 2027–2030). Association for Computing Machinery. https://doi.org/10.1145/3132847.3133047

Dawson, P. (2017). Assessment rubrics: Towards clearer and more replicable design, research and practice. Assessment & Evaluation in Higher Education, 42(3), 347–360. https://doi.org/10.1080/02602938.2015.1111294

Jebb, A. T., Ng, V., & Tay, L. (2021). A review of key Likert scale development advances: 1995–2019. Frontiers in Psychology, 12, 637547. https://doi.org/10.3389/fpsyg.2021.637547

Pecuchova, J., Benko, Ľ., & Drlik, M. (2025). Automated grading of open-ended questions in higher education using GenAI models. International Journal of Artificial Intelligence in Education, 35, 3813–3846. https://doi.org/10.1007/s40593-025-00517-2

Flodén, J. (2025). Grading exams using large language models: A comparison between human and AI grading of exams in higher education using ChatGPT. British Educational Research Journal, 51(1), 201–224. https://doi.org/10.1002/berj.4069

Hickman, L., Langer, M., Saef, R. M., & Tay, L. (2025). Automated speech recognition bias in personnel selection: The case of automatically scored job interviews. Journal of Applied Psychology, 110(6), 846–858. https://doi.org/10.1037/apl0001247

Downloads

Published

2026-09-30

How to Cite

Zhakypbekov, S., Bykov, A., Yermakova, V., & Sharshova, R. (2026). CRITERION-GUIDED AI ARCHITECTURE FOR LEARNING SUPPORT AND PRELIMINARY STUDENT ASSESSMENT. Scientific Journal of Astana IT University, 27(3), 208–227. https://doi.org/10.37943/LSGM5046

Issue

Section

Information Technologies