COMPARISON OF DIFFERENT OBJECT DETECTION TECHNIQUES FOR CENTRAL ASIAN FOOD RECOGNITION
DOI:
https://doi.org/10.37943/BZIF1573Keywords:
object detection , food recognition , Central Asian cuisine , dietary monitoring , Faster R-CNN , RetinaNet , YOLO , Detection Transformer , benchmark evaluationAbstract
Automated food detection is a major component in building reliable dietary monitoring systems, particularly in regions where cuisine-specific datasets are limited. Central Asian food recognition presents unique challenges due to strong intra-class variation, visually similar dishes, multi-item meal compositions, and long-tailed class distributions. To address this problem, we present a comprehensive benchmarking analysis of five modern object detection models across three groups: two-stage (Faster R-CNN), one-stage anchor-based (RetinaNet), one-stage anchor-free (YOLOv12, YOLO26), and transformer-based (RT-DETR). Experiments are conducted on a unified dataset constructed by merging the Global Gastronomic Culinary Dataset (GGCD) and the Food Portion Benchmark (FPB), resulting in 48,228 images across 258 food classes. The dataset includes controlled single-portion photographs and multi-item meals captured across various backgrounds and from different viewing angles. The merged dataset therefore covers both isolated food portions and more complex scenes containing several dishes. We compare the models in terms of detection accuracy, computational requirements, and inference speed, and examine AP by food class across all models. Overall, YOLOv12 achieves the highest accuracy, with an mAP50 of 75.80 and an mAP50-95 of 67.90, while using fewer parameters and less computation than the heavier baselines. RT-DETR reaches an mAP50 of 72.90 but requires more computation, whereas YOLO26 records the shortest single-image inference time on both GPU and CPU. Class-wise analysis further shows high performance on visually distinctive categories, with near-perfect AP for ten top-detected classes. These findings provide a practical baseline for Central Asian food detection and show that future progress depends on improving robustness for long-tail and visually ambiguous classes.
References
Liu, D., Zuo, E., Wang, D., He, L., Dong, L., & Lu, X. (2025). Deep Learning in Food Image Recognition: A Comprehensive review. Applied Sciences, 15(14), 7626. https://doi.org/10.3390/app15147626
Kaushal, S., Tammineni, D. K., Rana, P., Sharma, M., Sridhar, K., & Chen, H. (2024). Computer vision and deep learning-based approaches for detection of food nutrients/nutrition: New insights and advances. Trends in Food Science & Technology, 146, 104408. https://doi.org/10.1016/j.tifs.2024.104408
Min, W., Jiang, S., Liu, L., Rui, Y., & Jain, R. (2019). A survey on food computing. ACM Computing Surveys, 52(5), 1–36. https://doi.org/10.1145/3329168
Zhang, Y., Deng, L., Zhu, H., Wang, W., Ren, Z., Zhou, Q., Lu, S., Sun, S., Zhu, Z., Gorriz, J. M., & Wang, S. (2023). Deep learning in food category recognition. Information Fusion, 98, 101859. https://doi.org/10.1016/j.inffus.2023.101859
Bossard, L., Guillaumin, M., & Van Gool, L. (2014). Food-101 – Mining Discriminative Components with Random Forests. In Lecture Notes in Computer Science (pp. 446–461). https://doi.org/10.1007/978-3-319-10599-4_29
Min, W., Wang, Z., Liu, Y., Luo, M., Kang, L., Wei, X., Wei, X., & Jiang, S. (2023). Large scale visual food recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(8), 9932–9949. https://doi.org/10.1109/tpami.2023.3237871
Salvador, A., Hynes, N., Aytar, Y., Marin, J., Ofli, F., Weber, I., & Torralba, A. (2017). Learning Cross-Modal Embeddings for Cooking Recipes and Food Images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3020-3028. https://doi.org/10.1109/CVPR.2017.323
Chen, X., Zhu, Y., Zhou, H., Diao, L., & Wang, D. (2017). ChineseFoodNet: A large-scale Image Dataset for Chinese Food Recognition. arXiv (Cornell University). https://doi.org/10.48550/arxiv.1705.02743
Bolanos, M., & Radeva, P. (2016, December). Simultaneous food localization and recognition. In 2016 23rd International conference on pattern recognition (ICPR) (pp. 3140-3145). IEEE.
Jia, W., Li, Y., Qu, R., Baranowski, T., Burke, L. E., Zhang, H., Bai, Y., Mancino, J. M., Xu, G., Mao, Z., & Sun, M. (2018). Automatic food detection in egocentric images using artificial intelligence technology. Public Health Nutrition, 22(7), 1–12. https://doi.org/10.1017/s1368980018000538
Tan, S. W., Lee, C. P., Lim, K. M., & Lim, J. Y. (2023, August). Food detection and recognition with deep learning: A comparative study. In Proceedings of the International Conference on Information and Communication Technology (ICoICT) (pp. 283-288). IEEE.
Karabay, A., Bolatov, A., Varol, H.A., & Chan, M. (2023). A central Asian Food dataset for personalized dietary interventions. Nutrients, 15(7), 1728. https://doi.org/10.3390/nu15071728
Karabay, A., Varol, H. A., & Chan, M. Y. (2025). Improved food image recognition by leveraging deep learning and data-driven methods with an application to Central Asian Food Scene. Scientific Reports, 15(1), 14043. https://doi.org/10.1038/s41598-025-95770-9
Sanatbyek, A., Karabay, A., Varol, H. A., & Chan, M. Y. (2025, September). Deep Object Recognition-Based Analysis of Diverse Culinary Landscapes. In 2025 IEEE International Conference on Image Processing (ICIP) (pp. 1127-1132). IEEE.
Thames, Q., Karpur, A., Norris, W., Xia, F., Panait, L., Weyand, T., & Sim, J. (2021). Nutrition5K: towards automatic nutritional understanding of generic food. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2103.03375
Sanatbyek, A., Rakhimzhanova, T., Nurmanova, B., Omarova, Z., Rakhmankulova, A., Orazbayev, R., Varol, H. A., & Chan, M. Y. (2025). A multitask deep learning model for food scene recognition and portion estimation—the Food Portion Benchmark (FPB) dataset. IEEE Access, 13, 152033–152045. https://doi.org/10.1109/access.2025.3603287
Global Nutrition Report. (2025). Kazakhstan: Country nutrition profile. https://globalnutritionreport.org/resources/nutrition-profiles/asia/central-asia/kazakhstan/
Omarova, Z., Nurmanova, B., Sanatbyek, A., Varol, H. A., & Chan, M.-Y. (2025). Digital Mapping of Central Asian Foods: Towards a Standardized Visual Atlas for Nutritional Research. Nutrients, 17(21), 3315. https://doi.org/10.3390/nu17213315
Ren, S., He, K., Girshick, R., & Sun, J. (2016). Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(6), 1137–1149. https://doi.org/10.1109/tpami.2016.2577031
Lin, T. Y., Dollár, P., Girshick, R., He, K., Hariharan, B., & Belongie, S. (2017). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, 936-944. https://doi.org/10.1109/CVPR.2017.106
Lin, T., Goyal, P., Girshick, R., He, K., & Dollar, P. (2018). Focal loss for dense object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(2), 318-327. https://doi.org/10.1109/tpami.2018.2858826
Tian, Y., Ye, Q., & Doermann, D.S. (2025). YOLOv12: Attention-Centric Real-Time Object Detectors. ArXiv, abs/2502.12524.
Jocher, G., & Qiu, J. (2026). Ultralytics YOLO26 (Version 26.0.0) [Computer software]. https://github.com/ultralytics/ultralytics
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., & Zagoruyko, S. (2020). End-to-End Object Detection with Transformers. Lecture Notes in Computer Science, 213-229. https://doi.org/10.1007/978-3-030-58452-8_13
Zhu, X., Su, W., Lu, L., Li, B., Wang, X., & Dai, J. (2021). Deformable DETR: Deformable Transformers for End-to-End Object Detection. In Proceedings of the International Conference on Learning Representations (ICLR). OpenReview.
Zhao, Y., Lv, W., Xu, S., Wei, J., Wang, G., Dang, Q., … Chen, J. (2024). DETRs beat YOLOs on real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 16965-16974). https://doi.org/10.1109/CVPR52733.2024.01605
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., … Zitnick, C. L. (2014). Microsoft COCO: Common objects in context. In Computer Vision – ECCV 2014 (pp. 740-755). Springer. https://doi.org/10.1007/978-3-319-10602-1_48
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Articles are open access under the Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Authors who publish a manuscript in this journal agree to the following terms:
- The authors reserve the right to authorship of their work and transfer to the journal the right of first publication under the terms of the Creative Commons Attribution License, which allows others to freely distribute the published work with a mandatory link to the the original work and the first publication of the work in this journal.
- Authors have the right to conclude independent additional agreements that relate to the non-exclusive distribution of the work in the form in which it was published by this journal (for example, to post the work in the electronic repository of the institution or publish as part of a monograph), providing the link to the first publication of the work in this journal.
- Other terms stated in the Copyright Agreement.