COMPARISON OF YOLOV8N AND YOLOV12N MODELS FOR SCALABLE REAL TIME VIDEO MONITORING OF CONSTRUCTION SITES

Authors

DOI:

https://doi.org/10.37943/TSYT4497

Keywords:

computer vision , object detection , YOLO models , personal protective equipment , industrial safety monitoring , real-time inference , detection accuracy , neural network performance

Abstract

This study presents a comparative analysis of the object detection models YOLOv8n and YOLOv12n in their Nano configurations for real-time construction site monitoring tasks. Key model characteristics were evaluated, including detection accuracy, recall, inference speed, and computational resource consumption across various server configurations. Particular attention was given to the use of the Slicing Aided Hyper Inference (SAHI) method to improve the detection quality of small objects, as well as the impact of scene complexity on system performance.

YOLOv12n achieved significantly higher accuracy and recall than YOLOv8n. In addition, the integration of the SAHI method further enhanced small object detection performance. Parallel stream processing showed that servers equipped with Tesla L4 GPUs offered the best trade-off between inference speed and stability under multi-threaded workloads. The inclusion of data augmentation techniques simulating various weather conditions enhanced model robustness, especially for YOLOv8n, in noisy and complex scenes.

The findings underscore the clear advantage of YOLOv12n for applications requiring high recall in challenging environments, while YOLOv8n demonstrated strong efficiency on resource-constrained systems. The proposed approaches can be utilized in the development of intelligent monitoring systems for construction sites, industrial facilities, and logistics hubs, including automated access control and real-time employee activity tracking.

References

Wang, Z., Wu, Y., Yang, L., Thirunavukarasu, A., Evison, C., & Zhao, Y. (2021). Fast personal protective equipment detection for real construction sites using deep learning approaches. Sensors, 21(10), 3478. https://doi.org/10.3390/s21103478

Terven, J., Córdova-Esparza, D.-M., & Romero-González, J.-A. (2023). A comprehensive review of YOLO architectures in computer vision: From YOLOv1 to YOLOv8 and YOLO-NAS. Machine Learning and Knowledge Extraction, 5(4), 1680–1716. https://doi.org/10.3390/make5040083

Tian, Y., Ye, Q., & Doermann, D. (2025). YOLOv12: Attention-centric real-time object detectors. In Advances in Neural Information Processing Systems (Vol. 38, pp. 78433–78457). https://proceedings.neurips.cc/paper_files/paper/2025/hash/7103444259031cc58051f8c9a4868533-Abstract-Conference.html

Sapkota, R., Flores-Calero, M., Qureshi, R., Badgujar, C., Nepal, U., Poulose, A., Zeno, P., Vaddevolu, U. B. P., Khan, S., Shoman, M., Yan, H., & Karkee, M. (2025). YOLO advances to its genesis: a decadal and comprehensive review of the You Only Look Once (YOLO) series. Artificial Intelligence Review, 58(9). https://doi.org/10.1007/s10462-025-11253-3

Wang, C.-Y., Liao, H.-Y., Wu, Y.-H., Chen, P.-Y., Hsieh, J.-W., & Yeh, I.-H. (2020). CSPNet: A new backbone that can enhance learning capability of CNN. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) (pp. 1571–1580). https://doi.org/10.1109/CVPRW50498.2020.00203

Lin, T. Y., Dollár, P., Girshick, R., He, K., Hariharan, B., & Belongie, S. (2017). Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 2117-2125). https://doi.org/10.1109/CVPR.2017.106

Woo, S., Park, J., Lee, J.-Y., & Kweon, I. S. (2018). CBAM: Convolutional block attention module. In V. Ferrari, M. Hebert, C. Sminchisescu, & Y. Weiss (Eds.), Computer vision – ECCV 2018 (Lecture Notes in Computer Science, Vol. 11211). Springer. https://doi.org/10.1007/978-3-030-01234-2_1

Ultralytics. (2025). YOLO12: Attention-centric object detection. https://docs.ultralytics.com/models/yolo12/ (Accessed June 28, 2025)

Ultralytics. (2025). YOLOv10 vs YOLOv8: A technical comparison for object detection. https://docs.ultralytics.com/ru/compare/yolov10-vs-yolov8/3 (Accessed June 28, 2025)

Xu, L., Zhao, Y., Zhai, Y., et al. (2024). Small object detection in UAV images based on YOLOv8n. International Journal of Computational Intelligence Systems, 17, 223. https://doi.org/10.1007/s44196-024-00632-3

Yao, G., Zhu, S., Zhang, L., & Qi, M. (2024). HP-YOLOv8: High-precision small object detection algorithm for remote sensing images. Sensors, 24(15), 4858. https://doi.org/10.3390/s24154858

Wang, Q., Zhou, Z., & Zhang, Z. (2025). RSO-YOLO: A Real-Time Detector for Small and Occluded Objects in Autonomous Driving Scenarios. Sensors, 25(21), 6703. https://doi.org/10.3390/s25216703

Akyon, F. C., Altinuc, S. O., & Temizel, A. (2022, October). Slicing aided hyper inference and fine-tuning for small object detection. In 2022 IEEE international conference on image processing (ICIP) (pp. 966-970). IEEE. https://doi.org/10.1109/ICIP46576.2022.9897990

Lakens, D. (2013). Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs. Frontiers in psychology, 4, 863. https://doi.org/10.3389/fpsyg.2013.00863

Hu, J., Shen, L., & Sun, G. (2018). Squeeze-and-excitation networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 7132–7141). https://doi.org/10.1109/CVPR.2018.00745

Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 779–788). https://doi.org/10.1109/CVPR.2016.91

NVIDIA Corporation. (2022). NVIDIA Ada Lovelace architecture (Technical white paper). https://catalogone.com/wp-content/uploads/2024/06/NVIDIA-ADA-GPU-PROVIZ-Architecture-Whitepaper_1.1.pdf

Lee, H., Lee, J. S., & Choi, H. C. (2021). Parallelization of Non-Maximum Suppression. IEEE Access, 9, 166579-166587. https://doi.org/10.1109/ACCESS.2021.3134639

Huang, J., Rathod, V., Sun, C., Zhu, M., Korattikara, A., Fathi, A., ... & Murphy, K. (2017). Speed/accuracy trade-offs for modern convolutional object detectors. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 7310-7311). https://doi.org/10.1109/CVPR.2017.351

Bodla, N., Singh, B., Chellappa, R., & Davis, L. (2017). Soft-NMS—Improving object detection with one line of code. In Proceedings of the IEEE International Conference on Computer Vision (ICCV) (pp. 5562–5570). https://doi.org/10.1109/ICCV.2017.593

Tremblay, J., et al. (2018). Training deep networks with synthetic data: Bridging the reality gap by domain randomization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). https://doi.org/10.1109/CVPRW.2018.00143

Sapkota, R., Meng, Z., Churuvija, M., Du, X., Ma, Z., & Karkee, M. (2026). Comprehensive performance evaluation of yolov12, yolo11, yolov10, yolov9 and yolov8 on detecting and counting fruitlet in complex orchard environments. Agriculture Communications, 100125. https://doi.org/10.1016/j.agrcom.2026.100125

Gupta, H., Kotlyar, O., Andreasson, H., & Lilienthal, A. J. (2024). Robust object detection in challenging weather conditions. In Proceedings of the IEEE/CVF winter conference on applications of computer vision (pp. 7523-7532). https://doi.org/10.1109/WACV57701.2024.00735

Wojke, N., Bewley, A., & Paulus, D. (2017). Simple online and realtime tracking with a deep association metric. In Proceedings of the IEEE International Conference on Image Processing (ICIP) (pp. 3645–3649). https://doi.org/10.1109/ICIP.2017.8296962

Liu, S., Huang, D., & Wang, Y. (2019). Adaptive nms: Refining pedestrian detection in a crowd. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 6459-6468). https://doi.org/10.1109/CVPR.2019.00662

Downloads

Published

2026-06-30

How to Cite

Bektemyssova, G., Keresh, A., Nuralykyzy, S., Ziyada, M., Barlykbay, N. ., & Joldasbayev, S. (2026). COMPARISON OF YOLOV8N AND YOLOV12N MODELS FOR SCALABLE REAL TIME VIDEO MONITORING OF CONSTRUCTION SITES. Scientific Journal of Astana IT University, 26(2), 25–51. https://doi.org/10.37943/TSYT4497

Issue

Section

Information Technologies