COMPARISON OF YOLOV8N AND YOLOV12N MODELS FOR SCALABLE REAL TIME VIDEO MONITORING OF CONSTRUCTION SITES
DOI:
https://doi.org/10.37943/TSYT4497Keywords:
computer vision , object detection , YOLO models , personal protective equipment , industrial safety monitoring , real-time inference , detection accuracy , neural network performanceAbstract
This study presents a comparative analysis of the object detection models YOLOv8n and YOLOv12n in their Nano configurations for real-time construction site monitoring tasks. Key model characteristics were evaluated, including detection accuracy, recall, inference speed, and computational resource consumption across various server configurations. Particular attention was given to the use of the Slicing Aided Hyper Inference (SAHI) method to improve the detection quality of small objects, as well as the impact of scene complexity on system performance.
YOLOv12n achieved significantly higher accuracy and recall than YOLOv8n. In addition, the integration of the SAHI method further enhanced small object detection performance. Parallel stream processing showed that servers equipped with Tesla L4 GPUs offered the best trade-off between inference speed and stability under multi-threaded workloads. The inclusion of data augmentation techniques simulating various weather conditions enhanced model robustness, especially for YOLOv8n, in noisy and complex scenes.
The findings underscore the clear advantage of YOLOv12n for applications requiring high recall in challenging environments, while YOLOv8n demonstrated strong efficiency on resource-constrained systems. The proposed approaches can be utilized in the development of intelligent monitoring systems for construction sites, industrial facilities, and logistics hubs, including automated access control and real-time employee activity tracking.
References
Wang, Z., Wu, Y., Yang, L., Thirunavukarasu, A., Evison, C., & Zhao, Y. (2021). Fast personal protective equipment detection for real construction sites using deep learning approaches. Sensors, 21(10), 3478. https://doi.org/10.3390/s21103478
Terven, J., Córdova-Esparza, D.-M., & Romero-González, J.-A. (2023). A comprehensive review of YOLO architectures in computer vision: From YOLOv1 to YOLOv8 and YOLO-NAS. Machine Learning and Knowledge Extraction, 5(4), 1680–1716. https://doi.org/10.3390/make5040083
Tian, Y., Ye, Q., & Doermann, D. (2025). YOLOv12: Attention-centric real-time object detectors. In Advances in Neural Information Processing Systems (Vol. 38, pp. 78433–78457). https://proceedings.neurips.cc/paper_files/paper/2025/hash/7103444259031cc58051f8c9a4868533-Abstract-Conference.html
Sapkota, R., Flores-Calero, M., Qureshi, R., Badgujar, C., Nepal, U., Poulose, A., Zeno, P., Vaddevolu, U. B. P., Khan, S., Shoman, M., Yan, H., & Karkee, M. (2025). YOLO advances to its genesis: a decadal and comprehensive review of the You Only Look Once (YOLO) series. Artificial Intelligence Review, 58(9). https://doi.org/10.1007/s10462-025-11253-3
Wang, C.-Y., Liao, H.-Y., Wu, Y.-H., Chen, P.-Y., Hsieh, J.-W., & Yeh, I.-H. (2020). CSPNet: A new backbone that can enhance learning capability of CNN. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) (pp. 1571–1580). https://doi.org/10.1109/CVPRW50498.2020.00203
Lin, T. Y., Dollár, P., Girshick, R., He, K., Hariharan, B., & Belongie, S. (2017). Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 2117-2125). https://doi.org/10.1109/CVPR.2017.106
Woo, S., Park, J., Lee, J.-Y., & Kweon, I. S. (2018). CBAM: Convolutional block attention module. In V. Ferrari, M. Hebert, C. Sminchisescu, & Y. Weiss (Eds.), Computer vision – ECCV 2018 (Lecture Notes in Computer Science, Vol. 11211). Springer. https://doi.org/10.1007/978-3-030-01234-2_1
Ultralytics. (2025). YOLO12: Attention-centric object detection. https://docs.ultralytics.com/models/yolo12/ (Accessed June 28, 2025)
Ultralytics. (2025). YOLOv10 vs YOLOv8: A technical comparison for object detection. https://docs.ultralytics.com/ru/compare/yolov10-vs-yolov8/3 (Accessed June 28, 2025)
Xu, L., Zhao, Y., Zhai, Y., et al. (2024). Small object detection in UAV images based on YOLOv8n. International Journal of Computational Intelligence Systems, 17, 223. https://doi.org/10.1007/s44196-024-00632-3
Yao, G., Zhu, S., Zhang, L., & Qi, M. (2024). HP-YOLOv8: High-precision small object detection algorithm for remote sensing images. Sensors, 24(15), 4858. https://doi.org/10.3390/s24154858
Wang, Q., Zhou, Z., & Zhang, Z. (2025). RSO-YOLO: A Real-Time Detector for Small and Occluded Objects in Autonomous Driving Scenarios. Sensors, 25(21), 6703. https://doi.org/10.3390/s25216703
Akyon, F. C., Altinuc, S. O., & Temizel, A. (2022, October). Slicing aided hyper inference and fine-tuning for small object detection. In 2022 IEEE international conference on image processing (ICIP) (pp. 966-970). IEEE. https://doi.org/10.1109/ICIP46576.2022.9897990
Lakens, D. (2013). Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs. Frontiers in psychology, 4, 863. https://doi.org/10.3389/fpsyg.2013.00863
Hu, J., Shen, L., & Sun, G. (2018). Squeeze-and-excitation networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 7132–7141). https://doi.org/10.1109/CVPR.2018.00745
Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 779–788). https://doi.org/10.1109/CVPR.2016.91
NVIDIA Corporation. (2022). NVIDIA Ada Lovelace architecture (Technical white paper). https://catalogone.com/wp-content/uploads/2024/06/NVIDIA-ADA-GPU-PROVIZ-Architecture-Whitepaper_1.1.pdf
Lee, H., Lee, J. S., & Choi, H. C. (2021). Parallelization of Non-Maximum Suppression. IEEE Access, 9, 166579-166587. https://doi.org/10.1109/ACCESS.2021.3134639
Huang, J., Rathod, V., Sun, C., Zhu, M., Korattikara, A., Fathi, A., ... & Murphy, K. (2017). Speed/accuracy trade-offs for modern convolutional object detectors. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 7310-7311). https://doi.org/10.1109/CVPR.2017.351
Bodla, N., Singh, B., Chellappa, R., & Davis, L. (2017). Soft-NMS—Improving object detection with one line of code. In Proceedings of the IEEE International Conference on Computer Vision (ICCV) (pp. 5562–5570). https://doi.org/10.1109/ICCV.2017.593
Tremblay, J., et al. (2018). Training deep networks with synthetic data: Bridging the reality gap by domain randomization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). https://doi.org/10.1109/CVPRW.2018.00143
Sapkota, R., Meng, Z., Churuvija, M., Du, X., Ma, Z., & Karkee, M. (2026). Comprehensive performance evaluation of yolov12, yolo11, yolov10, yolov9 and yolov8 on detecting and counting fruitlet in complex orchard environments. Agriculture Communications, 100125. https://doi.org/10.1016/j.agrcom.2026.100125
Gupta, H., Kotlyar, O., Andreasson, H., & Lilienthal, A. J. (2024). Robust object detection in challenging weather conditions. In Proceedings of the IEEE/CVF winter conference on applications of computer vision (pp. 7523-7532). https://doi.org/10.1109/WACV57701.2024.00735
Wojke, N., Bewley, A., & Paulus, D. (2017). Simple online and realtime tracking with a deep association metric. In Proceedings of the IEEE International Conference on Image Processing (ICIP) (pp. 3645–3649). https://doi.org/10.1109/ICIP.2017.8296962
Liu, S., Huang, D., & Wang, Y. (2019). Adaptive nms: Refining pedestrian detection in a crowd. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 6459-6468). https://doi.org/10.1109/CVPR.2019.00662
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Articles are open access under the Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Authors who publish a manuscript in this journal agree to the following terms:
- The authors reserve the right to authorship of their work and transfer to the journal the right of first publication under the terms of the Creative Commons Attribution License, which allows others to freely distribute the published work with a mandatory link to the the original work and the first publication of the work in this journal.
- Authors have the right to conclude independent additional agreements that relate to the non-exclusive distribution of the work in the form in which it was published by this journal (for example, to post the work in the electronic repository of the institution or publish as part of a monograph), providing the link to the first publication of the work in this journal.
- Other terms stated in the Copyright Agreement.