AN EFFECTIVE METHOD FOR ANALYZING HUMAN MOVEMENT IN REAL TIME BASED ON YOLO-ROI
DOI:
https://doi.org/10.37943/TV013333%20Keywords:
YOLO , ROI , real-time processing , human motion analysis , computer vision , pose estimation , edge devices , latency optimizationAbstract
Real-time human movement analysis has become an essential component of modern computer vision applications, including sports performance assessment, rehabilitation monitoring, intelligent surveillance, and human–computer interaction. Although deep learning–based pose estimation methods provide accurate human body landmark detection, executing person detection on every video frame introduces considerable computational overhead, limiting their applicability on resource-constrained CPU-based systems.
This paper proposes an efficient YOLO–ROI framework for real-time human movement analysis that combines scheduled YOLO person detection, region-of-interest (ROI) generation, MediaPipe Pose estimation, and lightweight landmark smoothing. Instead of performing object detection on every frame, the proposed framework periodically executes the YOLO detector while reusing the most recently detected ROI between detection cycles. Consequently, MediaPipe Pose continuously estimates body landmarks within the ROI, significantly reducing redundant object detection operations and improving computational efficiency.
The proposed framework was evaluated using a self-collected dataset consisting of 200 videos recorded from 10 participants performing five physical exercises under controlled indoor conditions. Computational performance was analyzed under different ROI sizes, detection intervals, and MediaPipe Pose complexity levels using CPU-based execution.
Experimental results demonstrate that the proposed optimization strategy achieves a maximum processing speed of 61.32 FPS with an average processing latency of 16.31 ms, enabling stable real-time operation without GPU acceleration. The obtained results further show that scheduled YOLO detection, ROI reuse, and lightweight landmark smoothing effectively reduce computational workload while preserving continuous pose estimation throughout video processing.
The proposed framework provides a simple, computationally efficient, and easily deployable solution for real-time human movement analysis and can be readily integrated into sports analytics, rehabilitation systems, intelligent surveillance applications, and edge-based computer vision platforms.
References
Chen, H., Feng, R., Wu, S., Xu, H., Zhou, F., & Liu, Z. (2023). 2D human pose estimation: A survey. Multimedia Systems, 29(5), 3115–3138. https://doi.org/10.1007/s00530-022-01019-0
Gao, Z., Chen, J., Liu, Y., Jin, Y., & Tian, D. (2025). A systematic survey on human pose estimation: Upstream and downstream tasks, approaches, lightweight models, and prospects. Artificial Intelligence Review, 58(3), Article 68. https://doi.org/10.1007/s10462-024-11060-2
Cao, Z., Hidalgo, G., Simon, T., Wei, S.-E., & Sheikh, Y. (2021). OpenPose: Realtime multi-person 2D pose estimation using part affinity fields. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(1), 172–186. https://doi.org/10.1109/TPAMI.2019.2929257
Ahmadyan, A., Zhang, L., Ablavatski, A., Wei, J., & Grundmann, M. (2021). Objectron: A large scale dataset of object-centric videos in the wild with pose annotations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 7822–7831). https://openaccess.thecvf.com/content/CVPR2021/papers/Ahmadyan_Objectron_A_Large_Scale_Dataset_of_Object-Centric_Videos_in_the_CVPR_2021_paper.pdf
Lugaresi, C., Tang, J., Nash, H., McClanahan, C., Uboweja, E., Hays, M., Zhang, F., Chang, C.-L., Yong, M. G., Lee, J., Chang, W.-T., Hua, W., Georg, M., & Grundmann, M. (2019). MediaPipe: A framework for perceiving and processing reality [Workshop paper]. Third Workshop on Computer Vision for AR/VR at IEEE Computer Vision and Pattern Recognition (CVPR). https://research.google/Fpubs/mediapipe-a-framework-for-perceiving-and-processing-reality/
Redmon, J., & Farhadi, A. (2017). YOLO9000: Better, faster, stronger. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. https://openaccess.thecvf.com/content_cvpr_2017/papers/Redmon_YOLO9000_Better_Faster_CVPR_2017_paper.pdf
Wang, C.-Y., Bochkovskiy, A., & Liao, H.-Y. M. (2021). Scaled-YOLOv4: Scaling cross stage partial network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 13029–13038). https://openaccess.thecvf.com/content/CVPR2021/papers/Wang_Scaled-YOLOv4_Scaling_Cross_Stage_Partial_Network_CVPR_2021_paper.pdf
Wang, C.-Y., Bochkovskiy, A., & Liao, H.-Y. M. (2023). YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 7464–7475). https://doi.org/10.1109/CVPR52729.2023.00721
Jocher, G., Chaurasia, A., & Qiu, J. (2023). Ultralytics YOLOv8 [Computer software]. GitHub. https://github.com/ultralytics/ultralytics
Chu, X., Yang, W., Ouyang, W., Ma, C., Yuille, A. L., & Wang, X. (2017). Multi-context attention for human pose estimation. arXiv. https://doi.org/10.48550/arXiv.1702.07432
Pavllo, D., Feichtenhofer, C., Grangier, D., & Auli, M. (2019). 3D human pose estimation in video with temporal convolutions and semi-supervised training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 7753–7762). https://doi.org/10.1109/CVPR.2019.00794
Xiao, B., Wu, H., & Wei, Y. (2018). Simple baselines for human pose estimation and tracking. arXiv. https://doi.org/10.48550/arXiv.1804.06208
Wang, J., Sun, K., Cheng, T., Jiang, B., Deng, C., Zhao, Y., Liu, D., Mu, Y., Tan, M., Wang, X., Liu, W., & Xiao, B. (2021). Deep high-resolution representation learning for visual recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(10), 3349–3364. https://doi.org/10.1109/TPAMI.2020.2983686
Xu, Y., Zhang, J., Zhang, Q., & Tao, D. (2022). ViTPose: Simple vision transformer baselines for human pose estimation. arXiv. https://doi.org/10.48550/arXiv.2204.12484
Lu, P., Jiang, T., Li, Y., Li, X., Chen, K., & Yang, W. (2024). RTMO: Towards high-performance one-stage real-time multi-person pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 1491–1500). https://openaccess.thecvf.com/content/CVPR2024/papers/Lu_RTMO_Towards_High-Performance_One-Stage_Real-Time_Multi-Person_Pose_Estimation_CVPR_2024_paper.pdf
Kim, J.-W., Choi, J.-Y., Ha, E.-J., & Choi, J.-H. (2023). Human pose estimation using MediaPipe Pose and optimization method based on a humanoid model. Applied Sciences, 13(4), Article 2700. https://doi.org/10.3390/app13042700
Hii, C. S. T., Gan, K. B., Zainal, N., Mohamed Ibrahim, N., Azmin, S., Mat Desa, S. H., van de Warrenburg, B., & You, H. W. (2023). Automated gait analysis based on a marker-free pose estimation model. Sensors, 23(14), Article 6489. https://doi.org/10.3390/s23146489
Dill, S., Ahmadi, A., Grimmer, M., Haufe, D., Rohr, M., Zhao, Y., Sharbafi, M., & Hoog Antink, C. (2024). Accuracy evaluation of 3D pose reconstruction algorithms through stereo camera information fusion for physical exercises with MediaPipe Pose. Sensors, 24(23), Article 7772. https://doi.org/10.3390/s24237772
Zhang, Y., Sun, P., Jiang, Y., Yu, D., Weng, F., Yuan, Z., Luo, P., Liu, W., & Wang, X. (2022). ByteTrack: Multi-object tracking by associating every detection box. In S. Avidan, G. Brostow, M. Cissé, G. M. Farinella, & T. Hassner (Eds.), Computer vision – ECCV 2022 (pp. 1–21). Springer. https://doi.org/10.1007/978-3-031-20047-2_1
Fang, H.-S., Li, J., Tang, H., Xu, C., Zhu, H., Xiu, Y., Li, Y.-L., & Lu, C. (2023). AlphaPose: Whole-body regional multi-person pose estimation and tracking in real-time. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(6), 7157–7173. https://doi.org/10.1109/TPAMI.2022.3222784
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Articles are open access under the Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Authors who publish a manuscript in this journal agree to the following terms:
- The authors reserve the right to authorship of their work and transfer to the journal the right of first publication under the terms of the Creative Commons Attribution License, which allows others to freely distribute the published work with a mandatory link to the the original work and the first publication of the work in this journal.
- Authors have the right to conclude independent additional agreements that relate to the non-exclusive distribution of the work in the form in which it was published by this journal (for example, to post the work in the electronic repository of the institution or publish as part of a monograph), providing the link to the first publication of the work in this journal.
- Other terms stated in the Copyright Agreement.