TASK-AWARE EVALUATION OF JPEG RECOMPRESSION AND RGB BIT-DEPTH REDUCTION FOR OBJECT DETECTION

Authors

DOI:

https://doi.org/10.37943/TRMD1403%20

Keywords:

Image Compression; Quantization; JPEG Recompression; Bit-Depth Reduction; Ob¬ject Detection; YOLOv8; COCO Dataset; PSNR; SSIM; mAP.

Abstract

How does image storage affect object detection? Traditional image quality metrics typically focus on human vision, but computer vision models process images differently. To explore this gap, we evaluated the Microsoft COCO val2017 dataset under two distinct degradation methods: frequency-domain JPEG recompression (quality levels 94 to 25) and amplitude-domain uniform RGB bit-depth reduction (8 bits to 1 bit) stored in 24-bit PNG containers. We processed the degraded images using three distinct architectures: two convolutional networks (YOLOv8n, YOLOv8m) and a Vision Transformer (RT-DETR-L). The results demonstrate that JPEG recompression provides a highly efficient tradeoff; quality level 88 halves the dataset storage size (2.02× compression, reducing to 384 MB) with less than a 1.5% relative drop in mAP50 across all models. Conversely, storing uniformly quantized RGB data in lossless PNG containers proved highly inefficient. While 7-bit quantization preserved baseline accuracy, it artificially inflated the physical file size to 1912 MB. Furthermore, at 3 bits per channel, average mAP50 dropped by over 12%, and all networks effectively failed at the 2-bit and 1-bit thresholds. Crucially, scale-aware metrics revealed that small objects are disproportionately vulnerable to compression artifacts across both CNN and Transformer paradigms, suffering up to a 50% relative accuracy loss. Ultimately, we conclude that standard JPEG compression is significantly more robust for object detection than storing uniformly quantized RGB data in unoptimized containers. These empirical findings highlight that conventional human-centric quality metrics are insufficient for predicting downstream neural network performance. Future research should investigate whether these non-linear degradation trends persist across specialized domain-shifted datasets and modern learned compression algorithms.

Author Biographies

Amir Karatayev, Astana IT University, Kazakhstan

Master’s Degree, School of Software Engineering

Shynar Akhmetzhanova, Astana IT University, Kazakhstan

Candidate of Technical Sciences, Assistant Professor, School of Software Engineering

References

Wallace, G. K. (1992). The JPEG still picture compression standard. IEEE Transactions on Consumer Electronics, 38(1), xviii–xxxiv. https://doi.org/10.1109/30.125072

World Wide Web Consortium. (2003). Portable Network Graphics (PNG) specification (second edition) (W3C Recommendation 10 November 2003). https://www.w3.org/TR/PNG/

Wang, Z., Bovik, A. C., Sheikh, H. R., & Simoncelli, E. P. (2004). Image quality assessment: From error visibility to structural similarity. IEEE Transactions on Image Processing, 13(4), 600–612. https://doi.org/10.1109/TIP.2003.819861

Zhai, G., & Min, X. (2020). Perceptual image quality assessment: A survey. Science China Information Sciences, 63(11), 211301. https://doi.org/10.1007/s11432-019-2757-1

Horé, A., & Ziou, D. (2010). Image quality metrics: PSNR vs. SSIM. 2010 20th International Conference on Pattern Recognition, 2366–2369. https://doi.org/10.1109/ICPR.2010.579

Heckbert, P. S. (1982). Color image quantization for frame buffer display. Proceedings of the 9th Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH '82), 297–307. https://doi.org/10.1145/800064.801294

Floyd, R. W., & Steinberg, L. (1976). An adaptive algorithm for spatial grey scale. Proceedings of the Society of Information Display, 17(2), 75–77.

Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., & Kalenichenko, D. (2018). Quantization and training of neural networks for efficient integer-arithmetic-only inference. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2704–2713. https://doi.org/10.1109/CVPR.2018.00291

Han, S., Mao, H., & Dally, W. J. (2016). Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding. 4th International Conference on Learning Representations (ICLR 2016). https://doi.org/10.48550/arXiv.1510.00149

Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 779–788. https://doi.org/10.1109/CVPR.2016.91

Wang, C.-Y., Bochkovskiy, A., & Liao, H.-Y. M. (2023). YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7464–7475. https://doi.org/10.1109/CVPR52729.2023.00721

Jocher, G., Chaurasia, A., & Qiu, J. (2023). Ultralytics YOLOv8 (Version 8.0.0) [Computer software]. https://github.com/ultralytics/ultralytics

Lin, T.-Y., Maire, M., Belongie, S. J., Hays, J., Perona, P., Ramanan, D., Dollár, P., & Zitnick, C. L. (2014). Microsoft COCO: Common objects in context. In D. Fleet, T. Pajdla, B. Schiele, & T. Tuytelaars (Eds.), Computer vision – ECCV 2014 (pp. 740–755). Springer. https://doi.org/10.1007/978-3-319-10602-1_48

Song, M., Choi, J., & Han, B. (2021). Variable-rate deep image compression through spatially-adaptive feature transform. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2360–2369. https://doi.org/10.1109/ICCV48922.2021.00238

Ye, J., Yeo, H., Park, J., & Han, D. (2023). AccelIR: Task-aware image compression for accelerating neural restoration. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 17852–17861. https://doi.org/10.1109/CVPR52729.2023.01747

Gandor, T., & Nalepa, J. (2022). First gradually, then suddenly: Understanding the impact of image compression on object detection using deep learning. Sensors, 22(3), 1104. https://doi.org/10.3390/s22031104

Hao, Y., Pei, H., Lyu, Y., Yuan, Z., Rizzo, J.-R., Wang, Y., & Fang, Y. (2022). Understanding the impact of image quality and distance of objects to object detection performance. 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 5136-5142. https://doi.org/10.1109/IROS55552.2023.10342371

Janeiro, J. M., Frolov, S., El-Nouby, A., & Verbeek, J. (2023). Are visual recognition models robust to image compression? Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2510-2520. https://doi.org/10.1109/ICCV51070.2023.00236

Huang, C.-H., & Wu, J.-L. (2024). Unveiling the future of human and machine coding: A survey of end-to-end learned image compression. Entropy, 26(5), 357. https://doi.org/10.3390/e26050357

Jamil, S., Piran, M. J., Rahman, M. U., & Kwon, O. J. (2023). Learning-driven lossy image compression: A comprehensive survey. Engineering Applications of Artificial Intelligence, 123, 106176. https://doi.org/10.1016/j.engappai.2023.106176

Pavlov, I. (2013). LZMA SDK (Version 9.20) [Computer software]. https://www.7-zip.org/sdk.html

Downloads

Published

2026-09-30

How to Cite

Karatayev, A., & Akhmetzhanova, S. (2026). TASK-AWARE EVALUATION OF JPEG RECOMPRESSION AND RGB BIT-DEPTH REDUCTION FOR OBJECT DETECTION. Scientific Journal of Astana IT University, 27(3), 122–137. https://doi.org/10.37943/TRMD1403

Issue

Section

Information Technologies