TASK-AWARE EVALUATION OF JPEG RECOMPRESSION AND RGB BIT-DEPTH REDUCTION FOR OBJECT DETECTION
DOI:
https://doi.org/10.37943/TRMD1403%20Keywords:
Image Compression; Quantization; JPEG Recompression; Bit-Depth Reduction; Ob¬ject Detection; YOLOv8; COCO Dataset; PSNR; SSIM; mAP.Abstract
How does image storage affect object detection? Traditional image quality metrics typically focus on human vision, but computer vision models process images differently. To explore this gap, we evaluated the Microsoft COCO val2017 dataset under two distinct degradation methods: frequency-domain JPEG recompression (quality levels 94 to 25) and amplitude-domain uniform RGB bit-depth reduction (8 bits to 1 bit) stored in 24-bit PNG containers. We processed the degraded images using three distinct architectures: two convolutional networks (YOLOv8n, YOLOv8m) and a Vision Transformer (RT-DETR-L). The results demonstrate that JPEG recompression provides a highly efficient tradeoff; quality level 88 halves the dataset storage size (2.02× compression, reducing to 384 MB) with less than a 1.5% relative drop in mAP50 across all models. Conversely, storing uniformly quantized RGB data in lossless PNG containers proved highly inefficient. While 7-bit quantization preserved baseline accuracy, it artificially inflated the physical file size to 1912 MB. Furthermore, at 3 bits per channel, average mAP50 dropped by over 12%, and all networks effectively failed at the 2-bit and 1-bit thresholds. Crucially, scale-aware metrics revealed that small objects are disproportionately vulnerable to compression artifacts across both CNN and Transformer paradigms, suffering up to a 50% relative accuracy loss. Ultimately, we conclude that standard JPEG compression is significantly more robust for object detection than storing uniformly quantized RGB data in unoptimized containers. These empirical findings highlight that conventional human-centric quality metrics are insufficient for predicting downstream neural network performance. Future research should investigate whether these non-linear degradation trends persist across specialized domain-shifted datasets and modern learned compression algorithms.
References
Wallace, G. K. (1992). The JPEG still picture compression standard. IEEE Transactions on Consumer Electronics, 38(1), xviii–xxxiv. https://doi.org/10.1109/30.125072
World Wide Web Consortium. (2003). Portable Network Graphics (PNG) specification (second edition) (W3C Recommendation 10 November 2003). https://www.w3.org/TR/PNG/
Wang, Z., Bovik, A. C., Sheikh, H. R., & Simoncelli, E. P. (2004). Image quality assessment: From error visibility to structural similarity. IEEE Transactions on Image Processing, 13(4), 600–612. https://doi.org/10.1109/TIP.2003.819861
Zhai, G., & Min, X. (2020). Perceptual image quality assessment: A survey. Science China Information Sciences, 63(11), 211301. https://doi.org/10.1007/s11432-019-2757-1
Horé, A., & Ziou, D. (2010). Image quality metrics: PSNR vs. SSIM. 2010 20th International Conference on Pattern Recognition, 2366–2369. https://doi.org/10.1109/ICPR.2010.579
Heckbert, P. S. (1982). Color image quantization for frame buffer display. Proceedings of the 9th Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH '82), 297–307. https://doi.org/10.1145/800064.801294
Floyd, R. W., & Steinberg, L. (1976). An adaptive algorithm for spatial grey scale. Proceedings of the Society of Information Display, 17(2), 75–77.
Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., & Kalenichenko, D. (2018). Quantization and training of neural networks for efficient integer-arithmetic-only inference. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2704–2713. https://doi.org/10.1109/CVPR.2018.00291
Han, S., Mao, H., & Dally, W. J. (2016). Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding. 4th International Conference on Learning Representations (ICLR 2016). https://doi.org/10.48550/arXiv.1510.00149
Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 779–788. https://doi.org/10.1109/CVPR.2016.91
Wang, C.-Y., Bochkovskiy, A., & Liao, H.-Y. M. (2023). YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7464–7475. https://doi.org/10.1109/CVPR52729.2023.00721
Jocher, G., Chaurasia, A., & Qiu, J. (2023). Ultralytics YOLOv8 (Version 8.0.0) [Computer software]. https://github.com/ultralytics/ultralytics
Lin, T.-Y., Maire, M., Belongie, S. J., Hays, J., Perona, P., Ramanan, D., Dollár, P., & Zitnick, C. L. (2014). Microsoft COCO: Common objects in context. In D. Fleet, T. Pajdla, B. Schiele, & T. Tuytelaars (Eds.), Computer vision – ECCV 2014 (pp. 740–755). Springer. https://doi.org/10.1007/978-3-319-10602-1_48
Song, M., Choi, J., & Han, B. (2021). Variable-rate deep image compression through spatially-adaptive feature transform. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2360–2369. https://doi.org/10.1109/ICCV48922.2021.00238
Ye, J., Yeo, H., Park, J., & Han, D. (2023). AccelIR: Task-aware image compression for accelerating neural restoration. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 17852–17861. https://doi.org/10.1109/CVPR52729.2023.01747
Gandor, T., & Nalepa, J. (2022). First gradually, then suddenly: Understanding the impact of image compression on object detection using deep learning. Sensors, 22(3), 1104. https://doi.org/10.3390/s22031104
Hao, Y., Pei, H., Lyu, Y., Yuan, Z., Rizzo, J.-R., Wang, Y., & Fang, Y. (2022). Understanding the impact of image quality and distance of objects to object detection performance. 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 5136-5142. https://doi.org/10.1109/IROS55552.2023.10342371
Janeiro, J. M., Frolov, S., El-Nouby, A., & Verbeek, J. (2023). Are visual recognition models robust to image compression? Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2510-2520. https://doi.org/10.1109/ICCV51070.2023.00236
Huang, C.-H., & Wu, J.-L. (2024). Unveiling the future of human and machine coding: A survey of end-to-end learned image compression. Entropy, 26(5), 357. https://doi.org/10.3390/e26050357
Jamil, S., Piran, M. J., Rahman, M. U., & Kwon, O. J. (2023). Learning-driven lossy image compression: A comprehensive survey. Engineering Applications of Artificial Intelligence, 123, 106176. https://doi.org/10.1016/j.engappai.2023.106176
Pavlov, I. (2013). LZMA SDK (Version 9.20) [Computer software]. https://www.7-zip.org/sdk.html
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Articles are open access under the Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Authors who publish a manuscript in this journal agree to the following terms:
- The authors reserve the right to authorship of their work and transfer to the journal the right of first publication under the terms of the Creative Commons Attribution License, which allows others to freely distribute the published work with a mandatory link to the the original work and the first publication of the work in this journal.
- Authors have the right to conclude independent additional agreements that relate to the non-exclusive distribution of the work in the form in which it was published by this journal (for example, to post the work in the electronic repository of the institution or publish as part of a monograph), providing the link to the first publication of the work in this journal.
- Other terms stated in the Copyright Agreement.