BILATERAL VISION-LANGUAGE FRAMEWORK WITH CROSS-LATERAL ATTENTION FOR MAMMOGRAPHY

Authors

DOI:

https://doi.org/10.37943/HLVA7583%20

Keywords:

breast cancer, mammography, vision-language models, contrastive learning, cross-lateral attention

Abstract

Screening mammography is read comparatively: radiologists hold one breast against the other, and several finding categories have no meaning outside that comparison. Most automated pipelines nonetheless encode a single breast at a time. We describe Bilateral MV-CLIP, a contrastive image-text framework in which all four projections of a patient enter one forward pass and the two lateralities exchange information through a Cross-Lateral Attention block, with a learned gate and a placeholder token preserving behaviour when a projection or an entire side is unavailable. Training combines image-text and image-image objectives, graded similarity targets, and a symmetry term tying embedding geometry to recorded asymmetry status. We train on VinDr-Mammo with reports synthesised from structured annotations and evaluate by linear probing of frozen features, cross-modal retrieval and report decoding, against an independently trained single-breast model matched in backbone, optimiser, schedule and epoch budget, with checkpoints selected on a held-out validation partition. On a pre-specified composite endpoint pooling the three comparison-defined categories, fusion did not improve on that baseline: area under the curve 0.781 against 0.792, a difference of -1.1 points with a 95% confidence interval from -9.6 to +7.5. Per-category effects follow the direction the design predicts, but the partition holds 49 positive comparison-defined breasts and no difference survived correction. Retrieval is limited by weak instance-level image-text alignment rather than by report redundancy, with a pool-size-independent ranking accuracy of 0.691. On an external cohort of 11,913 patients labelled by biopsy the representations transfer at reduced discrimination, reaching 0.65 for any abnormality. We therefore report explicit bilateral comparison as plausible but, on this dataset, unconfirmed, and give confidence intervals throughout.

References

Sung, H., Ferlay, J., Siegel, R. L., Laversanne, M., Soerjomataram, I., Jemal, A., & Bray, F. (2021). Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA: A Cancer Journal for Clinicians, 71(3), 209–249. https://doi.org/10.3322/caac.21660

McKinney, S. M., Sieniek, M., Godbole, V., Godwin, J., Antropova, N., Ashrafian, H., Back, T., Chesus, M., Corrado, G. S., Darzi, A., Etemadi, M., Garcia-Vicente, F., Gilbert, F. J., Halling-Brown, M., Hassabis, D., Jansen, S., Karthikesalingam, A., Kelly, C. J., King, D., Ledsam, J. R., Melnick, D., Mostofi, H., Peng, L., Reicher, J. J., Romera-Paredes, B., Sidebottom, R., Suleyman, M., Tse, D., Young, K. C., De Fauw, J., & Shetty, S. (2020). International evaluation of an AI system for breast cancer screening. Nature, 577(7788), 89–94. https://doi.org/10.1038/s41586-019-1799-6

Spak, D. A., Plaxco, J. S., Santiago, L., Dryden, M. J., & Dogan, B. E. (2017). BI-RADS® fifth edition: A summary of changes. Diagnostic and Interventional Imaging, 98(3), 179–190. https://doi.org/10.1016/j.diii.2017.01.001

Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. Proceedings of Machine Learning Research, 139, 8748–8763. https://proceedings.mlr.press/v139/radford21a.html

Ghosh, S., Poynton, C. B., Visweswaran, S., & Batmanghelich, K. (2024). Mammo-CLIP: A vision language foundation model to enhance data efficiency and robustness in mammography. In M. G. Linguraru, Q. Dou, A. Feragen, S. Giannarou, B. Glocker, K. Lekadir, & J. A. Schnabel (Eds.), Medical image computing and computer assisted intervention – MICCAI 2024 (Lecture Notes in Computer Science, Vol. 15012, pp. 632–642). Springer. https://doi.org/10.1007/978-3-031-72390-2_59

Shen, L., Margolies, L. R., Rothstein, J. H., Fluder, E., McBride, R., & Sieh, W. (2019). Deep learning to improve breast cancer detection on screening mammography. Scientific Reports, 9(1), Article 12495. https://doi.org/10.1038/s41598-019-48995-4

Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021, May 3–7). An image is worth 16×16 words: Transformers for image recognition at scale [Conference paper]. International Conference on Learning Representations, Virtual. https://openreview.net/forum?id=YicbFdNTTy

Tian, Y., Krishnan, D., & Isola, P. (2020). Contrastive multiview coding. In A. Vedaldi, H. Bischof, T. Brox, & J.-M. Frahm (Eds.), Computer vision – ECCV 2020 (Lecture Notes in Computer Science, Vol. 12356, pp. 776–794). Springer. https://doi.org/10.1007/978-3-030-58621-8_45

Chen, T., Kornblith, S., Norouzi, M., & Hinton, G. (2020). A simple framework for contrastive learning of visual representations. Proceedings of Machine Learning Research, 119, 1597–1607. https://proceedings.mlr.press/v119/chen20j.html

He, K., Fan, H., Wu, Y., Xie, S., & Girshick, R. (2020). Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 9729–9738). IEEE. https://doi.org/10.1109/CVPR42600.2020.00975

Zhang, Y., Jiang, H., Miura, Y., Manning, C. D., & Langlotz, C. P. (2022). Contrastive learning of medical visual representations from paired images and text. Proceedings of Machine Learning Research, 182, 2–25. https://proceedings.mlr.press/v182/zhang22a.html

Boecking, B., Usuyama, N., Bannur, S., Castro, D. C., Schwaighofer, A., Hyland, S., Wetscherek, M., Naumann, T., Nori, A., Alvarez-Valle, J., Poon, H., & Oktay, O. (2022). Making the most of text semantics to improve biomedical vision–language processing. In S. Avidan, G. Brostow, M. Cissé, G. M. Farinella, & T. Hassner (Eds.), Computer vision – ECCV 2022 (Lecture Notes in Computer Science, Vol. 13696, pp. 1–21). Springer. https://doi.org/10.1007/978-3-031-20059-5_1

Huang, S.-C., Shen, L., Lungren, M. P., & Yeung, S. (2021). GLoRIA: A multimodal global-local representation learning framework for label-efficient medical image recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 3942–3951). IEEE. https://doi.org/10.1109/ICCV48922.2021.00391

Wang, Z., Wu, Z., Agarwal, D., & Sun, J. (2022). MedCLIP: Contrastive learning from unpaired medical images and text. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (pp. 3876–3887). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.emnlp-main.256

Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Vol. 1, pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423

Alsentzer, E., Murphy, J., Boag, W., Weng, W.-H., Jindi, D., Naumann, T., & McDermott, M. (2019). Publicly available clinical BERT embeddings. In Proceedings of the 2nd Clinical Natural Language Processing Workshop (pp. 72–78). Association for Computational Linguistics. https://doi.org/10.18653/v1/W19-1909

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008. https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html

Li, J., Li, D., Savarese, S., & Hoi, S. (2023). BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. Proceedings of Machine Learning Research, 202, 19730–19742. https://proceedings.mlr.press/v202/li23q.html

Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2002). BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (pp. 311–318). Association for Computational Linguistics. https://doi.org/10.3115/1073083.1073135

Lin, C.-Y. (2004). ROUGE: A package for automatic evaluation of summaries. In Text summarization branches out (pp. 74–81). Association for Computational Linguistics. https://aclanthology.org/W04-1013/

Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2020, April 26–30). BERTScore: Evaluating text generation with BERT [Conference paper]. International Conference on Learning Representations, Virtual. https://openreview.net/forum?id=SkeHuCVFDr

Khosla, P., Teterwak, P., Wang, C., Sarna, A., Tian, Y., Isola, P., Maschinot, A., Liu, C., & Krishnan, D. (2020). Supervised contrastive learning. Advances in Neural Information Processing Systems, 33, 18661–18673. https://proceedings.neurips.cc/paper/2020/hash/d89a66c7c80a29b1bdbab0f2a1a94af8-Abstract.html

Nguyen, H. T., Nguyen, H. Q., Pham, H. H., Lam, K., Le, L. T., Dao, M., & Vu, V. (2023). VinDr-Mammo: A large-scale benchmark dataset for computer-aided diagnosis in full-field digital mammography. Scientific Data, 10(1), Article 277. https://doi.org/10.1038/s41597-023-02100-7

Tan, M., & Le, Q. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. Proceedings of Machine Learning Research, 97, 6105–6114. https://proceedings.mlr.press/v97/tan19a.html

Loshchilov, I., & Hutter, F. (2019, May 6–9). Decoupled weight decay regularization [Conference paper]. International Conference on Learning Representations, New Orleans, LA, United States. https://openreview.net/forum?id=Bkg6RiCqY7

Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2022, April 25–29). LoRA: Low-rank adaptation of large language models [Conference paper]. International Conference on Learning Representations, Virtual. https://openreview.net/forum?id=nZeVKeeFYf9

Downloads

Published

2026-09-30

How to Cite

Abdikenov, B., Imasheva, A., & Zhaksylyk, T. (2026). BILATERAL VISION-LANGUAGE FRAMEWORK WITH CROSS-LATERAL ATTENTION FOR MAMMOGRAPHY. Scientific Journal of Astana IT University, 27(3), 187–207. https://doi.org/10.37943/HLVA7583

Issue

Section

Information Technologies