BILATERAL VISION-LANGUAGE FRAMEWORK WITH CROSS-LATERAL ATTENTION FOR MAMMOGRAPHY
DOI:
https://doi.org/10.37943/HLVA7583%20Keywords:
breast cancer, mammography, vision-language models, contrastive learning, cross-lateral attentionAbstract
Screening mammography is read comparatively: radiologists hold one breast against the other, and several finding categories have no meaning outside that comparison. Most automated pipelines nonetheless encode a single breast at a time. We describe Bilateral MV-CLIP, a contrastive image-text framework in which all four projections of a patient enter one forward pass and the two lateralities exchange information through a Cross-Lateral Attention block, with a learned gate and a placeholder token preserving behaviour when a projection or an entire side is unavailable. Training combines image-text and image-image objectives, graded similarity targets, and a symmetry term tying embedding geometry to recorded asymmetry status. We train on VinDr-Mammo with reports synthesised from structured annotations and evaluate by linear probing of frozen features, cross-modal retrieval and report decoding, against an independently trained single-breast model matched in backbone, optimiser, schedule and epoch budget, with checkpoints selected on a held-out validation partition. On a pre-specified composite endpoint pooling the three comparison-defined categories, fusion did not improve on that baseline: area under the curve 0.781 against 0.792, a difference of -1.1 points with a 95% confidence interval from -9.6 to +7.5. Per-category effects follow the direction the design predicts, but the partition holds 49 positive comparison-defined breasts and no difference survived correction. Retrieval is limited by weak instance-level image-text alignment rather than by report redundancy, with a pool-size-independent ranking accuracy of 0.691. On an external cohort of 11,913 patients labelled by biopsy the representations transfer at reduced discrimination, reaching 0.65 for any abnormality. We therefore report explicit bilateral comparison as plausible but, on this dataset, unconfirmed, and give confidence intervals throughout.
References
Sung, H., Ferlay, J., Siegel, R. L., Laversanne, M., Soerjomataram, I., Jemal, A., & Bray, F. (2021). Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA: A Cancer Journal for Clinicians, 71(3), 209–249. https://doi.org/10.3322/caac.21660
McKinney, S. M., Sieniek, M., Godbole, V., Godwin, J., Antropova, N., Ashrafian, H., Back, T., Chesus, M., Corrado, G. S., Darzi, A., Etemadi, M., Garcia-Vicente, F., Gilbert, F. J., Halling-Brown, M., Hassabis, D., Jansen, S., Karthikesalingam, A., Kelly, C. J., King, D., Ledsam, J. R., Melnick, D., Mostofi, H., Peng, L., Reicher, J. J., Romera-Paredes, B., Sidebottom, R., Suleyman, M., Tse, D., Young, K. C., De Fauw, J., & Shetty, S. (2020). International evaluation of an AI system for breast cancer screening. Nature, 577(7788), 89–94. https://doi.org/10.1038/s41586-019-1799-6
Spak, D. A., Plaxco, J. S., Santiago, L., Dryden, M. J., & Dogan, B. E. (2017). BI-RADS® fifth edition: A summary of changes. Diagnostic and Interventional Imaging, 98(3), 179–190. https://doi.org/10.1016/j.diii.2017.01.001
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. Proceedings of Machine Learning Research, 139, 8748–8763. https://proceedings.mlr.press/v139/radford21a.html
Ghosh, S., Poynton, C. B., Visweswaran, S., & Batmanghelich, K. (2024). Mammo-CLIP: A vision language foundation model to enhance data efficiency and robustness in mammography. In M. G. Linguraru, Q. Dou, A. Feragen, S. Giannarou, B. Glocker, K. Lekadir, & J. A. Schnabel (Eds.), Medical image computing and computer assisted intervention – MICCAI 2024 (Lecture Notes in Computer Science, Vol. 15012, pp. 632–642). Springer. https://doi.org/10.1007/978-3-031-72390-2_59
Shen, L., Margolies, L. R., Rothstein, J. H., Fluder, E., McBride, R., & Sieh, W. (2019). Deep learning to improve breast cancer detection on screening mammography. Scientific Reports, 9(1), Article 12495. https://doi.org/10.1038/s41598-019-48995-4
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021, May 3–7). An image is worth 16×16 words: Transformers for image recognition at scale [Conference paper]. International Conference on Learning Representations, Virtual. https://openreview.net/forum?id=YicbFdNTTy
Tian, Y., Krishnan, D., & Isola, P. (2020). Contrastive multiview coding. In A. Vedaldi, H. Bischof, T. Brox, & J.-M. Frahm (Eds.), Computer vision – ECCV 2020 (Lecture Notes in Computer Science, Vol. 12356, pp. 776–794). Springer. https://doi.org/10.1007/978-3-030-58621-8_45
Chen, T., Kornblith, S., Norouzi, M., & Hinton, G. (2020). A simple framework for contrastive learning of visual representations. Proceedings of Machine Learning Research, 119, 1597–1607. https://proceedings.mlr.press/v119/chen20j.html
He, K., Fan, H., Wu, Y., Xie, S., & Girshick, R. (2020). Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 9729–9738). IEEE. https://doi.org/10.1109/CVPR42600.2020.00975
Zhang, Y., Jiang, H., Miura, Y., Manning, C. D., & Langlotz, C. P. (2022). Contrastive learning of medical visual representations from paired images and text. Proceedings of Machine Learning Research, 182, 2–25. https://proceedings.mlr.press/v182/zhang22a.html
Boecking, B., Usuyama, N., Bannur, S., Castro, D. C., Schwaighofer, A., Hyland, S., Wetscherek, M., Naumann, T., Nori, A., Alvarez-Valle, J., Poon, H., & Oktay, O. (2022). Making the most of text semantics to improve biomedical vision–language processing. In S. Avidan, G. Brostow, M. Cissé, G. M. Farinella, & T. Hassner (Eds.), Computer vision – ECCV 2022 (Lecture Notes in Computer Science, Vol. 13696, pp. 1–21). Springer. https://doi.org/10.1007/978-3-031-20059-5_1
Huang, S.-C., Shen, L., Lungren, M. P., & Yeung, S. (2021). GLoRIA: A multimodal global-local representation learning framework for label-efficient medical image recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 3942–3951). IEEE. https://doi.org/10.1109/ICCV48922.2021.00391
Wang, Z., Wu, Z., Agarwal, D., & Sun, J. (2022). MedCLIP: Contrastive learning from unpaired medical images and text. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (pp. 3876–3887). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.emnlp-main.256
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Vol. 1, pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423
Alsentzer, E., Murphy, J., Boag, W., Weng, W.-H., Jindi, D., Naumann, T., & McDermott, M. (2019). Publicly available clinical BERT embeddings. In Proceedings of the 2nd Clinical Natural Language Processing Workshop (pp. 72–78). Association for Computational Linguistics. https://doi.org/10.18653/v1/W19-1909
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008. https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html
Li, J., Li, D., Savarese, S., & Hoi, S. (2023). BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. Proceedings of Machine Learning Research, 202, 19730–19742. https://proceedings.mlr.press/v202/li23q.html
Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2002). BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (pp. 311–318). Association for Computational Linguistics. https://doi.org/10.3115/1073083.1073135
Lin, C.-Y. (2004). ROUGE: A package for automatic evaluation of summaries. In Text summarization branches out (pp. 74–81). Association for Computational Linguistics. https://aclanthology.org/W04-1013/
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2020, April 26–30). BERTScore: Evaluating text generation with BERT [Conference paper]. International Conference on Learning Representations, Virtual. https://openreview.net/forum?id=SkeHuCVFDr
Khosla, P., Teterwak, P., Wang, C., Sarna, A., Tian, Y., Isola, P., Maschinot, A., Liu, C., & Krishnan, D. (2020). Supervised contrastive learning. Advances in Neural Information Processing Systems, 33, 18661–18673. https://proceedings.neurips.cc/paper/2020/hash/d89a66c7c80a29b1bdbab0f2a1a94af8-Abstract.html
Nguyen, H. T., Nguyen, H. Q., Pham, H. H., Lam, K., Le, L. T., Dao, M., & Vu, V. (2023). VinDr-Mammo: A large-scale benchmark dataset for computer-aided diagnosis in full-field digital mammography. Scientific Data, 10(1), Article 277. https://doi.org/10.1038/s41597-023-02100-7
Tan, M., & Le, Q. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. Proceedings of Machine Learning Research, 97, 6105–6114. https://proceedings.mlr.press/v97/tan19a.html
Loshchilov, I., & Hutter, F. (2019, May 6–9). Decoupled weight decay regularization [Conference paper]. International Conference on Learning Representations, New Orleans, LA, United States. https://openreview.net/forum?id=Bkg6RiCqY7
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2022, April 25–29). LoRA: Low-rank adaptation of large language models [Conference paper]. International Conference on Learning Representations, Virtual. https://openreview.net/forum?id=nZeVKeeFYf9
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Articles are open access under the Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Authors who publish a manuscript in this journal agree to the following terms:
- The authors reserve the right to authorship of their work and transfer to the journal the right of first publication under the terms of the Creative Commons Attribution License, which allows others to freely distribute the published work with a mandatory link to the the original work and the first publication of the work in this journal.
- Authors have the right to conclude independent additional agreements that relate to the non-exclusive distribution of the work in the form in which it was published by this journal (for example, to post the work in the electronic repository of the institution or publish as part of a monograph), providing the link to the first publication of the work in this journal.
- Other terms stated in the Copyright Agreement.