| 1 |
Amudhavalli, L. (2025). Exploring transformer-based architectures for large-scale multimodal information retrieval systems. International Journal of Computer Science and Engineering Innovations, 1(1):18–26. DOI: 10.64137/31079458/IJCSEI-V1I1P103.
|
|
| 2 |
Attallah, O. (2025). Multi-domain feature incorporation of lightweight convolutional neural networks and handcrafted features for lung and colon cancer diagnosis. Technologies, 13(5). DOI: 10.3390/technologies13050173.
|
|
| 3 |
Campos, L. O., Favoreto, G. S., Traina-Jr., C., Traina, A. J. M., and Cazzolato, M. T. (2026). Incomplete information retrieval without imputation: Exploiting the correlation among deep features. In Proceedings of the 39th IEEE International Symposium on Computer-Based Medical Systems (CBMS). To appear.
|
|
| 4 |
Dubey, S. R., Singh, S. K., and Chu, W.-T. (2022). Vision Transformer Hashing for Image Retrieval. In 2022 IEEE International Conference on Multimedia and Expo (ICME). IEEE. DOI: 10.1109/ICME52920.2022.9859900.
|
|
| 5 |
Fiel, S. and Sablatnig, R. (2024). Self-supervised vision transformers for writer retrieval. In Document Analysis and Recognition - ICDAR 2024, pages 380–396. Springer. DOI:
10.1007/978-3-031-70536-6_23.
|
|
| 6 |
Jiang, X., Tang, H., Pan, Y., and Li, Z. (2025). Rethinking vision transformer for largescale fine-grained image retrieval. arXiv preprint arXiv:2504.16691, 18(9). DOI: 10.48550/arXiv.2504.16691.
|
|
| 7 |
Johnson, A., Pollard, T., Mark, R., Berkowitz, S., and Horng, S. (2024). MIMIC-CXR
Database. PhysioNet. Version 2.1.0.
|
|
| 8 |
Kim, W., Son, B., and Kim, I. (2021). Vilt: Vision-and-language transfor-
mer without convolution or region supervision. In Proc. of the Internatio-
nal Conference on Machine Learning, volume 139 of PMLR. PMLR. DOI:
https://doi.org/10.48550/arXiv.2102.03334.
|
|
| 9 |
Liu, S. and Deng, W. (2015). Very deep convolutional neural network based image classi-
fication using small training sample size. In 2015 3rd IAPR Asian Conference on Pat-
tern Recognition (ACPR), pages 730–734. IEEE. DOI: 10.1109/ACPR.2015.7486599.
|
|
| 10 |
Peng, Y. (2025). A clip-based cross-modal matching model for image-text retrieval. In-
formation Technology and Control, 54(3):1030–1048.
|
|
| 11 |
Song, C. H., Yoon, J., Choi, S., and Avrithis, Y. (2023). Boosting vision transformers for
image retrieval. In 2023 IEEE/CVF Winter Conference on Applications of Computer
Vision (WACV), pages 107–117. DOI: 10.48550/arXiv.2210.11909.
|
|
| 12 |
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł.,
and Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Informa-
tion Processing Systems 30 (NIPS 2017). Curran Associates, Inc. DOI: 10.48550/ar-
Xiv.1706.03762.
|
|
| 13 |
Wang, X., Peng, Y., Lu, L., Lu, Z., Bagheri, M., and Summers, R. M. (2017). Chestx-
ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised clas-
sification and localization of common thorax diseases. In Proceedings of the IEEE
Conference on Computer Vision and Pattern Recognition (CVPR), pages 2097–2106.
DOI: 10.1109/CVPR.2017.369.
|
|
| 14 |
Xu, P., Zhu, X., and Clifton, D. A. (2023). Multimodal learning with transformers: A sur-
vey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(10):11993–
12013. DOI: 10.1109/TPAMI.2023.3274228.
|
|
| 15 |
Zhang, Z., Li, J., Zhu, L., Wang, T., and Shen, H. T. (2024). Cross-modal retrieval:
A review of methodologies, datasets, and future perspectives. IEEE Transactions on
Multimedia. DOI: 10.1109/ACCESS.2024.3444817.
|
|
| 16 |
Zheng, J., Liang, M., Yu, Y., Li, Y., and Xue, Z. (2024). Knowledge graph enhan-
ced multimodal transformer for image-text retrieval. In 2024 IEEE 40th Inter-
national Conference on Data Engineering (ICDE), pages 70–82. IEEE. DOI:
10.1109/ICDE60146.2024.00013.
|
|