Comprehensive review and analysis on multi modal image retrieval
*Pilli MounikaCorresponding authormounika260328@gmail.comDepartment of Computer Science and EngineeringJawaharlal Nehru Technological UniversityKakinada, Andhra Pradesh, 533003, IndiaView full profile → , K. Venkata Subba Reddykvsreddy2012@gmail.comDepartment of Computer Science & Engineering (AI&ML)Vidya Jyothi Institute of TechnologyJawaharlal Nehru Technological UniversityHyderabad, Telangana, 500075, IndiaView full profile → , N. Ramakrishnaiahnrkrishna27@gmail.comDepartment of Computer Science and EngineeringUniversity College of EngineeringJawaharlal Nehru Technological UniversityKakinada, Andhra Pradesh, 533003, IndiaView full profile →
* Corresponding author · click or hover a name for details
- Received:
- 12 Nov 2024
- Published Online:
- 17 Mar 2025
- Article type:
- Research Article
- Language:
- EN
- Article no.:
- JIOS-1923
- Pages:
- 403–413
Abstract
Keywords
Subject Classifications
References
[1] A. Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems, vol. 30, p. I (2017).
[2] T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 41, no. 2, pp. 423-443 (2018).
[3] P. Kaur, H. S. Pannu, and A. K. Malhi, “Comparative analysis on cross-modal information retrieval: A review,” Computer Science Review, vol. 39, p. 100336 (2021).
[4] Y. Huo, Q. Qin, J. Dai, L. Wang, W. Zhang, L. Huang, and C. Wang, Deep semantic-aware proxy hashing for multi-label cross-modal retrieval,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 1, pp. 576-589 (2023).
[5] F. Feng, X. Wang, and R. Li, “Cross-modal retrieval with correspondence autoencoder,” Proceedings of the 22nd ACM International Conference on Multimedia (2014).
[6] K. Zhou, F. H. Hassan, and G. K. Hoon, “The state of the art for cross-modal retrieval: A survey,” IEEE Access (2023).
[7] C. Zhang, J. Song, X. Zhu, L. Zhu, and S. Zhang, “HCMSL: Hybrid cross-modal similarity learning for cross-modal retrieval,” ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), vol. 17, no. 1s, pp. 1-22 (2021).
[8] L. I. U. Ying, G. U. O. Yingying, J. F. A. N. G., J. F. A. N., Y. H. A. O., and J. L. I. U., “Survey of research on deep learning image-text cross-modal retrieval,” Journal of Frontiers of Computer Science & Technology, vol. 16, no. 3 (2022).
[9] J. Chen, L. Zhang, C. Bai, and K. Kpalma, “Review of recent deep learning-based methods for image-text retrieval,” Proceedings of the 2020 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR), pp. 167-172, Aug. (2020).
[10] Z. Li, H. Lu, H. Fu, and G. Gu, “Image-text bidirectional learning network based cross-modal retrieval,” Neurocomputing, vol. 483, pp. 148-159 (2022).
[11] Z. Yuan, W. Zhang, C. Tian, X. Rong, Z. Zhang, H. Wang, and X. Sun, “Remote sensing cross-modal text-image retrieval based on global and local information,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-16 (2022).
[12] X. Tang, Y. Wang, J. Ma, X. Zhang, F. Liu, and L. Jiao, “Interacting-enhancing feature transformer for cross-modal remote-sensing image and text retrieval,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1-15 (2023).
[13] Q. Cheng, Y. Zhou, P. Fu, Y. Xu, and L. Zhang, “A deep semantic alignment network for the cross-modal image-text retrieval in remote sensing,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 14, pp. 4284-4297 (2021).
[14] G. Sucharitha, N. Arora, and S. C. Sharma, “Medical image retrieval using a novel local relative directional edge pattern and Zernike moments,” Multimedia Tools and Applications, vol. 82, no. 20, pp. 31737-31757 (2023).
[15] W. Zhou, H. Li, and Q. Tian, “Recent advance in content-based image retrieval: A literature survey,” arXiv preprint arXiv:1706.06064 (2017).
[16] G. Sucharitha and R. K. Senapati, “Local quantized edge binary patterns for colour texture image retrieval,” Journal of Theoretical & Applied Information Technology, vol. 96, no. 2 (2018).
[17] I. M. Hameed, S. H. Abdulhussain, and B. M. Mahmmod, “Content-based image retrieval: A review of recent trends,” Cogent Engineering, vol. 8, no. 1, p. 1927469 (2021).
[18] X. Li, J. Yang, and J. Ma, “Recent developments of content-based image retrieval (CBIR),” Neurocomputing, vol. 452, pp. 675-689 (2021).
[19] M. Alrahhal and K. P. Supreethi, “Content-Based Image Retrieval using Local Patterns and Supervised Machine Learning Techniques,” 2019 Amity International Conference on Artificial Intelligence (AICAI), Dubai, United Arab Emirates, pp. 118-124 (2019), doi: 10.1109/AICAI.2019.8701255.
[20] S. Qian, Y. Guo, J. Liu, X. Zhang, and Q. Wu, “Adaptive label-aware graph convolutional networks for cross-modal retrieval,” IEEE Transactions on Multimedia, vol. 24, pp. 3520-3532 (2021).
[21] N. Arora, G. Sucharitha, and S. C. Sharma, “MVM-LBP: Mean−Variance−Median based LBP for face recognition,” International Journal of Information Technology (2023).
[22] M. Hosseinzadeh and Y. Wang, “Composed query image retrieval using locally bounded features,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2020).
[23] D. Xie, X. Zhang, Y. Li, J. Chen, L. Wang, and W. Li, “Multi-task consistency-preserving adversarial hashing for cross-modal retrieval,” IEEE Transactions on Image Processing, vol. 29, pp. 3626-3637 (2020).
[24] F. Zhang, M. Xu, Q. Mao, and C. Xu, “Joint attribute manipulation and modality alignment learning for composing text and image to image retrieval,” Proceedings of the 28th ACM International Conference on Multimedia, pp. 3367-3376 (2020).
[25] S. Lee, D. Kim, and B. Han, “Cosmo: Content-style modulation for image retrieval with text feedback,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2021).
[26] S. Goenka, Z. Zheng, A. Jaiswal, R. Chada, Y. Wu, V. Hedau, and P. Natarajan, “FashionVLP: Vision language transformer for fashion retrieval with feedback,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14105-14115 (2022).
[27] N. Perveen, D. Roy, and C. K. Mohan, “Spontaneous expression recognition using universal attribute model,” IEEE Transactions on Image Processing, vol. 27, no. 11, pp. 5575-5584 (2018).
[28] M. Paolanti, C. Kaiser, R. Schallner, E. Frontoni, and P. Zingaretti, “Visual and textual sentiment analysis of brand-related social media pictures using deep convolutional neural networks,” in Image Analysis and Processing-ICIAP 2017: 19th International Conference, Catania, Italy, Sep. 11-15, Part I, Springer International Publishing, pp. 402-413 (2017).
[29] N. Medagoda, S. Shanmuganathan, and J. Whalley, “Sentiment lexicon construction using SentiWordNet 3.0,” in 2015 11th International Conference on Natural Computation (ICNC), IEEE (2015).
[30] Q. You, J. Luo, H. Jin, and J. Yang, “Cross-modality consistent regression for joint visual-textual sentiment analysis of social multimedia,” in Proceedings of the Ninth ACM International Conference on Web Search and Data Mining, pp. 13-22 (2016).
[31] D. Singh and C. K. Mohan, “Deep spatio-temporal representation for detection of road accidents using stacked autoencoder,” IEEE Transactions on Intelligent Transportation Systems, vol. 20, no. 3, pp. 879-887 (2019). doi: 10.1109/TMM.2018.2887021.
[32] J. Xu, F. Huang, X. Zhang, S. Wang, C. Li, Z. Li, and Y. He, “Visual-textual sentiment classification with bi-directional multi-level attention networks,” Knowledge-Based Systems, vol. 178, pp. 61-73 (2019).
[33] T. Zhou, H. Fu, G. Chen, J. Shen, and L. Shao, “Hi-net: Hybrid-fusion network for multi-modal MR image synthesis,” IEEE Transactions on Medical Imaging, vol. 39, no. 9, pp. 2772-2781 (2020).
[34] E. P. Ijjina and C. K. Mohan, “Human action recognition in RGB-D using motion sequence and deep learning,” Pattern Recognition, vol. 72, pp. 504-516 (2017). doi: 10.1016/j.patcog.2017.07.01.




