TARU PUBLICATIONS
Journal of Information and Optimization Sciences cover
Hybrid ·Peer-reviewed·ISSN (Online): 2169-0103·ISSN (Print): 0252-2667

WoS  JIF 2026 : 0.4 (Q4)

Powered by:Powered by

Monthly Journal: Publishes theoretical and applied research on topics in information and optimization sciences.

Issues up to 2022 co-published with and available at:Taylor & Francis
submissions@tarupublications.com
Open Access Research Article

Comprehensive review and analysis on multi modal image retrieval

* , ,

* Corresponding author · click or hover a name for details

pp. 403–413Vol. 46Issue 2March 2025DOI: 10.47974/JIOS-1923XML
Received:
12 Nov 2024
Published Online:
17 Mar 2025
Article type:
Research Article
Language:
EN
Article no.:
JIOS-1923
Pages:
403–413

Abstract

Multimodal image retrieval, which involves retrieving images using various modalities such as text, audio, or other images, has significant research importance due to its wide-ranging applications and potential to enhance user experiences across multiple domains. Traditional image retrieval systems, content-based image retrieval (CBIR) systems rely solely on visual features, which can be limiting. By integrating multiple modalities, multimodal retrieval systems can offer more robust and accurate results. A research problem is considered multimodal when it integrates information from more than one type of data source. In multimodal image retrieval (MMIR) systems, one form of data is used to search for outcomes in the same or different modalities. In contrast, cross-modal systems strictly retrieve information from a different modality. A significant challenge in these systems is the effective comparison of input-output queries from different data types, due to their basic forms and the subjective nature of content similarity. Researchers have proposed various techniques to address this challenge and to bridge the semantic gap in information retrieval across different modalities. This article presents comprehensive analysis of various research works pertained in this multimodal image retrieval. The results and comparative analysis of current research works on benchmark datasets have also been discussed. At the end of the paper, some of open issues are presented for future research directions.

Keywords

Subject Classifications

68P30

References

[1] A. Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems, vol. 30, p. I (2017).
[2] T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 41, no. 2, pp. 423-443 (2018).
[3] P. Kaur, H. S. Pannu, and A. K. Malhi, “Comparative analysis on cross-modal information retrieval: A review,” Computer Science Review, vol. 39, p. 100336 (2021).
[4] Y. Huo, Q. Qin, J. Dai, L. Wang, W. Zhang, L. Huang, and C. Wang, Deep semantic-aware proxy hashing for multi-label cross-modal retrieval,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 1, pp. 576-589 (2023).
[5] F. Feng, X. Wang, and R. Li, “Cross-modal retrieval with correspondence autoencoder,” Proceedings of the 22nd ACM International Conference on Multimedia (2014).
[6] K. Zhou, F. H. Hassan, and G. K. Hoon, “The state of the art for cross-modal retrieval: A survey,” IEEE Access (2023).
[7] C. Zhang, J. Song, X. Zhu, L. Zhu, and S. Zhang, “HCMSL: Hybrid cross-modal similarity learning for cross-modal retrieval,” ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), vol. 17, no. 1s, pp. 1-22 (2021).
[8] L. I. U. Ying, G. U. O. Yingying, J. F. A. N. G., J. F. A. N., Y. H. A. O., and J. L. I. U., “Survey of research on deep learning image-text cross-modal retrieval,” Journal of Frontiers of Computer Science & Technology, vol. 16, no. 3 (2022).
[9] J. Chen, L. Zhang, C. Bai, and K. Kpalma, “Review of recent deep learning-based methods for image-text retrieval,” Proceedings of the 2020 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR), pp. 167-172, Aug. (2020).
[10] Z. Li, H. Lu, H. Fu, and G. Gu, “Image-text bidirectional learning network based cross-modal retrieval,” Neurocomputing, vol. 483, pp. 148-159 (2022).
[11] Z. Yuan, W. Zhang, C. Tian, X. Rong, Z. Zhang, H. Wang, and X. Sun, “Remote sensing cross-modal text-image retrieval based on global and local information,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-16 (2022).
[12] X. Tang, Y. Wang, J. Ma, X. Zhang, F. Liu, and L. Jiao, “Interacting-enhancing feature transformer for cross-modal remote-sensing image and text retrieval,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1-15 (2023).
[13] Q. Cheng, Y. Zhou, P. Fu, Y. Xu, and L. Zhang, “A deep semantic alignment network for the cross-modal image-text retrieval in remote sensing,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 14, pp. 4284-4297 (2021).
[14] G. Sucharitha, N. Arora, and S. C. Sharma, “Medical image retrieval using a novel local relative directional edge pattern and Zernike moments,” Multimedia Tools and Applications, vol. 82, no. 20, pp. 31737-31757 (2023).
[15] W. Zhou, H. Li, and Q. Tian, “Recent advance in content-based image retrieval: A literature survey,” arXiv preprint arXiv:1706.06064 (2017).
[16] G. Sucharitha and R. K. Senapati, “Local quantized edge binary patterns for colour texture image retrieval,” Journal of Theoretical & Applied Information Technology, vol. 96, no. 2 (2018).
[17] I. M. Hameed, S. H. Abdulhussain, and B. M. Mahmmod, “Content-based image retrieval: A review of recent trends,” Cogent Engineering, vol. 8, no. 1, p. 1927469 (2021).
[18] X. Li, J. Yang, and J. Ma, “Recent developments of content-based image retrieval (CBIR),” Neurocomputing, vol. 452, pp. 675-689 (2021).
[19] M. Alrahhal and K. P. Supreethi, “Content-Based Image Retrieval using Local Patterns and Supervised Machine Learning Techniques,” 2019 Amity International Conference on Artificial Intelligence (AICAI), Dubai, United Arab Emirates, pp. 118-124 (2019), doi: 10.1109/AICAI.2019.8701255.
[20] S. Qian, Y. Guo, J. Liu, X. Zhang, and Q. Wu, “Adaptive label-aware graph convolutional networks for cross-modal retrieval,” IEEE Transactions on Multimedia, vol. 24, pp. 3520-3532 (2021).
[21] N. Arora, G. Sucharitha, and S. C. Sharma, “MVM-LBP: Mean−Variance−Median based LBP for face recognition,” International Journal of Information Technology (2023).
[22] M. Hosseinzadeh and Y. Wang, “Composed query image retrieval using locally bounded features,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2020).
[23] D. Xie, X. Zhang, Y. Li, J. Chen, L. Wang, and W. Li, “Multi-task consistency-preserving adversarial hashing for cross-modal retrieval,” IEEE Transactions on Image Processing, vol. 29, pp. 3626-3637 (2020).
[24] F. Zhang, M. Xu, Q. Mao, and C. Xu, “Joint attribute manipulation and modality alignment learning for composing text and image to image retrieval,” Proceedings of the 28th ACM International Conference on Multimedia, pp. 3367-3376 (2020).
[25] S. Lee, D. Kim, and B. Han, “Cosmo: Content-style modulation for image retrieval with text feedback,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2021).
[26] S. Goenka, Z. Zheng, A. Jaiswal, R. Chada, Y. Wu, V. Hedau, and P. Natarajan, “FashionVLP: Vision language transformer for fashion retrieval with feedback,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14105-14115 (2022).
[27] N. Perveen, D. Roy, and C. K. Mohan, “Spontaneous expression recognition using universal attribute model,” IEEE Transactions on Image Processing, vol. 27, no. 11, pp. 5575-5584 (2018).
[28] M. Paolanti, C. Kaiser, R. Schallner, E. Frontoni, and P. Zingaretti, “Visual and textual sentiment analysis of brand-related social media pictures using deep convolutional neural networks,” in Image Analysis and Processing-ICIAP 2017: 19th International Conference, Catania, Italy, Sep. 11-15, Part I, Springer International Publishing, pp. 402-413 (2017).
[29] N. Medagoda, S. Shanmuganathan, and J. Whalley, “Sentiment lexicon construction using SentiWordNet 3.0,” in 2015 11th International Conference on Natural Computation (ICNC), IEEE (2015).
[30] Q. You, J. Luo, H. Jin, and J. Yang, “Cross-modality consistent regression for joint visual-textual sentiment analysis of social multimedia,” in Proceedings of the Ninth ACM International Conference on Web Search and Data Mining, pp. 13-22 (2016).
[31] D. Singh and C. K. Mohan, “Deep spatio-temporal representation for detection of road accidents using stacked autoencoder,” IEEE Transactions on Intelligent Transportation Systems, vol. 20, no. 3, pp. 879-887 (2019). doi: 10.1109/TMM.2018.2887021.
[32] J. Xu, F. Huang, X. Zhang, S. Wang, C. Li, Z. Li, and Y. He, “Visual-textual sentiment classification with bi-directional multi-level attention networks,” Knowledge-Based Systems, vol. 178, pp. 61-73 (2019).
[33] T. Zhou, H. Fu, G. Chen, J. Shen, and L. Shao, “Hi-net: Hybrid-fusion network for multi-modal MR image synthesis,” IEEE Transactions on Medical Imaging, vol. 39, no. 9, pp. 2772-2781 (2020).
[34] E. P. Ijjina and C. K. Mohan, “Human action recognition in RGB-D using motion sequence and deep learning,” Pattern Recognition, vol. 72, pp. 504-516 (2017). doi: 10.1016/j.patcog.2017.07.01.

Views: 341Downloads: 85Citations: 0