TARU PUBLICATIONS
Journal of Information and Optimization Sciences cover
Hybrid ·Peer-reviewed·ISSN (Online): 2169-0103·ISSN (Print): 0252-2667

WoS  JIF 2026 : 0.4 (Q4)

Powered by:Powered by

Monthly Journal: Publishes theoretical and applied research on topics in information and optimization sciences.

Issues up to 2022 co-published with and available at:Taylor & Francis
submissions@tarupublications.com
Open Access Research Article

Optimized approach for efficient image retrieval using Resnet50 and CBAM

* , ,

* Corresponding author · click or hover a name for details

pp. 415–425Vol. 46Issue 2March 2025DOI: 10.47974/JIOS-1924XML
Received:
13 Nov 2024
Published Online:
17 Mar 2025
Article type:
Research Article
Language:
EN
Article no.:
JIOS-1924
Pages:
415–425

Abstract

The rapid expansion of multimedia repositories surpasses the capabilities of traditional data analysis. An advanced image retrieval (IR) model is essential for accurately retrieving relevant images. Numerous machine learning and deep learning models have been proposed to address this challenge effectively. Even though the deep learning models are so powerful in extracting the image features, they are getting failed in prioritizing the most informative parts of the feature maps and also ignored the focus on specific channels or spatial regions. This will lead to the less effective feature representations of images, which contain varying levels of importance across different regions or channels. As a result, the retrieval performance may suffer because the model may not effectively differentiate between similar images. To address this point, this paper concentrated on refining the feature maps by applying channel and spatial attention, which helps in highlighting the most informative features and suppressing less important ones. This has become possible by integrating the Convolutional block of attention module (CBAM) and ResNet-50 models. CBAM is integrated with ResNet-50 by adding attention modules after each residual block. These modules apply channel and spatial attention mechanisms to the feature maps, refining them before passing to the next layer. The proposed model is verified on benchmark databases Corel-10k, CIFAR. The significant performance of the proposed method shown over existing deep learning models.

Keywords

Subject Classifications

Primary 93A30Secondary 49K15

References

[1] D. Jiang and J. Kim, “Texture image retrieval using DTCWT-SVD and local binary pattern features,” Journal of Information Processing Systems, vol. 13, no. 6, pp. 1628–1639 (2017).
[2] G. Sucharitha and R. K. Senapati, “Biomedical image retrieval by using local directional edge binary patterns and Zernike moments,” Multimedia Tools and Applications, vol. 79, no. 3, pp. 1847–1864 (2020).
[3] Z. Zhu, C. Zhao, and Y. Hou, “Research on similarity measurement for texture image retrieval,” e45302 (2012).
[4] G. Sucharitha and R. K. Senapati, “Local extreme edge binary patterns for face recognition and image retrieval,” Journal of Advanced Research in Dynamical and Control Systems, vol. 10, pp. 644–654 (2018).
[5] D. Sudarvizhi, “Feature based image retrieval system using Zernike moments and Daubechies Wavelet Transform,” in 2016 International Conference on Recent Trends in Information Technology (ICRTIT), IEEE (2016).
[6] G. Sucharitha, R. K. Senapati, and A. B. Ranjan, “Secure and efficient content-based image retrieval using dominant local patterns and watermark encryption in cloud computing,” Cluster Computing, pp. 1–17 (2024).
[7] M. S. Ghaleb, A. H. Rahman, and L. C. Park, “Image retrieval based on deep learning,” Journal of System and Management Sciences, vol. 12, no. 2, pp. 477–496 (2022).
[8] S. R. Dubey, “A decade survey of content-based image retrieval using deep learning,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 5, pp. 2687–2704 (2021).
[9] Z. Rian, V. Christanti, and J. Hendryli, “Content-based image retrieval using convolutional neural networks,” in 2019 IEEE International Conference on Signals and Systems (ICSigSys), IEEE (2019).
[10] T.-Y. Lin, A. RoyChowdhury, and S. Maji, “Bilinear convolutional neural networks for fine-grained visual recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 6, pp. 1309–1322 (2017).
[11] M. S. Basha, S. K. Mouleeswaran, and K. Rajendra Prasad, “Hybrid visual computing models to discover the clusters assessment of high dimensional big data,” Soft Comput, vol. 27, pp. 4249–4262 (2023), doi: 10.1007/s00500-022-07092-x.
[12] S. A. Vassou, G. I. Papadopoulos, and C. D. Styliadis, “CoMo: a compact composite moment-based descriptor for image retrieval,” Proceedings of the 15th International Workshop on Content-Based Multimedia Indexing (2017).
[13] C. Zhang and J. Liu, “Content Based Deep Learning Image Retrieval: A Survey,” Proceedings of the 2023 9th International Conference on Communication and Information Processing (2023).
[14] S. Reddy K. and K. Rajendra Prasad, “An Extended Fuzzy C-Means Segmentation for an Efficient BTD with the Region of Interest of SCP,” IJITPM, vol. 12, no. 4, pp. 11–24 (2021), doi: 10.4018/IJITPM.2021100.
[15] Y. Luo, S. Liu, and X. Zhang, “LSTM pose machines,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2018).
[16] C. Vishnu, P. Jeripothula, R. Datla, S. Babu Ch., and C. Krishna Mohan, “mSODANet: A Network for Multi-Scale Object Detection in Aerial Images using Hierarchical Dilated Convolutions,” Pattern Recognition, pp. 108548 (2022), doi: 10.1016/j.patcog.2022.108548.
[17] K. Subba Reddy, K. Rajendra Prasad, G. R. Kamatam, and R. K. Gupta, “An extended visual methods to perform data cluster assessment in distributed data systems,” J Supercomput, vol. 78, pp. 8810–8829 (2022), doi: 10.1007/s11227-021-04243-z.
[18] S. Wilson and C. Krishna Mohan, “An information bottleneck approach to optimize the dictionary of visual data,” IEEE Transactions on Multimedia, vol. 20, no. 1, pp. 96–106 (2018), doi: 10.1109/TMM.2017.2716835.
[19] N. Perveen, D. Roy, and C. Krishna Mohan, “Facial Expression Recognition in Videos using Dynamic Kernels,” IEEE Transactions on Image Processing, vol. 29, pp. 8316–8325 (2020), doi: 10.1109/TIP.2020.3011846.
[20] K. Shaheed, M. A. Yaseen, and M. A. Aslam, “EfficientRMT-Net—An Efficient ResNet-50 and Vision Transformers Approach for Classifying Potato Plant Leaf Diseases,” Sensors, vol. 23, no. 23, pp. 9516 (2023).
[21] S. Woo, J. Park, and I. Y. Choi, “CBAM: Convolutional block attention module,” Proceedings of the European Conference on Computer Vision (ECCV) (2018).
[22] L. Deng, Y. Li, Z. Zhan, and X. Xie, “Multi-level attention network: Mixed time–frequency channel attention and multi-scale self-attentive standard deviation pooling for speaker recognition,” Engineering Applications of Artificial Intelligence, vol. 128, p. 107439 (2024).
[23] Corel, “Corel 10K Dataset,” Corel Corporation (2004).
[24] A. Krizhevsky and G. Hinton, “Learning Multiple Layers of Features from Tiny,” [Online]. Available: https://www.cs.toronto.edu/~kriz/cifar.html.

Views: 191Downloads: 78Citations: 0