TARU PUBLICATIONS
Journal of Discrete Mathematical Sciences and Cryptography cover
Hybrid ·Peer-reviewed·ISSN (Online): 2169-0065·ISSN (Print): 0972-0529

Monthly Journal: Publishes theoretical and applied research in all areas of Discrete Mathematical Sciences, Cryptography, Combinatorics, Elliptic Curves and Information Security.

Issues up to 2022 co-published with and available at:Taylor & Francis Online
submissions@tarupublications.com
Open Access Research Article

Double awareness mechanism based deep learning framework for image captioning

* ,

* Corresponding author · click or hover a name for details

pp. 1801–1817Vol. 26Issue 6September 2023DOI: 10.47974/JDMSC-1728 Crossmark XML
Received:
03 Feb 2023
Published Online:
20 Sep 2023
Article type:
Research Article
Language:
EN
Article no.:
JDMSC-1728
Pages:
1801–1817

Abstract

Understanding the qualities of an image and converting them into a phrase or sentence that makes sense is the process of image captioning. Neuroscience research has only recently made clear the connection between human vision and language formation. Although there have been many methods for captioning images, including content retrieval and template filling, the current trend is toward deep learning-based methods. Using an image encoder, feature vectors are created from an image through the deep learning process, and a language decoder converts these feature vectors into a string of words. Using encoder-decoder approach or simple attention based has not provided so much efficient results. In the proposed model, double awareness-based mechanism has been used. The primary goal of this study is to extract visual properties from the region of interest (RoI) of an image as well as text features using the glove embedding technique. Inception ResNet version of Convolutional neural network (CNN) is used as an encoder. As a decoder, a gated recurrent unit is employed. The proposed model is tested on Flickr8k dataset and it can be seen that the results achieved through double awareness mechanism are highly effective.

Keywords

Subject Classifications

68T07 Artificial neural networks and deep learning

References

[1] M. D. Zakir Hossain, F. Sohel, M. F. Shiratuddin, and H. Laga, “A comprehensive survey of deep learning for image captioning,” ACM Comput. Surv., vol. 51, no. 6 (2019), doi: 10.1145/3295748.
[2] S. Kalra and A. Leekha, “Survey of convolutional neural networks for image captioning,” J. Inf. Optim. Sci., vol. 41, no. 1, pp. 239–260 (2020), doi: 10.1080/02522667.2020.1715602.
[3] O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” Proc. IEEE Comput. Soc. Conf. Comput. Vis. Pattern Recognit., vol. 07-12-June, pp. 3156–3164 (2015), doi: 10.1109/CVPR.2015.7298935.
[4] X. Liu, Q. Xu, and N. Wang, “A survey on deep neural network-based image captioning,” Vis. Comput., vol. 35, no. 3, pp. 445–470 (2019), doi: 10.1007/s00371-018-1566-y.
[5] J. Wang, W. Wang, L. Wang, Z. Wang, D. D. Feng, and T. Tan, “Learning visual relationship and context-aware attention for image captioning,” Pattern Recognit., vol. 98, p. 107075 (2020), doi: 10.1016/j.patcog.2019.107075.
[6] X. Xiao, L. Wang, K. Ding, S. Xiang, and C. Pan, “Deep Hierarchical Encoder-Decoder Network for Image Captioning,” IEEE Trans. Multimed., vol. 21, no. 11, pp. 2942–2956 (2019), doi: 10.1109/TMM.2019.2915033.
[7] M. Liu, L. Li, H. Hu, W. Guan, and J. Tian, “Image caption generation with dual attention mechanism,” Inf. Process. Manag., vol. 57, no. 2, p. 102178 (2020), doi: 10.1016/j.ipm.2019.102178.
[8] J. Liu, K. Cheng, H. Jin, and Z. Wu, “An Image Captioning Algorithm Based on Combination Attention Mechanism,” Electron., vol. 11, no. 9 (2022), doi: 10.3390/electronics11091397.
[9] P. Anderson et al., “Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering,” Proc. IEEE Comput. Soc. Conf. Comput. Vis. Pattern Recognit., pp. 6077–6086 (2018), doi: 10.1109/CVPR.2018.00636.
[10] H. Sharma, M. Agrahari, S. K. Singh, M. Firoj, and R. K. Mishra, “Image Captioning: A Comprehensive Survey,” 2020 Int. Conf. Power Electron. IoT Appl. Renew. Energy its Control. PARC 2020, pp. 325–328 (2020), doi: 10.1109/PARC49193.2020.236619.
[11] S. Katiyar and S. K. Borgohain, “Image Captioning using Deep Stacked LSTMs, Contextual Word Embeddings and Data Augmentation” (2021), [Online]. Available: http://arxiv.org/abs/2102.11237.
[12] Y. H. Chang, Y. J. Chen, R. H. Huang, and Y. T. Yu, “Enhanced image captioning with color recognition using deep learning methods,” Appl. Sci., vol. 12, no. 1 (2022), doi: 10.3390/app12010209.
[13] J. H. Lim, C. S. Chan, K. W. Ng, L. Fan, and Q. Yang, “Protect, show, attend and tell: Empowering image captioning models with ownership protection,” Pattern Recognit., vol. 122, p. 108285 (2022), doi: 10.1016/j.patcog.2021.108285.
[14] X. Dong, C. Long, W. Xu, and C. Xiao, “Dual Graph Convolutional Networks with Transformer and Curriculum Learning for Image Captioning,” MM 2021 - Proc. 29th ACM Int. Conf. Multimed., pp. 2615–2624 (2021), doi: 10.1145/3474085.3475439.
[15] E. Cetinic, “Iconographic Image Captioning for Artworks,” Lect. Notes Comput. Sci. (including Subser. Lect. Notes Artif. Intell. Lect. Notes Bioinformatics), vol. 12663 LNCS, pp. 502–516 (2021), doi: 10.1007/978-3-030-68796-0_36.
[16] J. Ji et al., “Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer Network” (2020), [Online]. Available: http://arxiv.org/abs/2012.07061.
[17] S. Dubey, F. Olimov, M. A. Rafique, J. Kim, and M. Jeon, “Label-Attention Transformer with Geometrically Coherent Objects for Image Captioning,” pp. 1–14 (2021), [Online]. Available: http://arxiv.org/abs/2109.07799.
[18] A. U. Haque, S. Ghani, and M. Saeed, “Image Captioning with Positional and Geometrical Semantics,” IEEE Access, vol. 9, pp. 160917–160925 (2021), doi: 10.1109/ACCESS.2021.3131343.
[19] M. Kalimuthu, A. Mogadala, M. Mosbach, and D. Klakow, Fusion Models for Improved Image Captioning, vol. 12666 LNCS. Springer International Publishing (2021).
[20] K. Iwamura, J. Y. L. Kasahara, A. Moro, A. Yamashita, and H. Asama, “Potential of Incorporating Motion Estimation for Image Captioning,” 2021 IEEE/SICE Int. Symp. Syst. Integr. SII 2021, pp. 23–28 (2021), doi: 10.1109/IEEECONF49454.2021.9382725.
[21] J. Zhang, K. Li, Z. Wang, X. Zhao, and Z. Wang, “Visual enhanced gLSTM for image captioning,” Expert Syst. Appl., vol. 184, no. June, p. 115462 (2021), doi: 10.1016/j.eswa.2021.115462.
[22] I. Azhar, I. Afyouni, and A. Elnagar, “Facilitated deep learning models for image captioning,” 2021 55th Annu. Conf. Inf. Sci. Syst. CISS 2021 (2021), doi: 10.1109/CISS50987.2021.9400209.
[23] B. Wan, W. Jiang, Y. M. Fang, M. Zhu, Q. Li, and Y. Liu, “Revisiting image captioning via maximum discrepancy competition,” Pattern Recognit., vol. 122, p. 108358 (2022), doi: 10.1016/j.patcog.2021.108358.
[24] Gaurav and P. Mathur, “A Survey on Various Deep Learning Models for Automatic Image Captioning,” J. Phys. Conf. Ser., vol. 1950, no. 1 (2021), doi: 10.1088/1742-6596/1950/1/012045.
[25] Gaurav and P. Mathur, “An Attention Mechanism and GRU Based Deep Learning Model for Automatic Image Captioning,” Int. J. Eng. Trends Technol., vol. 70, no. 3, pp. 302–309 (2022), doi: 10.14445/22315381/IJETT-V70I3P234.
[26] Z. Yang and Q. Liu, “ATT-BM-SOM: A Framework of Effectively Choosing Image Information and Optimizing Syntax for Image Captioning,” IEEE Access, vol. 8, pp. 50565–50573 (2020), doi: 10.1109/ACCESS.2020.2980578.
[27] Z. Deng, Z. Jiang, R. Lan, W. Huang, and X. Luo, “Image captioning using DenseNet network and adaptive attention,” Signal Process. Image Commun., vol. 85, p. 115836 (2020), doi: 10.1016/j.image.2020.115836.
[28] X. Zhang, S. He, X. Song, R. W. H. Lau, J. Jiao, and Q. Ye, “Image captioning via semantic element embedding,” Neurocomputing, vol. 395, no. xxxx, pp. 212–221 (2020), doi: 10.1016/j.neucom.2018.02.112.
[29] J. Liu et al., “Interactive dual generative adversarial networks for image captioning,” AAAI 2020 - 34th AAAI Conf. Artif. Intell., pp. 11588–11595 (2020), doi: 10.1609/aaai.v34i07.6826.
[30] M. Chohan, A. Khan, M. S. Mahar, S. Hassan, A. Ghafoor, and M. Khan, “Image captioning using deep learning: A systematic literature review,” Int. J. Adv. Comput. Sci. Appl., vol. 11, no. 5, pp. 278–286 (2020), doi: 10.14569/IJACSA.2020.0110537.
[31] D. H. Fudholi, A. Zahra, and R. A. N. Nayoan, “A Study on Visual Understanding Image Captioning using Different Word Embeddings and CNN-Based Feature Extractions,” Kinet. Game Technol. Inf. Syst. Comput. Network, Comput. Electron. Control, vol. 4, no. 1, pp. 91–98 (2022), doi: 10.22219/kinetik.v7i1.1394.
[32] U. Sirisha and B. Sai Chandana, “Semantic interdisciplinary evaluation of image captioning models,” Cogent Eng., vol. 9, no. 1 (2022), doi: 10.1080/23311916.2022.2104333.
[33] D. Kumar, V. Srivastava, D. E. Popescu, and J. D. Hemanth, “Dual-Modal Transformer with Enhanced Inter-and Intra-Modality Interactions for Image Captioning,” Appl. Sci., vol. 12, no. 13 (2022), doi: 10.3390/app12136733.
[34] P. Girdhar, P. Johri, and D. Virmani, “Incept_LSTM : Accession for human activity concession in automatic surveillance,” J. Discret. Math. Sci. Cryptogr., vol. 25, no. 8, pp. 2259–2273 (2022), doi: 10.1080/09720529.2020.1804132.
[35] M. Rajani Shree and B. R. Shambhavi, “POS tagger model for Kannada text with CRF++ and deep learning approaches,” J. Discret. Math. Sci. Cryptogr., vol. 23, no. 2, pp. 485–493 (2020), doi: 10.1080/09720529.2020.1728902.
[36] Kumar, A., Singh, K., & Khan, T., L-RTAM: Logarithm based reliable trust assessment model for WBSNs. Journal of Discrete Mathematical Sciences and Cryptography, 24(6), 1701-1716 (2021).
[37] Khan, T., & Singh, K., Resource management based secure trust model for WSN. Journal of Discrete Mathematical Sciences and Cryptography, 22(8), 1453-1462 (2019).
[38] Umamaheswaran, S., John, R., Deepthi, S. S., & Dharmarajlu, S. M., Caption positioning structure for hard of hearing people using deep learning method. Journal of Discrete Mathematical Sciences and Cryptography, 25(3), 623-633 (2022).
[39] Amrutha, K., & Prabu, P., Effortless and beneficial processing of natural languages using transformers. Journal of Discrete Mathematical Sciences and Cryptography, 25(7), 1987-2005 (2022).
[40] Hore, P., & Sharma, A., Code-switched end-to-end Marathi speech recognition for especially abled people. Journal of Discrete Mathematical Sciences and Cryptography, 25(3), 771-784 (2022).

Views: 246Downloads: 81Citations: 0