Double awareness mechanism based deep learning framework for image captioning
*GauravCorresponding authorgauravsingla31@gmail.comDepartment of Information TechnologyManipal University JaipurJaipur, Rajasthan, IndiaView full profile → , Pratistha Mathurpratistha.mathur@jaipur.manipal.eduDepartment of Information TechnologyManipal University JaipurJaipur, Rajasthan, IndiaView full profile →
* Corresponding author · click or hover a name for details
- Received:
- 03 Feb 2023
- Published Online:
- 20 Sep 2023
- Article type:
- Research Article
- Language:
- EN
- Article no.:
- JDMSC-1728
- Pages:
- 1801–1817
Abstract
Keywords
Subject Classifications
References
[1] M. D. Zakir Hossain, F. Sohel, M. F. Shiratuddin, and H. Laga, “A comprehensive survey of deep learning for image captioning,” ACM Comput. Surv., vol. 51, no. 6 (2019), doi: 10.1145/3295748.
[2] S. Kalra and A. Leekha, “Survey of convolutional neural networks for image captioning,” J. Inf. Optim. Sci., vol. 41, no. 1, pp. 239–260 (2020), doi: 10.1080/02522667.2020.1715602.
[3] O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” Proc. IEEE Comput. Soc. Conf. Comput. Vis. Pattern Recognit., vol. 07-12-June, pp. 3156–3164 (2015), doi: 10.1109/CVPR.2015.7298935.
[4] X. Liu, Q. Xu, and N. Wang, “A survey on deep neural network-based image captioning,” Vis. Comput., vol. 35, no. 3, pp. 445–470 (2019), doi: 10.1007/s00371-018-1566-y.
[5] J. Wang, W. Wang, L. Wang, Z. Wang, D. D. Feng, and T. Tan, “Learning visual relationship and context-aware attention for image captioning,” Pattern Recognit., vol. 98, p. 107075 (2020), doi: 10.1016/j.patcog.2019.107075.
[6] X. Xiao, L. Wang, K. Ding, S. Xiang, and C. Pan, “Deep Hierarchical Encoder-Decoder Network for Image Captioning,” IEEE Trans. Multimed., vol. 21, no. 11, pp. 2942–2956 (2019), doi: 10.1109/TMM.2019.2915033.
[7] M. Liu, L. Li, H. Hu, W. Guan, and J. Tian, “Image caption generation with dual attention mechanism,” Inf. Process. Manag., vol. 57, no. 2, p. 102178 (2020), doi: 10.1016/j.ipm.2019.102178.
[8] J. Liu, K. Cheng, H. Jin, and Z. Wu, “An Image Captioning Algorithm Based on Combination Attention Mechanism,” Electron., vol. 11, no. 9 (2022), doi: 10.3390/electronics11091397.
[9] P. Anderson et al., “Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering,” Proc. IEEE Comput. Soc. Conf. Comput. Vis. Pattern Recognit., pp. 6077–6086 (2018), doi: 10.1109/CVPR.2018.00636.
[10] H. Sharma, M. Agrahari, S. K. Singh, M. Firoj, and R. K. Mishra, “Image Captioning: A Comprehensive Survey,” 2020 Int. Conf. Power Electron. IoT Appl. Renew. Energy its Control. PARC 2020, pp. 325–328 (2020), doi: 10.1109/PARC49193.2020.236619.
[11] S. Katiyar and S. K. Borgohain, “Image Captioning using Deep Stacked LSTMs, Contextual Word Embeddings and Data Augmentation” (2021), [Online]. Available: http://arxiv.org/abs/2102.11237.
[12] Y. H. Chang, Y. J. Chen, R. H. Huang, and Y. T. Yu, “Enhanced image captioning with color recognition using deep learning methods,” Appl. Sci., vol. 12, no. 1 (2022), doi: 10.3390/app12010209.
[13] J. H. Lim, C. S. Chan, K. W. Ng, L. Fan, and Q. Yang, “Protect, show, attend and tell: Empowering image captioning models with ownership protection,” Pattern Recognit., vol. 122, p. 108285 (2022), doi: 10.1016/j.patcog.2021.108285.
[14] X. Dong, C. Long, W. Xu, and C. Xiao, “Dual Graph Convolutional Networks with Transformer and Curriculum Learning for Image Captioning,” MM 2021 - Proc. 29th ACM Int. Conf. Multimed., pp. 2615–2624 (2021), doi: 10.1145/3474085.3475439.
[15] E. Cetinic, “Iconographic Image Captioning for Artworks,” Lect. Notes Comput. Sci. (including Subser. Lect. Notes Artif. Intell. Lect. Notes Bioinformatics), vol. 12663 LNCS, pp. 502–516 (2021), doi: 10.1007/978-3-030-68796-0_36.
[16] J. Ji et al., “Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer Network” (2020), [Online]. Available: http://arxiv.org/abs/2012.07061.
[17] S. Dubey, F. Olimov, M. A. Rafique, J. Kim, and M. Jeon, “Label-Attention Transformer with Geometrically Coherent Objects for Image Captioning,” pp. 1–14 (2021), [Online]. Available: http://arxiv.org/abs/2109.07799.
[18] A. U. Haque, S. Ghani, and M. Saeed, “Image Captioning with Positional and Geometrical Semantics,” IEEE Access, vol. 9, pp. 160917–160925 (2021), doi: 10.1109/ACCESS.2021.3131343.
[19] M. Kalimuthu, A. Mogadala, M. Mosbach, and D. Klakow, Fusion Models for Improved Image Captioning, vol. 12666 LNCS. Springer International Publishing (2021).
[20] K. Iwamura, J. Y. L. Kasahara, A. Moro, A. Yamashita, and H. Asama, “Potential of Incorporating Motion Estimation for Image Captioning,” 2021 IEEE/SICE Int. Symp. Syst. Integr. SII 2021, pp. 23–28 (2021), doi: 10.1109/IEEECONF49454.2021.9382725.
[21] J. Zhang, K. Li, Z. Wang, X. Zhao, and Z. Wang, “Visual enhanced gLSTM for image captioning,” Expert Syst. Appl., vol. 184, no. June, p. 115462 (2021), doi: 10.1016/j.eswa.2021.115462.
[22] I. Azhar, I. Afyouni, and A. Elnagar, “Facilitated deep learning models for image captioning,” 2021 55th Annu. Conf. Inf. Sci. Syst. CISS 2021 (2021), doi: 10.1109/CISS50987.2021.9400209.
[23] B. Wan, W. Jiang, Y. M. Fang, M. Zhu, Q. Li, and Y. Liu, “Revisiting image captioning via maximum discrepancy competition,” Pattern Recognit., vol. 122, p. 108358 (2022), doi: 10.1016/j.patcog.2021.108358.
[24] Gaurav and P. Mathur, “A Survey on Various Deep Learning Models for Automatic Image Captioning,” J. Phys. Conf. Ser., vol. 1950, no. 1 (2021), doi: 10.1088/1742-6596/1950/1/012045.
[25] Gaurav and P. Mathur, “An Attention Mechanism and GRU Based Deep Learning Model for Automatic Image Captioning,” Int. J. Eng. Trends Technol., vol. 70, no. 3, pp. 302–309 (2022), doi: 10.14445/22315381/IJETT-V70I3P234.
[26] Z. Yang and Q. Liu, “ATT-BM-SOM: A Framework of Effectively Choosing Image Information and Optimizing Syntax for Image Captioning,” IEEE Access, vol. 8, pp. 50565–50573 (2020), doi: 10.1109/ACCESS.2020.2980578.
[27] Z. Deng, Z. Jiang, R. Lan, W. Huang, and X. Luo, “Image captioning using DenseNet network and adaptive attention,” Signal Process. Image Commun., vol. 85, p. 115836 (2020), doi: 10.1016/j.image.2020.115836.
[28] X. Zhang, S. He, X. Song, R. W. H. Lau, J. Jiao, and Q. Ye, “Image captioning via semantic element embedding,” Neurocomputing, vol. 395, no. xxxx, pp. 212–221 (2020), doi: 10.1016/j.neucom.2018.02.112.
[29] J. Liu et al., “Interactive dual generative adversarial networks for image captioning,” AAAI 2020 - 34th AAAI Conf. Artif. Intell., pp. 11588–11595 (2020), doi: 10.1609/aaai.v34i07.6826.
[30] M. Chohan, A. Khan, M. S. Mahar, S. Hassan, A. Ghafoor, and M. Khan, “Image captioning using deep learning: A systematic literature review,” Int. J. Adv. Comput. Sci. Appl., vol. 11, no. 5, pp. 278–286 (2020), doi: 10.14569/IJACSA.2020.0110537.
[31] D. H. Fudholi, A. Zahra, and R. A. N. Nayoan, “A Study on Visual Understanding Image Captioning using Different Word Embeddings and CNN-Based Feature Extractions,” Kinet. Game Technol. Inf. Syst. Comput. Network, Comput. Electron. Control, vol. 4, no. 1, pp. 91–98 (2022), doi: 10.22219/kinetik.v7i1.1394.
[32] U. Sirisha and B. Sai Chandana, “Semantic interdisciplinary evaluation of image captioning models,” Cogent Eng., vol. 9, no. 1 (2022), doi: 10.1080/23311916.2022.2104333.
[33] D. Kumar, V. Srivastava, D. E. Popescu, and J. D. Hemanth, “Dual-Modal Transformer with Enhanced Inter-and Intra-Modality Interactions for Image Captioning,” Appl. Sci., vol. 12, no. 13 (2022), doi: 10.3390/app12136733.
[34] P. Girdhar, P. Johri, and D. Virmani, “Incept_LSTM : Accession for human activity concession in automatic surveillance,” J. Discret. Math. Sci. Cryptogr., vol. 25, no. 8, pp. 2259–2273 (2022), doi: 10.1080/09720529.2020.1804132.
[35] M. Rajani Shree and B. R. Shambhavi, “POS tagger model for Kannada text with CRF++ and deep learning approaches,” J. Discret. Math. Sci. Cryptogr., vol. 23, no. 2, pp. 485–493 (2020), doi: 10.1080/09720529.2020.1728902.
[36] Kumar, A., Singh, K., & Khan, T., L-RTAM: Logarithm based reliable trust assessment model for WBSNs. Journal of Discrete Mathematical Sciences and Cryptography, 24(6), 1701-1716 (2021).
[37] Khan, T., & Singh, K., Resource management based secure trust model for WSN. Journal of Discrete Mathematical Sciences and Cryptography, 22(8), 1453-1462 (2019).
[38] Umamaheswaran, S., John, R., Deepthi, S. S., & Dharmarajlu, S. M., Caption positioning structure for hard of hearing people using deep learning method. Journal of Discrete Mathematical Sciences and Cryptography, 25(3), 623-633 (2022).
[39] Amrutha, K., & Prabu, P., Effortless and beneficial processing of natural languages using transformers. Journal of Discrete Mathematical Sciences and Cryptography, 25(7), 1987-2005 (2022).
[40] Hore, P., & Sharma, A., Code-switched end-to-end Marathi speech recognition for especially abled people. Journal of Discrete Mathematical Sciences and Cryptography, 25(3), 771-784 (2022).




