Analysis of attention based deep learning approach for audio image descriptions
Achin Jainachin.mails@gmail.comDepartment of Information TechnologyBharati Vidyapeeth’s College of EngineeringPaschim Vihar, New Delhi, 110063, IndiaView full profile → , Sarita Yadavsarita1320@yahoo.co.inDepartment of Information TechnologyBharati Vidyapeeth’s College of EngineeringPaschim Vihar, New Delhi, 110063, IndiaView full profile → , Neetu Singhsinghneetu4bvcoe@gmail.comDepartment of Information TechnologyBharati Vidyapeeth’s College of EngineeringPaschim Vihar, New Delhi, 110063, IndiaView full profile → , Kajal Kaulkajalkaulphd@gmail.comDepartment of Information TechnologyBharati Vidyapeeth’s College of EngineeringPaschim Vihar, New Delhi, 110063, IndiaView full profile → , *Arun Kumar DubeyCorresponding authorarudubey@gmail.comDepartment of Information TechnologyBharati Vidyapeeth’s College of EngineeringPaschim Vihar, New Delhi, 110063, India0000-0002-6844-9213View full profile → , Prakhar Priyadarshiprakharpriya@gmail.comDepartment of Information TechnologyBharati Vidyapeeth’s College of EngineeringPaschim Vihar, New Delhi, 110063, IndiaView full profile → , Prabhav Sangaprabhav14.sanga@gmail.comDepartment of Information TechnologyBharati Vidyapeeth’s College of EngineeringPaschim Vihar, New Delhi, 110063, IndiaView full profile → , Surinder Kaurthisissurinderkaur1304@gmail.comDepartment of Information TechnologyBharati Vidyapeeth’s College of EngineeringPaschim Vihar, New Delhi, 110063, IndiaView full profile → , Ashima Airanashimajain046@gmail.comDepartment of Electrical and Electronics EngineeringBharati Vidyapeeth’s College of EngineeringPaschim Vihar, New Delhi, 110063, IndiaView full profile →
* Corresponding author · click or hover a name for details
- Received:
- 07 Aug 2024
- Published Online:
- 19 Feb 2025
- Article type:
- Research Article
- Language:
- EN
- Article no.:
- JIOS-1854
- Pages:
- 81–89
Abstract
Keywords
Subject Classifications
References
[1] A. Vaswani, S. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Information Processing Systems 30 (NeurIPS), Long Beach, USA, pp. 5998-6008 (2017).
[2] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” (2018), arXiv preprint arXiv:1810.04805.
[3] I. Sutskever, O. Vinyals, and Q. Le, “Sequence to sequence learning with neural networks,” (2014), arXiv preprint arXiv:1409.3215.
[4] K. Cho, B. van Merriënboer, D. Bahdanau, F. O. P. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using RNN encoder-decoder for statistical machine translation,” (2014), arXiv preprint arXiv:1406.1078.
[5] N. Kalchbrenner and P. Blunsom, “Recurrent continuous translation models,” Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing (2013).
[6] D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” (2014), arXiv preprint arXiv:1409.0473.
[7] A. S. Al-Shamayleh, O. Adwan, M. A. Alsharaiah, A. H. Hussein, Q. M. Kharma, and C. I. Eke, “A comprehensive literature review on image captioning methods and metrics based on deep learning technique,” Multimedia Tools and Applications, vol. 83, no. 12, pp. 34219-34268 (2024).
[8] I. Krejtz, A. Krejtz, W. Sienkiewicz, and M. Zubek, “Audio description as an aural guide of children’s visual attention: evidence from an eye-tracking study,” Proceedings of the Symposium on Eye Tracking Research and Applications (2012).
[9] M. A. Al-Malla, A. Jafar, and N. Ghneim, “Image captioning model using attention and object features to mimic human image understanding,” Journal of Big Data, vol. 9, no. 1, p. 20 (2022).
[10] A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3128-3137 (2015).
[11] Y. Luo, J. Li, Z. Li, and Y. Wang, “Natural language to visualization by neural machine translation,” IEEE Transactions on Visualization and Computer Graphics, vol. 28, no. 1, pp. 217-226 (2021).
[12] “Flickr8K dataset,” accessed from https://www.kaggle.com/datasets/adityajn105/flickr8k.




