Open Access
·Peer-reviewed·ISSN (Online): 2169-0103·ISSN (Print): 0252-2667
Powered by:DOICrossrefiThenticate
The Journal of Information and Optimization Sciences (JIOS) is a world leading journal publishing high quality, rigorously peer-reviewed original research in all mathematically-oriented theoretical and applied topics in information sciences, optimization sciences and related areas since 1980. Subjects include but are not limited to:
• Information Sciences
• Optimization Sciences
• Control Theory
• Operational Research
• Decision Sciences
• Information Theory
• Information Technology
• Computer Networks and Communications
• Mathematical Programming
• Modelling and Simulation
• Database Management
• Applications to Engineering Sciences
• Applications to Technology
Issues up to 2022 co-published with and available at:
Analysis of attention based deep learning approach for audio image descriptions
Achin Jainachin.mails@gmail.comDepartment of Information Technology Bharati Vidyapeeth’s College of EngineeringPaschim Vihar, New Delhi, 110063, IndiaView full profile →
, Sarita Yadavsarita1320@yahoo.co.inDepartment of Information Technology Bharati Vidyapeeth’s College of EngineeringPaschim Vihar, New Delhi, 110063, IndiaView full profile →
, Neetu Singhsinghneetu4bvcoe@gmail.comDepartment of Information Technology Bharati Vidyapeeth’s College of EngineeringPaschim Vihar, New Delhi, 110063, IndiaView full profile →
, Kajal Kaulkajalkaulphd@gmail.comDepartment of Information Technology Bharati Vidyapeeth’s College of EngineeringUniversity School of Information, Communication and Technology Guru Gobind Singh Indraprastha UniversityDwarka, New Delhi, 110078, IndiaView full profile →
, *Arun Kumar DubeyCorresponding authorarudubey@gmail.comDepartment of Information Technology Bharati Vidyapeeth’s College of EngineeringPaschim Vihar, New Delhi, 110063, India0000-0002-6844-9213View full profile →
, Prakhar Priyadarshiprakharpriya@gmail.comDepartment of Information Technology Bharati Vidyapeeth’s College of EngineeringPaschim Vihar, New Delhi, 110063, IndiaView full profile →
, Prabhav Sangaprabhav14.sanga@gmail.comDepartment of Information Technology Bharati Vidyapeeth’s College of EngineeringPaschim Vihar, New Delhi, 110063, IndiaView full profile →
, Surinder Kaurthisissurinderkaur1304@gmail.comDepartment of Information Technology Bharati Vidyapeeth’s College of EngineeringPaschim Vihar, New Delhi, 110063, IndiaView full profile →
, Ashima Airanashimajain046@gmail.comDepartment of Electrical and Electronics Engineering Bharati Vidyapeeth’s College of EngineeringPaschim Vihar, New Delhi, 110063, IndiaView full profile →
* Corresponding author · click or hover a name for details
The aim of this paper is to focus on utilizing deep learning approaches for the building of an intelligent image captioning system. To bridge the semantic gap between the literal depiction of the images and the languages, this combined architecture is employed. The structure of the model is of an encoder-decoder architecture which consists of InceptionV3 CNN for feature extraction and LSTMs with attention mechanism to generate the captions. The complexity of the model is determined based on the usage of the Flickr8K image dataset and it was revealed through the results that it can generate accurate, interesting and appropriate content describing the images. This study has good prospects for the development of framework for learning with multiple modalities especially for usage databases that require content-based searching, image-based data retrieval and usability for users with sight problems. The results show that InceptionV3 Model performs best with BLEU Score of 0.8504 followed by DenseNet 201 with score of 0.8274. Further analysis shows that the best BLEU score of 0.9172 is attained when the dataset of the best performing model, Inception V3, is split into 75:25 Train:Test and running the model for 50 epochs is done. This resulted in a loss function value of 0.70 for the above-mentioned configuration.
[1] A. Vaswani, S. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Information Processing Systems 30 (NeurIPS), Long Beach, USA, pp. 5998-6008 (2017). [2] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” (2018), arXiv preprint arXiv:1810.04805.[3] I. Sutskever, O. Vinyals, and Q. Le, “Sequence to sequence learning with neural networks,” (2014), arXiv preprint arXiv:1409.3215.[4] K. Cho, B. van Merriënboer, D. Bahdanau, F. O. P. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using RNN encoder-decoder for statistical machine translation,” (2014), arXiv preprint arXiv:1406.1078.[5] N. Kalchbrenner and P. Blunsom, “Recurrent continuous translation models,” Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing (2013).[6] D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” (2014), arXiv preprint arXiv:1409.0473.[7] A. S. Al-Shamayleh, O. Adwan, M. A. Alsharaiah, A. H. Hussein, Q. M. Kharma, and C. I. Eke, “A comprehensive literature review on image captioning methods and metrics based on deep learning technique,” Multimedia Tools and Applications, vol. 83, no. 12, pp. 34219-34268 (2024).[8] I. Krejtz, A. Krejtz, W. Sienkiewicz, and M. Zubek, “Audio description as an aural guide of children’s visual attention: evidence from an eye-tracking study,” Proceedings of the Symposium on Eye Tracking Research and Applications (2012).[9] M. A. Al-Malla, A. Jafar, and N. Ghneim, “Image captioning model using attention and object features to mimic human image understanding,” Journal of Big Data, vol. 9, no. 1, p. 20 (2022).[10] A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3128-3137 (2015).[11] Y. Luo, J. Li, Z. Li, and Y. Wang, “Natural language to visualization by neural machine translation,” IEEE Transactions on Visualization and Computer Graphics, vol. 28, no. 1, pp. 217-226 (2021).[12] “Flickr8K dataset,” accessed from https://www.kaggle.com/datasets/adityajn105/flickr8k.
Views: 227Downloads: 11Citations: 0
Install Journal of Information and Optimization SciencesFaster access from your home screen