TARU PUBLICATIONS
Journal of Information and Optimization Sciences cover
Open Access ·Peer-reviewed·ISSN (Online): 2169-0103·ISSN (Print): 0252-2667
Powered by:DOICrossrefiThenticate

The Journal of Information and Optimization Sciences (JIOS) is a world leading journal publishing high quality, rigorously peer-reviewed original research in all mathematically-oriented theoretical and applied topics in information sciences, optimization sciences and related areas since 1980. Subjects include but are not limited to: • Information Sciences • Optimization Sciences • Control Theory • Operational Research • Decision Sciences • Information Theory • Information Technology • Computer Networks and Communications • Mathematical Programming • Modelling and Simulation • Database Management • Applications to Engineering Sciences • Applications to Technology

Issues up to 2022 co-published with and available at:Taylor & Francis
submissions@tarupublications.com
Open Access Research Article

Analysis of attention based deep learning approach for audio image descriptions

, , , , * , , , ,

* Corresponding author · click or hover a name for details

pp. 81–89Vol. 46Issue 1January 2025DOI: 10.47974/JIOS-1854XML
Received:
07 Aug 2024
Published Online:
01 Jan 2025
Article type:
Research Article
Language:
EN
Article no.:
JIOS-1854
Pages:
81–89

Abstract

The aim of this paper is to focus on utilizing deep learning approaches for the building of an intelligent image captioning system. To bridge the semantic gap between the literal depiction of the images and the languages, this combined architecture is employed. The structure of the model is of an encoder-decoder architecture which consists of InceptionV3 CNN for feature extraction and LSTMs with attention mechanism to generate the captions. The complexity of the model is determined based on the usage of the Flickr8K image dataset and it was revealed through the results that it can generate accurate, interesting and appropriate content describing the images. This study has good prospects for the development of framework for learning with multiple modalities especially for usage databases that require content-based searching, image-based data retrieval and usability for users with sight problems. The results show that InceptionV3 Model performs best with BLEU Score of 0.8504 followed by DenseNet 201 with score of 0.8274. Further analysis shows that the best BLEU score of 0.9172 is attained when the dataset of the best performing model, Inception V3, is split into 75:25 Train:Test and running the model for 50 epochs is done. This resulted in a loss function value of 0.70 for the above-mentioned configuration. 

Keywords

Subject Classifications

68T07

References

[1] A. Vaswani, S. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Information Processing Systems 30 (NeurIPS), Long Beach, USA, pp. 5998-6008 (2017).    [2] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” (2018), arXiv preprint arXiv:1810.04805.[3] I. Sutskever, O. Vinyals, and Q. Le, “Sequence to sequence learning with neural networks,” (2014), arXiv preprint arXiv:1409.3215.[4] K. Cho, B. van Merriënboer, D. Bahdanau, F. O. P. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using RNN encoder-decoder for statistical machine translation,” (2014), arXiv preprint arXiv:1406.1078.[5] N. Kalchbrenner and P. Blunsom, “Recurrent continuous translation models,” Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing (2013).[6] D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” (2014), arXiv preprint arXiv:1409.0473.[7] A. S. Al-Shamayleh, O. Adwan, M. A. Alsharaiah, A. H. Hussein, Q. M. Kharma, and C. I. Eke, “A comprehensive literature review on image captioning methods and metrics based on deep learning technique,” Multimedia Tools and Applications, vol. 83, no. 12, pp. 34219-34268 (2024).[8] I. Krejtz, A. Krejtz, W. Sienkiewicz, and M. Zubek, “Audio description as an aural guide of children’s visual attention: evidence from an eye-tracking study,” Proceedings of the Symposium on Eye Tracking Research and Applications (2012).[9] M. A. Al-Malla, A. Jafar, and N. Ghneim, “Image captioning model using attention and object features to mimic human image understanding,” Journal of Big Data, vol. 9, no. 1, p. 20 (2022).[10] A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3128-3137 (2015).[11] Y. Luo, J. Li, Z. Li, and Y. Wang, “Natural language to visualization by neural machine translation,” IEEE Transactions on Visualization and Computer Graphics, vol. 28, no. 1, pp. 217-226 (2021).[12] “Flickr8K dataset,” accessed from https://www.kaggle.com/datasets/adityajn105/flickr8k.
Views: 227Downloads: 11Citations: 0