TARU PUBLICATIONS
Journal of Information and Optimization Sciences cover
Hybrid ·Peer-reviewed·ISSN (Online): 2169-0103·ISSN (Print): 0252-2667

WoS  JIF 2026 : 0.4 (Q4)

Powered by:Powered by

Monthly Journal: Publishes theoretical and applied research on topics in information and optimization sciences.

Issues up to 2022 co-published with and available at:Taylor & Francis
submissions@tarupublications.com
Open Access Research Article

Multimodal news document summarization

* , ,

* Corresponding author · click or hover a name for details

pp. 959–968Vol. 45Issue 4May 2024DOI: 10.47974/JIOS-1619XML
Published Online:
08 Jun 2024
Article type:
Research Article
Language:
EN
Article no.:
JIOS-1619
Pages:
959–968

Abstract

With the increase in multimedia content, the domain of multimodal processing is experiencing constant growth. The question of whether combining these modalities is beneficial may come up. In this work, we investigate this by working on multi-modal content for obtaining quality summaries. We have conducted several experiments on the extractive summarization process employing asynchronous text, audio, image,and video.  Information present in the multimedia content has been leveraged to bridge the semantic gaps between different modes. Vision Transformers and BERT have been used for the image-matching and similarity-checking tasks. Furthermore, audio transcriptions have been used for incorporating the audio information in the summaries. The obtained news summaries have been evaluated with Rouge Score and a comparative analysis has been done.

Keywords

Subject Classifications

Primary 68TxxSecondary 68Uxx

References

[1] Evangelopoulos, Georgios, et al. “Multimodal saliency and fusion for movie summarization based on aural, visual, and textual attention.” IEEE Transactions on Multimedia 15.7 : 1553-1568 (2013).
[2] Erol, Berna, D-S. Lee, and Jonathan Hull. “Multimodal summarization of meeting recordings.” 2003 International Conference on Multimedia and Expo. ICME’03. Proceedings (Cat. No. 03TH8698). Vol. 3. IEEE (2003).
[3] Tjondronegoro, Dian, et al. “Multi-modal summarization of key events and top players in sports tournament videos.” 2011 IEEE Workshop on Applications of Computer Vision (WACV). IEEE (2011).
[4] Zhang, Tao, et al. “Automatic generation of pattern-controlled product description in e-commerce.” The World Wide Web Conference (2019).
[5] Calixto, Iacer, Qun Liu, and Nick Campbell. “Incorporating global visual features into attention-based neural machine translation.” arXiv preprint arXiv:1701.06521 (2017).[6] D. R. Radev, H. Jing, M. Stys, and D. Tam, “Centroid-based´ summarization of multiple documents,” Information Processing Management, vol. 40, no. 6, pp. 919–938 (2004).
[7] Katragadda, Rahul, Prasad Pingali, and Vasudeva Varma. “Sentence Position revisited: A robust light-weight Update Summarization ‘baseline’ Algorithm.” Proceedings of the Third International Workshop on Cross Lingual Information Access: Addressing the Information Need of Multilingual Societies (CLIAWS3) (2009).
[8] Mihalcea, Rada, and Paul Tarau. “Textrank: Bringing order into text.” Proceedings of the 2004 conference on empirical methods in natural language processing (2004).
[9] Li, Haoran, et al. “Guiderank: A guided ranking graph model for multilingual multi-document summarization.” Natural Language Understanding and Intelligent Applications: 5th CCF Conference on Natural Language Processing and Chinese Computing, NLPCC 2016, and 24th International Conference on Computer Processing of Oriental Languages, ICCPOL 2016, Kunming, China, December 2–6, 2016, Proceedings 24. Springer International Publishing (2016).
[10] Goyal, Pawan, Laxmidhar Behera, and Thomas Martin McGinnity. “A context-based word indexing model for document summarization.” IEEE Transactions on Knowledge and Data Engineering 25.8 : 1693-1705 (2012).
[11]  X. Li, L. Du, and Y. D. Shen, “Update summarization via graph_based sentence ranking,” IEEE Transactions on Knowledge & Data Engineering, vol. 25, no. 5, pp. 1162–1174 (2013).
[12] Zhou, Xinjie, Xiaojun Wan, and Jianguo Xiao. “CMiner: Opinion extraction and summarization for chinese microblogs.” IEEE Transactions on Knowledge and Data Engineering 28.7 : 1650-1663 (2016).
[13] Erkan, Günes, and Dragomir R. Radev. “Lexrank: Graph-based lexical centrality as salience in text summarization.” Journal of Artificial Intelligence Research 22 : 457-479 (2004).
[14] Wan, Xiaojun, and Jianwu Yang. “Improved affinity graph-based multi-document summarization.” Proceedings of the human language technology conference of the NAACL, Companion volume: Short papers (2006).
[15] Brin, Sergey, and Lawrence Page. “The anatomy of a large-scale hypertextual web search engine.” Computer networks and ISDN systems 30.1-7 : 107-117 (1998).
[16] Bhatnagar, Vaibhav, et al. “Descriptive analysis of COVID-19 patients in the context of India.” Journal of Interdisciplinary Mathematics 24.3 : 489-504 (2021).
[17] Li, Haoran, et al. “Multi-modal summarization for asynchronous collection of text, image, audio and video.” Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (2017).
[18] Kumari, Rajani, et al. “Analysis and predictions of spread, recovery, and death caused by COVID-19 in India.” Big Data Mining and Analytics 4.2 : 65-75 (2021).
[19] Tjondronegoro, Dian, et al. “Multi-modal summarization of key events and top players in sports tournament videos.” 2011 IEEE Workshop on Applications of Computer Vision (WACV). IEEE (2011).
[20] Akhtar, Nadeem, et al. “Tiered sentence based topic model for multi-document summarization.” Journal of Information and Optimization Sciences 43.8 : 2131-2141 (2022).
[21] Akhtar, Nadeem, M. M. Sufyan Beg, and Md Muzakkir Hussain. “Sparse two level topic model for extraction of general summary words.” Journal of Interdisciplinary Mathematics 23.1 : 303-310 (2020).
[22] M. Del Fabro, A. Sobe, and L. Boszormenyi, “Summarization of ¨ real-life events based on community-contributed content,” in The Fourth International Conferences on Advances in Multimedia : 119–126 (2012).
[23] Schinas, Manos, et al. “Multimodal graph-based event detection and summarization in social media streams.” Proceedings of the 23rd ACM International Conference on Multimedia (2015).
[24] Vaswani, Ashish, et al. “Attention is all you need.” Advances in neural information processing systems 30 (2017).
[25] Li, Haoran, et al. “Multi-modal summarization for asynchronous collection of text, image, audio and video.” Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (2017).
[26] https://huggingface.co/nlpconnect/vit-gpt2-image-captioning Accessed on 20/2/24
[27] Lin, Chin-Yew, and Eduard Hovy. “Automatic evaluation of summaries using n-gram co-occurrence statistics.” Proceedings of the 2003 human language technology conference of the North American chapter of the association for computational linguistics (2003).

Views: 193Downloads: 60Citations: 1