<?xml version="1.0" encoding="UTF-8"?>
<article article-type="Research Article">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher">journal-of-information-and-optimization-sciences</journal-id>
      <journal-title-group>
        <journal-title>Journal of Information and Optimization Sciences</journal-title>
      </journal-title-group>
      <issn publication-format="electronic">2169-0103</issn>
      <issn publication-format="print">0252-2667</issn>
      <publisher>
        <publisher-name>Taru Publications</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.47974/JIOS-1619</article-id>
      <title-group>
        <article-title>Multimodal news document summarization</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes">
          <name>
            <surname>Javed</surname>
            <given-names>Hira</given-names>
          </name>
          <aff>Department of Computer Engineering, Aligarh Muslim University, Aligarh, Uttar Pradesh, India</aff>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Akhtar</surname>
            <given-names>Nadeem</given-names>
          </name>
          <aff>Department of Computer Engineering, Aligarh Muslim University, Aligarh, Uttar Pradesh, India</aff>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Beg</surname>
            <given-names>M. M. Sufyan</given-names>
          </name>
          <aff>Department of Computer Engineering, Aligarh Muslim University, Aligarh, Uttar Pradesh, India</aff>
        </contrib>
      </contrib-group>
      <volume>45</volume>
      <issue>4</issue>
      <fpage>959</fpage>
      <lpage>968</lpage>
      <pub-date date-type="pub">
        <day>08</day>
        <month>06</month>
        <year>2024</year>
      </pub-date>
      <abstract>
        <p>With the increase in multimedia content, the domain of multimodal processing is experiencing constant growth. The question of whether combining these modalities is beneficial may come up. In this work, we investigate this by working on multi-modal content for obtaining quality summaries. We have conducted several experiments on the extractive summarization process employing asynchronous text, audio, image,and video.  Information present in the multimedia content has been leveraged to bridge the semantic gaps between different modes. Vision Transformers and BERT have been used for the image-matching and similarity-checking tasks. Furthermore, audio transcriptions have been used for incorporating the audio information in the summaries. The obtained news summaries have been evaluated with Rouge Score and a comparative analysis has been done.</p>
      </abstract>
      <kwd-group>
        <kwd>Multimodal</kwd>
        <kwd>Summarization</kwd>
        <kwd>Machine learning</kwd>
        <kwd>Transformers</kwd>
      </kwd-group>
      <custom-meta-group>
        <custom-meta>
          <meta-name>access</meta-name>
          <meta-value>open</meta-value>
        </custom-meta>
        <custom-meta>
          <meta-name>retracted</meta-name>
          <meta-value>no</meta-value>
        </custom-meta>
      </custom-meta-group>
    </article-meta>
  </front>
</article>
