<?xml version="1.0" encoding="UTF-8"?>
<article article-type="Research Article">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher">journal-of-statistics-and-management-systems</journal-id>
      <journal-title-group>
        <journal-title> Journal of Statistics and Management Systems</journal-title>
      </journal-title-group>
      <issn publication-format="electronic">2169-0014</issn>
      <issn publication-format="print">0972-0510</issn>
      <publisher>
        <publisher-name>Taru Publications</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.47974/JSMS-1309</article-id>
      <title-group>
        <article-title>Managing token limitations with RoBERTa-large for enhanced plagiarism detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes">
          <name>
            <surname>Sohail</surname>
            <given-names>Md.</given-names>
          </name>
          <aff>Principal Architect, LTIMindtree Limited, Pune, Maharashtra, India</aff>
          <aff>Department of Computer Engineering, Savitribai Phule Pune University, Pune, Maharashtra, 411052, India</aff>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Thakre</surname>
            <given-names>Kalpana S.</given-names>
          </name>
          <aff>Department of Computer Engineering, Marathwada Mitra Mandal’s College of Engineering, Savitribai Phule Pune University, Pune, Maharashtra, 411052, India</aff>
        </contrib>
      </contrib-group>
      <volume>27</volume>
      <issue>5</issue>
      <fpage>1033</fpage>
      <lpage>1043</lpage>
      <pub-date date-type="pub">
        <day>05</day>
        <month>08</month>
        <year>2024</year>
      </pub-date>
      <abstract>
        <p>Plagiarism is a pervasive challenge in academic and digital landscapes, necessitating effective detection mechanisms. This paper presents an innovative approach utilizing advanced NLP models, specifically RoBERTa and RoBERTa-large, to enhance plagiarism detection capabilities. The seven-step methodology addresses data processing, model training, evaluation, storage, and detection, with a focus on the unique attributes of the RoBERTa-large model. Additionally, the proposed solution adeptly manages token limitations in RoBERTa-large, providing an efficient strategy for handling extended texts in plagiarism detection tasks.  This paper introduces a solution to overcome token limitations in the RoBERTa-large model, enhancing plagiarism detection accuracy. The approach involves systematic text segmentation, tokenization, and result amalgamation for efficient processing of longer texts. It tackles RoBERTa-large’s constraints, ensuring optimal performance in plagiarism detection tasks.</p>
      </abstract>
      <kwd-group>
        <kwd>Plagiarism detection</kwd>
        <kwd>NLP (Natural language processing)</kwd>
        <kwd>LLM (Large language processing)</kwd>
        <kwd>Roberta</kwd>
        <kwd>Machine learning algorithm</kwd>
      </kwd-group>
      <custom-meta-group>
        <custom-meta>
          <meta-name>access</meta-name>
          <meta-value>open</meta-value>
        </custom-meta>
        <custom-meta>
          <meta-name>retracted</meta-name>
          <meta-value>no</meta-value>
        </custom-meta>
      </custom-meta-group>
    </article-meta>
  </front>
</article>
