TARU PUBLICATIONS
 Journal of Statistics and Management Systems cover
Open Access ·Peer-reviewed·ISSN (Online): 2169-0014·ISSN (Print): 0972-0510
Powered by:Powered by

The Journal of Statistics and Management Systems (JSMS) is a world leading journal publishing high quality, rigorously peer-reviewed original research on theoretical and applied statistics and management systems since 1998. The scope is intentionally broad, but papers must make a novel contribution to the field to be considered for publication. Topics include, but are not limited to, the following: • Statistics • Applied Statistics • Industrial Statistics • Statistical Inference • Interdisciplinary role of Statistics • Actuarial Sciences • Decision Sciences • Managerial Aspects • Management Sciences • Management Information Systems

Issues up to 2022 co-published with and available at:Taylor & Francis Online
submissions@tarupublications.com
Open Access Research Article

Managing token limitations with RoBERTa-large for enhanced plagiarism detection

* ,

* Corresponding author · click or hover a name for details

pp. 1033–1043Vol. 27Issue 5July 2024DOI: 10.47974/JSMS-1309XML
Received:
07 May 2024
Published Online:
05 Aug 2024
Article type:
Research Article
Language:
EN
Article no.:
JSMS-1309
Pages:
1033–1043

Abstract

Plagiarism is a pervasive challenge in academic and digital landscapes, necessitating effective detection mechanisms. This paper presents an innovative approach utilizing advanced NLP models, specifically RoBERTa and RoBERTa-large, to enhance plagiarism detection capabilities. The seven-step methodology addresses data processing, model training, evaluation, storage, and detection, with a focus on the unique attributes of the RoBERTa-large model. Additionally, the proposed solution adeptly manages token limitations in RoBERTa-large, providing an efficient strategy for handling extended texts in plagiarism detection tasks.  This paper introduces a solution to overcome token limitations in the RoBERTa-large model, enhancing plagiarism detection accuracy. The approach involves systematic text segmentation, tokenization, and result amalgamation for efficient processing of longer texts. It tackles RoBERTa-large’s constraints, ensuring optimal performance in plagiarism detection tasks.

Keywords

Subject Classifications

(2010) 68M12

References

[1] Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization (2016). arXiv preprint arXiv:1607.06450.
[2] Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate (2014). CoRR, abs/1409.0473.
[3] Denny Britz, Anna Goldie, Minh-Thang Luong, and Quoc V. Le. Massive exploration of neural machine translation architectures (2017). CoRR, abs/1703.03906.
[4] Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding (2018). arXiv preprint arXiv:1810.04805.
[5] Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., ... & Stoyanov, V.  Roberta: A robustly optimized BERT approach (2019). arXiv preprint arXiv:1907.11692.
[6] Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. Convolutional sequence to sequence learning (2017). arXiv preprint arXiv:1705.03122v2.
[7] Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using RNN encoder-decoder for statistical machine translation (2014). CoRR, abs/1406.1078.
[8] Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. A structured self-attentive sentence embedding (2017). arXiv preprint arXiv:1703.03130. 
[9] Ofir Press and Lior Wolf. Using the output embedding to improve language models (2016). arXiv preprint arXiv:1608.05859.
[10] Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision (2015). CoRR, abs/1512.00567.
[11] Muttlak, Hassen A. “Estimation of parameters in a multiple regression model using rank set sampling.” Journal of Information and Optimization Sciences 17.3 : 521-533 (1996).
[12] Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. Exploring the limits of language modeling (2016). arXiv preprint arXiv:1602.02410.
[13] Altameem, Ayman, et al. “P-ROCK: a sustainable clustering algorithm for large categorical datasets.” Intell. Autom. Soft Comput 35.1 : 553-566 (2023).
[14] Wazalwar, Sampada S., and Urmila Shrawankar. “Interpretation of sign language into English using NLP techniques.” Journal of Information and Optimization Sciences 38.6 : 895-910 (2017).

Views: 160Downloads: 6Citations: 0