Managing token limitations with RoBERTa-large for enhanced plagiarism detection
*Md. SohailCorresponding authorSohailmd2007@gmail.comPrincipal Architect LTIMindtree Limited; Department of Computer Engineering Savitribai Phule Pune UniversityPrincipal Architect LTIMindtree LimitedPune, Maharashtra, IndiaView full profile → , Kalpana S. Thakrekalpanathakre@mmcoe.edu.inDepartment of Computer Engineering Marathwada Mitra Mandal’s College of Engineering Savitribai Phule Pune UniversityDepartment of Computer Engineering Marathwada Mitra Mandal’s College of EngineeringPune, Maharashtra, 411052, IndiaView full profile →
* Corresponding author · click or hover a name for details
- Received:
- 07 May 2024
- Published Online:
- 05 Aug 2024
- Article type:
- Research Article
- Language:
- EN
- Article no.:
- JSMS-1309
- Pages:
- 1033–1043
Abstract
Keywords
Subject Classifications
References
[1] Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization (2016). arXiv preprint arXiv:1607.06450.
[2] Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate (2014). CoRR, abs/1409.0473.
[3] Denny Britz, Anna Goldie, Minh-Thang Luong, and Quoc V. Le. Massive exploration of neural machine translation architectures (2017). CoRR, abs/1703.03906.
[4] Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding (2018). arXiv preprint arXiv:1810.04805.
[5] Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., ... & Stoyanov, V. Roberta: A robustly optimized BERT approach (2019). arXiv preprint arXiv:1907.11692.
[6] Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. Convolutional sequence to sequence learning (2017). arXiv preprint arXiv:1705.03122v2.
[7] Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using RNN encoder-decoder for statistical machine translation (2014). CoRR, abs/1406.1078.
[8] Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. A structured self-attentive sentence embedding (2017). arXiv preprint arXiv:1703.03130.
[9] Ofir Press and Lior Wolf. Using the output embedding to improve language models (2016). arXiv preprint arXiv:1608.05859.
[10] Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision (2015). CoRR, abs/1512.00567.
[11] Muttlak, Hassen A. “Estimation of parameters in a multiple regression model using rank set sampling.” Journal of Information and Optimization Sciences 17.3 : 521-533 (1996).
[12] Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. Exploring the limits of language modeling (2016). arXiv preprint arXiv:1602.02410.
[13] Altameem, Ayman, et al. “P-ROCK: a sustainable clustering algorithm for large categorical datasets.” Intell. Autom. Soft Comput 35.1 : 553-566 (2023).
[14] Wazalwar, Sampada S., and Urmila Shrawankar. “Interpretation of sign language into English using NLP techniques.” Journal of Information and Optimization Sciences 38.6 : 895-910 (2017).



