TARU PUBLICATIONS
Journal of Information and Optimization Sciences cover
Hybrid ·Peer-reviewed·ISSN (Online): 2169-0103·ISSN (Print): 0252-2667

WoS  JIF 2026 : 0.4 (Q4)

Powered by:Powered by

Monthly Journal: Publishes theoretical and applied research on topics in information and optimization sciences.

Issues up to 2022 co-published with and available at:Taylor & Francis
submissions@tarupublications.com
Open Access Research Article

Developing sentiment lexicon for Marathi : A comprehensive survey and analysis

* ,

* Corresponding author · click or hover a name for details

pp. 1141–1152Vol. 45Issue 4May 2024DOI: 10.47974/JIOS-1698XML
Published Online:
08 Jun 2024
Article type:
Research Article
Language:
EN
Article no.:
JIOS-1698
Pages:
1141–1152

Abstract

Sentiment Analysis plays an important role in developing AI applications involving human language and sense. Language-specific Sentiment Lexicon is important to accelerate Sentiment Analysis. Marathi is a morphologically rich but low-resource Indic Language. The language has a very strong grammatical base showing high resemblance to the human body. The paper elaborates on issues like Lexicon basics, word embedding, and Neural Network Models in the process of Lexicon Construction. Existing databases that are useful as seed databases are enlisted. A detailed study of various methods of Lexicon Construction is done. The methods highlight that Sentiment Neural Embedding is required and contextual information plays a vital role in detecting the polarity of a single word or document. The hybrid approach which uses lexicon-based features for deep learning provides better understanding of sentiment analysis. A method is proposed for MSL construction and sentiment analysis of Marathi text. Natural Language Processing Research for Indic Language is its boom. For Marathi, the efforts are seen. Considering the depth and scope of this language, more resource creation is a must for its revolutionization.

Keywords

Subject Classifications

68U15

References

[1] Busra Saylan , Songul Cinaroglu . Opinion mining and machine learning analysis : What emotions twitter data tell us about telemedicine?. Journal of Statistics & Management Systems ISSN 0972-0510 (Print), ISSN 2169-0014 (Online) Vol. 27, No. 3, pp. 605–633 (2024). DOI : 10.47974/JSMS-1028
[2] Huang, Minlie, et al. “Challenges in Building Intelligent Open-Domain Dialog Systems.” ACM Transactions on Information Systems, vol. 38, no. 3 (2020), https://doi.org/10.1145/3383123.
[3] Zhao, Deji, and Bo Ning. A Survey on Conversational Question-Answering Systems. Lecture Notes in Electrical Engineering, vol. 654 LNEE, no. 5, pp. 1857–61 (2021), https://doi.org/10.1007/978-981-15-8411-4_244.
[4] Pingle Aabha, et al. L3Cube-MahaSent-MD: A Multi-Domain Marathi Sentiment Analysis Dataset and Transformer Models (2023). arXiv:2306.13888v1https://doi.org/10.48550/arXiv.2306.13888.
[5] Yin Fulian, et al. “The Construction of Sentiment Lexicon Based on Context-Dependent Part-of-Speech Chunks for Semantic.” IEEE Access, vol. PP, p. 1 (2020), https://doi.org/10.1109/ACCESS.2020.2984284.
[6] Moeller Sarah, et al. “To POS Tag or Not to POS Tag: The Impact of POS Tags on Morphological Learning in Low-Resource Settings.” ACL-IJCNLP 2021 - 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, Proceedings of the Conference, pp. 966–78 (2021), https://doi.org/10.18653/v1/2021.acl-long.78.
[7] Shah Sonali Rajesh, and Abhishek Kaushik. Sentiment analysis on indian indigenous languages : a review on multilingual opinion mining. preprints 2019110338 https://doi.org/10.20944/preprints201911.0338.v1. 
[8] Ranathunga, Surangika, and Ranjiva Munasinghe. Indic Language Computing. no. June 2020 (2019), https://doi.org/10.1145/3343456.
[9] Lahoti, Pawan, et al. A Survey on NLP Resources , Tools , and Techniques For Marathi Language Processing. ACM Transactions on Asian and Low-Resource Language Information Processing, Volume 22 Issue 2 Article No 4 7 pp 1–34. https://doi.org/10.1145/3548457. 
[10] Kunchukuttan Anoop. GitHub - AI4Bharat/Indicnlp_catalog: A Collaborative Catalog of Resources for Indian Language NLP. https://github.com/AI4Bharat/indicnlp_catalog.30 March 2024.
[11] Mohammad Saif M., and Peter D. Turney. Crowdsourcing a Word-Emotion Association Lexicon. Computational Intelligence, vol. 29, no. 3,  pp. 436–65 (2013), https://doi.org/10.1111/j.1467-8640.2012.00460.x.
[12] Mohammad Saif M., and Peter D. Turney. Emotions Evoked by Common Words and Phrases: Using Mechanical Turk to Create an Emotion Lexicon. CAAGET ’10 Proceedings of the NAACL HLT 2010 Workshop on Computational Approaches to Analysis and Generation of Emotion in Text, no.June, pp. 26–34 (2010). 
[13] Baccianella, Stefano et al. SentiWordNet 3.0: An Enhanced Lexical Resource for Sentiment Analysis and Opinion Mining. International Conference on Language Resources and Evaluation (2010).
[14] Das Amitava, and Sivaji Bandyopadhyay. “SentiWordNet for Indian Languages.” The 8th Workshop on Asian Language Resources (ALR), August, no. August, pp. 56–63 (2010), http://www.aclweb.org/anthology/W/W10/W10-3208.pdf.
[15] Aditya Joshi, Sagar Ahire, Pushpak Bhattacharyya, ‘Sentiment Resources: Lexicons and Datasets’, Book Chapter, ‘A Practical Guide to Sentiment Analysis’, Editors: Dr. Dipankar Das, Dr. Erik Cambria, Springer Publication
[16] Kulkarni, Atharva, Meet Mandhane, Manali Likhitkar, Gayatri Kshirsagar, and Raviraj Joshi. L3CubeMahaSent: A Marathi Tweet-Based Sentiment Analysis Dataset. April (2021), http://arxiv.org/abs/2103.11408.
[17] Wijayanti Rini, and Andria Arisal. Automatic Indonesian Sentiment Lexicon Curation with Sentiment Valence Tuning for Social Media Sentiment Analysis. ACM Transactions on Asian and Low-Resource Language Information Processing. Volume 20 Issue 1Article No.:15 pp 1–16 https://doi.org/10.1145/3425632
[18] Gatti Lorenzo, et al. SentiWords : Deriving a High Precision and High Coverage Lexicon for Sentiment Analysis. arXiv:1510.09079v1 no. 4,  pp. 409–21 (2016). https://doi.org/10.1109/TAFFC.2015.2476456.
[19] Lai Siwei, et al. How to Generate a Good Word Embedding. arXiv:1507.05523v1 https://doi.org/10.48550/arXiv.1507.05523
[20] Li Minglei, et al. “Inferring Affective Meanings of Words from Word Embedding.” IEEE Transactions on Affective Computing, vol. 8, no. 4,  pp. 443–56 (2017), https://doi.org/10.1109/TAFFC.2017.2723012.
[21] Altameem, Ayman, et al. “P-ROCK: a sustainable clustering algorithm for large categorical datasets.” Intell. Autom. Soft Comput 35.1: 553-566 ((2023)). 
[22] D. Deng, L. Jing, J. Yu and S. Sun, “Sparse Self-Attention LSTM for Sentiment Lexicon Construction,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 27, no. 11, pp. 1777-1790, Nov. (2019), doi: 10.1109/TASLP.2019.2933326. 
[23] D. Deng, L. Jing, J. Yu, S. Sun and M. K. Ng, “Sentiment Lexicon Construction With Hierarchical Supervision Topic Model,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 27, no. 4, pp. 704-718, April (2019), doi: 10.1109/TASLP.2019.2892232.
[24] O. Wu, T. Yang, M. Li and M. Li, “Two-Level LSTM for Sentiment Analysis With Lexicon Embedding and Polar Flipping,” in IEEE Transactions on Cybernetics, vol. 52, no. 5, pp. 3867-3879, May (2022), doi: 10.1109/TCYB.2020.3017378.
[25] Khan H. T, Ridhorkar Sonali. A pragmatic analysis of text-based sentiment analysis models from an analytical perspective. Journal of Statistics & Management Systems ISSN 0972-0510 (Print), ISSN 2169-0014 (Online) Vol. 27, No. 2, pp. 285–294 (2024). DOI: 10.47974/JSMS-1254 https://doi.org/10.47974/JS

Views: 234Downloads: 84Citations: 0