TARU PUBLICATIONS
Journal of Information and Optimization Sciences cover
Hybrid ·Peer-reviewed·ISSN (Online): 2169-0103·ISSN (Print): 0252-2667

WoS  JIF 2026 : 0.4 (Q4)

Powered by:Powered by

Monthly Journal: Publishes theoretical and applied research on topics in information and optimization sciences.

Issues up to 2022 co-published with and available at:Taylor & Francis
submissions@tarupublications.com
Open Access Research Article

Optimized text generation using Markov models

* ,

* Corresponding author · click or hover a name for details

pp. 2031–2042Vol. 47Issue 5-BMay 2026DOI: 10.47974/JIOS-2294XML
Received:
01 Apr 2025
Published Online:
23 Apr 2026
Article type:
Research Article
Language:
EN
Article no.:
JIOS-2294
Pages:
2031–2042

Abstract

Many resource-poor and morphologically rich Indian languages struggle to benefit from recent advancements in feature representations in Natural Language Processing (NLP) because of lack of massive, annotated corpus and sturdy benchmarks. Natural Language Generation (NLG) or Textual content Generation is a subfield of NLP. A Markov chain is a stochastic model that forecasts future states based exclusively on the current state, utilizing a random probability distribution. It is computationally efficient, quick to execute, and requires minimal memory because its next state depends exclusively on the present state, with no influence from preceding states. In this paper, we have implemented 1) A word level Text Generation model with a Markov chain and 2) A character level Text Generation model using a Markov chain with the help of reinforcement-learning. We have used the Telugu dataset used in [1] for our experimentations. We generated only 20, 30 and 50 percent of the input text as summary with both our methods. These generated summaries used for text classification and evaluated the results. Our methods proved as computationally efficient. 

Keywords

Subject Classifications

68T5068T30

References

[1] M. Mareddy, “Am I a resource-poor language? Data sets, embeddings, models and analysis for four different NLP tasks in Telugu language,” ACM Trans. Asian Low-Resource Lang. Inf. Process., pp. 1–35 (2022).
[2] Y. Dang, Y. Zhang, and H. Chen, “A lexicon-enhanced method for sentiment classification: An experiment on online product reviews,” IEEE Intelligent Systems, vol. 25, no. 4, pp. 46–53 (2010).
[3] C. Ounis, I. MacDonald, and I. Soboroff, “Overview of the TREC-2008 Blog Track,” in Proc. 17th Text Retrieval Conf. (TREC), NIST (2008).
[4] M. Taboada, C. Anthony, and V. Kimberly, “Creating semantic orientation dictionaries,” in Proc. 5th Int. Conf. Language Resources and Evaluation (LREC), Genoa, Italy, pp. 427–432 (2006).
[5] T. Wilson, J. Wiebe, and P. Hoffmann, “Recognizing contextual polarity in phrase-level sentiment analysis,” in Proc. HLT/EMNLP, Vancouver, Canada, pp. 347–354 (2005).
[6] G. A. Miller, R. Beckwith, C. Fellbaum, D. Gross, and K. J. Miller, “WordNet: An on-line lexical database,” Oxford: Oxford University Press (1990).
[7] A. Esuli and F. Sebastiani, “SentiWordNet: A publicly available lexical resource for opinion mining,” in Proc. 5th Int. Conf. Language Resources and Evaluation (LREC), Genoa, Italy (2006).
[8] N. D. Gitari, Z. Zuping, H. Damien, and J. Long, “A lexicon-based approach for hate speech detection,” Int. J. Multimedia Ubiquitous Eng., vol. 10, no. 4, pp. 215–230 (2015).
[9] M. Hu and B. Liu, “Mining and summarizing customer reviews,” in Proc. 10th ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, pp. 168–177 (2004).
[10] C. Yang, K. H.-Y. Lin, and H.-H. Chen, “Building emotion lexicon from weblog corpora,” in Proc. 45th Annu. Meeting Assoc. for Computational Linguistics (ACL) Companion Volume, pp. 133–136 (2007).
[11] R. Jain, L. Raja, S. K. Sharma, and D. P. Bhatt, “Particle swarm optimization model for Hindi text summarization,” Journal of Information and Optimization Sciences, vol. 45, no. 4, pp. 839–850 (2024), doi : 10.47974/JIOS-1609.
[12] T. Choesang and B. R. Prathap, “Detection of Twitter hate tweets leading to crime using multiseries BERT model,” Journal of Information and Optimization Sciences, vol. 45, no. 4, pp. 885–896 (2024), doi : 10.47974/JIOS-1613.
[13] S. Abhijith, A. Poly, B. A. Jacob, G. Roy, and R. R. Cherish, “Decoding consumer voice: Sentiment analysis of web-scraped product reviews,” Journal of Information and Optimization Sciences, vol. 45, no. 4, pp. 913–923 (2024), doi : 10.47974/JIOS-1615.
[14] P. V. Kulkarni and K. S. Thakre, “Developing sentiment lexicon for Marathi: A comprehensive survey and analysis,” Journal of Information and Optimization Sciences, vol. 45, no. 4, pp. 1141–1152 (2024), doi : 10.47974/JIOS-1698.
[15] M. Tiwari, P. Aital, and P. Joshi, “A comprehensive survey and review of machine learning techniques in document processing: Industry applications and future directions,” Journal of Information and Optimization Sciences, vol. 45, no. 4, pp. 1177–1188 (2024).

Views: 77Downloads: 85Citations: 0