TARU PUBLICATIONS
Journal of Information and Optimization Sciences cover
Open Access ·Peer-reviewed·ISSN (Online): 2169-0103·ISSN (Print): 0252-2667
Powered by:DOICrossrefiThenticate

The Journal of Information and Optimization Sciences (JIOS) is a world leading journal publishing high quality, rigorously peer-reviewed original research in all mathematically-oriented theoretical and applied topics in information sciences, optimization sciences and related areas since 1980. Subjects include but are not limited to: • Information Sciences • Optimization Sciences • Control Theory • Operational Research • Decision Sciences • Information Theory • Information Technology • Computer Networks and Communications • Mathematical Programming • Modelling and Simulation • Database Management • Applications to Engineering Sciences • Applications to Technology

Issues up to 2022 co-published with and available at:Taylor & Francis
submissions@tarupublications.com
Open Access Research Article

Part of speech-based semantic similarity for RDF Predicate Selection : A benchmarking approach

, , *

* Corresponding author · click or hover a name for details

pp. 2179–2192Vol. 45Issue 8November 2024DOI: 10.47974/JIOS-1781XML
Received:
08 Apr 2024
Published Online:
30 Nov 2024
Article type:
Research Article
Language:
EN
Article no.:
JIOS-1781
Pages:
2179–2192

Abstract

RDF predicate selection and RDF subject identification are two key steps for effectively processing the natural language queries by the semantic question answering systems. For subject identification, techniques such as Named Entity Recognition are used. For the relatively complex task of predicate selection, Natural Language Processing techniques can be utilized along with knowledge bases to select the semantically matching relationships. Semantic features can be extracted from queries utilizing phrase embeddings of the transformer models. In this paper firstly, a novel Predicate Selection approach on RDF Knowledge Bases has been proposed and implemented utilizing BERT-based phrase embeddings and custom logic for key part-of-speech token combinations. Secondly, the implemented prototype is benchmarked over combinations of key part of speech tokens and stop words for predicate selection accuracy over YAGO Date Facts, demonstrating the effectiveness of our approach. Benchmarking results also highlight the importance of verb stop words for the task of predicate selection.

Keywords

Subject Classifications

(2010) 68T3068T3568T50

References

[1] T. Ahmad, M. Ahamad, S. U. Ahmed and N. Ahmad, “Short question-answers assessment using lexical and semantic similarity-based features”, Journal of Discrete Mathematical Sciences and Cryptography, vol. 25(7),, pp. 2057–2067 (2022).[2] A. Pereira, A. Trifan, R. Lopes and J. Oliveira, “Systematic review of question answering over knowledge bases”, IET Software, vol. 16(1), pp. 1-13 (2022).[3] Z. H. Amur, Y. Kwang Hooi, H. Bhanbhro, K. Dahri, and G. M. Soomro, “Short-Text Semantic Similarity (STSS): Techniques, Challenges and Future Perspectives”, Applied Sciences, vol. 13(6), pp. 3911 (2023).[4] P. Rathee, and S. K. Malik, “IWD towards Semantic similarity measure in ontology”, Journal of Information and Optimization Sciences, vol. 41(7), pp. 1561–1577 (2020).[5] P. Rathee, and S. K. Malik, “Ontology concept semantic similarity matching based on Ant Colony Optimization algorithm”, Journal of Information and Optimization Sciences, vol. 42(8), pp. 1987–2000 (2021).[6] P. J. Worth, “Word embeddings and semantic spaces in natural language processing”, International Journal of Intelligence Science, vol. 13(1), pp. 1-21 (2023).[7] M. Bilal, and A. A. Almazroi, A. A., “Effectiveness of fine-tuned BERT model in classification of helpful and unhelpful online customer reviews”, Electronic Commerce Research, vol. 23(4), pp. 2737-2757 (2023).[8] C. Antoniou and N. Bassiliades, “A survey on semantic question answering systems”, The Knowledge Engineering Review, vol. 37, pp. 1-42 (2022).[9] F. Yin, Y. Wang, J. Liu, and L. Lin, “The construction of sentiment lexicon based on context-dependent part-of-speech chunks for semantic disambiguation”, IEEE Access, vol. 8, pp. 63359-63367 (2020).[10] M. Wray, H. Doughty, and D. Damen, “On semantic similarity in video retrieval”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3650-3660 (2021).[11] I. Vulić, S. Baker, E. M. Ponti, U. Petti, I. Leviant, K. Wing, O. Majewska, E. Bar, M. Malone, T. Poibeau, and R. Reichart, “Multi-simlex: A large-scale evaluation of multilingual and crosslingual lexical semantic similarity”, Computational Linguistics, vol. 46(4), pp. 847-897 (2020).[12] F. S. Lievers, M. Bolognesi, and B. Winter, “The linguistic dimensions of concrete and abstract concepts: lexical category, morphological structure, countability, and etymology”, Cognitive Linguistics, vol. 32(4), pp. 641-670 (2021).[13] T. Rustamov, X. M. Jumanazarov, U. Almatova, Z. X. Mamaziyayev, and Z. A. Alibekova, “Classification symbols of words”, Asian Journal of Research in Social Sciences and Humanities, vol. 12(2), pp. 213-219 (2022).[14] S. Xu, S. Semnani, G. Campagna, and M. Lam, “AutoQA: From databases to QA semantic parsers with only synthetic training data”, In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (2020).[15] J. Reilly, A. M. Finley, C. P. Litovsky, and Y. N. Kenett, “Bigram semantic distance as an index of continuous semantic flow in natural language: Theory, tools, and applications”, Journal of Experimental Psychology: General, vol. 152(9), pp. 2578 (2023).[16] K. Yalcin, I. Cicekli, and G. Ercan, “An external plagiarism detection system based on part-of-speech (POS) tag n-grams and word embedding”, Expert Systems with Applications, vol. 197, pp. 116677 (2022).[17] X. Zhao, H. Zhang, and Y. Song, “PCR4ALL: A Comprehensive Evaluation Benchmark for Pronoun Coreference Resolution in English”, In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pp. 5963-5973 (2022).[18] R. Liu, R. Mao, A. T. Luu, and E. Cambria, “A brief survey on recent advances in coreference resolution”, Artificial Intelligence Review, vol. 56(12), pp. 14439-14481 (2023).[19] OpenAI. (2023). ChatGPT (Feb 13 version) [Large language model]. https://chat.openai.com/chat.[20] F. Mahdisoltani, J. Biega, and F. M. Suchanek, “Yago3: A knowledge base from multilingual wikipedias”, In CIDR (2013).
Views: 117Downloads: 8Citations: 0