TARU PUBLICATIONS
 Journal of Statistics and Management Systems cover
Hybrid ·Peer-reviewed·ISSN (Online): 2169-0014·ISSN (Print): 0972-0510

Monthly Journal: Publishes peer-reviewed aticles on theoretical and applied statistics and management systems, expoloring industrial statistics, actuarial and decision sciences.

Issues up to 2022 co-published with and available at:Taylor & Francis Online
submissions@tarupublications.com
Open Access Research Article

Designing a structured approach for consistency verification in AI systems for heartbeat disorder-related conversations

* ,

* Corresponding author · click or hover a name for details

pp. 1733–1741Vol. 27Issue 8November 2024DOI: 10.47974/JSMS-1402XML
Received:
12 Jun 2024
Published Online:
18 Dec 2024
Article type:
Research Article
Language:
EN
Article no.:
JSMS-1402
Pages:
1733–1741

Abstract

Large language models (LLMs) have shown mind-boggling question-answering capabilities. Most of their responses make one feel that they understand and remember context and are able to retrieve relevant content to generate almost accurate, human-like responses. While LLMs seem to mimic human intelligence, their responses can sometimes be inconsistent. Although LLMs are getting acceptance in healthcare for tasks like diagnosis support or patient communication, their inaccurate, inconsistent, or misleading information can be dangerous and can lead to harmful medical decisions if not properly validated by experts. Due to these limitations, LLM-powered full automation is still a dream. This paper presents a structured approach that can be used as basis to check LLMs level of “understanding” of medical knowledge and their “inferential capabilities” by evaluation of their responses to the word problems related to human heartbeat. The crafted word problems designed in the light of the proposed approach can be very helpful in identifying inherent weaknesses stemming from their fundamental nature. The study can be highly beneficial for healthcare professionals in determining the appropriate level of adoption of modern artificial intelligence (AI) technologies to identify heartbeat related disorders like Arrhythmia.

Keywords

Subject Classifications

Primary 00A05Secondary 68T50

References

[1] W. X. Zhao, X. Y. Li, Z. H. Wang, and A. B. Chen, “A survey of large language models,” arXiv, Oct. 13 (2024), arXiv:2303.18223. doi: 10.48550/arXiv.2303.18223.
[2] I. Mirzadeh, K. Alizadeh, H. Shahrokhi, O. Tuzel, S. Bengio, and M. Farajtabar, “GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models,” Oct. 07 (2024), arXiv: arXiv:2410.05229. doi: 10.48550/arXiv.2410.05229.
[3] K. Cobbe, A. S. Patel, J. S. Brown, and L. D. Miller, “Training verifiers to solve math word problems,” arXiv, Nov. 18 (2021), arXiv:2110.14168. doi: 10.48550/arXiv.2110.14168.
[4] N. Dziri, S. J. Lee, P. M. Sharma, and R. T. Zhang, “Faith and fate: Limits of transformers on compositionality,” arXiv, Oct. 31 (2023), arXiv:2305.18654. doi: 10.48550/arXiv.2305.18654.
[5] M. Sallam, “ChatGPT Utility in Healthcare Education, Research, and Practice: Systematic Review on the Promising Perspectives and Valid Concerns,” Healthc. Basel Switz., vol. 11, no. 6, p. 887, Mar. (2023), doi: 10.3390/healthcare11060887.
[6] M. Cascella, J. Montomoli, V. Bellini, and E. Bignami, “Evaluating the Feasibility of ChatGPT in Healthcare: An Analysis of Multiple Clinical and Research Scenarios,” J. Med. Syst., vol. 47, no. 1, p. 33 (2023), doi: 10.1007/s10916-023-01925-4.
[7] Y. Li, Z. Li, K. Zhang, R. Dan, S. Jiang, and Y. Zhang, “ChatDoctor: A Medical Chat Model Fine-Tuned on a Large Language Model Meta-AI (LLaMA) Using Medical Domain Knowledge,” Jun. 24 (2023), arXiv: arXiv:2303.14070. doi: 10.48550/arXiv.2303.14070.
[8] Y. Chang, Z. X. Li, M. P. Kim, and A. B. Zhao, “A survey on evaluation of large language models,” ACM Trans. Intell. Syst. Technol., vol. 15, no. 3, pp. 39:1-39:45, Mar. (2024), doi: 10.1145/3641289.
[9] Y. Liu, S. M. Zhao, X. T. Zhang, and Q. W. Huang, “Trustworthy LLMs: A survey and guideline for evaluating large language models’ alignment,” arXiv, Mar. 21 (2024), arXiv:2308.05374. doi: 10.48550/arXiv.2308.05374.
[10] Y. Sun, J. K. Wang, L. Z. Xu, and F. G. Yu, “LLM4Vuln: A unified evaluation framework for decoupling and enhancing LLMs’ vulnerability reasoning,” arXiv, Sep. 5 (2024), arXiv:2401.16185. doi: 10.48550/arXiv.2401.16185.

Views: 144Downloads: 47Citations: 0