Diagnostic performance of a specific local Turkish large language model in detecting alcohol use disorder across internal medicine practice: A 42-patient pilot study
Authors
İlker Boğa, Mete Tuğcan ÜçdalFiles
Abstract
Objective: Alcohol use disorder (AUD) is underdiagnosed in internal medicine. We evaluated a local large language model (LLM), fine-tuned on a Turkish medical corpus, for detecting AUD across outpatient, emergency and inpatient settings.
Method: We retrospectively analyzed 42 adult patients at a tertiary state hospital (January–May 2024), in a single-center de-identified dataset modeled on the MIMIC-IV schema. The reference standard was the clinician-confirmed DSM-5 AUD diagnosis; comparators were the Alcohol Use Disorders Identification Test, Consumption version (AUDIT-C), gamma-glutamyl transferase (GGT), mean corpuscular volume and the aspartate aminotransferase/alanine aminotransferase ratio. A Llama-3-8B-Instruct model, fine-tuned via low-rank adaptation on an in-hospital GPU server, produced binary predictions, probabilities and severity categories. Analyses used Wilson 95% confidence intervals, bootstrap AUC, Cohen's κ and the
Results: Mean age was 53.0 ± 15.0 years; 50.0% were female. AUD prevalence was 47.6% (20/42). The LLM achieved sensitivity 85.0% (95% CI, 64.0–94.8), specificity 90.9% (72.2–97.5), accuracy 88.1%, area under the curve (AUC) 0.864 (0.720–0.977) and κ = 0.761. Within the 17 true positives, severity matched the reference exactly (quadratic-weighted κ = 1.000), conditional on correct detection in a small subgroup. Within this enriched pilot cohort, AUDIT-C and GGT showed perfect separation (AUC = 1.000) and the LLM AUC was lower (DeLong p = 0.038). This ordering is a ceiling effect of the present cohort and is not a general property of these markers.
Conclusion: The local LLM achieved clinically acceptable accuracy and exact severity classification within correctly flagged cases. Its principal value lies where AUDIT-C is not routine...
References
American Psychiatric Association. (2013). Diagnostic and statistical manual of mental disorders (5th ed.). https://doi.org/10.1176/appi.books.9780890425596
Bossuyt, P. M., Reitsma, J. B., Bruns, D. E., Gatsonis, C. A., Glasziou, P. P., Irwig, L., Lijmer, J. G., Moher, D., Rennie, D., de Vet, H. C. W., Kressel, H. Y., Rifai, N., Golub, R. M., Altman, D. G., Hooft, L., Korevaar, D. A., & Cohen, J. F. (2015). STARD 2015: An updated list of essential items for reporting diagnostic accuracy studies. BMJ, 351, h5527. https://doi.org/10.1136/bmj.h5527
Bush, K., Kivlahan, D. R., McDonell, M. B., Fihn, S. D., & Bradley, K. A. (1998). The AUDIT alcohol consumption questions (AUDIT-C): An effective brief screening test for problem drinking. Archives of Internal Medicine, 158(16), 1789–1795. https://doi.org/10.1001/archinte.158.16.1789
Carvalho, A. F., Heilig, M., Perez, A., Probst, C., & Rehm, J. (2019). Alcohol use disorders. The Lancet, 394(10200), 781–792. https://doi.org/10.1016/S0140-6736(19)31775-1
Conigrave, K. M., Davies, P., Haber, P., & Whitfield, J. B. (2003). Traditional markers of excessive alcohol use. Addiction, 98(s2), 31–43. https://doi.org/10.1046/j.1359-6357.2003.00581.x
Connor, J. P., Haber, P. S., & Hall, W. D. (2016). Alcohol use disorders. The Lancet, 387(10022), 988–998. https://doi.org/10.1016/S0140-6736(15)00122-1
DeLong, E. R., DeLong, D. M., & Clarke-Pearson, D. L. (1988). Comparing the areas under two or more correlated receiver operating characteristic curves: A nonparametric approach. Biometrics, 44(3), 837–845. https://doi.org/10.2307/2531595
Goh, E., Gallo, R., Hom, J., Strong, E., Weng, Y., Kerman, H., Cool, J. A., Kanjee, Z., Parsons, A. S., Ahuja, N., Horvitz, E., Yang, D., Milstein, A., Olson, A. P. J., Rodman, A., & Chen, J. H. (2024). Large language model influence on diagnostic reasoning: A randomized clinical trial. JAMA Network Open, 7(10), e2440969. https://doi.org/10.1001/jamanetworkopen.2024.40969
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., Yang, A., Fan, A., Goyal, A., Hartshorn, A., Yang, A., Mitra, A., Sravankumar, A., Korenev, A., Hinsvark, A., … Ma, Z. (2024). The Llama 3 herd of models (arXiv:2407.21783). arXiv. https://doi.org/10.48550/arXiv.2407.21783
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2022). LoRA: Low-rank adaptation of large language models [Conference paper]. Tenth International Conference on Learning Representations (ICLR 2022). https://openreview.net/forum?id=nZeVKeeFYf9
Johnson, A. E. W., Bulgarelli, L., Shen, L., Gayles, A., Shammout, A., Horng, S., Pollard, T. J., Hao, S., Moody, B., Gow, B., Lehman, L. H., Celi, L. A., & Mark, R. G. (2023). MIMIC-IV, a freely accessible electronic health record dataset. Scientific Data, 10, Article 1. https://doi.org/10.1038/s41597-022-01899-x
Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174. https://doi.org/10.2307/2529310
Mitchell, A. J., Meader, N., Bird, V., & Rizzo, M. (2012). Clinical recognition and recording of alcohol disorders by clinicians in primary and secondary care: Meta-analysis. The British Journal of Psychiatry, 201(2), 93–100. https://doi.org/10.1192/bjp.bp.110.091199
Niemelä, O. (2016). Biomarker-based approaches for assessing alcohol use disorders. International Journal of Environmental Research and Public Health, 13(2), Article 166. https://doi.org/10.3390/ijerph13020166
Rehm, J., Anderson, P., Manthey, J., Shield, K. D., Struzzo, P., Wojnar, M., & Gual, A. (2016). Alcohol use disorders in primary health care: What do we know and where do we go? Alcohol and Alcoholism, 51(4), 422–427. https://doi.org/10.1093/alcalc/agv127
Saatcioğlu, Ö., Evren, C., & Çakmak, D. (2002). Alkol kullanım bozuklukları tanıma testinin (AKBT) geçerlik ve güvenirlik çalışması [Validity and reliability study of the Alcohol Use Disorders Identification Test]. Türkiye'de Psikiyatri, 4(2–3), 107–113.
Saunders, J. B., Aasland, O. G., Babor, T. F., De La Fuente, J. R., & Grant, M. (1993). Development of the alcohol use disorders identification test (AUDIT): WHO collaborative project on early detection of persons with harmful alcohol consumption-II. Addiction, 88(6), 791–804. https://doi.org/10.1111/j.1360-0443.1993.tb02093.x
Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., Payne, P., Seneviratne, M., Gamble, P., Kelly, C., Babiker, A., Schärli, N., Chowdhery, A., Mansfield, P., Demner-Fushman, D., … Natarajan, V. (2023). Large language models encode clinical knowledge. Nature, 620(7972), 172–180. https://doi.org/10.1038/s41586-023-06291-2
Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F., & Ting, D. S. W. (2023). Large language models in medicine. Nature Medicine, 29(8), 1930–1940. https://doi.org/10.1038/s41591-023-02448-8
Topol, E. J. (2019). High-performance medicine: The convergence of human and artificial intelligence. Nature Medicine, 25(1), 44–56. https://doi.org/10.1038/s41591-018-0300-7
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Canton Ferrer, C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., … Scialom, T. (2023). Llama 2: Open foundation and fine-tuned chat models (arXiv:2307.09288). arXiv. https://doi.org/10.48550/arXiv.2307.09288
Türkiye İstatistik Kurumu. (2023). Türkiye sağlık araştırması, 2022 [Turkish Health Survey, 2022] (Haber Bülteni No. 49747). https://data.tuik.gov.tr/Bulten/Index?p=Turkiye-Saglik-Arastirmasi-2022-49747
Wilson, E. B. (1927). Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association, 22(158), 209–212. https://doi.org/10.1080/01621459.1927.10502953
World Health Organization. (2018). Global status report on alcohol and health 2018. https://www.who.int/publications/i/item/9789241565639
