Top
Journal of Dependence

Diagnostic performance of a specific local Turkish large language model in detecting alcohol use disorder across internal medicine practice: A 42-patient pilot study

Authors

Files

pdf

Abstract

Objective: Alcohol use disorder (AUD) is underdiagnosed in internal medicine. We evaluated a local large language model (LLM), fine-tuned on a Turkish medical corpus, for detecting AUD across outpatient, emergency and inpatient settings.

Method: We retrospectively analyzed 42 adult patients at a tertiary state hospital (January–May 2024), in a single-center de-identified dataset modeled on the MIMIC-IV schema. The reference standard was the clinician-confirmed DSM-5 AUD diagnosis; comparators were the Alcohol Use Disorders Identification Test, Consumption version (AUDIT-C), gamma-glutamyl transferase (GGT), mean corpuscular volume and the aspartate aminotransferase/alanine aminotransferase ratio. A Llama-3-8B-Instruct model, fine-tuned via low-rank adaptation on an in-hospital GPU server, produced binary predictions, probabilities and severity categories. Analyses used Wilson 95% confidence intervals, bootstrap AUC, Cohen's κ and the 

Results: Mean age was 53.0 ± 15.0 years; 50.0% were female. AUD prevalence was 47.6% (20/42). The LLM achieved sensitivity 85.0% (95% CI, 64.0–94.8), specificity 90.9% (72.2–97.5), accuracy 88.1%, area under the curve (AUC) 0.864 (0.720–0.977) and κ = 0.761. Within the 17 true positives, severity matched the reference exactly (quadratic-weighted κ = 1.000), conditional on correct detection in a small subgroup. Within this enriched pilot cohort, AUDIT-C and GGT showed perfect separation (AUC = 1.000) and the LLM AUC was lower (DeLong p = 0.038). This ordering is a ceiling effect of the present cohort and is not a general property of these markers.

Conclusion: The local LLM achieved clinically acceptable accuracy and exact severity classification within correctly flagged cases. Its principal value lies where AUDIT-C is not routine...

pdf

References

American Psychiatric Association. (2013). Diagnostic and statistical manual of mental disorders (5th ed.). https://doi.org/10.1176/appi.books.9780890425596

Bossuyt, P. M., Reitsma, J. B., Bruns, D. E., Gatsonis, C. A., Glasziou, P. P., Irwig, L., Lijmer, J. G., Moher, D., Rennie, D., de Vet, H. C. W., Kressel, H. Y., Rifai, N., Golub, R. M., Altman, D. G., Hooft, L., Korevaar, D. A., & Cohen, J. F. (2015). STARD 2015: An updated list of essential items for reporting diagnostic accuracy studies. BMJ, 351, h5527. https://doi.org/10.1136/bmj.h5527

Bush, K., Kivlahan, D. R., McDonell, M. B., Fihn, S. D., & Bradley, K. A. (1998). The AUDIT alcohol consumption questions (AUDIT-C): An effective brief screening test for problem drinking. Archives of Internal Medicine, 158(16), 1789–1795. https://doi.org/10.1001/archinte.158.16.1789

Carvalho, A. F., Heilig, M., Perez, A., Probst, C., & Rehm, J. (2019). Alcohol use disorders. The Lancet, 394(10200), 781–792. https://doi.org/10.1016/S0140-6736(19)31775-1

Conigrave, K. M., Davies, P., Haber, P., & Whitfield, J. B. (2003). Traditional markers of excessive alcohol use. Addiction, 98(s2), 31–43. https://doi.org/10.1046/j.1359-6357.2003.00581.x

Connor, J. P., Haber, P. S., & Hall, W. D. (2016). Alcohol use disorders. The Lancet, 387(10022), 988–998. https://doi.org/10.1016/S0140-6736(15)00122-1

DeLong, E. R., DeLong, D. M., & Clarke-Pearson, D. L. (1988). Comparing the areas under two or more correlated receiver operating characteristic curves: A nonparametric approach. Biometrics, 44(3), 837–845. https://doi.org/10.2307/2531595

Goh, E., Gallo, R., Hom, J., Strong, E., Weng, Y., Kerman, H., Cool, J. A., Kanjee, Z., Parsons, A. S., Ahuja, N., Horvitz, E., Yang, D., Milstein, A., Olson, A. P. J., Rodman, A., & Chen, J. H. (2024). Large language model influence on diagnostic reasoning: A randomized clinical trial. JAMA Network Open, 7(10), e2440969. https://doi.org/10.1001/jamanetworkopen.2024.40969

Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., Yang, A., Fan, A., Goyal, A., Hartshorn, A., Yang, A., Mitra, A., Sravankumar, A., Korenev, A., Hinsvark, A., … Ma, Z. (2024). The Llama 3 herd of models (arXiv:2407.21783). arXiv. https://doi.org/10.48550/arXiv.2407.21783

Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2022). LoRA: Low-rank adaptation of large language models [Conference paper]. Tenth International Conference on Learning Representations (ICLR 2022). https://openreview.net/forum?id=nZeVKeeFYf9

Johnson, A. E. W., Bulgarelli, L., Shen, L., Gayles, A., Shammout, A., Horng, S., Pollard, T. J., Hao, S., Moody, B., Gow, B., Lehman, L. H., Celi, L. A., & Mark, R. G. (2023). MIMIC-IV, a freely accessible electronic health record dataset. Scientific Data, 10, Article 1. https://doi.org/10.1038/s41597-022-01899-x

Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174. https://doi.org/10.2307/2529310

Mitchell, A. J., Meader, N., Bird, V., & Rizzo, M. (2012). Clinical recognition and recording of alcohol disorders by clinicians in primary and secondary care: Meta-analysis. The British Journal of Psychiatry, 201(2), 93–100. https://doi.org/10.1192/bjp.bp.110.091199

Niemelä, O. (2016). Biomarker-based approaches for assessing alcohol use disorders. International Journal of Environmental Research and Public Health, 13(2), Article 166. https://doi.org/10.3390/ijerph13020166

Rehm, J., Anderson, P., Manthey, J., Shield, K. D., Struzzo, P., Wojnar, M., & Gual, A. (2016). Alcohol use disorders in primary health care: What do we know and where do we go? Alcohol and Alcoholism, 51(4), 422–427. https://doi.org/10.1093/alcalc/agv127

Saatcioğlu, Ö., Evren, C., & Çakmak, D. (2002). Alkol kullanım bozuklukları tanıma testinin (AKBT) geçerlik ve güvenirlik çalışması [Validity and reliability study of the Alcohol Use Disorders Identification Test]. Türkiye'de Psikiyatri, 4(2–3), 107–113.

Saunders, J. B., Aasland, O. G., Babor, T. F., De La Fuente, J. R., & Grant, M. (1993). Development of the alcohol use disorders identification test (AUDIT): WHO collaborative project on early detection of persons with harmful alcohol consumption-II. Addiction, 88(6), 791–804. https://doi.org/10.1111/j.1360-0443.1993.tb02093.x

Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., Payne, P., Seneviratne, M., Gamble, P., Kelly, C., Babiker, A., Schärli, N., Chowdhery, A., Mansfield, P., Demner-Fushman, D., … Natarajan, V. (2023). Large language models encode clinical knowledge. Nature, 620(7972), 172–180. https://doi.org/10.1038/s41586-023-06291-2

Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F., & Ting, D. S. W. (2023). Large language models in medicine. Nature Medicine, 29(8), 1930–1940. https://doi.org/10.1038/s41591-023-02448-8

Topol, E. J. (2019). High-performance medicine: The convergence of human and artificial intelligence. Nature Medicine, 25(1), 44–56. https://doi.org/10.1038/s41591-018-0300-7

Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Canton Ferrer, C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., … Scialom, T. (2023). Llama 2: Open foundation and fine-tuned chat models (arXiv:2307.09288). arXiv. https://doi.org/10.48550/arXiv.2307.09288

Türkiye İstatistik Kurumu. (2023). Türkiye sağlık araştırması, 2022 [Turkish Health Survey, 2022] (Haber Bülteni No. 49747). https://data.tuik.gov.tr/Bulten/Index?p=Turkiye-Saglik-Arastirmasi-2022-49747

Wilson, E. B. (1927). Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association, 22(158), 209–212. https://doi.org/10.1080/01621459.1927.10502953

World Health Organization. (2018). Global status report on alcohol and health 2018. https://www.who.int/publications/i/item/9789241565639

Details