Evolutionary perspective of large language models on shaping research insights into healthcare disparities

Authors

  • David An Harvard University

DOI:

https://doi.org/10.47989/ir31260845

Keywords:

LLM, h-index, AI Evolution, Healthcare disparities

Abstract

Introduction. Advances in large language models (LLMs) offer a chance to act as scientific assistants, helping people grasp complex research areas. This study examines how LLMs evolve in healthcare disparities research, with attention to public access to relevant information.

Methods. We studied three well-known LLMs: ChatGPT, Copilot, and Gemini. Each week, we asked them a consistent prompt about research themes in healthcare disparities and tracked how their answers changed over a one-month period.

Analysis. The themes produced by the LLMs were categorised and cross-checked against H-index values from the Web of Science to verify relevance. This dual approach shows how LLMs’ outputs develop over time and how such progress could help researchers navigate trends.

Results. The outputs aligned with actual scientific impact and trends in the field, indicating that LLMs can help people understand the healthcare disparities landscape. Time-series comparisons showed differences among the models in how broadly and deeply they identified and classified themes.

Conclusion. The study offers a framework that uses the evolution of multiple LLMs to illuminate AI tools for studying healthcare disparities, informing future research and public engagement strategies.

References

Abràmoff, M. D., Tarver, M. E., Loyo-Berrios, N., Trujillo, S., Char, D., Obermeyer, Z., & Eydelman, M. B., Foundational Principles of Ophthalmic Imaging and Algorithmic Interpretation Working Group of the Collaborative Community for Ophthalmic Imaging Foundation, W., DC, and Maisel, W. H. (2023). Considerations for addressing bias in artificial intelligence for health equity. NPJ Digital Medicine, 6(1), 170. https://www.nature.com/articles/s41746-023-00913-9

Aldoseri, A., Al-Khalifa, K. N., & Hamouda, A. M. (2023). Re-thinking data strategy and integration for artificial intelligence: concepts, opportunities, and challenges. Applied Sciences, 13(12), 7082. https://doi.org/10.3390/app13127082

Amirova, A., Fteropoulli, T., Ahmed, N., Cowie, M. R., & Leibo, J. Z. (2024). Framework-based qualitative analysis of free responses of large language models: algorithmic fidelity. PloS One, 19(3), e0300024. https://doi.org/10.1371/journal.pone.0300024

Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., & Wingate, D. (2023). Out of one, many: using language models to simulate human samples. Political Analysis, 31(3), 337-351. https://doi.org/10.1017/pan.2023.2

Bailey, R., Sharpe, D., Kwiatkowski, T., Watson, S., Dexter Samuels, A., & Hall, J. (2018). Mental health care disparities now and in the future. Journal of Racial and Ethnic Health Disparities, 5, 351-356. https://doi.org/10.1007/s40615-017-0377-6

Beheshti, M., Toubal, I. E., Alaboud, K., Almalaysha, M., Ogundele, O. B., Turabieh, H., Abdalnabi, N., Boren, S. A., Scott, G. J., & Dahu, B. M. (2025). Evaluating the reliability of ChatGPT for health-related questions: a systematic review. Informatics, 12(1), 9. https://doi.org/10.3390/informatics12010009

Berkman, N. D., Davis, T. C., & McCormack, L. (2010). Health literacy: what is it? Journal of Health Communication, 15(S2), 9-19. https://doi.org/10.1080/10810730.2010.499985

Braveman, P. (2006). Health disparities and health equity: concepts and measurement. Annual Review of Public Health, 27, 167-194. https://doi.org/10.1146/annurev.publhealth.27.021405.102103

Brondolo, E., Gallo, L. C., & Myers, H. F. (2009). Race, racism and health: disparities, mechanisms, and interventions. Journal of Behavioral Medicine, 32, 1-8. https://doi.org/10.1007/s10865-008-9190-3

Busch, F., Hoffmann, L., Rueger, C., van Dijk, E. H. C., Kader, R., Ortiz-Prado, E., Makowski, M. R., Saba, L., Hadamitzky, M., Kather, J. N., Truhn, D., Cuocolo, R., Adams, L. C., & Bressem, K. K. (2025). Current applications and challenges in large language models for patient care: a systematic review. Communications Medicine, 5, 26. https://doi.org/10.1038/s43856-024-00717-2

Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y., & Ye, W. (2024). A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology, 15(3), 1-45. https://doi.org/10.1145/3641289

Clusmann, J., Kolbinger, F. R., Muti, H. S., Carrero, Z. I., Eckardt, J. N., Laleh, N. G., Löffler, C. M. L., Schwarzkopf, S. C., Unger, M., Veldhuizen, G. P., & Wagner, S. J. (2023). The future landscape of large language models in medicine. Communications Medicine, 3(1), 141. https://doi.org/10.1038/s43856-023-00370-1

Cook, B. L., Hou, S. S. Y., Lee-Tauler, S. Y., Progovac, A. M., Samson, F., & Sanchez, M. J. (2019). A review of mental health and mental health care disparities research: 2011-2014. Medical Care Research and Review, 76(6), 683-710. https://doi.org/10.1177/1077558718780592

Cook, B. L., McGuire, T., & Miranda, J. (2007). Measuring trends in mental health care disparities, 2000–2004. Psychiatric Services, 58(12), 1533-1540. https://doi.org/10.1176/ps.2007.58.12.1533

Costas, R., & Bordons, M. (2007). The H-index: advantages, limitations and its relation with other bibliometric indicators at the micro level. Journal of Informetrics, 1(3), 193-203. https://doi.org/10.1016/j.joi.2007.02.001

Fiscella, K., & Sanders, M. R. (2016). Racial and ethnic disparities in the quality of health care. Annual Review of Public Health, 37(1), 375-394. https://doi.org/10.1146/annurev-publhealth-032315-021439

FitzGerald, C., & Hurst, S. (2017). Implicit bias in healthcare professionals: a systematic review. BMC Medical Ethics, 18, 1-18. https://doi.org/10.1186/s12910-017-0179-8

Fullman, N., Yearwood, J., Abay, S. M., Abbafati, C., Abd-Allah, F., Abdela, J., Abdelalim, A., Abebe, Z., Abebo, T. A., Aboyans, V., & Abraha, H. N. (2018). Measuring performance on the Healthcare Access and Quality Index for 195 countries and territories and selected subnational locations: a systematic analysis from the Global Burden of Disease Study 2016. The Lancet, 391(10136), 2236-2271. https://doi.org/10.1016/S0140-6736(18)30994-2

Google. (2023). Bard updates from Google I/O 2023: Images, new features. Retrieved 15 June 2024, from https://blog.google/technology/ai/google-bard-updates-io-2023/

Kirubarajan, A., Patel, P., Leung, S., Prethipan, T., & Sierra, S. (2021). Barriers to fertility care for racial/ethnic minority groups: a qualitative systematic review. F&S Reviews, 2(2), 150-159. https://doi.org/10.1016/j.xfnr.2021.01.001

Maity, S., & Saikia, M. J. (2025). Large language models in healthcare and medical applications: a review. Bioengineering, 12, 631. https://doi.org/10.3390/bioengineering12060631

Meduri, K., Gonaygunta, H., Nadella, G. S., Pawar, P. P., & Kumar, D. (2024). Adaptive intelligence: GPT-powered language models for dynamic responses to emerging healthcare challenges. International Journal of Advanced Research in Computer and Communication Engineering, 13(1), 104-109. https://doi.org/10.17148/IJARCCE.2024.13114

Mehdi, Y. (2023). Announcing Microsoft Copilot, your everyday AI companion. Retrieved 15 June 2024, from https://blogs.microsoft.com/blog/2023/09/21/announcing-microsoft-copilot-your-everyday-ai-companion/

Miranda, J., McGuire, T. G., Williams, D. R., & Wang, P. (2008). Mental health in the context of health disparities. American Journal of Psychiatry, 165(9), 1102-1108. https://doi.org/10.1176/appi.ajp.2008.08030333

Nazer, L. H., Zatarah, R., Waldrip, S., Ke, J. X. C., Moukheiber, M., Khanna, A. K., Hicklen, R. S., Moukheiber, L., Moukheiber, D., Ma, H., & Mathur, P. (2023). Bias in artificial intelligence algorithms and recommendations for mitigation. PLOS Digital Health, 2(6), e0000278. https://doi.org/10.1371/journal.pdig.0000278

Nazi, Z. A., & Peng, W. (2024). Large language models in healthcare and medical domain: A review. Informatics, 11(3), 57. https://doi.org/10.3390/informatics11030057

Nogueira, L., White, K. E., Bell, B., Alegria, K. E., Bennett, G., Edmondson, D., Epel, E., Holman, E. A., Kronish, I. M., & Thayer, J. (2022). The role of behavioral medicine in addressing climate change-related health inequities. Translational Behavioral Medicine, 12(4), 526-534. https://doi.org/10.1093/tbm/ibac005

Norris, M., & Oppenheim, C. (2010). The H‐index: a broad review of a new bibliometric indicator. Journal of Documentation, 66(5), 681-705. https://doi.org/10.1108/00220411011066790

Ong, J. C. L., Seng, B. J. J., Law, J. Z. F., Low, L. L., Kwa, A. L. H., Giacomini, K. M., & Ting, D. S. W. (2024). Artificial intelligence, ChatGPT, and other large language models for social determinants of health: current state and future directions. Cell Reports Medicine, 5(1), 101356. https://doi.org/10.1016/j.xcrm.2023.101356

OpenAI. (2024). ChatGPT4.0 [Large language model]. Retrieved 15 June 2024, from https://openai.com/chatgpt/

Ray, P. P. (2023). ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope. Internet of Things and Cyber-Physical Systems, 3, 121-154. https://doi.org/10.1016/j.iotcps.2023.04.003

Riley, W. J. (2012). Health disparities: gaps in access, quality and affordability of medical care. Transactions of the American Clinical and Climatological Association, 123, 167. https://pmc.ncbi.nlm.nih.gov/articles/PMC3540621/

Ruprecht, M. M., Wang, X., Johnson, A. K., Xu, J., Felt, D., Ihenacho, S., Stonehouse, P., Curry, C. W., DeBroux, C., Costa, D., & Phillips Ii, G. (2021). Evidence of social and structural COVID-19 disparities by sexual orientation, gender identity, and race/ethnicity in an urban environment. Journal of Urban Health, 98(1), 27-40. https://doi.org/10.1007/s11524-020-00497-9

Saeed, S. A., & Masters, R. M. (2021). Disparities in health care and the digital divide. Current Psychiatry Reports, 23, 1-6. https://doi.org/10.1007/s11920-021-01274-4

Schlicht, I. B., Zhao, Z., Sayin, B., Flek, L., & Rosso, P. (2025). Do LLMs provide consistent answers to health-related questions across languages? In European Conference on Information Retrieval (pp. 314-322). https://doi.org/10.1007/978-3-031-88714-7_30

Shah, F. A., & Jawaid, S. A. (2023). The H-index: an indicator of research and publication output. Pakistan Journal of Medical Sciences, 39(2), 315-316. https://doi.org/10.12669/pjms.39.2.7398

Singh, N., Lawrence, K., Richardson, S., & Mann, D. M. (2023). Centering health equity in large language model deployment. PLOS Digital Health, 2(10), e0000367. https://doi.org/10.1371/journal.pdig.0000367

Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., & Payne, P. (2023). Large language models encode clinical knowledge. Nature, 620(7972), 172-180. https://doi.org/10.1038/s41586-023-06291-2

Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F., & Ting, D. S. W. (2023). Large language models in medicine. Nature Medicine, 29(8), 1930-1940. https://doi.org/10.1038/s41591-023-02448-8

Veinot, T. C., Mitchell, H., & Ancker, J. S. (2018). Good intentions are not enough: how informatics interventions can worsen inequality. Journal of the American Medical Informatics Association, 25(8), 1080-1088. https://doi.org/10.1093/jamia/ocy052

Wang, D., & Zhang, S. (2024). Large language models in medical and healthcare fields: applications, advances, and challenges. Artificial Intelligence Review, 57(11), 299. https://doi.org/10.1007/s10462-024-10921-0

Wang, L., Wan, Z., Ni, C., Song, Q., Li, Y., Clayton, E., Malin, B., & Yin, Z. (2024). Applications and concerns of ChatGPT and other conversational large language models in health care: systematic review. Journal of Medical Internet Research, 26, e22769. https://doi.org/10.2196/22769

Yang, R., Tan, T. F., Lu, W., Thirunavukarasu, A. J., Ting, D. S. W., & Liu, N. (2023). Large language models in health care: Development, applications, and challenges. Health Care Science, 2(4), 255-263. https://doi.org/10.1002/hcs2.61

Yu, E., Chu, X., Zhang, W., Meng, X., Yang, Y., Ji, X., & Wu, C. (2025). Large language models in medicine: applications, challenges, and future directions. International Journal of Medical Sciences, 22(11), 2792-2801. https://doi.org/10.7150/ijms.111780

Published

2026-05-15

How to Cite

An, D. (2026). Evolutionary perspective of large language models on shaping research insights into healthcare disparities. Information Research an International Electronic Journal, 31(2), 482–494. https://doi.org/10.47989/ir31260845