How Large Language Models Are Reported in Qualitative Research: A Scoping Review

- 333332
Paper Abstract
Favorite this paper
How to cite this paper?
Abstract

Introduction: Large language models (LLMs) are increasingly embedded in qualitative research workflows. Their use promises efficiency and scale, while raising concerns about transparency, reproducibility, and methodological fit. A scoping review was conducted to map how studies describe the use of LLMs in qualitative research and to identify strengths and gaps in reporting practices.

Goals and Methods: The review aimed to describe study characteristics, document roles assigned to LLMs across the research pipeline, and assess the clarity of reporting on models, prompts, configuration, validation, ethics, and reflexivity. Peer-reviewed literature published from 2020 to 2025 was searched. Records were screened by multiple reviewers with conflicts resolved through discussion. Extraction captured methods, design, sample, model type and version, access pathway, parameter settings, prompting strategies, human AI interaction, validation against human analysis, outcomes attributed to model use, and indicators aligned with qualitative reporting norms.

Results: A total of 5,049 records were identified. After duplicate removal, 4,201 titles and abstracts were screened. In total, 149 full-text articles were assessed, and 78 studies met the inclusion criteria. Data sources were interviews, focus groups, documents, and social media posts. Applications included coding, theme development, topic identification, summarization, sentiment analysis, and support for instrument design. Widely used models included GPT variants, Claude, and Llama. Reporting quality varied markedly, ranging from very detailed descriptions of model version, access pathway, parameters, prompts, and validation procedures to minimal statements that provided only a model name. Parameter settings were often omitted. Validation involved comparing model outputs with human coding. Ethical issues and reflexivity were addressed inconsistently.

Conclusions: The current literature reveals a rapid and diverse adoption of large language models in qualitative research, accompanied by highly variable reporting. Clear guidance on model specification, prompting, validation, ethics, and reflexivity is required. These findings directly inform the development of COREQ + LLM to support transparent and trustworthy reporting.

Share your ideas or questions with the authors!

Did you know that the greatest stimulus in scientific and cultural development is curiosity? Leave your questions or suggestions to the author!

Sign in to interact

Have a question or suggestion? Share your feedback with the authors!

Institutions
  • 1 Health Care Informatics, Faculty of Health, School of Medicine, Witten/Herdecke University, Witten, Germany
  • 2 Health Services Research, Faculty of Health, School of Medicine, Witten/Herdecke University, Witten, Germany
  • 3 Chair of Clinical Pharmacology, Faculty of Health, School of Medicine, Witten/Herdecke University, Witten, Germany
Track
  • 1. Qualitative Research in Health
Keywords
Qualitative Research
Artificial Intelligence
Large Language Models
Scoping Review
COREQ+LLM