Exploring the Use of Language Models for Open Coding: A Comparative Study of Fine-Tuned T5 and GPT-5

- 333299
Project Abstract
Favorite this paper
How to cite this paper?
Abstract

Introduction
AI is increasingly applied in qualitative research, yet the tension between procedural reproducibility and interpretive richness remains a challenge. Open coding, a core procedure in qualitative analysis, is central to this issue. However, systematic studies on applying language models to open coding remain limited.

Goals and Methods
This study explored how language models may support open coding by comparing two representative approaches: fine-tuning an open-source model, and prompt engineering with a commercial model. For the former, we fine-tuned the Japanese-pretrained open-source model Retrieva-jp/t5-base-long (hereafter, T5) to generate codes directly from text. For the latter, we instructed OpenAI GPT-5 to follow the sequential steps of SCAT (Steps for Coding and Theorization, a structured method for coding qualitative data developed in Japan) and produce a final code after intermediate outputs.

The dataset comprised 1,977 Japanese text–label pairs from 152 published SCAT studies, where labels had been assigned by the original authors. These labels are not definitive “correct answers” but served as reference points. 200 pairs were randomly selected for testing, and the remainder used for training. Outputs were evaluated quantitatively, using BERTScore, and qualitatively through interpretive assessment. Given architectural differences, the comparison is exploratory rather than a strict performance test.

Results
T5 model achieved a mean BERTScore of 0.727 (SD = 3.89×10⁻³), while GPT-5 scored 0.696 (SD = 3.30×10⁻³). T5 model frequently reproduced surface-level terms, yielding higher similarity scores. GPT-5 often produced paraphrases and abstractions, resulting in lower scores but outputs that more closely resembled interpretive researcher coding.

Conclusion
T5 model was effective for reproducible, summary-oriented coding, whereas GPT-5 better approximated interpretive, process-oriented coding. This exploratory comparison suggests that model choice should align with research goals. Future work should involve controlled comparisons within the same model family and multiple evaluation criteria to clarify AI’s role in qualitative analysis.

Share your ideas or questions with the authors!

Did you know that the greatest stimulus in scientific and cultural development is curiosity? Leave your questions or suggestions to the author!

Sign in to interact

Have a question or suggestion? Share your feedback with the authors!

Institutions
  • 1 Institute of Science Tokyo
Track
  • 3. Qualitative Research in Social Science
Keywords
Open Coding
AI
fine-tuning
prompt engineering
BERTScore