Evaluating Command-Line Interface-Based Agentic Large Language Model Coding Tools for Non-English Thematic Analysis: A 3 x 3 x 3 Factorial Study of Models, Prompts and Stochastic Variability
Keywords:
Agentic coding tools; thematic analysis; large language models; prompt engineering; natural language processingAbstract
Command-line interface (CLI)-based agentic coding tools enable large language models (LLMs) to autonomously read local data files, write and execute analysis scripts, and generate outputs, but whether they can reliably perform thematic analysis of non-English educational data remains understudied. This 3 × 3 × 3 factorial experiment compared three frontier LLMs (Claude Opus 4.6, GPT-5.3-Codex, and Gemini 3 Pro Preview) across three prompt tiers (detailed, structured, and exploratory) with three repetitions per condition, yielding 27 runs. Each run processed 89 anonymized student evaluations (Likert ratings and open-ended Thai comments) from a Year 3 preventive dentistry course over three academic years. Outputs were evaluated against 10 required items covering statistical computation, thematic analysis, and visualization, with the generated scripts inspected to classify categorization as direct reading or keyword matching. Statistical computation and overview charts each passed 93% (25/27), and detail charts passed in all 27 runs, with failures confined to exploratory prompts. Thematic reading was the primary difference. Opus and Gemini read Thai directly while GPT usually generated keyword-matching scripts. Structured prompts achieved the highest thematic reading pass rate at 78% (7/9), followed by detailed at 67% (6/9) and exploratory at 33% (3/9). Identical configurations produced different outputs across repetitions, with theme counts ranging from 4 to 31. CLI-based agentic coding tools can perform thematic analyses of Thai-language educational feedback, but natural language processing output reliability depends on model selection, prompt specificity, and stochastic variability. Multiple independent runs are recommended to assess consistency.
https://doi.org/10.26803/ijlter.25.8.5
References
Almulla, M., & Ali, S. I. (2024). The changing educational landscape for sustainable online experiences: Implications of ChatGPT in Arab students' learning experience. International Journal of Learning, Teaching and Educational Research, 23(9), 285–306. https://doi.org/10.26803/ijlter.23.9.15
Arreerard, R., Mander, S., & Piao, S. (2022). Survey on Thai NLP language resources and tools. Proceedings of the Thirteenth Language Resources and Evaluation Conference, 6495–6505. https://aclanthology.org/2022.lrec-1.697/
Atil, B., Aykent, S., Chittams, A., Fu, L., Passonneau, R. J., Radcliffe, E., Rajagopal, G. R., Sloan, A., Tudrej, T., Ture, F., Wu, Z., Xu, L., & Baldwin, B. (2024). Non-determinism of "deterministic" LLM settings. arXiv. https://doi.org/10.48550/arXiv.2408.04667
Bano, M., Hoda, R., Zowghi, D., & Treude, C. (2024). Large language models for qualitative research in software engineering: Exploring opportunities and challenges. Automated Software Engineering, 31(1), Article 8. https://doi.org/10.1007/s10515-023-00407-8
Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77–101. https://doi.org/10.1191/1478088706qp063oa
Braun, V., & Clarke, V. (2019). Reflecting on reflexive thematic analysis. Qualitative Research in Sport, Exercise and Health, 11(4), 589–597. https://doi.org/10.1080/2159676X.2019.1628806
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., ... Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901. https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf
Cai, C., Hong, S., Ma, M., Feng, H., Du, S., Chow, M., Teo, W. L., Liu, S., & Fan, X. (2025). Analyzing the teaching and learning environments through student feedback at scale: A multi-agent LLMs framework. Education and Information Technologies, 30(15), 21815–21847. https://doi.org/10.1007/s10639-025-13633-2
Denny, P., Prather, J., Becker, B. A., Finnie-Ansley, J., Hellas, A., Leinonen, J., Luxton-Reilly, A., Reeves, B. N., Santos, E. A., & Sarsa, S. (2024). Computing education in the era of generative AI. Communications of the ACM, 67(2), 56–67. https://doi.org/10.1145/3624720
El-Hakim, M., Anthonappa, R., & Fawzy, A. (2025). Artificial intelligence in dental education: A scoping review of applications, challenges, and gaps. Dentistry Journal, 13(9), Article 384. https://doi.org/10.3390/dj13090384
Gridach, M., Nanavati, J., Zine El Abidine, K., Mendes, L., & Mack, C. (2025). Agentic AI for scientific discovery: A survey of progress, challenges, and future directions. arXiv. https://doi.org/10.48550/arXiv.2503.08979
Herrera-Poyatos, D., Pelaez-Gonzalez, C., Zuheros, C., Herrera-Poyatos, A., Tejedor, V., Herrera, F., & Montes, R. (2025). An overview of model uncertainty and variability in LLM-based sentiment analysis: Challenges, mitigation strategies, and the role of explainability. Frontiers in Artificial Intelligence, 8, Article 1609097. https://doi.org/10.3389/frai.2025.1609097
Jiang, J., Wang, F., Shen, J., Kim, S., & Kim, S. (2026). A survey on large language models for code generation. ACM Transactions on Software Engineering and Methodology, 35(2), Article 58, 1–72. https://doi.org/10.1145/3747588
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. R. (2024). SWE-bench: Can language models resolve real-world GitHub issues? Proceedings of the Twelfth International Conference on Learning Representations (ICLR). https://doi.org/10.48550/arXiv.2310.06770
Jita, T., & Thaanyane, M. E. (2025). Evaluating the validity and impact of teaching practice assessment tools in higher education. International Journal of Learning, Teaching and Educational Research, 24(7), 486–503. https://doi.org/10.26803/ijlter.24.7.24
Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., ... Kasneci, G. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, Article 102274. https://doi.org/10.1016/j.lindif.2023.102274
Kim, D., Lee, S., Kim, Y., Rutherford, A., & Park, C. (2025). Representing the under-represented: Cultural and core capability benchmarks for developing Thai large language models. In *Proceedings of the 31st International Conference on Computational Linguistics* (pp. 4114–4129). Association for Computational Linguistics. https://aclanthology.org/2025.coling-main.278/
Li, Q., Peng, H., Li, J., Xia, C., Yang, R., Sun, L., Yu, P. S., & He, L. (2022). A survey on text classification: From traditional to deep learning. ACM Transactions on Intelligent Systems and Technology, 13(2), 1–41. https://doi.org/10.1145/3495162
Limkonchotiwat, P., Masuk, K., Nonesung, S., Mai-On, C., Nutanong, S., Ponwitayarat, W., & Manakul, P. (2025). Assessing Thai dialect performance in LLMs with automatic benchmarks and human evaluation. arXiv. https://doi.org/10.48550/arXiv.2504.05898
Lucas, H. C., Upperman, J. S., & Robinson, J. R. (2024). A systematic review of large language models and their implications in medical education. Medical Education, 58(11), 1276–1285. https://doi.org/10.1111/medu.15402
Martinez Montes, C., Feldt, R., Miguel Martos, C., Ouhbi, S., Premanandan, S., & Graziotin, D. (2025). Large language models in thematic analysis: Prompt engineering evaluation and guidelines for qualitative software engineering research. arXiv. https://doi.org/10.48550/arXiv.2510.18456
Miller, J. K., & Tang, W. (2025). Evaluating LLM metrics through real-world capabilities. arXiv. https://doi.org/10.48550/arXiv.2505.08253
Parker, M. J., Anderson, C., Stone, C., & Oh, Y. R. (2025). A large language model approach to educational survey feedback analysis. International Journal of Artificial Intelligence in Education, 35(2), 444–481. https://doi.org/10.1007/s40593-024-00414-0
Phatthiyaphaibun, W., Chaovavanich, K., Polpanumas, C., Suriyawongkul, A., Lowphansirikul, L., Chormai, P., Limkonchotiwat, P., Suntorntip, T., & Udomcharoenchaikit, C. (2023). PyThaiNLP: Thai natural language processing in Python. In *Proceedings of the 3rd Workshop for Natural Language Processing Open Source Software (NLP-OSS 2023)* (pp. 25–36). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.nlposs-1.4
Qin, L., Chen, Q., Zhou, Y., Chen, Z., Li, Y., Liao, L., Li, M., Che, W., & Yu, P. S. (2025). A survey of multilingual large language models. Patterns, 6(1), Article 101118. https://doi.org/10.1016/j.patter.2024.101118
Sahoo, P., Singh, A. K., Saha, S., Jain, V., Mondal, S., & Chadha, A. (2024). A systematic survey of prompt engineering in large language models: Techniques and applications. arXiv. https://doi.org/10.48550/arXiv.2402.07927
Saleh, Y., Abu Talib, M., Nasir, Q., & Dakalbab, F. (2025). Evaluating large language models: A systematic review of efficiency, applications, and future directions. Frontiers in Computer Science, 7, Article 1523699. https://doi.org/10.3389/fcomp.2025.1523699
Schulhoff, S., Ilie, M., Balepur, N., Kahadze, K., Liu, A., Si, C., Li, Y., Gupta, A., Han, H., Dulepet, P. S., Vidyadhara, S., Ki, D., Agrawal, S., Pham, C., Kroiz, G., Li, F., Tao, H., Srivastava, A., Da Costa, H., ... Resnik, P. (2024). The prompt report: A systematic survey of prompting techniques. arXiv. https://doi.org/10.48550/arXiv.2406.06608
Sun, M., Han, R., Jiang, B., Qi, H., Sun, D., Yuan, Y., & Huang, J. (2025). A survey on large language model-based agents for statistics and data science. The American Statistician, 1–14. https://doi.org/10.1080/00031305.2025.2561140
Tai, R. H., Bentley, L. R., Xia, X., Sitt, J. M., Fankhauser, S. C., Chicas-Mosier, A. M., & Monteith, B. G. (2024). An examination of the use of large language models to aid analysis of textual data. International Journal of Qualitative Methods, 23, Article 16094069241231168. https://doi.org/10.1177/16094069241231168
Wang, J. J., & Wang, V. X. (2025). Assessing consistency and reproducibility in the outputs of large language models: Evidence across diverse finance and accounting tasks. arXiv. https://doi.org/10.48550/arXiv.2503.16974
Wei, X., Cui, X., Cheng, N., Wang, X., Zhang, X., Huang, S., Xie, P., Xu, J., Chen, Y., Zhang, M., Jiang, Y., & Han, W. (2023). ChatIE: Zero-shot information extraction via chatting with ChatGPT. arXiv. https://doi.org/10.48550/arXiv.2302.10205
Wen, C., Clough, P., Paton, R., & Middleton, R. (2026). Leveraging large language models for thematic analysis: A case study in the charity sector. AI & Society, 41(1), 731–748. https://doi.org/10.1007/s00146-025-02487-4
Wijnen-Meijer, M., van den Broek, S., Koens, F., & ten Cate, O. (2020). Vertical integration in medical education: The broader perspective. BMC Medical Education, 20(1), Article 509. https://doi.org/10.1186/s12909-020-02433-6
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Chawalit Chanintonsongkhla, Thananya Chongcharoenkit

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
All articles published by IJLTER are licensed under a Creative Commons Attribution Non-Commercial No-Derivatives 4.0 International License (CCBY-NC-ND4.0).