
10
J Gandhara Med Dent Sci
and equity.
LIMITATIONS
Limitations include the single-institution context,
limiting generalisability; the nested structure of
reections within students; the modest sample size for
subgroup analyses; and the proprietary nature of AI
systems, which limits transparency regarding training
data and potential sources of bias. Future studies should
employ larger and more diverse datasets, incorporate
formal measures of language prociency, and use
multilevel modelling to enhance analytic precision.
CONCLUSIONS
GenAI oers promising opportunities to enhance
efciency and consistency in assessing reective
writing, particularly in terms of structural and linguistic
features. However, reective writing is fundamentally a
human, relational, and interpretive practice, deeply tied
to learner identity, emotional insight, and contextual
meaning-making. In this study, GenAI consistently
underperformed on higher-order reective constructs
and raised concerns about linguistic fairness among
participants. These ndings underscore that GenAI
should supplement, not supplant, human evaluators.
Ethical and educationally sound integration requires
hybrid models, transparent governance, fairness
monitoring, and explicit protection of teacher agency.
As AI becomes more embedded in health professions
education, maintaining human judgment at the core of
reective assessment remains essential for both the
integrity of learning and the development of reective
practitioners.
CONFLICT OF INTEREST: None
FUNDING SOURCES: None
REFERENCES
1. Schön DA. The reective practitioner: How professionals think
in action. London: Temple Smith; 1983.
2. Mann K, Gordon J, MacLeod A. Reection and reective
practice in health professions education: a systematic review.
Adv Health Sci Educ Theory Pract. 2009;14(4):595–621.
https://doi.org/10.1007/s10459-007-9090-2. PMID: 18274874
3. Sandars J. The use of reflection in medical education: AMEE
Guide No. 44. Med Teach. 2009;31(8):685–95.
https://doi.org/10.1080/01421590903050374. PMID: 19811128
4. Adeani IS, Febriani RB, Syafryadin S. Using Gibbs’ reective
cycle in making reections of literary analysis. Indonesian EFL
J. 2020;6(2):139–48. https://doi.org/10.25134/iej.v6i2.3385.
5. Jasper M, Rosser M. Reection and reective practice. In:
Jasper M, Rosser M, editors. Professional development,
reection, and decision-making in nursing and healthcare.
Chichester: Wiley-Blackwell; 2013. p. 41–82.
6. Moon J. Using reective learning to improve the impact of short
courses and workshops. J Contin Educ Health Prof.
2004;24(1):4–11.
https://doi.org/10.1002/chp.1340240103.PMID: 15069907
7. Ryan M, Ryan M. Theorising a model for teaching and
assessing reective learning in higher education. High Educ Res
Dev.2013;32(2):244–57.
https://doi.org/10.1080/07294360.2012.661704.
8. Wald HS, Borkan JM, Taylor JS, Anthony D, Reis SP.
Fostering and evaluating reective capacity in medical
education: developing the REFLECT rubric for assessing
reective writing. Acad Med. 2012;87(1):41–50.
https://doi.org/10.1097/ACM.0b013e31823b55fa.PMID:221040
58
9. Williamson S, Seewoodhary R. A review and reection on the
visual rehabilitation progress of an older person following
cataract surgery two years on. Int J Ther Rehabil.
2016;23(5):242–6. https://doi.org/10.12968/ijtr.2016.23.5.242.
10. Kumar P. Large language models (LLMs): survey, technical
frameworks, and future challenges. Artif Intell Rev.
2024;57(10):260. https://doi.org/10.1007/s10462-024-10874-9.
11. Gallegos IO, Rossi RA, Barrow J, Tanjim MM, Kim S,
Dernoncourt F, et al. Bias and fairness in large language
models: a survey. Comput Linguist. 2024;50(3):1097–179.
https://doi.org/10.1162/coli_a_00523.
12. Liang J-C, Hwang G-J, Chen M-RA, Darmawansah D. Roles
and research foci of artificial intelligence in language education:
an integrated bibliographic analysis and systematic review
approach. Interact Learn Environ. 2023;31(7):4270–96.
https://doi.org/10.1080/10494820.2021.1982642.
13. Williamson B, Macgilchrist F, Potter J. Re-examining AI,
automation and datacation in education. Abingdon:
Routledge/Taylor & Francis; 2023. p. 1–5.
14. Lee D, Arnold M, Srivastava A, Plastow K, Strelan P, Ploeckl
F, et al. The impact of generative AI on higher education
learning and teaching: a study of educators’ perspectives.
Comput Educ Artif Intell. 2024;6:100221.
https://doi.org/10.1016/j.caeai.2024.100221.
15. Creswell JW, Plano Clark VL. Revisiting mixed methods
research designs twenty years later. Handb Mixed Methods Res
Des. 2023;1(1):21–36.
16. Hox J, de Leeuw E, Klausch T. Mixed-mode research: issues in
design and analysis. In: Biemer PP, de Leeuw ED, Eckman S,
Kreuter F, Lyberg LE, Tucker C, West BT, editors. Total survey
error in practice. Hoboken (NJ): Wiley; 2017. p. 511–30.
17. Raudenbush SW, Bryk AS. Hierarchical linear models:
applications and data analysis methods. 2nd ed. Thousand Oaks
(CA): Sage Publications; 2002.
18. Kember D, McKay J, Sinclair K, Wong FKY. A four-category
scheme for coding and assessing the level of reection in
written work. Assess Eval High Educ. 2008;33(4):369–79.
https://doi.org/10.1080/02602930701293355.
19. Polit DF, Beck CT, Owen SV. Is the CVI an acceptable
indicator of content validity? Appraisal and recommendations.
Res Nurs Health. 2007;30(4):459–67.
https://doi.org/10.1002/nur.20199. PMID: 17654487
20. Jonsson A, Svingby G. The use of scoring rubrics: reliability,
validity and educational consequences. Educ Res Rev.
2007;2(2):130-44.https:doi.org/10.1016/j.edurev.2007.05.002.
21. Chinta SV, Wang Z, Yin Z, Hoang N, Gonzalez M, Quy TL, et
al. FairAIED: Navigating fairness, bias, and ethics in
educational AI applications. arXiv preprint arXiv:2407.18745.
2024. Available from: https://arxiv.org/abs/2407.18745.
Can Articial Intelligence (AI) Judge Reection
January - March 2026