sözaltı news Science
Science
EN AZ

Modeling hierarchical response accuracy of large language models using a Bayesian generalized linear mixed model

nature.com 23.09.2026 02:00 4 views

Large Language Models (LLMs) demonstrate strong performance in Natural Language Processing. In neuroanatomy, their response accuracy across hierarchical data structures (responses nested within questions) has not yet been investigated. Evaluating how these models perform under varying cognitive demands is important for their effective integration into medical education.

A total of 183 neuroanatomy questions from the Turkish Medical Specialization Examination (2013–2021) were reviewed, and 176 single correct answer questions were obtained by excluding seven figure-based questions. Each question was reformulated into prompts representing three role-based prompt levels: medical student, recently graduated physician, and doctoral-level specialist. Responses from the three LLMs were scored as correct or incorrect.

Data were analyzed using a Bayesian Generalized Linear Mixed Model (GLMM) with a mixture prior for the random intercept, with model type, prompt level, and their interaction as fixed effects. The distribution of the question-level intercept was highly skewed, showing a ceiling for the items answered correctly by all LLMs under all prompt levels (141 of 176 questions, 80.1%). To deal with this excess of one, a spike and slab prior was specified for the random intercept, separating a ceiling component (spike) from a variable-difficulty component (slab).

Posterior estimates were obtained using Monte Carlo sampling with CmdStan in Python. ChatGPT-4o and Microsoft Copilot showed higher overall accuracy than Google Gemini, with ChatGPT-4o performing best across all prompt levels. The spike and slab model showed better out-of-sample predictive performance than the standard model specifications (ΔELPD-WAIC = 22.64).

The posterior global spike probability (\(\:\varphi\:\)=0.706, 95% HDI [0.516, 0.826]) indicated that 141 questions (80.1%) were assigned to the ceiling component, while the remaining 35 questions (19.9%) showed true variability in item difficulty. The question level random intercept standard deviation for the slab component (\(\:_\)=2.158, 95% HDI [1.289, 3.312]) reflected substantial between item variability in baseline easiness among non-ceiling neuroanatomical questions. The Bayesian GLMM approach implemented with a spike and slab prior distribution modeled the hierarchical structure of the data where responses were nested within questions.

The ceiling effect was addressed by decomposing question level random intercepts into separate components representing the ceiling and variable difficulty providing an estimate of performance variability. It was obtained between-item variability in LLM performance among non-ceiling questions. Role-based prompt differentiation alone was insufficient to generate systematic differences in model accuracy.

Extract — continue reading at the source.

Read full story