Boosting AI-based speech performance assessment through pairwise comparative scoring of crowd data
Open-ended performance assessments can capture intellectual status that is difficult to measure using highly constrained response formats. However, their practical use is limited by the difficulty of obtaining reliable, scalable, and interpretable scores. Recent advances in large language models provide new opportunities for automated scoring, but the measurement quality of AI-generated scores remains an important empirical question.
We evaluated AI models for complex speech-based assessment, and proposed a pairwise comparative scoring framework designed to improve reliability, while examining validity evidence through associations with working memory. Four speech tasks, each with three difficulty levels, were administered to 122 participants. We applied three AI scoring approaches: (a) individual scoring, where each speech response was assessed separately; (b) pairwise scoring with absolute scores, where paired responses were presented together and each participant received an absolute score; and (c) pairwise scoring with relative scores, where each score was converted into a within-pair difference.
Human ratings provided a conventional benchmark. Working memory was measured to examine theoretically relevant validity evidence. Pairwise comparative scoring showed higher reliability than individual scoring, reaching levels comparable to human scores averaged across multiple raters, while providing theoretically consistent validity evidence.
These results support AI-based comparative scoring for complex open-ended assessment. This work was partially supported by Seed Funds from the University of Hong Kong (2402101405, 2502251399, 2503251476). The funders have/had no role in study design, data collection and analysis, decision to publish or preparation of the manuscript.
Complex Neural Signals Decoding Lab, Faculty of Education, The University of Hong Kong, Pokfulam, Hong Kong Island, Hong Kong The authors declare no competing interests. Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Below is the link to the electronic supplementary material.
Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material.
Extract — continue reading at the source.