sözaltı news Science
Science
EN AZ

Designing fairness: best practices for gender-sensitive development of cognitive ability tests in recruitment

nature.com 25.09.2026 02:00 3 views

Psychological tests play a pivotal role in high-stakes decisions such as recruitment, yet traditional development guidelines concentrate fairness work downstream—at test administration and post hoc statistical bias correction—while offering little concrete guidance for the design stage, where constructs are defined and operationalized. Drawing on constructivist and organizational justice theories, we argue that a “bias by design” arises when unexamined default assumptions narrow construct representation, systematically disadvantaging groups before any item is administered. We develop this for cognitive ability tests, which exhibit the largest subgroup differences among common selection instruments and routinely rely on figural matrices as proxies for general ability, thereby disadvantaging women.

We contrast procedural justice (standardized administration) with distributive justice (equitable score distributions), clarifying that the latter applies to a principled class of constructs—latent and content-general, such as fluid intelligence—for which content-linked group differences signal construct-irrelevant variance rather than true differences. To address this gap, we adapt Stanford’s Gendered Innovations framework and demonstrate its application through the Modularer Kurzintelligenztest (M-KIT; Dantlgraber et al., 2015). Fairness can be addressed at three hierarchically ordered levels, where higher levels constrain lower ones: the theoretical model (broadening construct representation), the task format (removing construct-irrelevant demands), and the item (differential item functioning as final refinement, not primary remedy).

We distill practical recommendations for embedding distributive justice early, including questioning default assumptions and documenting fairness deliberations. Our approach reframes fairness not as a trade-off but as integral to validity, offering a roadmap for more inclusive assessments. Its impact (like that of ideology in general) is always indirect: in the formation and selection of preferred goals, values, methodologies, and explanations” (Keller, 1995, p. 137).

Psychological tests are widely used to inform recruitment, educational, and other aptitude-related decisions. For example, tests of cognitive ability (Armoneit et al., 2020; Kulikowski et al., 2025), situational judgement tests (Cabrera and Nguyen, 2001), or achievement motivation tests (Schmidt-Atzert, 2004) could indicate whether a person fits into a certain degree program or job position. Indeed, test scores in the past have been documented to play a crucial role in predicting important life outcomes such as academic (Rohde and Thompson, 2007) and job performance (Schmidt and Hunter, 1998; Steel and Faroborzi, 2024).

Thus, because of the wide-reaching consequences a test-based decision could have for an applicant’s or student’s future, it is important to ensure that no group of individuals systematically scores higher or lower than other groups for reasons that cannot be reasonably justified – that is, to ensure the criterion of test fairness is fulfilled. Not surprisingly, there are lively discussions about fairness regarding an applicant’s cultural background (Reynolds and Suzuki, 2012; Stumpf et al., 2017; Wilson et al., 2023), ethnicity (Chung-Yan and Cronshaw, 2002; Reynolds et al., 2021; Schmidt and Hunter, 1974), and gender (Herzberg, 2024; Sparfeldt and Schult, 2024; Striebing et al., 2024). The guidelines typically used as a basis for developing psychological tests – and, still more so, the degree to which developers adhere to them – concentrate fairness work at the stages of test administration and post-hoc statistical evaluation, while offering comparatively little concrete guidance for the design stage; even those few recommendations that do apply to design are followed inconsistently in practice (e.g., Jonson et al., 2019).

In this paper, we argue that different conceptions of fairness play a crucial role in producing this asymmetry. Where emphasis is placed on the test criterion of objectivity, attention gravitates toward procedural justice, that is, toward process-related aspects of testing. We propose that greater attention should be given to distributive justice—how test results are allocated across subgroups—and, crucially, that this attention must be paid at the point where constructs are defined and operationalized in recruitment settings.

Extract — continue reading at the source.

Read full story