Abstract
The 10-item Big Five Inventory (BFI-10) is a personality measure widely used in numerous contexts, including large cohort studies. We report that its openness subscale (two items) is flawed due to a linguistically ambiguous item: ‘I see myself as someone who has few artistic interests’. The word ‘few’ is often misinterpreted as ‘some’ instead of its intended meaning of ‘almost none’, leading to inconsistent responses. Across nine datasets (N = 6,640), this item shows weak or no inter-item correlations with the other openness item, contrasting sharply with those observed within the other four personality dimensions. To confirm that the issue stems from wording rather than semantics, we tested a revised item (‘…has hardly any artistic interests’), which improves inter-item correlations, test–retest reliability, criterion validity, and statistical power. Simulation analyses show that the original item substantially reduces statistical power to detect true changes in openness. The problematic item also appears in the 44-item BFI, suggesting possible misinterpretations in the aesthetics facet. To address existing datasets, we evaluated adjustment methods through additional simulations, showing that regression-based adjustment approaches can partially but meaningfully reduce bias. We recommend re-evaluating prior findings, applying the correction strategies to existing datasets, and updating future data collection.
Plain Language Summary
The 10-item Big Five Inventory (BFI-10) is a short and popular test used to measure personality. We found that one of its two questions assessing ‘openness’ (receptiveness to new ideas, experiences, and variety) is confusing and may lead respondents to answer incorrectly. The question reads, ‘I see myself as someone who has few artistic interests’. Many people interpret the word ‘few’ to mean ‘some’, when it actually means ‘almost none’. We analyzed data from over 6,600 participants across nine datasets and consistently found evidence for this misinterpretation. Specifically, answers to this question do not match well with answers to the other openness question, making the openness score unreliable. Moreover, when we replaced the word ‘few’ with ‘hardly any’, the results became more consistent and reliable, and the change improved several indicators of measurement quality. For researchers working with existing datasets that include the original question, we also show that statistical adjustments can partially correct for the bias. Because this flawed question has been used for many years and also appears in the longer 44-item version of the questionnaire, some previous findings about openness might need re-evaluation. We suggest using the revised version and paying closer attention to how subtle wording choices can affect personality research.
Keywords
Get full access to this article
View all access options for this article.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
