How much data is needed for predictive models of consumer choice?
Decision making often occurs under bounded cognitive ressources. Building on data from Amasino et al., 2025 a large study about consumer choices, we examine how well machine learning methods and cognitive models can predict the study participants' decisions. While our results demonstrate that individual product choices can be predicted from past behavior with high accuracy, we further aim to identify the number of trials and participants required to ensure stable predictive performance. We examine the choice task through an analysis that varies both the number of trials per subject and the number of subjects used for model training. Using data from 302 participants with up to 42 trials each, we train Random Forest and CNN-Transformer models to establish an upper performance bound as data increases. Predictive performance saturates at moderate sample sizes, indicating limits to behavioural predictability under bounded cognitive resources. Random Forest models achieve accuracies above 80% with fewer than 20 trials per subject and roughly 50 subjects. Contrary to the common assumption that neural networks require large datasets, our CNN-based model needed only slightly more data and trained considerably faster. A cognitive decision tree capturing the key determinants of participants’ choices was derived from the identified features.
Keywords
There is nothing here yet. Be the first to create a thread.
Cite this as: