Features, Labels, and the Politics of Measurement: How Data Definitions Shape Machine Learning
Features, labels, and the politics of measurement examines how machine-learning systems transform messy social, institutional, and technical realities into variables that algorithms can process. This article explains why features are not neutral inputs and labels are not simple truths: both depend on definitions, measurement practices, institutional priorities, historical data, and human judgment. It explores construct validity, proxy variables, annotation, classification, target definition, missing data, measurement error, bias, fairness, documentation, and governance. The article shows how predictive systems can misread the world when data categories flatten context, encode institutional history, or mistake what is measurable for what matters. By connecting machine learning with measurement theory, social classification, and algorithmic accountability, it argues that responsible computational reasoning must examine how data are defined before models are trained, evaluated, deployed, or trusted in public, scientific, commercial, educational, and administrative settings alike today.









