[edit]
Data Quality as a Causal Variable in Smartphone Health Monitoring
Proceedings of the Fourth UK AI Conference 2026, PMLR 348:1-11, 2026.
Abstract
Smartphone-based health monitoring studies commonly treat data quality as a preprocessing step: flag or remove low-quality recordings, then proceed with the remainder. This places data quality outside the model when it belongs inside, as part of the data-generating process if it is associated with the outcome of interest, filtering on it changes what the remaining data represent. We test this on the mPower 2.0 Parkinson’s disease dataset (81 participants, 15,398 resting tremor recordings), applying an automated assessment framework at three levels – study-level session completion, task-level recording duration, and signal-level flat-segment detection – and testing each check against diagnosis at the participant level. The three behave differently. Duration filtering removes recordings from Parkinson’s disease participants (clustered odds ratio 2.23, 95% CI 1.50–3.32); flat signal filtering removes control recordings (OR 0.52, 0.23–1.19), because on a resting tremor task low signal variation is the expected observation for a participant without tremor rather than a measurement failure; session completion shows no association once follow-up length is accounted for (OR 1.13, 0.53–2.37). Filtering on data quality is therefore not neutral, and its direction cannot be assumed. The main classification results are nonetheless stable after duration filtering. Data quality belongs inside the causal model as separate variables, each tested against the outcome before it is filtered on.