Data Quality as a Causal Variable in Smartphone Health Monitoring

Abdulrahman A. Aloyayri, Max A. Little, Nawfal A. Zakar
Proceedings of the Fourth UK AI Conference 2026, PMLR 348:1-11, 2026.

Abstract

Smartphone-based health monitoring studies commonly treat data quality as a preprocessing step: flag or remove low-quality recordings, then proceed with the remainder. This places data quality outside the model when it belongs inside, as part of the data-generating process if it is associated with the outcome of interest, filtering on it changes what the remaining data represent. We test this on the mPower 2.0 Parkinson’s disease dataset (81 participants, 15,398 resting tremor recordings), applying an automated assessment framework at three levels – study-level session completion, task-level recording duration, and signal-level flat-segment detection – and testing each check against diagnosis at the participant level. The three behave differently. Duration filtering removes recordings from Parkinson’s disease participants (clustered odds ratio 2.23, 95% CI 1.50–3.32); flat signal filtering removes control recordings (OR 0.52, 0.23–1.19), because on a resting tremor task low signal variation is the expected observation for a participant without tremor rather than a measurement failure; session completion shows no association once follow-up length is accounted for (OR 1.13, 0.53–2.37). Filtering on data quality is therefore not neutral, and its direction cannot be assumed. The main classification results are nonetheless stable after duration filtering. Data quality belongs inside the causal model as separate variables, each tested against the outcome before it is filtered on.

Cite this Paper


BibTeX
@InProceedings{pmlr-v348-aloyayri26a, title = {Data Quality as a Causal Variable in Smartphone Health Monitoring}, author = {Aloyayri, Abdulrahman A. and Little, Max A. and Zakar, Nawfal A.}, booktitle = {Proceedings of the Fourth UK AI Conference 2026}, pages = {1--11}, year = {2026}, editor = {Benford, Alistair and Büyükateş, Baturalp and Cabrera, Christian and Kiden, Sarah and Salili-James, Arianna and Zakka, Vincent and Zhou, Feng}, volume = {348}, series = {Proceedings of Machine Learning Research}, month = {29--30 Sep}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v348/main/assets/aloyayri26a/aloyayri26a.pdf}, url = {https://proceedings.mlr.press/v348/aloyayri26a.html}, abstract = {Smartphone-based health monitoring studies commonly treat data quality as a preprocessing step: flag or remove low-quality recordings, then proceed with the remainder. This places data quality outside the model when it belongs inside, as part of the data-generating process if it is associated with the outcome of interest, filtering on it changes what the remaining data represent. We test this on the mPower 2.0 Parkinson’s disease dataset (81 participants, 15,398 resting tremor recordings), applying an automated assessment framework at three levels – study-level session completion, task-level recording duration, and signal-level flat-segment detection – and testing each check against diagnosis at the participant level. The three behave differently. Duration filtering removes recordings from Parkinson’s disease participants (clustered odds ratio 2.23, 95% CI 1.50–3.32); flat signal filtering removes control recordings (OR 0.52, 0.23–1.19), because on a resting tremor task low signal variation is the expected observation for a participant without tremor rather than a measurement failure; session completion shows no association once follow-up length is accounted for (OR 1.13, 0.53–2.37). Filtering on data quality is therefore not neutral, and its direction cannot be assumed. The main classification results are nonetheless stable after duration filtering. Data quality belongs inside the causal model as separate variables, each tested against the outcome before it is filtered on.} }
Endnote
%0 Conference Paper %T Data Quality as a Causal Variable in Smartphone Health Monitoring %A Abdulrahman A. Aloyayri %A Max A. Little %A Nawfal A. Zakar %B Proceedings of the Fourth UK AI Conference 2026 %C Proceedings of Machine Learning Research %D 2026 %E Alistair Benford %E Baturalp Büyükateş %E Christian Cabrera %E Sarah Kiden %E Arianna Salili-James %E Vincent Zakka %E Feng Zhou %F pmlr-v348-aloyayri26a %I PMLR %P 1--11 %U https://proceedings.mlr.press/v348/aloyayri26a.html %V 348 %X Smartphone-based health monitoring studies commonly treat data quality as a preprocessing step: flag or remove low-quality recordings, then proceed with the remainder. This places data quality outside the model when it belongs inside, as part of the data-generating process if it is associated with the outcome of interest, filtering on it changes what the remaining data represent. We test this on the mPower 2.0 Parkinson’s disease dataset (81 participants, 15,398 resting tremor recordings), applying an automated assessment framework at three levels – study-level session completion, task-level recording duration, and signal-level flat-segment detection – and testing each check against diagnosis at the participant level. The three behave differently. Duration filtering removes recordings from Parkinson’s disease participants (clustered odds ratio 2.23, 95% CI 1.50–3.32); flat signal filtering removes control recordings (OR 0.52, 0.23–1.19), because on a resting tremor task low signal variation is the expected observation for a participant without tremor rather than a measurement failure; session completion shows no association once follow-up length is accounted for (OR 1.13, 0.53–2.37). Filtering on data quality is therefore not neutral, and its direction cannot be assumed. The main classification results are nonetheless stable after duration filtering. Data quality belongs inside the causal model as separate variables, each tested against the outcome before it is filtered on.
APA
Aloyayri, A.A., Little, M.A. & Zakar, N.A.. (2026). Data Quality as a Causal Variable in Smartphone Health Monitoring. Proceedings of the Fourth UK AI Conference 2026, in Proceedings of Machine Learning Research 348:1-11 Available from https://proceedings.mlr.press/v348/aloyayri26a.html.

Related Material