Evidence Against Homogeneity: Identifying Process Mismatch for Copy Number Variation Detection

Austin Talbot, Alex V. Kotlar, Yue Ke
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:1950-1975, 2026.

Abstract

Batch effects represent a major confounder in genomic diagnostics. In copy number variant (CNV) detection from next-generation sequencing, many algorithms compare read depth between test samples and a reference derived from the processing batch, assuming samples are process-matched. When this assumption is violated, with causes ranging from reagent lot changes and sample quality differences to multi-site processing, the reference becomes inappropriate, introducing false CNV calls or masking true pathogenic variants. Detecting such heterogeneity before downstream analysis is critical for reliable clinical interpretation. Existing batch effect detection methods either cluster samples based on raw features, risking conflation of biological signal with technical variation, or require known batch labels that are frequently unavailable. We introduce a method that addresses both limitations by clustering samples according to their Bayesian model evidence. The central insight is that evidence quantifies compatibility between data and model assumptions, technical artifacts violate assumptions and reduce evidence, whereas biological variation, including CNV status, is anticipated by the model and yields high evidence. This asymmetry provides a discriminative signal that separates batch effects from biology. We formalize heterogeneity detection as a likelihood ratio test for mixture structure in evidence space, using parametric bootstrap calibration to ensure conservative false positive rates. We validate our approach on synthetic data demonstrating proper Type I error control, three clinical targeted sequencing panels (liquid biopsy, BRCA, and thalassemia) exhibiting distinct batch effect mechanisms, and mouse electrophysiology recordings demonstrating cross-modality generalization. Our method achieves superior clustering accuracy compared to standard correlation-based and dimensionality-reduction approaches while maintaining the conservativeness required for clinical usage.

Cite this Paper


BibTeX
@InProceedings{pmlr-v340-talbot26a, title = {Evidence Against Homogeneity: Identifying Process Mismatch for Copy Number Variation Detection}, author = {Talbot, Austin and Kotlar, Alex V. and Ke, Yue}, booktitle = {Proceedings of the 11th Machine Learning for Healthcare Conference}, pages = {1950--1975}, year = {2026}, editor = {Krishnan, Rahul G. and van Amsterdam, Wouter A. C. and Chopra, Sumit and Overgaard, Shauna and Hughes, Michael and Ötleş, Erkin and Shen, Yiqiu and Shanmugam, Divya and Nayan, Madhur and Engelhard, Matthew and Fackler, Jim and Oberst, Michael}, volume = {340}, series = {Proceedings of Machine Learning Research}, month = {12--14 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v340/main/assets/talbot26a/talbot26a.pdf}, url = {https://proceedings.mlr.press/v340/talbot26a.html}, abstract = {Batch effects represent a major confounder in genomic diagnostics. In copy number variant (CNV) detection from next-generation sequencing, many algorithms compare read depth between test samples and a reference derived from the processing batch, assuming samples are process-matched. When this assumption is violated, with causes ranging from reagent lot changes and sample quality differences to multi-site processing, the reference becomes inappropriate, introducing false CNV calls or masking true pathogenic variants. Detecting such heterogeneity before downstream analysis is critical for reliable clinical interpretation. Existing batch effect detection methods either cluster samples based on raw features, risking conflation of biological signal with technical variation, or require known batch labels that are frequently unavailable. We introduce a method that addresses both limitations by clustering samples according to their Bayesian model evidence. The central insight is that evidence quantifies compatibility between data and model assumptions, technical artifacts violate assumptions and reduce evidence, whereas biological variation, including CNV status, is anticipated by the model and yields high evidence. This asymmetry provides a discriminative signal that separates batch effects from biology. We formalize heterogeneity detection as a likelihood ratio test for mixture structure in evidence space, using parametric bootstrap calibration to ensure conservative false positive rates. We validate our approach on synthetic data demonstrating proper Type I error control, three clinical targeted sequencing panels (liquid biopsy, BRCA, and thalassemia) exhibiting distinct batch effect mechanisms, and mouse electrophysiology recordings demonstrating cross-modality generalization. Our method achieves superior clustering accuracy compared to standard correlation-based and dimensionality-reduction approaches while maintaining the conservativeness required for clinical usage.} }
Endnote
%0 Conference Paper %T Evidence Against Homogeneity: Identifying Process Mismatch for Copy Number Variation Detection %A Austin Talbot %A Alex V. Kotlar %A Yue Ke %B Proceedings of the 11th Machine Learning for Healthcare Conference %C Proceedings of Machine Learning Research %D 2026 %E Rahul G. Krishnan %E Wouter A. C. van Amsterdam %E Sumit Chopra %E Shauna Overgaard %E Michael Hughes %E Erkin Ötleş %E Yiqiu Shen %E Divya Shanmugam %E Madhur Nayan %E Matthew Engelhard %E Jim Fackler %E Michael Oberst %F pmlr-v340-talbot26a %I PMLR %P 1950--1975 %U https://proceedings.mlr.press/v340/talbot26a.html %V 340 %X Batch effects represent a major confounder in genomic diagnostics. In copy number variant (CNV) detection from next-generation sequencing, many algorithms compare read depth between test samples and a reference derived from the processing batch, assuming samples are process-matched. When this assumption is violated, with causes ranging from reagent lot changes and sample quality differences to multi-site processing, the reference becomes inappropriate, introducing false CNV calls or masking true pathogenic variants. Detecting such heterogeneity before downstream analysis is critical for reliable clinical interpretation. Existing batch effect detection methods either cluster samples based on raw features, risking conflation of biological signal with technical variation, or require known batch labels that are frequently unavailable. We introduce a method that addresses both limitations by clustering samples according to their Bayesian model evidence. The central insight is that evidence quantifies compatibility between data and model assumptions, technical artifacts violate assumptions and reduce evidence, whereas biological variation, including CNV status, is anticipated by the model and yields high evidence. This asymmetry provides a discriminative signal that separates batch effects from biology. We formalize heterogeneity detection as a likelihood ratio test for mixture structure in evidence space, using parametric bootstrap calibration to ensure conservative false positive rates. We validate our approach on synthetic data demonstrating proper Type I error control, three clinical targeted sequencing panels (liquid biopsy, BRCA, and thalassemia) exhibiting distinct batch effect mechanisms, and mouse electrophysiology recordings demonstrating cross-modality generalization. Our method achieves superior clustering accuracy compared to standard correlation-based and dimensionality-reduction approaches while maintaining the conservativeness required for clinical usage.
APA
Talbot, A., Kotlar, A.V. & Ke, Y.. (2026). Evidence Against Homogeneity: Identifying Process Mismatch for Copy Number Variation Detection. Proceedings of the 11th Machine Learning for Healthcare Conference, in Proceedings of Machine Learning Research 340:1950-1975 Available from https://proceedings.mlr.press/v340/talbot26a.html.

Related Material