A Goemans-Williamson type algorithm for identifying subcohorts in clinical trials

Pratik Worah
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:3394-3402, 2026.

Abstract

We design an efficient algorithm that outputs tests for identifying predominantly homogeneous subcohorts of patients from large in-homogeneous datasets. Our theoretical contribution is a rounding technique, similar to that of Goemans and Wiliamson (1995), which approximates the optimal solution within a factor of $0.82$. As an application, we use our algorithm to trade-off sensitivity for specificity to systematically identify clinically interesting homogeneous subcohorts of patients in the RNA microarray data set for breast cancer from Curtis et al. (2012). One identified subcohort suggests a link between LXR over-expression and BRCA2 and MSH6 methylation levels for patients in that subcohort.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-worah26a, title = { A Goemans-Williamson type algorithm for identifying subcohorts in clinical trials }, author = {Worah, Pratik}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {3394--3402}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/worah26a/worah26a.pdf}, url = {https://proceedings.mlr.press/v300/worah26a.html}, abstract = { We design an efficient algorithm that outputs tests for identifying predominantly homogeneous subcohorts of patients from large in-homogeneous datasets. Our theoretical contribution is a rounding technique, similar to that of Goemans and Wiliamson (1995), which approximates the optimal solution within a factor of $0.82$. As an application, we use our algorithm to trade-off sensitivity for specificity to systematically identify clinically interesting homogeneous subcohorts of patients in the RNA microarray data set for breast cancer from Curtis et al. (2012). One identified subcohort suggests a link between LXR over-expression and BRCA2 and MSH6 methylation levels for patients in that subcohort. } }
Endnote
%0 Conference Paper %T A Goemans-Williamson type algorithm for identifying subcohorts in clinical trials %A Pratik Worah %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-worah26a %I PMLR %P 3394--3402 %U https://proceedings.mlr.press/v300/worah26a.html %V 300 %X We design an efficient algorithm that outputs tests for identifying predominantly homogeneous subcohorts of patients from large in-homogeneous datasets. Our theoretical contribution is a rounding technique, similar to that of Goemans and Wiliamson (1995), which approximates the optimal solution within a factor of $0.82$. As an application, we use our algorithm to trade-off sensitivity for specificity to systematically identify clinically interesting homogeneous subcohorts of patients in the RNA microarray data set for breast cancer from Curtis et al. (2012). One identified subcohort suggests a link between LXR over-expression and BRCA2 and MSH6 methylation levels for patients in that subcohort.
APA
Worah, P.. (2026). A Goemans-Williamson type algorithm for identifying subcohorts in clinical trials . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:3394-3402 Available from https://proceedings.mlr.press/v300/worah26a.html.

Related Material