Conditional Diffusion Models for Imbalanced Tabular Regression

Nathaniel Kang
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:2688-2712, 2026.

Abstract

Imbalanced regression, where certain continuous target ranges are severely underrepresented, poses a fundamental challenge for predictive modeling under uncertainty. Existing oversampling methods rely on local interpolation, which fails to capture complex conditional distributions in rare target regions. We propose TabOversample, a conditional diffusion framework that generates high-fidelity synthetic tabular samples for underrepresented targets via a relevance-weighted denoising objective. We establish four theoretical results: equivalence of relevance weighting to maximum likelihood under a reweighted measure, targeted probability mass amplification in rare regions, quality guarantees for generate-then-filter sampling through order statistics, and a formal connection to distributionally robust optimization over a chi-squared uncertainty set. Across nine benchmark imbalanced regression datasets—four numerical-dominant and five categorical-rich—and against thirteen oversampling and reweighting baselines evaluated over ten seeds, TabOversample attains the best relevance-weighted error (SERA) on eight of nine datasets and the best rare-region RMSE on all nine, substantially improving prediction accuracy in rare target regions while maintaining overall performance.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-kang26a, title = {Conditional Diffusion Models for Imbalanced Tabular Regression}, author = {Kang, Nathaniel}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {2688--2712}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/kang26a/kang26a.pdf}, url = {https://proceedings.mlr.press/v337/kang26a.html}, abstract = {Imbalanced regression, where certain continuous target ranges are severely underrepresented, poses a fundamental challenge for predictive modeling under uncertainty. Existing oversampling methods rely on local interpolation, which fails to capture complex conditional distributions in rare target regions. We propose TabOversample, a conditional diffusion framework that generates high-fidelity synthetic tabular samples for underrepresented targets via a relevance-weighted denoising objective. We establish four theoretical results: equivalence of relevance weighting to maximum likelihood under a reweighted measure, targeted probability mass amplification in rare regions, quality guarantees for generate-then-filter sampling through order statistics, and a formal connection to distributionally robust optimization over a chi-squared uncertainty set. Across nine benchmark imbalanced regression datasets—four numerical-dominant and five categorical-rich—and against thirteen oversampling and reweighting baselines evaluated over ten seeds, TabOversample attains the best relevance-weighted error (SERA) on eight of nine datasets and the best rare-region RMSE on all nine, substantially improving prediction accuracy in rare target regions while maintaining overall performance.} }
Endnote
%0 Conference Paper %T Conditional Diffusion Models for Imbalanced Tabular Regression %A Nathaniel Kang %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-kang26a %I PMLR %P 2688--2712 %U https://proceedings.mlr.press/v337/kang26a.html %V 337 %X Imbalanced regression, where certain continuous target ranges are severely underrepresented, poses a fundamental challenge for predictive modeling under uncertainty. Existing oversampling methods rely on local interpolation, which fails to capture complex conditional distributions in rare target regions. We propose TabOversample, a conditional diffusion framework that generates high-fidelity synthetic tabular samples for underrepresented targets via a relevance-weighted denoising objective. We establish four theoretical results: equivalence of relevance weighting to maximum likelihood under a reweighted measure, targeted probability mass amplification in rare regions, quality guarantees for generate-then-filter sampling through order statistics, and a formal connection to distributionally robust optimization over a chi-squared uncertainty set. Across nine benchmark imbalanced regression datasets—four numerical-dominant and five categorical-rich—and against thirteen oversampling and reweighting baselines evaluated over ten seeds, TabOversample attains the best relevance-weighted error (SERA) on eight of nine datasets and the best rare-region RMSE on all nine, substantially improving prediction accuracy in rare target regions while maintaining overall performance.
APA
Kang, N.. (2026). Conditional Diffusion Models for Imbalanced Tabular Regression. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:2688-2712 Available from https://proceedings.mlr.press/v337/kang26a.html.

Related Material