[edit]
Conditional Diffusion Models for Imbalanced Tabular Regression
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:2688-2712, 2026.
Abstract
Imbalanced regression, where certain continuous target ranges are severely underrepresented, poses a fundamental challenge for predictive modeling under uncertainty. Existing oversampling methods rely on local interpolation, which fails to capture complex conditional distributions in rare target regions. We propose TabOversample, a conditional diffusion framework that generates high-fidelity synthetic tabular samples for underrepresented targets via a relevance-weighted denoising objective. We establish four theoretical results: equivalence of relevance weighting to maximum likelihood under a reweighted measure, targeted probability mass amplification in rare regions, quality guarantees for generate-then-filter sampling through order statistics, and a formal connection to distributionally robust optimization over a chi-squared uncertainty set. Across nine benchmark imbalanced regression datasets—four numerical-dominant and five categorical-rich—and against thirteen oversampling and reweighting baselines evaluated over ten seeds, TabOversample attains the best relevance-weighted error (SERA) on eight of nine datasets and the best rare-region RMSE on all nine, substantially improving prediction accuracy in rare target regions while maintaining overall performance.