[edit]
When to Trust Simulated Data: Conformal Prediction for Selective Synthetic Data Integration
Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, PMLR 329:867-886, 2026.
Abstract
Integrating simulated data into object detection training is a scalable strategy for alleviating data scarcity, yet simulated samples may exhibit geometric inaccuracies and domain artifacts that can degrade model performance when naively combined with real data. Rather than treating all synthetic samples as equally beneficial, we frame synthetic data inclusion as a reliability-aware selection problem and propose a conformal prediction-based framework to address it. The key insight is that a well-behaved synthetic sample should produce consistent object localizations across detectors trained on real and mixed data; violations of this consistency are captured as a nonconformity score, calibrated on a held-out set drawn from both real and synthetic data to derive a principled acceptance threshold. We evaluate on two benchmark datasets spanning autonomous driving and underwater sonar imaging, demonstrating consistent improvements in precision and recall across multiple model architectures. Explainability analysis further indicates that conformal selection yields more spatially coherent model attention, providing interpretable evidence that filtering acts on genuine domain artifacts rather than spurious correlations.