When to Trust Simulated Data: Conformal Prediction for Selective Synthetic Data Integration

Sakir Furkan Yondem, Seyda Yoncaci, Gizem Karagoz, Benedikt Schlereth-Groh, Engin Avci, Ramin Tavakoli Kolagari, Fatima Rabia Yapicioglu
Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, PMLR 329:867-886, 2026.

Abstract

Integrating simulated data into object detection training is a scalable strategy for alleviating data scarcity, yet simulated samples may exhibit geometric inaccuracies and domain artifacts that can degrade model performance when naively combined with real data. Rather than treating all synthetic samples as equally beneficial, we frame synthetic data inclusion as a reliability-aware selection problem and propose a conformal prediction-based framework to address it. The key insight is that a well-behaved synthetic sample should produce consistent object localizations across detectors trained on real and mixed data; violations of this consistency are captured as a nonconformity score, calibrated on a held-out set drawn from both real and synthetic data to derive a principled acceptance threshold. We evaluate on two benchmark datasets spanning autonomous driving and underwater sonar imaging, demonstrating consistent improvements in precision and recall across multiple model architectures. Explainability analysis further indicates that conformal selection yields more spatially coherent model attention, providing interpretable evidence that filtering acts on genuine domain artifacts rather than spurious correlations.

Cite this Paper


BibTeX
@InProceedings{pmlr-v329-yondem26a, title = {When to Trust Simulated Data: Conformal Prediction for Selective Synthetic Data Integration}, author = {Yondem, Sakir Furkan and Yoncaci, Seyda and Karagoz, Gizem and Schlereth-Groh, Benedikt and Avci, Engin and Kolagari, Ramin Tavakoli and Yapicioglu, Fatima Rabia}, booktitle = {Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications}, pages = {867--886}, year = {2026}, editor = {Ahlberg, Ernst and Johansson, Ulf and Boström, Henrik and Carlevaro, Alberto and Hallberg Szabadváry, Johan and Carlsson, Lars}, volume = {329}, series = {Proceedings of Machine Learning Research}, month = {02--04 Sep}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v329/main/assets/yondem26a/yondem26a.pdf}, url = {https://proceedings.mlr.press/v329/yondem26a.html}, abstract = {Integrating simulated data into object detection training is a scalable strategy for alleviating data scarcity, yet simulated samples may exhibit geometric inaccuracies and domain artifacts that can degrade model performance when naively combined with real data. Rather than treating all synthetic samples as equally beneficial, we frame synthetic data inclusion as a reliability-aware selection problem and propose a conformal prediction-based framework to address it. The key insight is that a well-behaved synthetic sample should produce consistent object localizations across detectors trained on real and mixed data; violations of this consistency are captured as a nonconformity score, calibrated on a held-out set drawn from both real and synthetic data to derive a principled acceptance threshold. We evaluate on two benchmark datasets spanning autonomous driving and underwater sonar imaging, demonstrating consistent improvements in precision and recall across multiple model architectures. Explainability analysis further indicates that conformal selection yields more spatially coherent model attention, providing interpretable evidence that filtering acts on genuine domain artifacts rather than spurious correlations.} }
Endnote
%0 Conference Paper %T When to Trust Simulated Data: Conformal Prediction for Selective Synthetic Data Integration %A Sakir Furkan Yondem %A Seyda Yoncaci %A Gizem Karagoz %A Benedikt Schlereth-Groh %A Engin Avci %A Ramin Tavakoli Kolagari %A Fatima Rabia Yapicioglu %B Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications %C Proceedings of Machine Learning Research %D 2026 %E Ernst Ahlberg %E Ulf Johansson %E Henrik Boström %E Alberto Carlevaro %E Johan Hallberg Szabadváry %E Lars Carlsson %F pmlr-v329-yondem26a %I PMLR %P 867--886 %U https://proceedings.mlr.press/v329/yondem26a.html %V 329 %X Integrating simulated data into object detection training is a scalable strategy for alleviating data scarcity, yet simulated samples may exhibit geometric inaccuracies and domain artifacts that can degrade model performance when naively combined with real data. Rather than treating all synthetic samples as equally beneficial, we frame synthetic data inclusion as a reliability-aware selection problem and propose a conformal prediction-based framework to address it. The key insight is that a well-behaved synthetic sample should produce consistent object localizations across detectors trained on real and mixed data; violations of this consistency are captured as a nonconformity score, calibrated on a held-out set drawn from both real and synthetic data to derive a principled acceptance threshold. We evaluate on two benchmark datasets spanning autonomous driving and underwater sonar imaging, demonstrating consistent improvements in precision and recall across multiple model architectures. Explainability analysis further indicates that conformal selection yields more spatially coherent model attention, providing interpretable evidence that filtering acts on genuine domain artifacts rather than spurious correlations.
APA
Yondem, S.F., Yoncaci, S., Karagoz, G., Schlereth-Groh, B., Avci, E., Kolagari, R.T. & Yapicioglu, F.R.. (2026). When to Trust Simulated Data: Conformal Prediction for Selective Synthetic Data Integration. Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, in Proceedings of Machine Learning Research 329:867-886 Available from https://proceedings.mlr.press/v329/yondem26a.html.

Related Material