Cure-SFT: Diagnostic-Guided Data Curation for Instruction Tuning

Yuankang Fu, Xinrong Gong, Chen Gong, Tong Zhang, Kaixiang Yang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:31833-31853, 2026.

Abstract

Instruction data curation is central to improving the instruction-following ability of large language models. However, existing approaches often struggle to simultaneously maintain data quality, diversity, and distributional consistency, largely because they do not explicitly distinguish semantic redundancy from quality defects and rely on coarse-grained modeling of instruction data quality. To address this issue, we propose Cure-SFT, a coarse-to-fine, diagnostic-guided method for instruction data curation that explicitly disentangles semantic redundancy from quality defects. Specifically, Cure-SFT removes redundant samples via stratified semantic-geometric sampling, applies teacher models for diagnostic triage, and performs targeted defect remediation on fixable samples. Our experiments show that Cure-SFT can surpass full-data instruction tuning using only 10% of the data budget. Moreover, Cure-SFT consistently outperforms strong selection-based and rewriting-based baselines across data budgets, supporting the effectiveness of diagnostic-guided data curation.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-fu26f, title = {Cure-{SFT}: Diagnostic-Guided Data Curation for Instruction Tuning}, author = {Fu, Yuankang and Gong, Xinrong and Gong, Chen and Zhang, Tong and Yang, Kaixiang}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {31833--31853}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/fu26f/fu26f.pdf}, url = {https://proceedings.mlr.press/v306/fu26f.html}, abstract = {Instruction data curation is central to improving the instruction-following ability of large language models. However, existing approaches often struggle to simultaneously maintain data quality, diversity, and distributional consistency, largely because they do not explicitly distinguish semantic redundancy from quality defects and rely on coarse-grained modeling of instruction data quality. To address this issue, we propose Cure-SFT, a coarse-to-fine, diagnostic-guided method for instruction data curation that explicitly disentangles semantic redundancy from quality defects. Specifically, Cure-SFT removes redundant samples via stratified semantic-geometric sampling, applies teacher models for diagnostic triage, and performs targeted defect remediation on fixable samples. Our experiments show that Cure-SFT can surpass full-data instruction tuning using only 10% of the data budget. Moreover, Cure-SFT consistently outperforms strong selection-based and rewriting-based baselines across data budgets, supporting the effectiveness of diagnostic-guided data curation.} }
Endnote
%0 Conference Paper %T Cure-SFT: Diagnostic-Guided Data Curation for Instruction Tuning %A Yuankang Fu %A Xinrong Gong %A Chen Gong %A Tong Zhang %A Kaixiang Yang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-fu26f %I PMLR %P 31833--31853 %U https://proceedings.mlr.press/v306/fu26f.html %V 306 %X Instruction data curation is central to improving the instruction-following ability of large language models. However, existing approaches often struggle to simultaneously maintain data quality, diversity, and distributional consistency, largely because they do not explicitly distinguish semantic redundancy from quality defects and rely on coarse-grained modeling of instruction data quality. To address this issue, we propose Cure-SFT, a coarse-to-fine, diagnostic-guided method for instruction data curation that explicitly disentangles semantic redundancy from quality defects. Specifically, Cure-SFT removes redundant samples via stratified semantic-geometric sampling, applies teacher models for diagnostic triage, and performs targeted defect remediation on fixable samples. Our experiments show that Cure-SFT can surpass full-data instruction tuning using only 10% of the data budget. Moreover, Cure-SFT consistently outperforms strong selection-based and rewriting-based baselines across data budgets, supporting the effectiveness of diagnostic-guided data curation.
APA
Fu, Y., Gong, X., Gong, C., Zhang, T. & Yang, K.. (2026). Cure-SFT: Diagnostic-Guided Data Curation for Instruction Tuning. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:31833-31853 Available from https://proceedings.mlr.press/v306/fu26f.html.

Related Material