[edit]
$f_BGD$: Learning Embeddings From Positive Unlabeled Data with BGD
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:197-206, 2018.
Abstract
Learning sparse features from only positive and unlabeled (PU) data is a fundamental task for problems of several domains, such as natural lan- guage processing (NLP), computer vision (CV), information retrieval (IR). Considering the nu- merous amount of unlabeled data, most prevalent methods rely on negative sampling (NS) to in- crease computational efficiency. However, sam- pling a fraction of unlabeled data as negative for training may ignore other important examples, and thus lead to non-optimal prediction perfor- mance. To address this, we present a fast and generic batch gradient descent optimizer (fBGD) to learn from all training examples without sam- pling. By leveraging sparsity in PU data, we ac- celerate fBGD by several magnitudes, making its time complexity the same level as the NS- based stochastic gradient descent method. Mean- while, we observe that the standard batch gradi- ent method suffers from gradient instability is- sues due to the sparsity property. Driven by a theoretical analysis for this potential cause, an in- tuitive solution arises naturally. To verify its effi- cacy, we perform experiments on multiple tasks with PU data across domains, and show that fBGD consistently outperforms NS-based mod- els on all tasks with comparable efficiency.