New Feature Splitting Criteria for Co-training Using Genetic Algorithm Optimization

Often in real world applications only a small number of labeled data is available while unlabeled data is abundant. Therefore, it is important to make use of unlabeled data. Co-training is a popular semi-supervised learning technique that uses a small set of labeled data and enough unlabeled data to...

Full description

Saved in:

Bibliographic Details
Published in	Multiple Classifier Systems pp. 22 - 32
Main Authors	Salaheldin, Ahmed, El Gayar, Neamat
Format	Book Chapter
Language	English
Published	Berlin, Heidelberg Springer Berlin Heidelberg
Series	Lecture Notes in Computer Science
Subjects	Genetic Algorithm Individual View Label Data Support Vector Machine Unlabeled Data
Online Access	Get full text

Cover

Loading…

More Information
Summary:	Often in real world applications only a small number of labeled data is available while unlabeled data is abundant. Therefore, it is important to make use of unlabeled data. Co-training is a popular semi-supervised learning technique that uses a small set of labeled data and enough unlabeled data to create more accurate classification models. A key feature for successful co-training is to split the features among more than one view. In this paper we propose new splitting criteria based on the confidence of the views, the diversity of the views, and compare them to random and natural splits. We also examine a previously proposed artificial split that maximizes the independence between the views, and propose a mixed criterion for splitting features based on both the confidence and the independence of the views. Genetic algorithms are used to choose the splits which optimize the independence of the views given the class, the confidence of the views in their predictions, and the diversity of the views. We demonstrate that our proposed splitting criteria improve the performance of co-training.
ISBN:	9783642121265 3642121268
ISSN:	0302-9743 1611-3349
DOI:	10.1007/978-3-642-12127-2_3