Persistent Laplacian-enhanced algorithm for scarcely labeled data classification

The success of many machine learning (ML) methods depends crucially on having large amounts of labeled data. However, obtaining enough labeled data can be expensive, time-consuming, and subject to ethical constraints for many applications. One approach that has shown tremendous value in addressing t...

Full description

Saved in:

Bibliographic Details
Published in	Machine learning Vol. 113; no. 10; pp. 7267 - 7292
Main Authors	Bhusal, Gokul, Merkurjev, Ekaterina, Wei, Guo-Wei
Format	Journal Article
Language	English
Published	New York Springer US 01.10.2024 Springer Nature B.V
Subjects	Algorithms Artificial Intelligence Classification Computer Science Control Cost analysis Graph theory Machine Learning Mechatronics Natural language processing Natural Language Processing (NLP) Robotics Semi-supervised learning Simulation and Modeling Speech recognition Topology Scarcely labeled data Topology-based framework Persistent Laplacian Graph MBO technique
Online Access	Get full text

Cover

Loading…

More Information
Summary:	The success of many machine learning (ML) methods depends crucially on having large amounts of labeled data. However, obtaining enough labeled data can be expensive, time-consuming, and subject to ethical constraints for many applications. One approach that has shown tremendous value in addressing this challenge is semi-supervised learning (SSL); this technique utilizes both labeled and unlabeled data during training, often with much less labeled data than unlabeled data, which is often relatively easy and inexpensive to obtain. In fact, SSL methods are particularly useful in applications where the cost of labeling data is especially expensive, such as medical analysis, natural language processing, or speech recognition. A subset of SSL methods that have achieved great success in various domains involves algorithms that integrate graph-based techniques. These procedures are popular due to the vast amount of information provided by the graphical framework. In this work, we propose an algebraic topology-based semi-supervised method called persistent Laplacian-enhanced graph MBO by integrating persistent spectral graph theory with the classical Merriman–Bence–Osher (MBO) scheme. Specifically, we use a filtration procedure to generate a sequence of chain complexes and associated families of simplicial complexes, from which we construct a family of persistent Laplacians. Overall, it is a very efficient procedure that requires much less labeled data to perform well compared to many ML techniques, and it can be adapted for both small and large datasets. We evaluate the performance of our method on classification, and the results indicate that the technique outperforms other existing semi-supervised algorithms.
ISSN:	0885-6125 1573-0565
DOI:	10.1007/s10994-024-06616-w