Characteristics of publicly available skin cancer image datasets: a systematic review

Publicly available skin image datasets are increasingly used to develop machine learning algorithms for skin cancer diagnosis. However, the total number of datasets and their respective content is currently unclear. This systematic review aimed to identify and evaluate all publicly available skin im...

Full description

Saved in:
Bibliographic Details
Published inThe Lancet. Digital health Vol. 4; no. 1; pp. e64 - e74
Main Authors Wen, David, Khan, Saad M, Ji Xu, Antonio, Ibrahim, Hussein, Smith, Luke, Caballero, Jose, Zepeda, Luis, de Blas Perez, Carlos, Denniston, Alastair K, Liu, Xiaoxuan, Matin, Rubeta N
Format Journal Article
LanguageEnglish
Published England Elsevier Ltd 01.01.2022
Subjects
Online AccessGet full text

Cover

Loading…
More Information
Summary:Publicly available skin image datasets are increasingly used to develop machine learning algorithms for skin cancer diagnosis. However, the total number of datasets and their respective content is currently unclear. This systematic review aimed to identify and evaluate all publicly available skin image datasets used for skin cancer diagnosis by exploring their characteristics, data access requirements, and associated image metadata. A combined MEDLINE, Google, and Google Dataset search identified 21 open access datasets containing 106 950 skin lesion images, 17 open access atlases, eight regulated access datasets, and three regulated access atlases. Images and accompanying data from open access datasets were evaluated by two independent reviewers. Among the 14 datasets that reported country of origin, most (11 [79%]) originated from Europe, North America, and Oceania exclusively. Most datasets (19 [91%]) contained dermoscopic images or macroscopic photographs only. Clinical information was available regarding age for 81 662 images (76·4%), sex for 82 848 (77·5%), and body site for 79 561 (74·4%). Subject ethnicity data were available for 1415 images (1·3%), and Fitzpatrick skin type data for 2236 (2·1%). There was limited and variable reporting of characteristics and metadata among datasets, with substantial under-representation of darker skin types. This is the first systematic review to characterise publicly available skin image datasets, highlighting limited applicability to real-life clinical settings and restricted population representation, precluding generalisability. Quality standards for characteristics and metadata reporting for skin image datasets are needed.
Bibliography:ObjectType-Article-1
SourceType-Scholarly Journals-1
ObjectType-Feature-2
content type line 23
ObjectType-Undefined-3
ISSN:2589-7500
2589-7500
DOI:10.1016/S2589-7500(21)00252-1