A benchmark for neural network robustness in skin cancer classification

Roman C Maron; Justin G Schlager; Sarah Haggenmüller; Christof von Kalle; Jochen S Utikal; Friedegund Meier; Frank F Gellrich; Sarah Hobelsberger; Axel Hauschild; Lars French; Lucie Heinzerling; Max Schlaak; Kamran Ghoreschi; Franz J Hilke; Gabriela Poch; Markus V Heppt; Carola Berking; Sebastian Haferkamp; Wiebke Sondermann; Dirk Schadendorf; Bastian Schilling; Matthias Goebeler; Eva Krieghoff-Henning; Achim Hekler; Stefan Fröhling; Daniel B Lipka; Jakob N Kather; Titus J Brinker

doi:10.1016/j.ejca.2021.06.047

A benchmark for neural network robustness in skin cancer classification

Eur J Cancer. 2021 Sep:155:191-199. doi: 10.1016/j.ejca.2021.06.047. Epub 2021 Aug 11.

Authors

Roman C Maron¹, Justin G Schlager², Sarah Haggenmüller¹, Christof von Kalle³, Jochen S Utikal⁴, Friedegund Meier⁵, Frank F Gellrich⁵, Sarah Hobelsberger⁵, Axel Hauschild⁶, Lars French², Lucie Heinzerling², Max Schlaak⁷, Kamran Ghoreschi⁷, Franz J Hilke⁷, Gabriela Poch⁷, Markus V Heppt⁸, Carola Berking⁸, Sebastian Haferkamp⁹, Wiebke Sondermann¹⁰, Dirk Schadendorf¹⁰, Bastian Schilling¹¹, Matthias Goebeler¹¹, Eva Krieghoff-Henning¹, Achim Hekler¹, Stefan Fröhling¹², Daniel B Lipka¹³, Jakob N Kather¹⁴, Titus J Brinker¹⁵

Affiliations

¹ Digital Biomarkers for Oncology Group, National Center for Tumor Diseases (NCT), German Cancer Research Center (DKFZ), Heidelberg, Germany.
² Department of Dermatology and Allergy, University Hospital, LMU Munich, Munich, Germany.
³ Department of Clinical-Translational Sciences, Charité University Medicine and Berlin Institute of Health (BIH), Berlin, Germany.
⁴ Department of Dermatology, Heidelberg University, Mannheim, Germany; Skin Cancer Unit, German Cancer Research Center (DKFZ), Heidelberg, Germany.
⁵ Skin Cancer Center at the University Cancer Centre and National Center for Tumor Diseases Dresden, Department of Dermatology, University Hospital Carl Gustav Carus, Technische Universität Dresden, Dresden, Germany.
⁶ Department of Dermatology, University Hospital of Kiel, Kiel, Germany.
⁷ Charité - Universitätsmedizin Berlin, Department of Dermatology, Venereology and Allergology, Berlin, Germany.
⁸ Department of Dermatology, University Hospital Erlangen, Erlangen, Germany.
⁹ Department of Dermatology, University Hospital Regensburg, Regensburg, Germany.
¹⁰ Department of Dermatology, University Hospital Essen, Essen, Germany.
¹¹ Department of Dermatology, Venereology and Allergology, University Hospital Würzburg, Würzburg, Germany.
¹² Division of Translational Medical Oncology, German Cancer Research Center (DKFZ) & National Center for Tumor Diseases (NCT), Heidelberg, Germany.
¹³ Section of Translational Cancer Epigenomics, Division of Translational Medical Oncology, German Cancer Research Center (DKFZ) & National Center for Tumor Diseases (NCT), Heidelberg, Germany.
¹⁴ Department of Medicine III, University Hospital RWTH Aachen, Aachen, Germany.
¹⁵ Digital Biomarkers for Oncology Group, National Center for Tumor Diseases (NCT), German Cancer Research Center (DKFZ), Heidelberg, Germany. Electronic address: titus.brinker@dkfz.de.

PMID: 34388516
DOI: 10.1016/j.ejca.2021.06.047

Abstract

Background: One prominent application for deep learning-based classifiers is skin cancer classification on dermoscopic images. However, classifier evaluation is often limited to holdout data which can mask common shortcomings such as susceptibility to confounding factors. To increase clinical applicability, it is necessary to thoroughly evaluate such classifiers on out-of-distribution (OOD) data.

Objective: The objective of the study was to establish a dermoscopic skin cancer benchmark in which classifier robustness to OOD data can be measured.

Methods: Using a proprietary dermoscopic image database and a set of image transformations, we create an OOD robustness benchmark and evaluate the robustness of four different convolutional neural network (CNN) architectures on it.

Results: The benchmark contains three data sets-Skin Archive Munich (SAM), SAM-corrupted (SAM-C) and SAM-perturbed (SAM-P)-and is publicly available for download. To maintain the benchmark's OOD status, ground truth labels are not provided and test results should be sent to us for assessment. The SAM data set contains 319 unmodified and biopsy-verified dermoscopic melanoma (n = 194) and nevus (n = 125) images. SAM-C and SAM-P contain images from SAM which were artificially modified to test a classifier against low-quality inputs and to measure its prediction stability over small image changes, respectively. All four CNNs showed susceptibility to corruptions and perturbations.

Conclusions: This benchmark provides three data sets which allow for OOD testing of binary skin cancer classifiers. Our classifier performance confirms the shortcomings of CNNs and provides a frame of reference. Altogether, this benchmark should facilitate a more thorough evaluation process and thereby enable the development of more robust skin cancer classifiers.

Keywords: Artificial intelligence; Benchmarking; Deep learning; Dermatology; Melanoma; Nevus; Skin neoplasms.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Benchmarking / standards*
Humans
Neural Networks, Computer*
Skin Neoplasms / classification*