Simultaneous truth and performance level estimation (STAPLE): an algorithm for the validation of image segmentation.

There is no author summary for this article yet. Authors can add summaries to their articles on ScienceOpen to make them more accessible to a non-specialist audience.

Abstract

Characterizing the performance of image segmentation approaches has been a persistent challenge. Performance analysis is important since segmentation algorithms often have limited accuracy and precision. Interactive drawing of the desired segmentation by human raters has often been the only acceptable approach, and yet suffers from intra-rater and inter-rater variability. Automated algorithms have been sought in order to remove the variability introduced by raters, but such algorithms must be assessed to ensure they are suitable for the task. The performance of raters (human or algorithmic) generating segmentations of medical images has been difficult to quantify because of the difficulty of obtaining or estimating a known true segmentation for clinical data. Although physical and digital phantoms can be constructed for which ground truth is known or readily estimated, such phantoms do not fully reflect clinical images due to the difficulty of constructing phantoms which reproduce the full range of imaging characteristics and normal and pathological anatomical variability observed in clinical data. Comparison to a collection of segmentations by raters is an attractive alternative since it can be carried out directly on the relevant clinical imaging data. However, the most appropriate measure or set of measures with which to compare such segmentations has not been clarified and several measures are used in practice. We present here an expectation-maximization algorithm for simultaneous truth and performance level estimation (STAPLE). The algorithm considers a collection of segmentations and computes a probabilistic estimate of the true segmentation and a measure of the performance level represented by each segmentation. The source of each segmentation in the collection may be an appropriately trained human rater or raters, or may be an automated segmentation algorithm. The probabilistic estimate of the true segmentation is formed by estimating an optimal combination of the segmentations, weighting each segmentation depending upon the estimated performance level, and incorporating a prior model for the spatial distribution of structures being segmented as well as spatial homogeneity constraints. STAPLE is straightforward to apply to clinical imaging data, it readily enables assessment of the performance of an automated image segmentation algorithm, and enables direct comparison of human rater and algorithm performance.

Related collections

Author and article information

Journal

Journal ID (iso-abbrev): IEEE Trans Med Imaging

Title: IEEE transactions on medical imaging

Publisher: Institute of Electrical and Electronics Engineers (IEEE)

ISSN: 0278-0062

ISSN (Print): 0278-0062

Publication date (Electronic): Jul 2004

Volume: 23

Issue: 7

Affiliations

[1 ] Harvard Medical School and the Department of Radiology of Brigham and Women's Hospital, 75 Francis St, Boston, MA 02115, USA. warfield@bwh.harvard.edu

Article

Mid ID: NIHMS2330

DOI: 10.1109/TMI.2004.828354

PMC ID: 1283110

PubMed ID: 15250643

SO-VID: b5ab5c7f-8189-4982-acd7-c1f94130b39d

History

Data availability:

Comments

Comment on this article

scite_

Cited by 378

See all cited by

- Version 1
- Version 1