Investigating the potential of deep learning for patient-specific quality assurance of salivary gland contours using EORTC-1219-DAHANCA-29 clinical trial data

Authors

  • Hanne Nijhuis Department of Radiation Oncology, Amsterdam UMC, Vrije Universiteit Amsterdam, Amsterdam, The Netherlands
  • Ward van Rooij Department of Radiation Oncology, Amsterdam UMC, Vrije Universiteit Amsterdam, Amsterdam, The Netherlands
  • Vincent Gregoire Department of Radiation Oncology, Centre Leon Berard, Lyon, France
  • Jens Overgaard Department of Clinical Medicine – Department of Experimental Clinical Oncology, Aarhus University, Aarhus N, Denmark
  • Berend J. Slotman Department of Radiation Oncology, Amsterdam UMC, Vrije Universiteit Amsterdam, Amsterdam, The Netherlands
  • Wilko F. Verbakel Department of Radiation Oncology, Amsterdam UMC, Vrije Universiteit Amsterdam, Amsterdam, The Netherlands
  • Max Dahele Department of Radiation Oncology, Amsterdam UMC, Vrije Universiteit Amsterdam, Amsterdam, The Netherlands

DOI:

https://doi.org/10.1080/0284186X.2020.1863463

Keywords:

Deep learning, Radiotherapy, Clinical trial, Quality assurance, Segmentation, Salivary glands

Abstract

Introduction

Manual quality assurance (QA) of radiotherapy contours for clinical trials is time and labor intensive and subject to inter-observer variability. Therefore, we investigated whether deep-learning (DL) can provide an automated solution to salivary gland contour QA.

Material and methods

DL-models were trained to generate contours for parotid (PG) and submandibular glands (SMG). Sørensen–Dice coefficient (SDC) and Hausdorff distance (HD) were used to assess agreement between DL and clinical contours and thresholds were defined to highlight cases as potentially sub-optimal. 3 types of deliberate errors (expansion, contraction and displacement) were gradually applied to a test set, to confirm that SDC and HD were suitable QA metrics. DL-based QA was performed on 62 patients from the EORTC-1219-DAHANCA-29 trial. All highlighted contours were visually inspected.

Results

Increasing the magnitude of all 3 types of errors resulted in progressively severe deterioration/increase in average SDC/HD. 19/124 clinical PG contours were highlighted as potentially sub-optimal, of which 5 (26%) were actually deemed clinically sub-optimal. 2/19 non-highlighted contours were false negatives (11%). 15/69 clinical SMG contours were highlighted, with 7 (47%) deemed clinically sub-optimal and 2/15 non-highlighted contours were false negatives (13%). For most incorrectly highlighted contours causes for low agreement could be identified.

Conclusion

Automated DL-based contour QA is feasible but some visual inspection remains essential. The substantial number of false positives were caused by sub-optimal performance of the DL-model. Improvements to the model will increase the extent of automation and reliability, facilitating the adoption of DL-based contour QA in clinical trials and routine practice.

Downloads

Download data is not yet available.

Downloads

Published

2021-05-04

How to Cite

Nijhuis, H., van Rooij, W., Gregoire, V., Overgaard, J., Slotman, B. J., Verbakel, W. F., & Dahele, M. (2021). Investigating the potential of deep learning for patient-specific quality assurance of salivary gland contours using EORTC-1219-DAHANCA-29 clinical trial data. Acta Oncologica, 60(5), 575–581. https://doi.org/10.1080/0284186X.2020.1863463