Prediction of clinical trial enrollment rates

Cameron Bieganek, Constantin Aliferis, Sisi Ma

Research output: Contribution to journalArticlepeer-review


Clinical trials represent a critical milestone of translational and clinical sciences. However, poor recruitment to clinical trials has been a long standing problem affecting institutions all over the world. One way to reduce the cost incurred by insufficient enrollment is to minimize initiating trials that are most likely to fall short of their enrollment goal. Hence, the ability to predict which proposed trials will meet enrollment goals prior to the start of the trial is highly beneficial. In the current study, we leveraged a data set extracted from that consists of 46,724 U.S. based clinical trials from 1990 to 2020. We constructed 4,636 candidate predictors based on data collected by and external sources for enrollment rate prediction using various state-of-the-art machine learning methods. Taking advantage of a nested time series cross-validation design, our models resulted in good predictive performance that is generalizable to future data and stable over time. Moreover, information content analysis revealed the study design related features to be the most informative feature type regarding enrollment. Compared to the performance of models built with all features, the performance of models built with study design related features is only marginally worse (AUC = 0.78 ± 0.03 vs. AUC = 0.76 ± 0.02). The results presented can form the basis for data-driven decision support systems to assess whether proposed clinical trials would likely meet their enrollment goal.

Original languageEnglish (US)
Article numbere0263193
JournalPloS one
Issue number2
StatePublished - Feb 2022

Bibliographical note

Publisher Copyright:
© 2022 Bieganek et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.


  • Algorithms
  • Censuses
  • Clinical Trials, Phase I as Topic
  • Clinical Trials, Phase III as Topic
  • Forecasting
  • Humans
  • Machine Learning
  • Models, Theoretical
  • Natural Language Processing
  • Patient Selection
  • Translational Science, Biomedical

PubMed: MeSH publication types

  • Journal Article
  • Research Support, N.I.H., Extramural


Dive into the research topics of 'Prediction of clinical trial enrollment rates'. Together they form a unique fingerprint.

Cite this