Clinical Trials Analytics at Scale
Apache SparkSpark SQLPySparkDatabricksData qualityHealthcare
Four business questions answered with Spark SQL on a January 2025 extract of more than 520,000 registered clinical trials, framed around long-term planning and market evaluation for the pharmaceutical sector. The work combines typed schema design, defensive parsing of messy fields and clearly commented, reproducible SQL.
521KTrials analysed
36.5 moMean trial length
76.6%Interventional studies
755Diabetes trials completed in 2019 (peak)
The problem
Strategy teams in pharma need to know what is being studied, how often and for how long. The registry extract is large, multi-valued (conditions are stored as delimited lists) and inconsistent (mixed date formats and shifted records), so each question starts with data engineering.
The data
- Loaded with an explicit StructType schema instead of inferred types, then exposed as a temporary SQL view so every query is repeatable.
- A frequency check on Study Type exposed malformed rows (numbers and free text shifted into the column). The analysis whitelists the three valid study types.
- Start and completion dates mix yyyy-MM and MM/dd/yyyy formats. They are parsed with CASE logic and regex guards, and negative durations are excluded.
Approach
- Study typesCount trials per study type, excluding null and malformed values.
- Most common conditionsNormalise the multi-valued Conditions field with regexp_replace, split and LATERAL VIEW explode, lower-case and trim, then rank the top ten.
- Average trial lengthA CTE parses both date formats; MONTHS_BETWEEN gives the duration in months, averaged over trials with valid start and completion dates.
- Diabetes research over timeCompleted studies that mention diabetes, grouped by completion year (1989 to 2025) and charted in Databricks.
Results
Trials by study type
View as table
| Interventional | 399,654 |
| Observational | 120,816 |
| Expanded access | 966 |
Ten most-studied conditions (number of trials)
View as table
| Healthy | 10,786 |
| Obesity | 8,616 |
| Breast cancer | 8,031 |
| Diabetes mellitus | 6,759 |
| Pain | 6,657 |
| Stroke | 5,328 |
| Depression | 4,728 |
| Hypertension | 4,719 |
| Prostate cancer | 4,084 |
| Cancer | 3,866 |
Completed diabetes studies by completion year
View as table
| Completed studies | |
|---|---|
| 1989 | 2 |
| 1990 | 1 |
| 1991 | 1 |
| 1992 | 3 |
| 1993 | 3 |
| 1994 | 2 |
| 1995 | 2 |
| 1996 | 2 |
| 1997 | 3 |
| 1998 | 10 |
| 1999 | 11 |
| 2000 | 20 |
| 2001 | 26 |
| 2002 | 51 |
| 2003 | 76 |
| 2004 | 124 |
| 2005 | 190 |
| 2006 | 244 |
| 2007 | 316 |
| 2008 | 415 |
| 2009 | 459 |
| 2010 | 531 |
| 2011 | 514 |
| 2012 | 589 |
| 2013 | 583 |
| 2014 | 595 |
| 2015 | 665 |
| 2016 | 633 |
| 2017 | 716 |
| 2018 | 685 |
| 2019 | 755 |
| 2020 | 559 |
| 2021 | 577 |
| 2022 | 654 |
| 2023 | 616 |
| 2024 | 427 |
Key findings
- Interventional studies make up about three quarters of the registry (399,654 of 521,436), observational studies a further 23%, and expanded-access programmes under 0.2%.
- Healthy volunteers are the most common category (10,786 trials), followed by obesity (8,616), breast cancer (8,031), diabetes mellitus (6,759) and pain (6,657).
- Cleaning changed the answer. A naive split on the pipe character leaves diabetes mellitus (6,759 trials once normalised, fourth overall) out of the top ten altogether, and case and delimiter normalisation lifts obesity from 6,949 to 8,616.
- The average trial runs 36.5 months, just over three years from start to completion.
- Completed diabetes studies grew from 51 in 2002 to a peak of 755 in 2019, then settled between roughly 430 and 660 a year. The most recent years may be understated, and 2025 is left out of the chart, because the extract dates from January 2025.
Next steps
- Replace the delimiter heuristic with a controlled vocabulary such as MeSH so compound condition names stay intact.
- Cache the cleaned DataFrame and partition it by year to speed up repeated queries.
- Extend the analysis to sponsors and funder types.