Selected work

Data science projects

Two sides of my work: data science projects with their results, and apps I have built.

Seven projects from my MSc in Data Science, each with the question, the method, the results and what I would do next.

Machine learning

Supervised, unsupervised and NLP models on real datasets
0%25%50%75%100%Logistic regressionRandom forest
Machine learning
71%Buyer recall after tuning (from 35%)

Online Shopper Purchase-Intent Prediction

Predicts which e-commerce sessions end in a purchase. Tuning for the minority class lifted logistic-regression buyer recall from 35% to 71% on 12,330 sessions.

Pythonscikit-learnClassificationImbalanced dataGridSearchCV
Word cloud of positive review vocabulary
NLP
90.1%Accuracy on held-out reviews

IMDB Review Sentiment Classifier

Classifies 50,000 movie reviews as positive or negative with TF-IDF features and linear models, reaching 90.1% accuracy and 0.966 ROC-AUC on 10,000 held-out reviews.

PythonNLPTF-IDFscikit-learnLogistic regression
K-Means clusters in PCA space
Machine learning
2,111Records

Lifestyle and Obesity Segmentation

Unsupervised segmentation of 2,111 people by eating habits, activity and physical profile. K-Means outperformed hierarchical clustering on all three validity metrics.

PythonUnsupervised learningK-MeansHierarchical clusteringPCA

Big data

Spark and Databricks on large datasets
00.0750.150.2250.3λ 0.01 · 10 itλ 0.01 · 20 itλ 0.1 · 10 itλ 0.1 · 20 it
Big data
200KInteractions modelled

Steam Game Recommender

A collaborative-filtering recommender built with Spark MLlib on 200,000 Steam purchase and playtime records, tuned across eight ALS configurations and tracked with MLflow.

PySparkSpark MLlibALSDatabricksMLflow
020040060080019891994199920042009201420192024
Big data
521KTrials analysed

Clinical Trials Analytics at Scale

Spark SQL analysis of more than 520,000 registered clinical trials on Databricks: study-type mix, most-researched conditions, average trial duration and the rise of diabetes research.

Apache SparkSpark SQLPySparkDatabricksData quality

SQL

Database design and T-SQL in SQL Server
PassengersPK PassengerIDUQ PNR · EmailName · DoB · MealFlightsPK FlightIDUQ FlightNumberOrigin · DestinationReservationsPK ReservationIDFK PNR → PassengersFK FlightID → FlightsEmployeesPK EmployeeIDUQ Username · EmailRole · PasswordHashTicketsPK TicketIDFK ReservationIDFK EmployeeIDBaggagePK BaggageIDFK TicketIDWeight · Status · Fee
Databases
6Tables

Airport Ticketing System in SQL Server

A relational back end for airline ticketing: six tables, five stored procedures, two functions, two views, a trigger and role-based security.

T-SQLSQL ServerDatabase designStored proceduresTriggers
CustomersKEY customer_idname · emailcountryOrdersKEY order_idFK customer_idOrder_itemsFK order_idFK product_idquantity · line totalsProductsKEY product_idproduct_namecategoryPaymentsFK order_idpayment_id · methodamount_paid
Databases
5Tables

Online Shopping Analytics in SQL Server

Foreign keys over five imported CSV tables, then ten T-SQL answers to business questions: subqueries, joins, aggregation, TOP-N ranking, VAT arithmetic and a discount procedure.

T-SQLSQL ServerSubqueriesJoinsAggregation

My Fildata dissertation and internship work is covered by confidentiality agreements, so it is described on the Experience page rather than shown here.