All projects
Machine learningMSc Data Science · University of Salford

Online Shopper Purchase-Intent Prediction

Pythonscikit-learnClassificationImbalanced dataGridSearchCVRandom forest

Only 15% of online shopping sessions end in a purchase, so a model can reach about 85% accuracy by predicting "no purchase" every time. This project treats it as an imbalance problem: stratified splitting, class weighting, F1-driven hyperparameter search and evaluation on the minority class. The tuned logistic regression doubles the share of real buyers it catches.

71%Buyer recall after tuning (from 35%)
0.65Best buyer-class F1 (tuned forest)
12,330Sessions analysed
15.5%Sessions ending in purchase

The problem

A shop that can spot high-intent visitors can time offers and support where they matter. But with 85 sessions in 100 ending without a sale, plain accuracy rewards a model that never predicts a buyer, so the evaluation has to focus on the buyers themselves.

The data

  • 12,330 sessions described by page-visit counts and durations, bounce and exit rates, page value, special-day proximity, month, visitor type, weekend flag and technical attributes. No missing values.
  • Returning visitors dominate (10,551 sessions, 86%) and May is the busiest month (3,364 sessions). Bounce and exit rates are strongly correlated with each other, as are product-page counts and their durations.

Approach

  1. ExploreDistributions, boxplots by outcome, categorical splits and a correlation heatmap.
  2. PreprocessBooleans to integers, one-hot encoding for month and visitor type, standardised numeric features.
  3. Split with stratificationA 75/25 split that keeps the 15.5% purchase rate in both train and test sets.
  4. BaselinesLogistic regression and a 200-tree random forest, judged by confusion matrix and per-class precision, recall and F1.
  5. Tune for the minority classGridSearchCV with 3-fold cross-validation and F1 scoring: 40 logistic-regression candidates (penalty, C, solver, class weight) and 216 forest candidates (trees, depth, split and leaf size, class weight).
  6. CompareBaseline and tuned models side by side on the same untouched test set.

Results

Test-set results (3,083 sessions; precision, recall and F1 refer to the buyer class)
ModelAccuracyPrecisionRecallF1
Logistic regression (baseline)87.8%71.7%35.0%0.47
Logistic regression (tuned)87.4%57.6%71.1%0.64
Random forest (baseline)89.9%74.3%53.2%0.62
Random forest (tuned)88.7%62.5%66.9%0.65

Buyer recall: share of real purchases the model catches

0%25%50%75%100%Logistic regressionRandom forest
BaselineTuned
View as table
BaselineTuned
Logistic regression35%71.1%
Random forest53.2%66.9%

Buyer-class F1 score

0%25%50%75%100%Logistic regressionRandom forest
BaselineTuned
View as table
BaselineTuned
Logistic regression47%63.6%
Random forest62%64.6%

Confusion matrices on the test set

Logistic regression, baseline
Logistic regression, tuned
Random forest, baseline
Random forest, tuned

Colour shows the share of each true class. The bottom-right cell is the buyers the model caught.

Correlation heatmap of numeric features
Correlation heatmap of the numeric session features.

Key findings

  • Accuracy is the wrong yardstick here: the untuned logistic regression scores 87.8% yet finds only 35% of buyers (167 of 477).
  • Class weighting plus F1-driven tuning moved the operating point. The tuned logistic regression catches 339 of 477 buyers (71%), at the price of more false alarms (precision 58%).
  • The tuned random forest gives the best balance (F1 0.65, recall 67%, precision 63%). A strongly regularised linear model (L1 penalty, C = 0.01) comes close at 0.64 and is far easier to explain.
  • The right choice depends on the cost of missing a buyer versus the cost of a wasted intervention, and the two tuned models give the business a clear precision–recall trade-off to choose from.

Next steps

  • Add precision–recall AUC and probability calibration to pick a decision threshold by business cost.
  • Try gradient boosting (XGBoost, LightGBM) and cost-sensitive or SMOTE-based approaches.
  • Wrap preprocessing and the model in one scikit-learn Pipeline so scaling is fitted only on training folds.