Online Shopper Purchase-Intent Prediction
Pythonscikit-learnClassificationImbalanced dataGridSearchCVRandom forest
Only 15% of online shopping sessions end in a purchase, so a model can reach about 85% accuracy by predicting "no purchase" every time. This project treats it as an imbalance problem: stratified splitting, class weighting, F1-driven hyperparameter search and evaluation on the minority class. The tuned logistic regression doubles the share of real buyers it catches.
71%Buyer recall after tuning (from 35%)
0.65Best buyer-class F1 (tuned forest)
12,330Sessions analysed
15.5%Sessions ending in purchase
The problem
A shop that can spot high-intent visitors can time offers and support where they matter. But with 85 sessions in 100 ending without a sale, plain accuracy rewards a model that never predicts a buyer, so the evaluation has to focus on the buyers themselves.
The data
- 12,330 sessions described by page-visit counts and durations, bounce and exit rates, page value, special-day proximity, month, visitor type, weekend flag and technical attributes. No missing values.
- Returning visitors dominate (10,551 sessions, 86%) and May is the busiest month (3,364 sessions). Bounce and exit rates are strongly correlated with each other, as are product-page counts and their durations.
Approach
- ExploreDistributions, boxplots by outcome, categorical splits and a correlation heatmap.
- PreprocessBooleans to integers, one-hot encoding for month and visitor type, standardised numeric features.
- Split with stratificationA 75/25 split that keeps the 15.5% purchase rate in both train and test sets.
- BaselinesLogistic regression and a 200-tree random forest, judged by confusion matrix and per-class precision, recall and F1.
- Tune for the minority classGridSearchCV with 3-fold cross-validation and F1 scoring: 40 logistic-regression candidates (penalty, C, solver, class weight) and 216 forest candidates (trees, depth, split and leaf size, class weight).
- CompareBaseline and tuned models side by side on the same untouched test set.
Results
| Model | Accuracy | Precision | Recall | F1 |
|---|---|---|---|---|
| Logistic regression (baseline) | 87.8% | 71.7% | 35.0% | 0.47 |
| Logistic regression (tuned) | 87.4% | 57.6% | 71.1% | 0.64 |
| Random forest (baseline) | 89.9% | 74.3% | 53.2% | 0.62 |
| Random forest (tuned) | 88.7% | 62.5% | 66.9% | 0.65 |
Buyer recall: share of real purchases the model catches
BaselineTuned
View as table
| Baseline | Tuned | |
|---|---|---|
| Logistic regression | 35% | 71.1% |
| Random forest | 53.2% | 66.9% |
Buyer-class F1 score
BaselineTuned
View as table
| Baseline | Tuned | |
|---|---|---|
| Logistic regression | 47% | 63.6% |
| Random forest | 62% | 64.6% |
Confusion matrices on the test set
predicted no purchasepredicted purchase
actual no purchase
2,54097.5% of row
662.5% of row
actual purchase31065.0% of row
16735.0% of row
predicted no purchasepredicted purchase
actual no purchase
2,35690.4% of row
2509.6% of row
actual purchase13828.9% of row
33971.1% of row
predicted no purchasepredicted purchase
actual no purchase
2,51896.6% of row
883.4% of row
actual purchase22346.8% of row
25453.2% of row
predicted no purchasepredicted purchase
actual no purchase
2,41592.7% of row
1917.3% of row
actual purchase15833.1% of row
31966.9% of row
Colour shows the share of each true class. The bottom-right cell is the buyers the model caught.
Key findings
- Accuracy is the wrong yardstick here: the untuned logistic regression scores 87.8% yet finds only 35% of buyers (167 of 477).
- Class weighting plus F1-driven tuning moved the operating point. The tuned logistic regression catches 339 of 477 buyers (71%), at the price of more false alarms (precision 58%).
- The tuned random forest gives the best balance (F1 0.65, recall 67%, precision 63%). A strongly regularised linear model (L1 penalty, C = 0.01) comes close at 0.64 and is far easier to explain.
- The right choice depends on the cost of missing a buyer versus the cost of a wasted intervention, and the two tuned models give the business a clear precision–recall trade-off to choose from.
Next steps
- Add precision–recall AUC and probability calibration to pick a decision threshold by business cost.
- Try gradient boosting (XGBoost, LightGBM) and cost-sensitive or SMOTE-based approaches.
- Wrap preprocessing and the model in one scikit-learn Pipeline so scaling is fitted only on training folds.