Steam Game Recommender
PySparkSpark MLlibALSDatabricksMLflowCollaborative filtering
A scalable recommender that learns what 12,393 Steam players are likely to enjoy from implicit signals alone, purchases and hours played, and returns a ranked top-10 list for every user and the ten most likely players for every game. Built end to end in PySpark on Databricks, with every experiment logged in MLflow.
200KInteractions modelled
12,393Users
5,155Games
0.209Best held-out RMSE
The problem
A player faces thousands of titles, and Steam gives no star ratings, only whether a user bought a game and how long they played it. A recommender has to infer taste from behaviour, and the signal is noisy and heavily skewed: a handful of titles absorb most of the playtime.
The data
- 129,511 purchase events and 70,489 play events across 12,393 users and 5,155 games, with no missing values.
- Playtime is strongly right-skewed (mean 48.9 hours, standard deviation 229, maximum 999). Dota 2 alone logs 981,685 hours, about three times the runner-up.
Approach
- Explore at scaleSpark DataFrame aggregations for the most-played, most-purchased and most-popular games, and the heaviest users.
- Engineer implicit ratingsPurchase = 1.0; play = log(1 + hours) to tame the long tail; both summed per user–game pair, then min–max scaled to [0, 1].
- Encode identifiersStringIndexer converts user and game identifiers to the integer IDs that ALS requires.
- Train a baselineALS with implicit preferences on an 80/20 split, with cold-start rows dropped from evaluation.
- Tune systematicallyGrid over rank (10, 20), regularisation (0.01, 0.1) and iterations (10, 20): eight models compared by held-out RMSE.
- Track and serveParameters, RMSE and the model logged to MLflow (experiment Steam_Optimized_Recommender); top-10 games per user and top-10 users per game mapped back to readable game names.
Results
Held-out RMSE by ALS configuration (lower is better)
Rank 10Rank 20
View as table
| Rank 10 | Rank 20 | |
|---|---|---|
| λ 0.01 · 10 it | 0.238 | 0.257 |
| λ 0.01 · 20 it | 0.242 | 0.264 |
| λ 0.1 · 10 it | 0.209 | 0.217 |
| λ 0.1 · 20 it | 0.209 | 0.217 |
Most-played games (total hours played)
View as table
| Dota 2 | 981,685 |
| Counter-Strike: GO | 322,772 |
| Team Fortress 2 | 173,673 |
| Counter-Strike | 134,261 |
| Civilization V | 99,821 |
| CS: Source | 96,076 |
| Skyrim | 70,889 |
| Garry's Mod | 49,725 |
| CoD: MW2 Multiplayer | 42,010 |
| Left 4 Dead 2 | 33,597 |
Key findings
- Regularisation was the deciding hyperparameter: a value of 0.1 held RMSE at 0.209 to 0.217, while 0.01 pushed it to 0.238 to 0.264, a clear sign of over-fitting on sparse interactions.
- A smaller latent space generalised better: rank 10 beat rank 20 at the same regularisation (0.2094 against 0.2171).
- The search confirmed that Spark’s default settings (rank 10, regularisation 0.1) were already effectively optimal, with the winning configuration improving RMSE by less than 0.00001, and showed exactly which settings to avoid.
- Every user receives ten ranked game suggestions with predicted preference scores, and every game receives ten candidate players.
Next steps
- Add ranking metrics (precision@k, MAP, NDCG), which suit top-N recommendation better than RMSE on implicit scores.
- Benchmark against a popularity baseline and item-based neighbours.
- Bring in side information such as genre and price to ease the cold-start problem.