Lifestyle and Obesity Segmentation
PythonUnsupervised learningK-MeansHierarchical clusteringPCAscikit-learn
Can lifestyle data alone reveal natural groups related to weight? Using 16 survey features, I built a preprocessing pipeline, chose the number of clusters with the elbow method, compared K-Means with Ward hierarchical clustering and judged both with three internal metrics and a PCA projection. The scores are candid: the segments overlap, which is itself a finding about how continuous lifestyle behaviour is.
2,111Records
31Features after encoding
k = 5Clusters (elbow method)
3 of 3Metrics won by K-Means
The problem
Weight is shaped by many habits at once: diet, water intake, activity, screen time, transport. Rather than predicting a label, this project asks whether people fall into distinct lifestyle profiles, and how clear those profiles really are.
The data
- 2,111 records and 17 columns, 16 of them features (gender, age, height, weight, family history, eating habits, smoking, water intake, activity, screen time, alcohol, transport). No missing values; 24 duplicate rows noted.
- Height and weight are the most correlated numeric pair (r = 0.46). The behavioural features are mostly weakly correlated with each other (|r| of 0.3 or less), so they add complementary information.
Approach
- Audit and exploreDescriptive statistics, histograms, boxplots, categorical counts, a correlation heatmap and a pairplot.
- Keep it unsupervisedThe obesity-level label is removed from the clustering input.
- PreprocessStandardise numeric features and one-hot encode categorical ones, giving 31 dimensions.
- Choose kElbow curve of within-cluster sum of squares for k = 1 to 10, pointing to five clusters.
- Cluster two waysK-Means (k-means++ initialisation) and agglomerative clustering with Ward linkage, guided by a dendrogram on a 400-record sample.
- ValidateSilhouette, Davies–Bouldin and Calinski–Harabasz scores, plus a 2-D PCA projection of each result.
Results
| Method | Silhouette ↑ | Davies–Bouldin ↓ | Calinski–Harabasz ↑ |
|---|---|---|---|
| K-Means | 0.141 | 1.987 | 260.2 |
| Hierarchical (Ward) | 0.120 | 2.066 | 220.4 |
Key findings
- K-Means wins on all three internal metrics (silhouette 0.141 against 0.120, Davies–Bouldin 1.99 against 2.07, Calinski–Harabasz 260 against 220).
- Silhouette scores near 0.14 indicate overlapping clusters. The PCA views show a continuous cloud rather than separate islands, which suggests that lifestyle varies along a continuum and that the five segments are best read as soft profiles, not hard categories.
- The elbow curve bends gradually rather than breaking sharply, another sign that there are no strongly separated groups. Five clusters is a reasonable choice, not a definitive one.
Next steps
- Profile each cluster against the withheld obesity label to see whether segments line up with clinical categories.
- Try Gaussian mixture models, which give soft assignments for overlapping groups.
- Use Gower distance or k-prototypes to handle mixed numeric and categorical data natively.