All projects
Machine learningMSc Data Science · University of Salford

Lifestyle and Obesity Segmentation

PythonUnsupervised learningK-MeansHierarchical clusteringPCAscikit-learn

Can lifestyle data alone reveal natural groups related to weight? Using 16 survey features, I built a preprocessing pipeline, chose the number of clusters with the elbow method, compared K-Means with Ward hierarchical clustering and judged both with three internal metrics and a PCA projection. The scores are candid: the segments overlap, which is itself a finding about how continuous lifestyle behaviour is.

2,111Records
31Features after encoding
k = 5Clusters (elbow method)
3 of 3Metrics won by K-Means

The problem

Weight is shaped by many habits at once: diet, water intake, activity, screen time, transport. Rather than predicting a label, this project asks whether people fall into distinct lifestyle profiles, and how clear those profiles really are.

The data

  • 2,111 records and 17 columns, 16 of them features (gender, age, height, weight, family history, eating habits, smoking, water intake, activity, screen time, alcohol, transport). No missing values; 24 duplicate rows noted.
  • Height and weight are the most correlated numeric pair (r = 0.46). The behavioural features are mostly weakly correlated with each other (|r| of 0.3 or less), so they add complementary information.

Approach

  1. Audit and exploreDescriptive statistics, histograms, boxplots, categorical counts, a correlation heatmap and a pairplot.
  2. Keep it unsupervisedThe obesity-level label is removed from the clustering input.
  3. PreprocessStandardise numeric features and one-hot encode categorical ones, giving 31 dimensions.
  4. Choose kElbow curve of within-cluster sum of squares for k = 1 to 10, pointing to five clusters.
  5. Cluster two waysK-Means (k-means++ initialisation) and agglomerative clustering with Ward linkage, guided by a dendrogram on a 400-record sample.
  6. ValidateSilhouette, Davies–Bouldin and Calinski–Harabasz scores, plus a 2-D PCA projection of each result.

Results

Internal validity metrics (k = 5)
MethodSilhouette ↑Davies–Bouldin ↓Calinski–Harabasz ↑
K-Means0.1411.987260.2
Hierarchical (Ward)0.1202.066220.4
Elbow curve of within-cluster sum of squares
Elbow curve used to choose five clusters.
K-Means clusters in PCA space
K-Means clusters in a 2-D PCA projection.
Hierarchical clusters in PCA space
Hierarchical clusters in the same projection.

Key findings

  • K-Means wins on all three internal metrics (silhouette 0.141 against 0.120, Davies–Bouldin 1.99 against 2.07, Calinski–Harabasz 260 against 220).
  • Silhouette scores near 0.14 indicate overlapping clusters. The PCA views show a continuous cloud rather than separate islands, which suggests that lifestyle varies along a continuum and that the five segments are best read as soft profiles, not hard categories.
  • The elbow curve bends gradually rather than breaking sharply, another sign that there are no strongly separated groups. Five clusters is a reasonable choice, not a definitive one.

Next steps

  • Profile each cluster against the withheld obesity label to see whether segments line up with clinical categories.
  • Try Gaussian mixture models, which give soft assignments for overlapping groups.
  • Use Gower distance or k-prototypes to handle mixed numeric and categorical data natively.