IMDB Review Sentiment Classifier
PythonNLPTF-IDFscikit-learnLogistic regressionLinear SVM
A text-classification pipeline from raw reviews to evaluated models: HTML and noise removal, TF-IDF unigram and bigram features, and a head-to-head between logistic regression and a linear SVM, finished with ROC-AUC and a manual look at the reviews each model gets wrong.
90.1%Accuracy on held-out reviews
0.966ROC-AUC
50,000Reviews
10,000Test reviews
The problem
Free-text opinions are the richest feedback a business gets and the hardest to read at scale. The goal is a fast, explainable classifier that turns raw reviews into a sentiment label, and a clear picture of where it still goes wrong.
The data
- 50,000 reviews with no missing values; the classes are perfectly balanced, so accuracy is a fair headline metric here.
- Reviews average 231 words (median 173, longest 2,470) and contain HTML line breaks that must be removed before modelling.
Approach
- AuditShape, missing values, duplicates and review-length statistics in characters and words.
- CleanStrip HTML tags, lower-case, keep letters and apostrophes, collapse whitespace.
- ExploreClass balance, length distribution by sentiment and word clouds for each class.
- VectoriseTF-IDF with unigrams and bigrams, capped at 10,000 features, ignoring terms seen in fewer than five reviews; fitted on the training split only to avoid leakage.
- ModelLogistic regression and a linear SVM trained on the same features.
- EvaluateAccuracy, precision, recall, F1, confusion matrices and ROC-AUC, followed by a manual inspection of misclassified reviews.
Results
| Model | Accuracy | Precision | Recall | F1 | ROC-AUC |
|---|---|---|---|---|---|
| Logistic regression | 90.1% | 89.5% | 90.8% | 90.2% | 0.9655 |
| Linear SVM | 89.8% | 89.3% | 90.3% | 89.8% | 0.9639 |
Confusion matrices on the test set
predicted negativepredicted positive
actual negative
4,46989.4% of row
53110.6% of row
actual positive4599.2% of row
4,54190.8% of row
predicted negativepredicted positive
actual negative
4,46189.2% of row
53910.8% of row
actual positive4859.7% of row
4,51590.3% of row
Colour shows the share of each true class.
Key findings
- Logistic regression edges the linear SVM on every headline metric (accuracy 90.1% against 89.8%, ROC-AUC 0.9655 against 0.9639), a gap of 34 reviews in 10,000, so the two are effectively tied.
- Errors are balanced across classes (531 false positives against 459 false negatives for logistic regression), so there is no systematic lean towards either sentiment.
- A sparse linear model on TF-IDF bigrams is a strong, fast and explainable baseline for this task.
- Reading the misclassified reviews shows where bag-of-words features reach their limit: negation, mixed opinions and hedged verdicts.
Next steps
- Remove exact duplicate reviews (418 exist) before splitting.
- Fine-tune a transformer such as DistilBERT and compare it against this baseline.
- Publish the most influential words and phrases per class from the model coefficients.