All projects
NLPMSc Data Science · University of Salford

IMDB Review Sentiment Classifier

PythonNLPTF-IDFscikit-learnLogistic regressionLinear SVM

A text-classification pipeline from raw reviews to evaluated models: HTML and noise removal, TF-IDF unigram and bigram features, and a head-to-head between logistic regression and a linear SVM, finished with ROC-AUC and a manual look at the reviews each model gets wrong.

90.1%Accuracy on held-out reviews
0.966ROC-AUC
50,000Reviews
10,000Test reviews

The problem

Free-text opinions are the richest feedback a business gets and the hardest to read at scale. The goal is a fast, explainable classifier that turns raw reviews into a sentiment label, and a clear picture of where it still goes wrong.

The data

  • 50,000 reviews with no missing values; the classes are perfectly balanced, so accuracy is a fair headline metric here.
  • Reviews average 231 words (median 173, longest 2,470) and contain HTML line breaks that must be removed before modelling.

Approach

  1. AuditShape, missing values, duplicates and review-length statistics in characters and words.
  2. CleanStrip HTML tags, lower-case, keep letters and apostrophes, collapse whitespace.
  3. ExploreClass balance, length distribution by sentiment and word clouds for each class.
  4. VectoriseTF-IDF with unigrams and bigrams, capped at 10,000 features, ignoring terms seen in fewer than five reviews; fitted on the training split only to avoid leakage.
  5. ModelLogistic regression and a linear SVM trained on the same features.
  6. EvaluateAccuracy, precision, recall, F1, confusion matrices and ROC-AUC, followed by a manual inspection of misclassified reviews.

Results

Held-out results (10,000 reviews; precision, recall and F1 refer to the positive class)
ModelAccuracyPrecisionRecallF1ROC-AUC
Logistic regression90.1%89.5%90.8%90.2%0.9655
Linear SVM89.8%89.3%90.3%89.8%0.9639

Confusion matrices on the test set

Logistic regression
Linear SVM

Colour shows the share of each true class.

Word cloud of positive reviews
Most frequent words in positive reviews.
Word cloud of negative reviews
Most frequent words in negative reviews.

Key findings

  • Logistic regression edges the linear SVM on every headline metric (accuracy 90.1% against 89.8%, ROC-AUC 0.9655 against 0.9639), a gap of 34 reviews in 10,000, so the two are effectively tied.
  • Errors are balanced across classes (531 false positives against 459 false negatives for logistic regression), so there is no systematic lean towards either sentiment.
  • A sparse linear model on TF-IDF bigrams is a strong, fast and explainable baseline for this task.
  • Reading the misclassified reviews shows where bag-of-words features reach their limit: negation, mixed opinions and hedged verdicts.

Next steps

  • Remove exact duplicate reviews (418 exist) before splitting.
  • Fine-tune a transformer such as DistilBERT and compare it against this baseline.
  • Publish the most influential words and phrases per class from the model coefficients.