Data scientist in Manchester.
I'm Abdullah Mir, a data scientist based in Manchester, UK, with an MSc in Data Science from the University of Salford (January 2025 to April 2026). My work covers machine learning, statistical analysis, SQL and big data, applied to real datasets in an industry placement with Fildata.
Data science internship with Fildata.
My MSc dissertation was an industry-partnered project with Fildata, where I spent four months as a remote data science intern (January to April 2026). I analysed datasets of 300K+ records in Python with Pandas and NumPy, benchmarked five classification models (Logistic Regression, Random Forest, XGBoost, LightGBM and CatBoost) on an imbalanced binary problem, and built Power BI dashboards that turned model outputs into clear insight for non-technical stakeholders.
The work is covered by confidentiality agreements, so it is described on the Experience page rather than shown in detail.
Imbalance-aware machine learning.
- Classification with Logistic Regression, Random Forest, XGBoost, LightGBM and CatBoost, using cross-validation and hyperparameter tuning.
- Evaluation with ROC-AUC, F1 and recall instead of accuracy alone, so a model is judged on the minority class it needs to find.
- Reproducible preprocessing and feature engineering pipelines in Python and scikit-learn.
- SQL and NoSQL: querying and optimising databases, and T-SQL in SQL Server.
- Big data with Spark, Databricks and MLflow.
- Power BI dashboards for non-technical stakeholders.
Seven MSc data science projects, with results.
Each project sets out the question, the data, the method, the results and what I would do next.
- Online Shopper Purchase-Intent Prediction: buyer recall lifted from 35% to 71% on 12,330 sessions.
- IMDB Review Sentiment Classifier: 90.1% accuracy and 0.966 ROC-AUC on 10,000 held-out reviews.
- Lifestyle and Obesity Segmentation: K-Means and hierarchical clustering compared on 2,111 people.
- Steam Game Recommender: collaborative filtering with Spark MLlib on 200,000 records.
- Clinical Trials Analytics at Scale: Spark SQL on more than 520,000 registered clinical trials.
- Airport Ticketing System in SQL Server: a relational back end with stored procedures, views and a trigger.
- Online Shopping Analytics in SQL Server: ten T-SQL answers to business questions.
Certified in data science and machine learning.
The IBM Data Science Professional Certificate (12 courses) and the Machine Learning Specialisation from DeepLearning.AI and Stanford University (three courses), plus my Fildata internship certificate. All are on the Certificates page.
Taking analysis into a product.
Before data science I spent 3.5+ years as a Flutter developer, so I am comfortable taking analysis out of a notebook and into something people use. See my Flutter developer profile.