Juliet AsantewaaSarpong

Data scientist and statistician. I find out what actually drives results, not just what moves with them.

Four years turning messy, real-world data into decisions for Microsoft, Unilever, P&G and BIC, from pricing and market analysis to audience and media strategy. Now a PhD researcher in Statistics at the University of Edinburgh, specialising in causal inference, and bringing that rigour to customers, campaigns and clinical data alike.

Experience

My career, plotted the way a statistician would: each bar is a role or degree, from start to finish.

2025 – present

Tutor, Mathematics & Statistics

University of Edinburgh

Tutorials and academic support for undergraduate and postgraduate students. Also a Postgraduate Research Ambassador since late 2025.

2022 – 2023

Project Supervisor

Polymorph Labs

Built analytics systems to identify learning gaps and automate performance reporting, improving student engagement and comprehension by about 30%.

2021 – 2022

Data Scientist

Maverick Research

Market research for consumer goods brands including BIC, Unilever and Fanmilk across 15+ cities.

  • Cleaned and validated field-agent retail data, improving accuracy by 20%, and designed the SQL database behind it.
  • Grouped products into categories by use and tracked pricing, supply, demand and distribution within each client's category, such as stationery for BIC.
  • Identified the regions with most growth potential to shape where campaigns ran and what they said, contributing to over 25% sales growth.
2020 – 2021

Data Analyst & Media Executive

Carat (Dentsu)

Media and audience analytics for Microsoft, P&G, Beiersdorf and Betway.

  • Ran demographic and audience analysis to match segments to each product, and presented findings in client strategy meetings to choose media channels.
  • Built automated media-planning templates that cut operational workload by 20%, and optimised campaigns that saved up to 30% in cost.

Projects

Applied work in Python and R, each built on real data and written up so you can check every number.

As a campaign targets high-value customers more strongly, the naive estimate climbs away from the true effect while TMLE stays on it

Did the email campaign actually work?

When a campaign targets its best customers, a naive comparison overstated its impact by about 60%. Checked against a randomised experiment on 64,000 customers, AIPW and TMLE, implemented from scratch, recover the true effect with honest confidence intervals.

Python, scikit-learn, causal inference, simulation

Cumulative gains curve: calling the top-scored 30% of bank clients reaches 75% of subscribers

Who should the bank call?

On 41,188 real telemarketing calls, targeting the model's top 30% of clients wins 2.5 times as many term-deposit subscriptions as random calling on the same budget, and keeps 93% of the profit with 70% fewer calls.

Python, XGBoost, random forest, SHAP, business case modelling

Interactive app predicting a board game's BoardGameGeek rating and what drives it

What makes a board game highly rated?

An analysis of 15,249 BoardGameGeek games, with an app that predicts how a new game would be rated. Release year is the strongest predictor, and many popular mechanics lose their advantage once year and length are held constant.

R, tidyverse, glmnet, ranger, Shiny running in the browser

Research

Methods for separating cause from correlation when you can't run an experiment.

C-TMLE for causal inference in high-dimensional genomic data

PhD thesis, University of Edinburgh

Estimating the causal effects of genetic variants on disease from observational UK Biobank data, including settings where data are missing in non-random ways. The focus is on doubly robust, targeted estimators that stay reliable when models are misspecified.

Targeted Learning in Genomics and Molecular BiomedicineUnder review

Co-author, International Journal of Biostatistics

Contributed the C-TMLE implementation, integrating it into the open-source Julia package TMLE.jl.

Benchmarking causal machine learning estimators

Simulation study in R on the Eddie HPC cluster

A large-scale comparison of TMLE, C-TMLE variants and scalable C-TMLE against GLM and GLMnet baselines under correctly specified and misspecified models, with cross-fitted nuisance models and jackknife inference chosen for its coverage.

Toolkit

Methods
Causal inference, statistical modelling, machine learning, market and audience analysis
Languages
R, Python, Julia, SQL
Visualisation
Tableau, Power BI, Looker
Data & infrastructure
Data cleaning and validation, SQL database design, Git, HPC, cloud platforms

Education

PhD Statistics
University of Edinburgh, September 2024 – present. School of Mathematics funding package.
MSc Mathematical Sciences
African Institute for Mathematical Sciences, 2023 – 2024. Bending Spoons Scholar.
BSc Mathematics, First Class
Kwame Nkrumah University of Science and Technology, 2016 – 2020. Ato-Dadzie Foundation Scholar.

Outside work

I'm usually at a table somewhere: playing tabletop RPGs like the Cosmere RPG and City of Mist, or working through a stack of board games. The rest of the time I'm reading, comic books very much included, and I review most of what I read on Goodreads.