Research · open data
E-commerce retention model on 500,000+ transactions
Which customers come back? A retention model on a public online-retail dataset that beats the benchmark by 9 points of AUC.
The task
Predict which customers of an online retailer will buy again, and turn the answer into budget advice: how much to spend on winning new customers versus keeping current ones.
The data
UCI Online Retail II: 500,000+ real transactions of a UK online gift retailer, customers in 38 countries, December 2010 – December 2011. Public dataset, CC BY 4.0.
The approach
- Cleaned cancellations, returns and service lines; built average order value and RFM features (recency, frequency, monetary) per customer.
- Tested the difference between customer groups for significance: Welch's t-test, t = 15.42, p < 0.001.
- Trained a Random Forest classifier for retention and compared it with a 0.70 ROC-AUC benchmark.
- Turned model output into recommendations on acquisition budget and retention offers.
Results in numbers
Source: UCI Online Retail II (Chen, 2019), sheet Year 2010-2011, CC BY 4.0; our calculation
- Dec 2010 cohort
- All cohorts, average
Source: UCI Online Retail II (Chen, 2019), sheet Year 2010-2011, CC BY 4.0; our calculation
Charts are recalculated on the same public dataset for our RFM segmentation article; the model metrics above are from the original project.
Where this applies
In a typical online store a minority of customers brings most of the money. We find them in your order data, track whether they come back, and put the numbers that matter on one weekly dashboard.
Read the method
Have a similar question about your data?
Start with a free 20-minute call. You get a written quote with fixed scope and price.
More case studies: Running retail on numbers: 12 years of pricing, purchasing and stock · Staking fund portfolio: the range of outcomes, not one forecast