End-to-End Data Science & Product Analytics

Product User Churn Prediction & Analytics

An end-to-end portfolio project combining Python, SQL, ETL, product analytics, machine learning, model evaluation, automated testing, and an interactive Streamlit dashboard.

5,000 Users
229K+ Product Events
17 Predictive Features
2.72% Churn Rate

Project Overview

The project demonstrates how raw product event data can be transformed into actionable product analytics and a machine-learning based churn risk-ranking system.

Business Problem

Identify users showing behavioral signals associated with future churn so product and retention teams can prioritize analysis and potential interventions.

Data Engineering

Raw event data is cleaned, validated, transformed through an ETL pipeline, and loaded into SQLite for analytical SQL workflows.

Machine Learning

Logistic Regression and Random Forest models are evaluated using metrics appropriate for highly imbalanced churn classification.

Results at a Glance

The project combines product analytics with machine learning to identify behavioral signals associated with future churn and support retention prioritization.

Dataset

5,000 users and more than 229,000 product events were generated and processed through the ETL pipeline.

ML Dataset

3,307 eligible users with 17 predictive features were used for churn modeling.

Churn Rate

90 churners and 3,217 non-churners produced a churn rate of 2.72%, creating a highly imbalanced classification problem.

Model Performance โ€” 80/20 Holdout

Logistic Regression ROC-AUC 0.931
Logistic Regression PR-AUC 0.168
Random Forest ROC-AUC 0.953
Random Forest PR-AUC 0.402

Robustness โ€” 5-Fold Cross-Validation

Mean ROC-AUC 0.919
ROC-AUC Std. Dev. 0.013
Mean PR-AUC 0.244
PR-AUC Std. Dev. 0.054
What the analysis found

Churned users showed lower average engagement and substantially higher activity recency. Subscription history, activity recency, and purchase-related features were important model signals. The subscription_count result has an important synthetic-data caveat because the churn definition is tied to cancellation after subscription.

Project Architecture

The workflow connects data generation, ETL, SQL analytics, feature engineering, machine learning, evaluation, and dashboarding.

Raw Events
Product activity data
โ†’
Python ETL
Clean & validate
โ†’
SQLite
Store clean events
โ†’
SQL Features
User-level metrics
โ†’
ML Dataset
17 predictive features
โ†’
ML Models
LR + Random Forest
โ†’
Risk Score
Prioritize users
โ†’
Dashboard
Product insights

Product Analytics

User behavior is summarized through engagement, session, conversion, and recency features.

Engagement Features

Total events Count
Total sessions Count
Active days Count
Logins / product views Count

Conversion Features

Search rate Rate
Cart rate Rate
Purchase rate Rate

Recency Features

Days since last activity Days
Days since last login Days
Days since last product view Days
Events per active day Rate
Sessions per active day Rate

Machine Learning Results

The target is highly imbalanced, with churn representing only 2.72% of eligible users. The project therefore emphasizes ROC-AUC, PR-AUC, precision, recall, and threshold analysis rather than accuracy alone.

Holdout Evaluation

Logistic Regression ROC-AUC 0.931
Logistic Regression PR-AUC 0.168
Random Forest ROC-AUC 0.953
Random Forest PR-AUC 0.402

5-Fold Cross-Validation

Mean ROC-AUC 0.919
ROC-AUC Std. Dev. 0.013
Mean PR-AUC 0.244
PR-AUC Std. Dev. 0.054
Model interpretation

Subscription history was the strongest feature, but it has an important synthetic-data caveat: the churn definition is based on cancellation after subscription. Activity recency and purchase-related behavior also provide predictive signal. Feature importance indicates model contribution, not causation.

Churn Risk Score

The dashboard uses the Random Forest output as a relative risk score rather than a literal calibrated probability. Thresholds can be adjusted to explore the trade-off between identifying potential churners and targeting additional users.

Interactive Dashboard

The Streamlit dashboard brings together product KPIs, churn analysis, model performance, feature importance, risk segmentation, and user-level risk exploration.

Product Churn Analytics Streamlit dashboard

Engineering Quality

Automated Tests

Eight automated tests cover ETL behavior and ML dataset validation.

Continuous Integration

GitHub Actions automatically runs the project test suite.

Reproducibility

The project contains scripts, requirements, SQL, notebooks, documentation, and validation workflows.

Limitations & Production Roadmap

This is a production-style portfolio prototype using synthetic data. A production implementation would require real product data, a business-approved churn definition, point-in-time feature generation, time-based validation, leakage auditing, probability calibration, model monitoring, data-quality monitoring, and controlled experiments to measure whether retention interventions actually reduce churn.