DP
DHRUV PATEL
AI & DATA SCIENCE · GEC RAJKOT
STATUS
● ONLINE
DATE
VERSION
v 2.6
SYSTEM BOOT PROGRESS0%
INITIALIZING SYSTEMS...
PORTFOLIO v2.6
Back to Projects
Machine Learning

Cricket Score Prediction System

An ML-powered cricket match score predictor built with Python, Scikit-learn, and Streamlit. Predicts final innings score with min/max range and upcoming overs bar chart using IPL historical data.

Interactive Showcase Gallery (3 Slides)

Cricket Score Prediction System - Predictive Engine
Slide 1 of 3

Predictive Engine

Project Specifications

CategoryMachine Learning
Statuscompleted
Author / ArchitectDhruv Patel

Ready to explore this project in action?

Check out the live interactive deployment or inspect the full source code on GitHub.

Technologies Used

Python 3.9+Scikit-learnStreamlitPandasNumPyMatplotlibGradient Boosting RegressorOne-Hot EncodingStandard ScalerPickle

About The Project

The Cricket Score Prediction System is a machine learning web application that predicts the realistic final score of a cricket innings in real time. It is powered by a Gradient Boosting Regressor trained on a cleaned IPL historical dataset (ipl.csv).

The model goes beyond simple run-rate projections by engineering smart features that capture match context: current_run_rate (how fast runs are scoring), projected_score (linear extrapolation), pressure_factor (derived from wickets and overs remaining), and is_death_overs (binary flag for overs 17-20 where scoring accelerates). Teams and venues are encoded with One-Hot Encoding; numerical features are Standard Scaled.

The Streamlit app allows users to input current match state (runs, overs, wickets, batting/bowling team, venue) and instantly receive a prediction showing the average score, a minimum and maximum range, and a bar chart of projected scores for each upcoming over. Sample dataset viewing is built in.

The project is fully separated: train_model.py (data prep, feature engineering, training, saving .pkl files) and Cricket-Score-Prediction.py (Streamlit UI, model loading, inference). This separation of training and serving is a production ML best practice. Deployed live on Streamlit Cloud.

Key Features & Architecture Highlights

Gradient Boosting Regressor trained on cleaned IPL historical dataset

Smart feature engineering: run_rate, projected_score, pressure_factor, is_death_overs

One-Hot Encoding for teams/venue + Standard Scaling for numerical features

Predicts final score as average with min/max realistic range

Bar chart of projected scores for each remaining over

Clean separation: train_model.py (training) vs app file (serving) — production ML pattern

Deployed live on Streamlit Cloud with interactive web UI

In-app dataset viewer for transparency and exploration

Interested in building something extraordinary?

Whether you have an ambitious AI project, full-stack application, or collaboration in mind, I'd love to connect.