Codebasics Statistical ML Build Framework

Given any prediction problem with labelled numeric data, the user can select the right ML algorithm, clean the data, train and evaluate a production-ready model using scikit-learn, and interpret results with precision, recall, and confusion matrices.

// TL;DR

The Codebasics Statistical ML Build Framework is a step-by-step method for building production-ready machine learning models on labelled numeric data using scikit-learn. You use it whenever you have a tabular dataset with clear input features and a defined target — and a lightweight statistical model (regression or classification) beats overkill generative AI. It walks you from defining the task type, cleaning messy data, and splitting train/test sets, through selecting the right algorithm, training with model.fit, running inference, and evaluating with R², precision, recall, F1, and confusion matrices. Use it for delivery-time prediction, fraud detection, churn, and any prediction problem where cheap, fast, and interpretable wins.

// When should you use the Codebasics Statistical ML Build Framework?

Use this skill whenever you have a labelled dataset and need a lightweight, fast, cost-effective prediction solution — especially when generative AI is overkill and a statistical model (regression, classification) will do the job. Ideal for numeric tabular data with clear input features and a defined target variable.

// What do you need before you start building a statistical ML model?

  • Problem statementrequired
    What are you trying to predict? State the business context clearly (e.g. 'predict delivery time for a food platform').
  • Dataset descriptionrequired
    What columns/features are available? What is the target (dependent) variable? Is it continuous (regression) or categorical (classification)?
  • Task typerequired
    Is this regression, binary classification, or multiclass classification?
  • Data quality notes
    Known issues: missing values, inconsistent formats, outliers, non-numeric columns that need encoding.
  • Evaluation priority
    Which matters more — overall accuracy, precision, or recall? (e.g. fraud detection demands high recall)

// What core principles guide statistical machine learning with scikit-learn?

Car vs. Bike Principle

Machine learning models are the 'bike' in your AI toolkit — lightweight, cheap to train, and faster to reach the destination than a large generative AI 'car'. In any AI project, use generative AI for some parts and statistical machine learning for others, based on the situation. Don't use a sword to cut an apple.

Input-Output vs. Input-Logic

In traditional software you supply input + logic (code) and get output. In machine learning you supply input + output (labelled data) and the training process derives the logic (the equation or model). This is the core difference — ML trains machines on data so they can make predictions without explicit programming.

Training vs. Inference

All ML work has two phases. Training: feed labelled data into an algorithm to find the best-fitting parameters. Inference: give a new unseen input to the trained model and it produces a prediction. Never confuse evaluating on training data with evaluating on unseen data.

Data Matters More Than the Model

In any statistical ML project, 80% of effort goes into cleaning and preparing data. The model training and fine-tuning is a small fraction. Firsthand reality in enterprise projects: clients describe their data as clean but it is always messy. Accept this and prioritise data quality above algorithm selection.

Train-Test Split Principle

Never evaluate a model on the data it was trained on — that measures memorisation, not learning. Always split data (e.g. 80/20) into a training set and a held-out test set. Use test-set performance as your confidence benchmark before deploying to production.

Accuracy Is Not Enough

Relying on accuracy alone is dangerous for imbalanced datasets (e.g. fraud: 10,000 transactions, 5 fraud cases — a dummy model that always says 'not fraud' scores 99.9% accuracy). Always examine the classification report (precision, recall, F1) and confusion matrix alongside accuracy.

Ground Truth / Labelled Dataset

The foundation of supervised learning is a labelled dataset — records where both the input features and the correct output label are known. This is called ground truth. Without it, supervised training cannot occur.

// How do you build and evaluate a production-ready ML model step by step?

  1. 1

    Define the ML task type

    Ask: Is the target variable a continuous number (regression) or a category (classification)? If classification, is it binary (2 classes) or multiclass (3+)? This determines which algorithm family to use. Examples: predicting delivery time → regression; predicting spam/not-spam → binary classification; predicting flower species → multiclass classification.

  2. 2

    Identify independent variables (features) and the dependent variable (target)

    Features = independent variables = the inputs your model will learn from. Target = dependent variable = what you are predicting. In code: X = df[feature_columns] (DataFrame, capital X because it can be multi-column). y = df[target_column] (Series, lowercase y). Confirm feature relevance to the prediction task before proceeding.

  3. 3

    Audit and clean the data

    Scroll the raw data with your eyeballs first — you will immediately spot patterns. Then programmatically check for: (a) missing values / NaN, (b) inconsistent category labels (e.g. 'Male', 'M', 'm', 'MALE' all meaning the same thing), (c) outliers that are physically impossible (e.g. age = 200), (d) columns that are numeric but stored as strings (need type conversion), (e) question marks or placeholder values masking nulls. Fix these before splitting. Remember: data matters more than the model.

  4. 4

    Split data into training set and test set using train_test_split

    Use sklearn.model_selection.train_test_split(X, y, test_size=0.2, random_state=42). A common split is 80% train / 20% test. Set random_state for reproducibility — without it, each run produces a different split, making benchmarking unreliable. Never touch the test set during training or feature engineering decisions.

  5. 5

    Select and instantiate the correct algorithm

    Regression problems → LinearRegression (start here, especially for continuous numeric targets with linear relationships). Binary classification → LogisticRegression (despite the name, it is a classifier; applies a sigmoid/logit function on top of a linear model). Multiclass classification → LogisticRegression also handles this natively. Later in the course: Decision Trees, Random Forest, XGBoost, SVM, KNN, K-Means (unsupervised). Rule of thumb: start simple (linear/logistic), then upgrade if performance is insufficient.

  6. 6

    Train the model by calling model.fit(X_train, y_train)

    This single call is where training happens. For LinearRegression it finds the best-fitting line by minimising Mean Squared Error (MSE) using gradient descent — an iterative mathematical technique that adjusts slope (m) and intercept (b) until error is minimised. For LogisticRegression it fits a linear model then applies the sigmoid function to squash output to [0, 1]. After fitting, inspect model.coef_ (the learned slopes/weights) and model.intercept_ to understand what the model learned.

  7. 7

    Run inference using model.predict(X_new)

    Always pass a 2D array or DataFrame — sklearn expects this format even for a single prediction. For regression: output is a continuous number. For classification: output is the predicted class label. For probability scores in classification: use model.predict_proba(X_new) — returns probability per class, which is more informative than a hard label.

  8. 8

    Evaluate model performance on the test set

    For regression: use model.score(X_test, y_test) → returns R² score (0 to 1; closer to 1 is better). Also compare train R² vs test R² — a large gap signals overfitting. For classification: (a) model.score(X_test, y_test) → overall accuracy, but do not stop here. (b) Print classification_report(y_test, y_pred) → shows precision, recall, and F1 per class. (c) Plot confusion_matrix(y_test, y_pred) → diagonal = correct predictions; off-diagonal = mistakes. Understand which type of error matters most for your use case before declaring the model ready.

  9. 9

    Interpret precision, recall, and F1 against the business objective

    Precision (for a class): of all records the model labelled as this class, what fraction were truly that class? Start from predictions. Recall (for a class): of all records that truly belong to this class, what fraction did the model correctly identify? Start from ground truth. F1 = 2 × (precision × recall) / (precision + recall) — a single number balancing both. For fraud/medical/safety use cases, recall is typically more important (you cannot afford to miss true positives). For spam filtering, precision may matter more (false positives annoy users).

  10. 10

    Iterate and practice with exercises before declaring completion

    Do not just watch — practice on a new dataset with a fresh problem statement. Watching without coding is like watching a pickleball coaching video and expecting to become an expert. Write the code yourself. Only submit your model to production once you have a test-set benchmark you are confident in.

// What are real examples of applying this ML framework?

A food delivery platform wants to estimate delivery time for new orders. Historical data includes distance, traffic level, and weather conditions for past deliveries.

Task type: regression (continuous target — minutes). Features (independent variables): distance, traffic level, raining. Target (dependent variable): delivery time in minutes. Use LinearRegression. After training, model.coef_ gives minutes-per-unit for each feature; model.intercept_ gives the fixed preparation time baseline. Evaluate with R² on the held-out test set. The trained equation is: time = m1*distance + m2*traffic + m3*rain + b.

An insurance company wants to predict whether a prospect will purchase insurance based on their age, to prioritise outreach.

Task type: binary classification (0 = no insurance, 1 = has insurance). Feature: age. Target: has_insurance. A straight linear regression line gets skewed by outliers (e.g. a 90-year-old without insurance). Use LogisticRegression instead — it applies the sigmoid function to produce a probability between 0 and 1, making the decision boundary robust to outliers. Use predict_proba() to rank prospects by likelihood rather than using a hard cut-off. Evaluate with classification_report and confusion_matrix; focus on recall if missing a buyer is costly.

A botanical research tool needs to classify flower species (3 types) based on physical measurements of their leaf structures.

Task type: multiclass classification (3 output categories). Features: 4 numeric measurements. Target: species name. Use LogisticRegression (handles multiclass natively). Visualise feature distributions with scatter plots per class before training — some classes may be trivially separable (reducing expected error). After training, compare predicted vs actual labels in a DataFrame. Confusion matrix diagonal shows correct predictions per class; off-diagonal reveals which species pairs are most confused. Report precision/recall per class — easy-to-separate classes should score near 1.0.

A bank wants to detect fraudulent credit card transactions. Out of 10,000 transactions, only 5 are fraud.

Task type: binary classification with severe class imbalance. Do NOT rely on accuracy — a dummy model always predicting 'not fraud' scores 99.95% accuracy while being completely useless. Focus on recall for the fraud class (missing a fraud is catastrophic) and examine the confusion matrix. A model with 70% recall for fraud is far more valuable than a 99% accurate model with 0% fraud recall. Use classification_report to expose this distinction clearly.

// What mistakes should you avoid when building statistical ML models?

  • Evaluating your model on training data instead of a held-out test set — this measures memorisation, not generalisation. Always use train_test_split and evaluate on X_test, y_test only.
  • Relying on accuracy alone for imbalanced datasets. A dummy model that always predicts the majority class can score 99%+ accuracy while being completely useless. Always check precision, recall, F1, and confusion matrix.
  • Passing a 1D array to model.predict() — sklearn expects a 2D array. Wrap single predictions in a list-of-lists or a DataFrame.
  • Using LinearRegression for binary classification targets. The straight best-fit line is vulnerable to outliers that skew the decision boundary. Use LogisticRegression (sigmoid function) for classification tasks.
  • Skipping data cleaning before training. Real-world data will have inconsistent category labels, impossible values, missing entries, and wrong types. Feeding messy data to a model produces unreliable output regardless of algorithm quality.
  • Treating the model as more important than the data. 80% of time in real ML projects is data cleaning and preparation; model training is a small fraction. Underinvesting in data quality is the primary source of poor model performance in practice.
  • Not setting random_state in train_test_split — without it, each run produces a different split, making benchmark comparisons meaningless and results non-reproducible.
  • Confusing the two phases: training (model.fit) must always precede inference (model.predict). Never call predict before fit.
  • Applying only Simple Linear Regression when multiple factors affect the target. Real-world problems almost always require Multiple Linear Regression with several features (independent variables) simultaneously.

// What are the key terms in statistical machine learning?

Statistical Machine Learning
The discipline of training machines on labelled data using mathematical/statistical algorithms (linear regression, logistic regression, decision trees, SVM, etc.) so they can make predictions without explicit programming. Lightweight, cheap to train, and highly relevant alongside generative AI.
Ground Truth / Labelled Dataset
A dataset where each record has both input features and the correct output label already assigned. This is what you feed to a supervised ML algorithm during training.
Training
The phase where you run model.fit(X_train, y_train) — the algorithm finds the best-fitting parameters (e.g. slope m and intercept b) by minimising error across the labelled dataset.
Inference
The phase where you call model.predict(X_new) on a trained model to produce predictions for unseen data.
Independent Variable (Feature)
An input variable used to make a prediction. Called a 'feature' in ML contexts. There can be many features simultaneously (Multiple Linear Regression). Represented as capital X — a DataFrame.
Dependent Variable (Target)
The variable you are trying to predict. It 'depends on' the features. Represented as lowercase y — a Series.
Gradient Descent
The core mathematical optimisation technique used during training to iteratively adjust model parameters (m and b) until the Mean Squared Error is minimised and the best-fitting line or boundary is found.
Mean Squared Error (MSE)
The loss function used in linear regression: average of squared differences between actual values (y) and predicted values (ŷ). Squaring prevents positive and negative errors from cancelling out and penalises large errors more heavily than Mean Absolute Error (MAE).
R² Score
The evaluation metric for regression models, ranging 0 to 1. Closer to 1 means the model explains the variance in the target variable well. Used via model.score(X_test, y_test).
Train-Test Split
The practice of dividing a labelled dataset into a training portion (e.g. 80%) used to fit the model, and a test portion (e.g. 20%) held out for unbiased evaluation. Implemented via sklearn's train_test_split.
Sigmoid / Logit Function
The mathematical function applied in logistic regression: f(x) = 1 / (1 + e^(-x)). It compresses any input number into a probability between 0 and 1, enabling classification. Makes the decision boundary robust to outliers that would skew a straight linear regression line.
Binary Classification
A classification task where the output has exactly two categories (e.g. fraud / not fraud, spam / not spam, has insurance / no insurance).
Multiclass Classification
A classification task where the output has three or more categories (e.g. flower species: setosa, versicolor, virginica; news categories: business, technology, entertainment).
Precision
Of all records the model predicted as a given class, what fraction were truly that class? Start from predictions. High precision means few false positives.
Recall
Of all records that truly belong to a given class, what fraction did the model correctly identify? Start from ground truth. High recall means few false negatives — critical for fraud, medical, or safety use cases.
F1 Score
A single metric combining precision and recall: F1 = 2 × (precision × recall) / (precision + recall). Useful when you need one number that balances both concerns.
Confusion Matrix
A grid showing correct and incorrect predictions per class. Diagonal entries = correct predictions; off-diagonal entries = mistakes. Gives an immediate visual clue about which types of errors the model is making.
Classification Report
The sklearn output from classification_report(y_true, y_pred) showing precision, recall, F1, and support for every class plus overall accuracy. Always use this instead of accuracy alone.
Multiple Linear Regression
Linear regression with more than one independent variable: y = m1*x1 + m2*x2 + ... + mn*xn + b. This is what real-world regression models look like — a single feature is rarely sufficient.
Coefficient (model.coef_)
The learned slope values (m) for each feature after training. Represents the contribution of each feature to the prediction — e.g. extra minutes per kilometre in a delivery time model.
Intercept (model.intercept_)
The learned constant (b) after training. Represents the baseline prediction when all features are zero — e.g. the fixed preparation time before a delivery driver even leaves.

// FREQUENTLY ASKED QUESTIONS

What is the Codebasics Statistical ML Build Framework?

It's a repeatable workflow for building production-ready machine learning models on labelled numeric data using scikit-learn. It covers defining the task type (regression, binary, or multiclass classification), cleaning messy data, splitting train/test sets, selecting the right algorithm, training with model.fit, running inference, and evaluating with R², precision, recall, F1, and confusion matrices — all interpreted against your business objective.

What is statistical machine learning and how is it different from generative AI?

Statistical machine learning trains mathematical algorithms (linear regression, logistic regression, decision trees) on labelled data to make predictions without explicit programming. Unlike generative AI, it's lightweight, cheap to train, and fast. The framework calls it the 'bike' versus generative AI's 'car' — use the bike for tabular prediction tasks where a big generative model would be overkill. Don't use a sword to cut an apple.

How do I choose between regression and classification for my problem?

Look at your target variable. If it's a continuous number (like delivery time in minutes), use regression with LinearRegression. If it's a category, use classification: binary (two classes like fraud/not-fraud) or multiclass (three or more like flower species). Both classification types use LogisticRegression, which handles multiclass natively. This single decision determines your algorithm family before you write any model code.

How do I train and evaluate a model in scikit-learn step by step?

Split your data with train_test_split(X, y, test_size=0.2, random_state=42), instantiate your algorithm (LinearRegression or LogisticRegression), call model.fit(X_train, y_train) to train, then model.predict(X_new) for inference. Evaluate only on the held-out test set: use model.score for R² or accuracy, then classification_report for precision/recall/F1 and confusion_matrix to see which errors the model makes.

How does this framework compare to just throwing data at a neural network?

This framework prioritises the simplest model that works — start with linear or logistic regression, then upgrade only if performance is insufficient. Neural networks are heavier, slower, harder to interpret, and often unnecessary for tabular numeric data. The framework's 'Car vs. Bike' principle argues statistical ML reaches the destination faster and cheaper. It also enforces disciplined data cleaning, which matters more than algorithm choice.

When should I use this framework instead of generative AI?

Use it whenever you have a labelled dataset with clear input features and a defined target variable, and you need a fast, cheap, interpretable prediction. It's ideal for numeric tabular data — delivery-time estimation, churn prediction, fraud detection, or lead scoring. If a statistical model (regression or classification) solves the problem, generative AI is overkill and wastes compute, money, and latency.

Why shouldn't I trust accuracy alone when evaluating my model?

Accuracy is dangerous for imbalanced datasets. In fraud detection with 10,000 transactions and 5 fraud cases, a dummy model that always predicts 'not fraud' scores 99.95% accuracy while being completely useless. Always examine the classification report (precision, recall, F1) and confusion matrix. For fraud, medical, or safety use cases, recall matters most — you cannot afford to miss true positives.

What's the difference between precision and recall?

Precision starts from predictions: of all records the model labelled as a class, what fraction were truly that class? High precision means few false positives. Recall starts from ground truth: of all records that truly belong to a class, what fraction did the model catch? High recall means few false negatives. Fraud and medical use cases prioritise recall; spam filtering often prioritises precision.

What results can I expect after applying this framework?

You'll produce a trained model with a trustworthy test-set benchmark you can confidently deploy — an R² score for regression, or precision, recall, F1, and a confusion matrix for classification. You'll understand what the model learned via model.coef_ and model.intercept_, know which errors it makes, and be able to justify readiness against your specific business objective rather than a misleading accuracy number.

Why does data cleaning matter more than the algorithm?

In real enterprise ML projects, 80% of effort goes into cleaning and preparing data; model training is a small fraction. Clients always describe their data as clean, but it's always messy — inconsistent category labels, impossible values, missing entries, and wrong types. Feeding messy data to any algorithm produces unreliable output. Prioritising data quality above algorithm selection is the biggest driver of model performance.

What is a train-test split and why is random_state important?

A train-test split divides labelled data into a training portion (typically 80%) to fit the model and a held-out test portion (20%) for unbiased evaluation. Never evaluate on training data — that measures memorisation, not learning. Setting random_state=42 fixes the split so every run is reproducible; without it, each run produces a different split, making benchmark comparisons meaningless.

Why use LogisticRegression instead of LinearRegression for classification?

LinearRegression fits a straight best-fit line that gets skewed by outliers, corrupting the decision boundary for classification. LogisticRegression fits a linear model then applies the sigmoid function to squash output into a probability between 0 and 1, making the boundary robust to outliers. Despite its name, it's a classifier — and predict_proba() lets you rank predictions by likelihood instead of a hard cut-off.

// GET THIS SKILL — FREE

Use this skill in your AI

Every skill on SkillForge is free. Drop your email and copy this skill straight into Claude, ChatGPT, or any LLM.

We'll email you when new skills drop. Unsubscribe anytime.