Frequently Asked Questions About Codebasics Statistical ML Build Framework
22 answers covering everything from basics to advanced usage.
// Basics
What kind of data do I need to use this framework?
You need a labelled dataset — records where both the input features and the correct output label are known (called ground truth). The features should be numeric or convertible to numeric, arranged as tabular data with clear columns. You also need a defined target variable: a continuous number for regression, or a category for classification. Without labelled data, supervised training cannot occur.
What does 'input-output vs input-logic' mean?
In traditional software you supply input plus logic (code) and get output. In machine learning you supply input plus output (labelled data), and the training process derives the logic itself — the equation or model. This is the core difference: ML trains machines on examples so they discover the rules automatically, instead of you programming every rule explicitly.
What is the difference between training and inference?
Training is when you call model.fit(X_train, y_train) — the algorithm finds the best-fitting parameters by minimising error across labelled data. Inference is when you call model.predict(X_new) on the trained model to produce predictions for unseen data. Training must always come first; never call predict before fit. And never confuse evaluating on training data with evaluating on unseen data.
How much should I trust my model before deploying to production?
Trust it only after you have a test-set benchmark you're genuinely confident in — evaluated on held-out data, not training data — and after you've interpreted precision, recall, F1, and the confusion matrix against your specific business objective. Understand which error type matters most for your use case. And practise on a fresh dataset first; watching without coding is like watching a coaching video and expecting expertise.
// How To
How do I identify my features and target in code?
Features are your independent variables (inputs) and go into X = df[feature_columns] — a DataFrame, capital X because it can be multi-column. Target is your dependent variable (what you predict) and goes into y = df[target_column] — a Series, lowercase y. Confirm each feature is actually relevant to the prediction task before proceeding to the split.
How do I audit and clean my data before training?
First scroll the raw data with your eyeballs to spot patterns. Then programmatically check for missing values/NaN, inconsistent category labels ('Male', 'M', 'm', 'MALE'), physically impossible outliers (age = 200), numeric columns stored as strings needing conversion, and placeholder values like question marks masking nulls. Fix all of these before splitting the data — cleaning after the split risks leaking test-set information.
How do I make a single prediction with a trained model?
Always pass a 2D array or DataFrame to model.predict — scikit-learn expects this format even for a single prediction, so wrap it in a list-of-lists or DataFrame. For regression you get a continuous number; for classification you get a predicted class label. For probability scores, use model.predict_proba(X_new), which returns a probability per class — more informative than a hard label.
How do I interpret model.coef_ and model.intercept_?
After training a linear model, model.coef_ gives the learned slope (weight) for each feature — the contribution of that feature to the prediction, like extra minutes per kilometre in a delivery-time model. model.intercept_ gives the baseline prediction when all features are zero, like the fixed preparation time before a driver even leaves. Together they reveal what the model actually learned.
// Troubleshooting
My model scores high on training data but poorly on new data. What's wrong?
You're likely overfitting — the model memorised the training data instead of learning general patterns. Compare train R² versus test R²; a large gap signals overfitting. The fix: always evaluate on the held-out test set only, never on training data. If the gap persists, try a simpler model, more data, or fewer features. High training accuracy alone means nothing.
Why do I get an error when calling model.predict on a single row?
You probably passed a 1D array. Scikit-learn expects a 2D array even for a single prediction. Wrap your input in a list-of-lists like model.predict([[value1, value2]]) or pass a single-row DataFrame. This reshapes the input into the samples-by-features format sklearn requires, resolving the dimension error.
My accuracy is 99% but the model seems useless. Why?
You almost certainly have an imbalanced dataset. If one class dominates — like 5 fraud cases in 10,000 transactions — a model that always predicts the majority class scores 99%+ accuracy while catching zero fraud. Stop relying on accuracy. Print classification_report to see per-class precision and recall, and plot the confusion matrix. A 70% recall on the minority class beats 99% useless accuracy.
My results change every time I run the code. How do I fix that?
Set random_state in train_test_split, for example train_test_split(X, y, test_size=0.2, random_state=42). Without it, each run produces a different train/test split, making benchmark comparisons meaningless and results non-reproducible. Fixing random_state guarantees the same split every time so you can reliably compare models and track improvements.
// Comparisons
How does statistical ML compare to using generative AI for prediction tasks?
Statistical ML is the 'bike' — lightweight, cheap to train, fast to run, and interpretable. Generative AI is the 'car' — powerful but heavy, expensive, slower, and often overkill for tabular prediction. For a labelled numeric dataset with a clear target, a regression or classification model reaches the destination faster and cheaper. The principle: in any AI project, use each tool where it fits the situation.
How does LinearRegression compare to LogisticRegression?
LinearRegression predicts continuous numbers by fitting a best-fit line that minimises Mean Squared Error via gradient descent — use it for regression. LogisticRegression, despite its name, is a classifier: it fits a linear model then applies the sigmoid function to output a probability between 0 and 1. Use LinearRegression for continuous targets and LogisticRegression for binary or multiclass classification, especially where outliers would skew a straight line.
When should I use precision versus recall as my main metric?
Use recall when missing a true positive is costly — fraud detection, medical diagnosis, and safety systems all demand high recall because a missed case is catastrophic. Use precision when false positives are the bigger problem — spam filtering prioritises precision because wrongly flagging real emails annoys users. When you need to balance both in one number, use the F1 score, which is the harmonic mean of precision and recall.
Why start with simple models instead of the most powerful algorithm?
The framework's rule of thumb is start simple — LinearRegression or LogisticRegression — then upgrade only if performance is insufficient. Simple models are faster to train, easier to interpret, and often good enough. Jumping straight to Decision Trees, Random Forest, XGBoost, or SVM adds complexity and overfitting risk before you've established a baseline. A simple model's test-set score tells you whether upgrading is even worth it.
// Advanced
What is gradient descent and why does it matter?
Gradient descent is the core optimisation technique used during training. It iteratively adjusts model parameters — the slope (m) and intercept (b) — until the Mean Squared Error is minimised and the best-fitting line or boundary is found. You don't code it directly; model.fit runs it under the hood. Understanding it explains how the model 'learns' the logic from labelled data.
Why is Mean Squared Error used instead of Mean Absolute Error?
Mean Squared Error averages the squared differences between actual and predicted values. Squaring prevents positive and negative errors from cancelling out and penalises large errors more heavily than Mean Absolute Error does. This makes MSE especially sensitive to big mistakes, which is often desirable — you'd rather a model avoid occasional huge errors than many tiny ones. It's the default loss function in linear regression.
When should I move beyond Simple Linear Regression to Multiple Linear Regression?
Almost always. Real-world problems rarely depend on a single factor. Multiple Linear Regression uses several independent variables: y = m1*x1 + m2*x2 + ... + mn*xn + b. A delivery-time model needs distance, traffic, and weather together, not just distance. Applying only Simple Linear Regression when multiple factors affect the target is a common pitfall that underfits the real relationship.
How do I read a confusion matrix effectively?
A confusion matrix is a grid of predicted versus actual classes. Diagonal entries are correct predictions; off-diagonal entries are mistakes. Reading which off-diagonal cells are largest tells you exactly which classes the model confuses — for example, which two flower species get mixed up most. Pair it with the classification report to connect those errors to precision and recall per class before deciding the model is ready.
What does predict_proba give me that predict doesn't?
model.predict returns a hard class label; model.predict_proba returns the probability for each class. Probabilities let you rank predictions by likelihood instead of using a fixed cut-off — invaluable for prioritising outreach, like ranking insurance prospects by purchase probability. You can also tune the decision threshold to favour recall or precision depending on your business objective, rather than accepting the default 0.5 boundary.
Should I clean the test set the same way as the training set?
Apply consistent transformations, but never make feature-engineering or cleaning decisions based on the test set. The test set must stay untouched during training and decision-making to give an unbiased evaluation. Fit cleaning logic and encoders on training data, then apply them to the test data. Peeking at the test set leaks information and inflates your benchmark, making it untrustworthy for production.