Frequently Asked Questions About Simplilearn AI Capstone Project Architect

22 answers covering everything from basics to advanced usage.

// Basics

What is transfer learning and why does this methodology require it?

Transfer learning takes a pretrained model like VGG16 or ResNet, removes its top classification layers with include_top=False, inserts custom Dense layers matching your classes, and retrains on your data. The methodology requires it because pretrained architectures already encode general visual features and consistently outperform custom CNNs built from scratch on standard image tasks — saving both training time and accuracy.

What is the Two-Part Project Structure?

Every capstone track has exactly two parts: Part 1 is always a modelling or detection task, and Part 2 is always an analysis, recommendation, or forecasting task. For example, Track 1 pairs YOLO detection with accident-data analysis. You must complete both parts of your chosen track and should never conflate them when scoping or sequencing work.

What does the Generate-Then-Aggregate Sales Column pattern mean?

It's a two-step process for Track 3. First create a row-level sales column as unit_price × item_count. Then group by date and sum that column to produce a daily total sales series. You never model raw price or item_count alone — the derived total sales figure is the actual forecasting target, and all resampling flows from this aggregated daily series.

What is YOLO and how is it used in Track 1?

YOLO (You Only Look Once) is a real-time object detection model used via a cloned repository. It outputs bounding box coordinates plus class labels for every detected object. In Track 1 you verify data.yaml points to your images and labels folders, update the class names, run train.py with GPU enabled, then run detect.py on test images and visualize the boxes.

// How To

How do I set up data.yaml correctly for YOLO training?

Point data.yaml to the correct local or Colab directory paths for your training and validation image folders, set the number of classes, and update the 'names' field to match your exact vehicle class labels. Verify these paths before running train.py — a misconfigured path causes the script to fail silently or train on the wrong data.

How do I run the augmentation with-and-without comparison?

Train the same architecture twice. First train it plain with no augmentation and record validation accuracy. Then prepend a Sequential block of RandomFlip, RandomRotation, and RandomZoom to the same architecture, retrain, and record validation accuracy again. Compare the two curves and explicitly state which run generalized better and why — augmentation should reduce overfitting.

How do I build the item-based collaborative filter in Track 2?

Use tourism_rating (user_id, place_id, rating) as your primary data, then merge with tourism_with_id on place_id to add place names and city metadata. Build the model so the input is a location and the output is a ranked list of similar locations sharing rating patterns across users. Finally show top-N recommendations for a sample input location.

How do I extract time features for the regression models?

Convert your date column to datetime with pd.to_datetime(), then use the .dt accessor to extract day_of_week, month, quarter, year, and day_of_month. These become the features that Linear Regression, Random Forest, and XGBoost use to predict the daily sales column. Sort by date before splitting so the last 6 months form your test set.

How do I fill missing values in accident data?

Fill nulls in numeric count columns — like deaths, occupants, cyclists, or other_vehicle — with zero. In accident and event data, a missing value means no recorded event of that type, so zero is semantically correct. This also prevents nulls from breaking downstream aggregations like value_counts() and groupby() summaries.

// Troubleshooting

Why is my custom output layer throwing a shape mismatch?

Two common causes: you forgot include_top=False, so VGG16's original head is still attached, or your final Dense layer's neuron count doesn't equal your exact number of output classes. Set include_top=False, then make the last Dense layer's neuron count match your class count precisely with a softmax activation.

My YOLO training seems to do nothing or uses wrong images — what's wrong?

Your data.yaml almost certainly points to the wrong directory. YOLO training scripts fail silently or use incorrect data when paths are misconfigured. Re-verify that data.yaml references your actual images and labels folders, that the class count is correct, and that GPU is enabled in the runtime.

Why is my forecasting model performing suspiciously well?

You likely shuffled or randomly split the time-series data, which leaks future information into training. Sort by date ascending and hold out only the last 6 months as the test set. Also confirm you're predicting the derived total sales column, not raw price or item_count, and that your date column is proper datetime.

My three-way DataFrame merge is dropping rows or throwing errors — how do I fix it?

Don't merge three DataFrames in one call. Use the Merge-in-Two-Steps pattern: merge two DataFrames on a shared key like store_id, capture the result, then merge that result with the third on another shared key like item_id. Confirm each merge key exists in both DataFrames being joined at that step.

// Comparisons

How does this methodology compare to a generic Kaggle notebook approach?

Generic notebooks often skip disciplined scoping and mix concerns. This methodology enforces a two-part structure per track, mandates transfer learning over scratch builds, prescribes time-ordered splits, two-step merges, and null-filling rules, and requires explicit model comparison with the right metric. That structure prevents the silent errors — data leakage, shape mismatches, broken aggregations — that ad-hoc notebooks routinely introduce.

Should I use accuracy or RMSE to evaluate my model?

Use accuracy for classification tasks (Track 1 detection outputs and Track 2 image classification) and RMSE for regression tasks (Track 3 sales forecasting). For Track 3, evaluate all three regression models with RMSE and select the lowest as champion. For Track 2 Part 1, compare validation accuracy across the augmented and non-augmented runs.

Is item-based collaborative filtering better than a content-based recommender here?

For Track 2, item-based collaborative filtering is preferred because the ratings CSV (user_id, place_id, rating) is the primary data source and captures collective preference patterns. Content-based approaches rely on item metadata as the primary signal, which is secondary here. Since the input is a location and the goal is finding similarly-rated locations, item-based collaborative filtering fits the problem directly.

Random Forest vs XGBoost vs Linear Regression — which wins for sales forecasting?

There's no fixed winner — that's why the methodology trains all three in parallel and compares them by RMSE. Linear Regression is a fast baseline; Random Forest and XGBoost capture nonlinear patterns better on tabular time features. Select whichever produces the lowest RMSE on your last-6-months test set as the champion for the next-year forecast.

// Advanced

Can I combine two tracks in a single project?

The methodology says complete exactly one track but both of its parts. Combining tracks isn't the intended design and dilutes scope. If your real project genuinely needs both detection and forecasting, treat them as separate applications of the methodology rather than merging them, since each track has its own dataset audit, modelling, and evaluation requirements.

How should I detect and handle outliers in the ratings or sales data?

For Track 2 ratings, use z-score to identify ratings far from the mean and drop those rows, along with nulls and duplicates. For Track 3 sales, check for outliers using z-score or Isolation Forest before modelling. Handle these during the data-type and missing-value step so aggregations and model training aren't skewed by anomalies.

How do I produce multi-scale trend views of sales data?

After building your datetime-indexed daily_sales DataFrame, use pandas resample: .resample('W') for weekly, .resample('M') for monthly, and .resample('Q') for quarterly aggregations. Plot each view to reveal trends at different scales. You can also group by restaurant_id to rank stores and group by item_id plus store_id to find the most popular item per store.

What reference notebooks support each track?

Track 1 references the Deep Learning Lesson 10 notebook (object detection with YOLO). Track 2 references Deep Learning Lesson 9 (transfer learning, section 9.04) and Lesson 8.08 (image loading from directories), plus Machine Learning Lesson 7 for recommendation systems. Track 3 relies on standard pandas merging, resampling, and scikit-learn/XGBoost regression workflows.

How do I document and validate model performance at the end?

Report the metric appropriate to your task — accuracy for classification, RMSE for regression. For Track 3, present a summary table comparing all three regression models before declaring the champion. For Track 2 Part 1, explicitly compare the with- and without-augmentation runs and state which generalized better and why. Always pair numbers with the required visualizations.