Module 6 β Deep Learning with PyTorchΒΆ
Goal: Open the black box under every model in the second half of this course. Youβll write the training loop that scikit-learn hides, train neural networks honestly (validation curves, regularisation, early stopping), give categorical data learned embeddings, race your network against gradient boosting with no thumb on the scale β and wrap the result behind a scikit-learn API, ready to serve.
Estimated time: ~3.5β4.5 hours (three lessons of ~65β80 min each, plus exercises).
Prerequisites: Module 5 is essential β NB 8 (NumPy shapes), NB 17β18 (sklearn workflow, ROC-AUC, the βnever tune on testβ discipline), NB 19 (encoding & the golden rule). NB 6 (classes) for nn.Module. Nothing later than Module 5 is assumed.
scikit-learn gave you model.fit(X, y) β this module hands you the inside:
NB 20 β the machine NB 21 β the craft NB 22 β the payoff
tensors β autograd β DataLoaders Β· overfitting embeddings for categories
nn.Module β the 5-step dropout Β· weight decay the honest bake-off vs GBM
training loop Β· churn MLP early stopping Β· schedules sklearn wrapper Β· TorchScript
vs logistic baseline save/load Β· seeds Β· 4 bugs β serving
βββββββββββββββββββ one dataset throughout: the NB 17/18 SaaS churn customers ββ
Everything in this module runs 100% offline β PyTorch is a local library, not a service: no API key, no network, and every model trains on a laptop CPU in seconds (the datasets are small on purpose). Google Colab has PyTorch preinstalled, so the Colab badges just work; locally, one pip install torch (CPU wheel) is the only extra. Without torch installed, every runtime cell checks HAS_TORCH and skips gracefully β the code stays readable either way.
πΊοΈ How this relates to the Module 5 appendix track (A1βA3). The appendices are the condensed reference tour of PyTorch β βwatch the library at workβ, lighter scaffolding. This module is the full classroom treatment: slower, on the courseβs own churn data, with the complete checkpoint/exercise/debug-me apparatus. Do either first; they reinforce each other. Beyond this module, A2 (vision & sequences) and A3 (fine-tuning transformers) remain the breadth extensions for images, sequences and transfer learning β and Module 12 (DeepTab) is NB 22βs architecture, industrial-strength.
Notebooks at a glanceΒΆ
# |
Notebook |
β± Time |
Difficulty |
What youβll build |
|---|---|---|---|---|
50 |
|
~80 min |
Intermediate β Advanced |
Tensors, autograd, |
51 |
|
~70 min |
Intermediate β Advanced |
The training craft: |
52 |
|
~65 min |
Advanced |
|
Notebook guidesΒΆ
20 Β· PyTorch Fundamentals β 20_pytorch_fundamentals.ipynbΒΆ
Everything AI-shaped in this course β the LLMs of Module 8, the transformers behind BERTopic, DeepTab β is a neural network trained with gradient descent in PyTorch (or a framework on top of it). This lesson builds the four load-bearing ideas in order β tensors β autograd β nn.Module β the training loop β each verified with your own eyes (autograd is checked against calculus once, so you trust it forever). The mental model: training is a thermostat loop β forward (guess), loss (how wrong?), backward (which way is downhill?), step (nudge every knob), zero (clean slate). It ends on familiar ground: an MLP churn classifier on the NB 17/18 SaaS dataset that has to face a logistic-regression baseline β and roughly ties it, which is the honest result that frames the whole module.
Learning objectives:
Manipulate
torch.Tensors (shapes, dtypes, broadcasting, NumPy interop) and name the two things they add over NumPy (autograd + devices).Use autograd (
requires_gradβ.backward()β.grad) and explain why gradients must be zeroed every step.Write gradient descent by hand, then rebuild it idiomatically with
nn.Module,nn.MSELoss/nn.BCEWithLogitsLossandSGD/Adamβ mapping every idiom back to the by-hand version.Explain what hidden layers + ReLU buy (learned feature engineering) and read logits β sigmoid β probabilities correctly.
Train a churn MLP and compare it honestly (ROC-AUC) against
LogisticRegression.
Sections: 1 Where scikit-learn stops Β· 2 Setup & smoke test Β· 3 Tensors Β· 4 Autograd Β· 5 Gradient descent, watched live (fit a line to ad spend) Β· 6 The same fit, the PyTorch way Β· 7 From line to network β why hidden layers Β· 8 The real thing: a churn classifier (logits, BCE, the baseline scoreboard).
Practice: 4 β checkpoints Β· 5 π§ͺ practice exercises (ββββ, incl. a π debug-me on the missing zero_grad) Β· 4 π§ stretch exercises (be-the-optimizer, lr sweep, multiclass CrossEntropyLoss, logistic-regression-is-a-one-layer-net) Β· π bonus mini-project: your own fit() with train/val history Β· β
self-assessment.
Offline: synthetic data generated in-notebook (same seed as NB 17/18); HAS_TORCH guard throughout.
21 Β· Training Neural Networks Well β 21_training_neural_networks.ipynbΒΆ
βThe loss went downβ is a low bar: a network will happily memorise the training set and call it learning. This lesson is the craft of catching and preventing that. The mental model: capacity is a loan; the validation curve is the repayment schedule. A deliberately oversized network on deliberately tiny data makes overfitting visible (the unforgettable two-curve plot), then each treatment is applied and watched working: dropout + weight decay, early stopping with patience and best-weight restore, learning-rate schedules. Around that core, the professional habits: proper train/val/test three-way splits, Dataset/DataLoader batching, state_dict save/load, seeds, device-agnostic code β and a bug clinic that runs the four classic failures live (exploding/NaN loss, frozen loss, silent shape broadcasting, forgotten train()/eval() modes).
Learning objectives:
Batch with
TensorDataset+DataLoader; define step vs. epoch; say what mini-batches buy.Read overfitting off the two-curve diagnostic and treat it with dropout, weight decay and early stopping.
Replace lr-fiddling with
StepLR/ReduceLROnPlateau.Save/load via
state_dict(withmap_location), snapshot best weights, and write reproducible, device-agnostic code.Diagnose the four classic bugs from their symptoms, fast.
Sections: 1 Setup β a proper three-way split Β· 2 Mini-batches (Dataset/DataLoader) Β· 3 The two-curve diagnostic β watch a network memorise Β· 4 Dropout & weight decay Β· 5 Early stopping Β· 6 LR schedules Β· 7 Saving & loading Β· 8 Reproducibility & devices Β· 9 The bug clinic.
Practice: 4 β checkpoints Β· 4 π§ͺ practice exercises (ββββ, incl. a π debug-me: validating in train() mode with dropout firing) Β· 4 π§ stretch exercises (gradient clipping, 5-fold CV for a net, a custom Dataset, optimizer bake-off) Β· π bonus mini-project: a mini-Trainer (batching + schedule + early stopping + restore β Lightning in miniature) Β· β
self-assessment.
Offline: same churn data rebuilt in-notebook; HAS_TORCH guard throughout.
22 Β· PyTorch in Practice: Embeddings, Baselines & Serving β 22_pytorch_in_practice.ipynbΒΆ
The business lesson. First, deep learningβs genuinely new gift to tabular data: nn.Embedding turns categories into learned vectors β the notebook trains a 2-D region embedding and plots the map the model drew for itself (high-churn markets end up neighbours; nobody told it). Those embeddings join scaled numerics in a tabular network β the entity-embeddings architecture behind Module 12βs DeepTab β complete with the UNK-at-index-0 convention that survives unseen categories. Then the reckoning: an honest bake-off against logistic regression, random forest and HistGradientBoosting (spoiler: on 600 near-linear rows the thirty-second baseline wins β and the lesson is why, and what would flip it). Finally, shipping: a TorchTabularClassifier that speaks fluent sklearn (fit/predict_proba, clone-able, survives cross_val_score), the three serving artifacts (weights + scaler + vocab), and TorchScript export β with ONNX/quantization pointers into appendix A3. Closes with the moduleβs decision guide: boosting vs. embeddings vs. TabPFN vs. transfer learning, chosen by problem, not ideology.
Learning objectives:
Explain
nn.Embeddingas a trainable lookup table; pick dimensions (min(50, (card+1)//2)); handle unseen categories with the UNK convention.Build the embeddings β numerics tabular architecture.
Run a fair bake-off and state from evidence when deep learning earns its keep on tables.
Wrap a torch model behind the sklearn API so NB 18βs whole evaluation toolkit applies untouched.
Package a model for serving (three artifacts; TorchScript) and route new problems through the decision guide.
Sections: 1 The tabular reality check Β· 2 Setup β churn data with a second categorical (region) Β· 3 nn.Embedding β a lookup table that learns (the region map) Β· 4 The tabular network Β· 5 The bake-off Β· 6 A sklearn-style wrapper Β· 7 Serving β TorchScript & the three artifacts Β· 8 The decision guide.
Practice: 4 β checkpoints Β· 4 π§ͺ practice exercises (ββββ, incl. a π debug-me: the LATAM customer that crashes a vocab with no UNK slot) Β· 4 π§ stretch exercises (high-cardinality one-hot vs embedding, embeddings-as-features for simpler models, a two-head multi-task net, EmbeddingBag text classification) Β· π bonus mini-project: DeepTab-lite certified via cross_val_score, raced against boosting with error bars Β· β
self-assessment.
Offline: synthetic data in-notebook; HAS_TORCH guard throughout; TorchScript demo cleans up after itself.
The discipline this module trainsΒΆ
The loop is the whole game. Forward β loss β backward β step β zero trains everything from a 2-parameter line to an LLM; every framework is packaging around it.
Validation curves over vibes. Capacity is a loan; every dial (epochs, width, dropout, lr) is set against data the model hasnβt seen β and the test set is opened once.
Baselines are brutal, on purpose. Both bake-offs in this module end with the humble model tying or winning on small tabular data. Deep learningβs rent comes from embeddings, modalities, transfer and scale β reasons, or no network.
Models are artifacts.
state_dict+ fitted scaler + vocab travel together; reload and verify before anything ships. A PyTorch model in production is a file plus a function, subject to Module 13/12 discipline like everything else.
How these notebooks workΒΆ
Each lesson follows the course rhythm: short teaching sections punctuated by β Quick exercise (~2 min) checkpoints with collapsible solutions (executed in a fresh kernel), then a graded block β π§ͺ practice exercises (β-rated, each set with a π debug-me), π§ stretch exercises, and a π bonus mini-project β closing with key takeaways, a β self-assessment and the π next step. One dataset (the NB 17/18 SaaS churn customers) carries through all three lessons, so every new mechanism lands on familiar ground.
Where nextΒΆ
β Module 7 β Industry Applications (../07_industry_applications/) β the spine continues: churn value, fraud, segmentation and demand, with your new deep-learning judgment in your back pocket.
β Module 12 β DeepTab (../12_deeptab/) β NB 22βs entity-embedding architecture, industrial-strength (FT-Transformer, SAINT, Mamba) behind the same sklearn API you just built by hand.
β Module 5 appendices A2βA3 (../05_machine_learning/) β CNNs, RNNs, a tiny Transformer, then fine-tuning pretrained transformers: the modalities and transfer learning this module kept pointing at.
β Module 13 β Production (../13_production/) β your wrapper is an artifact; ship it like one.
π Finished this module? Test yourself with the Module 6 quiz β five questions, ~10 minutes.