Loading saved session · W1 · D1

Learn it here.
Prove it to yourself.

Read the lesson, work one exercise, explain it in your own words, and return for two short recall checks. AI chat is optional. Your answers stay in this browser.

Delayed recall

Return without rereading.

Due reviews unlock on their date. Write the answer before checking your notes.

Loading recall schedule…

Evidence

Checkpoint history

Loading checkpoints…

Inference systems · 28 Sep–27 Dec 2026

The 13-week route

Week 1 has seven complete lessons. Weeks 2–13 have scoped contracts; their daily lessons are planned.

Week 1 · Day 1 · 120 min · ready

From token IDs to next-token scores

Predict four tensor shapes and explain why this toy's last-position output ignores earlier tokens.

Week 1 · Day 2 · 120 min · ready

Token bytes, tensor axes, and position

Trace byte IDs through lookup, learned position addition, and a batched projection.

Week 1 · Day 3 · 120 min · ready

Single-head attention from three tokens

Compute one attention row and explain why its key weights sum to one.

Week 1 · Day 4 · 120 min · ready

Causal masks and multiple heads

Mask future keys, verify causal isolation, and trace multi-head shapes.

Week 1 · Day 5 · 120 min · ready

A transformer block and output head

Annotate one reference block from token vectors to logits and locate where loss would enter.

Week 1 · Day 6 · 300 min · ready

Trace the reference decoder on CPU

Trace token IDs to logits in python_ref/model.py and retain one reproducible correctness baseline.

Week 1 · Day 7 · 300 min · ready

Reconstruct, debug, and defend the baseline

Reconstruct the forward path closed-book, fix one causal bug, and state what the Week 1 evidence does and does not show.

View Weeks 2–13 contracts · detailed lessons planned

Week 2 · Oct 5–11 · planned

Generation and KV cache

Before this: Week 1 causal forward trace

Progression: Generation loop → prefill/decode → cached state → cache shapes/bytes → cached versus uncached correctness.

Reading: Learn Inference: LLM mechanics

Weekend evidence: A tested local comparison with cache arithmetic and matching relevant logits under declared tolerance.

Pass criterion: Show cache-size arithmetic and matching relevant cached versus uncached logits with fixed model, inputs, and settings; explain reused work.

Boundary: Inspect a compatible implementation first; this repo's reference cache path is not assumed. Detailed lessons remain to be authored.

Week 3 · Oct 12–18 · planned

Measurement

Before this: Week 2 generation and cache model

Progression: Timing boundaries → TTFT/inter-token latency → throughput/concurrency → memory/dtype → warmup, variance, profiling.

Reading: Learn Inference: latency and throughput

Weekend evidence: Reproducible output and a prediction-versus-observation note with pinned workload, versions, and repetitions.

Pass criterion: Pin workload, precision, versions, warmups, repetitions, and metric definitions; diagnose one misleading setup.

Boundary: Audit an existing harness first. One forward time is insufficient. Detailed lessons remain to be authored.

Week 4 · Oct 19–25 · planned

GPU execution and CUDA/Triton

Before this: Week 3 measurement definitions and reference comparisons

Progression: CPU/GPU roles → grid/block/thread/warp → access patterns → tiny CUDA kernel → introductory Triton fusion.

Reading: Wafer AI: GPU fundamentals and kernels

Weekend evidence: Reference-checked vector operation and fused softmax, including edge cases and memory-traffic explanation.

Pass criterion: Compare edge cases with a reference and explain indexing, masking, coalescing, and memory traffic; NVIDIA execution remains pending without compatible hardware.

Boundary: NVIDIA execution requires separately approved compatible hardware, pinned tooling, and cost cap. Mac/source practice cannot pass that gate. Detailed lessons and host verification remain.

Week 5 · Oct 26–Nov 1 · planned

Attention performance

Before this: Week 4 GPU execution model and correctness checks

Progression: Materialized attention → IO/tiling → numerical stability → optimized path → controlled comparison.

Weekend evidence: Reference-versus-optimized comparison with shapes, dtypes, correctness tolerance, timing, and transfer limits.

Pass criterion: Explain correctness, measured bottleneck, speed/space tradeoff, and where the result does not transfer.

Boundary: Mac measurements do not establish NVIDIA performance. Detailed lessons and compatible path verification remain.

Week 6 · Nov 2–8 · planned

Trace vLLM

Before this: Weeks 2–5 inference, measurement, and attention concepts

Progression: Engine overview → request entry → scheduling → model runner → output/KV lifecycle.

Reading: Wafer AI: inference engines

Weekend evidence: Source-linked request trace at a pinned revision and a hand-simulated request.

Pass criterion: Trace a request end to end at a pinned source revision; identify state ownership and where time can accumulate.

Boundary: Mac source review is code-understanding evidence, not a vLLM runtime result. Detailed lessons and source revision remain to be pinned.

Week 7 · Nov 9–15 · planned

Scheduling and KV ownership

Before this: Week 6 request-to-output source trace

Progression: Arrivals → token budgets → KV allocation/release → batching/prefill scheduling → fairness/failure.

Weekend evidence: Three-request simulation and one investigated scheduling choice with allocation invariants.

Pass criterion: Check KV allocation invariants for three requests and predict fairness and latency/throughput tradeoffs.

Boundary: Simulation is not a production benchmark. Detailed lessons remain to be authored.

Week 8 · Nov 16–22 · planned

Bounded runtime change

Before this: Week 7 scheduling trace and an executable baseline

Progression: Choose issue → reproduce → test → patch → before/after → review.

Weekend evidence: A reviewable patch, correctness tests, controlled measurements, and a regression-risk explanation.

Pass criterion: Reproduce the issue, pass correctness tests, compare controlled before/after evidence, and name a regression risk.

Boundary: Requires an executable baseline; no upstream publication is implied. Detailed lessons depend on the chosen issue.

Week 9 · Nov 23–29 · planned

Serving behavior

Before this: Weeks 3 and 8 controlled measurement and bounded runtime change

Progression: Request mix → service targets → concurrency/queueing → tail latency → overload tradeoffs.

Weekend evidence: Bounded serving-load report separating queue and model time, plus changed-mix diagnosis.

Pass criterion: Report latency distribution, throughput, queue versus model time, and the cost of an intervention under a changed request mix.

Boundary: Label local simulation separately from runtime measurement. Detailed lessons and environment remain to be specified.

Week 10 · Nov 30–Dec 6 · planned

Single device to cluster

Before this: Week 9 serving limits and Week 2 memory arithmetic

Progression: Capacity → sharding → communication volume → topology → design comparison.

Reading: Wafer AI: distributed inference

Weekend evidence: Declared model/workload memory and communication calculation with topology-aware assumptions.

Pass criterion: Defend declared memory, communication, and topology assumptions; separate calculation from execution.

Boundary: Multi-GPU reasoning is required; a multi-GPU benchmark is not. Detailed lessons remain to be authored.

Week 11 · Dec 7–13 · planned

Capstone investigation

Before this: One bounded question supported by Weeks 1–10 evidence

Progression: One hypothesis → frozen workload → controls → bounded intervention → predicted failure.

Weekend evidence: Raw evidence, code, config, stop rule, and competing explanations; a supported negative result counts.

Pass criterion: Preserve raw evidence, code, config, controls, stop rule, and competing explanations, including a supported negative result.

Boundary: Use one investigation from prior work, not a new project. Detailed lessons depend on the selected hypothesis.

Week 12 · Dec 14–20 · planned

Capstone review

Before this: Week 11 frozen capstone workload and raw evidence

Progression: Reproduce → vary workload → find regressions → challenge mechanism → conclude.

Weekend evidence: Write-up with reproduction steps, limitations, uncertainty, changed-workload evidence, and ship/reject/revise decision.

Pass criterion: Reproduce and vary workload, identify regressions and uncertainty, then defend ship/reject/revise from evidence.

Boundary: One favorable run is not a universal speed claim. Detailed lessons remain to be authored.

Week 13 · Dec 21–27 · planned

Unfamiliar-problem assessment

Before this: Weeks 1–12 traces, measurements, and reviewed limitations

Progression: Tensor/correctness → cache/memory → measurement trap → serving diagnosis → review.

Weekend evidence: Predict, diagnose, interpret evidence, and defend the next test on unfamiliar inputs.

Pass criterion: On unfamiliar inputs, predict, choose a diagnostic, interpret evidence, and defend the next test without claiming schedule-based mastery.

Boundary: Finishing the schedule does not itself establish professional seniority. Detailed assessment prompts remain to be authored.

Public course · indexable

Browse the complete learning spine

10 modules · choose any module as your current session
01

Functions, data, and parameters

A model is a function whose adjustable parameters are learned from examples.

A function maps an input to an output. In y = mx + b, x is the input and the predicted y is the output.

The examples stay fixed while fitting. The parameters m and b change because they control the function's behavior.

Learning means choosing parameter values that make predictions agree with the observed targets. The same distinction scales from a two-parameter line to a transformer with billions of parameters.

Exercise: For (x, y) = (0,1), (1,3), (2,5), (3,7), (4,9), choose m and b for y = mx + b. Show all five predictions and label inputs, targets, parameters, and predictions.

Mastery gate: Explain a parameter without LLM jargon and predict the effect of changing m versus b on a new input.

Read the complete lesson →
02

Loss and gradient descent

Loss measures error; gradients indicate how parameter changes affect it.

A loss function compresses many prediction errors into one number that can guide training.

A gradient is a local slope. Gradient descent moves parameters against that slope, scaled by the learning rate.

Exercise: Compute MSE for two line fits, then take one gradient-descent step by hand.

Mastery gate: Predict the symptoms of a learning rate that is too high or too low.

Read the complete lesson →
03

Vectors, matrices, and tensors

Neural networks organize multiply-and-add operations over arrays with explicit shapes.

Vectors hold features, matrices transform them, and tensors generalize these arrays to more axes.

Shape annotations expose which dimensions represent batch, sequence, embedding, heads, and vocabulary.

Exercise: Trace the shapes through one linear layer and diagnose one deliberately mismatched axis.

Mastery gate: Read a shape error and identify which axis is wrong.

Read the complete lesson →
04

Non-linear networks and backpropagation

Activations make depth expressive; backpropagation assigns credit through the computation graph.

Stacking only linear transformations still produces one linear transformation. Non-linear activations let networks fit curved decision boundaries.

Backpropagation applies the chain rule from the loss back to every parameter that influenced it.

Exercise: Inspect or train a two-layer MLP on a non-linear toy dataset and explain the gradient path.

Mastery gate: Explain why an activation enables behavior that stacked linear layers cannot.

Read the complete lesson →
05

ML paradigms and scaling

Pretraining, supervision, preferences, reinforcement, and scaling solve different problems.

Self-supervision builds broad predictive capability; supervised and preference methods shape behavior for particular tasks.

More scale can add capacity and knowledge, but cannot repair invalid data or an unreliable evaluation ruler.

Exercise: Classify retained posttrainllm attempts by learning paradigm and name the signal each used.

Mastery gate: Explain why scale cannot rescue a broken evaluation or contaminated dataset.

Read the complete lesson →
06

Tokenization, embeddings, and language modeling

Text becomes token IDs, embeddings, contextual states, and next-token probabilities.

A tokenizer chooses the discrete units a model sees. Embeddings map each token ID into a learned vector.

A causal language model learns to predict the next token from the tokens before it.

Exercise: Tokenize three prompts and inspect how punctuation and identifiers split.

Mastery gate: Explain how tokenization can affect SQL and tool-call reliability.

Read the complete lesson →
07

Attention and transformer blocks

Attention routes information between positions; repeated blocks refine contextual representations.

Queries express what a position seeks, keys describe what positions offer, and values carry the information to combine.

A transformer block combines attention, an MLP, normalization, residual paths, and positional information.

Exercise: Work one tiny query/key/value attention calculation with explicit shapes.

Mastery gate: Describe what attention can copy or route that a position-wise MLP cannot.

Read the complete lesson →
08

Training mechanics

Batches, optimizers, schedules, precision, and correctness gates determine whether training works.

A loss curve only becomes useful when paired with data checks, held-out behavior, gradient health, and reproducible settings.

Tiny-overfit is the first gate: if a model cannot memorize a tiny repeated set, scaling the run hides the bug rather than fixing it.

Exercise: Inspect a tiny-overfit run and diagnose data, learning-rate, capacity, and precision failures separately.

Mastery gate: Distinguish a data bug, optimization bug, and capacity limit from their symptoms.

Read the complete lesson →
09

SFT, LoRA, and preference tuning

Post-training changes behavior using labeled demonstrations, compact updates, or preferences.

SFT imitates desired outputs. LoRA learns a low-rank update while keeping the base weights frozen.

Preference methods compare better and worse responses; weak pair construction or missing reference control can collapse useful behavior.

Exercise: Compare the successful SQL SFT run with the failed hygiene preference run and isolate the changed variables.

Mastery gate: Design a bounded recipe with data, target, regression slices, budget, and stop rule.

Read the complete lesson →
10

Evals, rewards, and self-improvement

A frozen ruler, verifiable feedback, and retained failures make improvement measurable.

An evaluation must first pass a frontier-ceiling check: if a capable model cannot satisfy the ruler, the ruler is measuring noise.

Held-out data, regression slices, provenance, and cost measurements turn a score into a decision.

Exercise: Draft a target, frozen baseline, held-out gate, regressions, resource budget, and ship/retry/reject rule.

Mastery gate: Reject an attractive score when its ruler, holdout, provenance, or regression gate is invalid.

Read the complete lesson →
Applied reading

Open the industry case studies

5 cases · each turned into a practical exercise

Savante / Aryabhata: specialist data and evaluation

Study how curated domain data, filtered reasoning traces, SFT, and verifiable-reward RL combine in a JEE mathematics specialist. The published result does not isolate the contribution of each stage.

Exercise: Write a recipe teardown separating initialization, data, filtering, SFT, reward, evaluation, missing ablations, and leakage boundaries.

Open the case study and sources →

Bonsai 2 27B: capability retained per deployment cost

Distinguish ternary representation, packed artifact size, peak runtime memory, kernel support, latency, and capability retention relative to the source model.

Exercise: Draft a same-Mac comparison sheet covering exact revisions, task gates, peak RAM, TTFT, prefill, decode throughput, and total task latency.

Open the case study and sources →

QORL: parameter-aware query optimization

Compare native PostgreSQL planning, structured hint search, and LLM proposals under equal budgets, then evaluate a parameter-to-plan dispatcher on fresh held-out measurements.

Exercise: Freeze parameter splits, cache policy, timeouts, SQL-equivalence checks, unfamiliar-case fallback, routing overhead, and tuning-cost break-even.

Open the retained PRD and sources →

Inside vLLM: inference systems and scaling boundaries

Trace admission, scheduling, paged KV allocation, prefill, decode, batching, caching, and the boundary from one device to distributed serving.

Exercise: Hand-simulate three requests under fixed token and KV-block budgets, then predict effects on TTFT, inter-token latency, throughput, and memory.

Open the article study guide →

Splash: model-specific Mac inference

Study how shape-specific Metal kernels, a model-specific speculative draft, packed weights, cache reuse, and a startup memory plan trade generic model support for Mac-local serving performance.

Exercise: Design a same-Mac comparison with a general engine covering cold and cached TTFT, prefill, decode, concurrency, memory, output validity, task completion, and per-model tuning cost.

Open the launch analysis and source guide →