Week 2 · Oct 5–11 · planned
Generation and KV cache
Before this: Week 1 causal forward trace
Progression: Generation loop → prefill/decode → cached state → cache shapes/bytes → cached versus uncached correctness.
Reading: Learn Inference: LLM mechanics
Weekend evidence: A tested local comparison with cache arithmetic and matching relevant logits under declared tolerance.
Pass criterion: Show cache-size arithmetic and matching relevant cached versus uncached logits with fixed model, inputs, and settings; explain reused work.
Boundary: Inspect a compatible implementation first; this repo's reference cache path is not assumed. Detailed lessons remain to be authored.
Week 3 · Oct 12–18 · planned
Measurement
Before this: Week 2 generation and cache model
Progression: Timing boundaries → TTFT/inter-token latency → throughput/concurrency → memory/dtype → warmup, variance, profiling.
Reading: Learn Inference: latency and throughput
Weekend evidence: Reproducible output and a prediction-versus-observation note with pinned workload, versions, and repetitions.
Pass criterion: Pin workload, precision, versions, warmups, repetitions, and metric definitions; diagnose one misleading setup.
Boundary: Audit an existing harness first. One forward time is insufficient. Detailed lessons remain to be authored.
Week 4 · Oct 19–25 · planned
GPU execution and CUDA/Triton
Before this: Week 3 measurement definitions and reference comparisons
Progression: CPU/GPU roles → grid/block/thread/warp → access patterns → tiny CUDA kernel → introductory Triton fusion.
Reading: Wafer AI: GPU fundamentals and kernels
Weekend evidence: Reference-checked vector operation and fused softmax, including edge cases and memory-traffic explanation.
Pass criterion: Compare edge cases with a reference and explain indexing, masking, coalescing, and memory traffic; NVIDIA execution remains pending without compatible hardware.
Boundary: NVIDIA execution requires separately approved compatible hardware, pinned tooling, and cost cap. Mac/source practice cannot pass that gate. Detailed lessons and host verification remain.
Week 5 · Oct 26–Nov 1 · planned
Attention performance
Before this: Week 4 GPU execution model and correctness checks
Progression: Materialized attention → IO/tiling → numerical stability → optimized path → controlled comparison.
Weekend evidence: Reference-versus-optimized comparison with shapes, dtypes, correctness tolerance, timing, and transfer limits.
Pass criterion: Explain correctness, measured bottleneck, speed/space tradeoff, and where the result does not transfer.
Boundary: Mac measurements do not establish NVIDIA performance. Detailed lessons and compatible path verification remain.
Week 6 · Nov 2–8 · planned
Trace vLLM
Before this: Weeks 2–5 inference, measurement, and attention concepts
Progression: Engine overview → request entry → scheduling → model runner → output/KV lifecycle.
Reading: Wafer AI: inference engines
Weekend evidence: Source-linked request trace at a pinned revision and a hand-simulated request.
Pass criterion: Trace a request end to end at a pinned source revision; identify state ownership and where time can accumulate.
Boundary: Mac source review is code-understanding evidence, not a vLLM runtime result. Detailed lessons and source revision remain to be pinned.
Week 7 · Nov 9–15 · planned
Scheduling and KV ownership
Before this: Week 6 request-to-output source trace
Progression: Arrivals → token budgets → KV allocation/release → batching/prefill scheduling → fairness/failure.
Weekend evidence: Three-request simulation and one investigated scheduling choice with allocation invariants.
Pass criterion: Check KV allocation invariants for three requests and predict fairness and latency/throughput tradeoffs.
Boundary: Simulation is not a production benchmark. Detailed lessons remain to be authored.
Week 8 · Nov 16–22 · planned
Bounded runtime change
Before this: Week 7 scheduling trace and an executable baseline
Progression: Choose issue → reproduce → test → patch → before/after → review.
Weekend evidence: A reviewable patch, correctness tests, controlled measurements, and a regression-risk explanation.
Pass criterion: Reproduce the issue, pass correctness tests, compare controlled before/after evidence, and name a regression risk.
Boundary: Requires an executable baseline; no upstream publication is implied. Detailed lessons depend on the chosen issue.
Week 9 · Nov 23–29 · planned
Serving behavior
Before this: Weeks 3 and 8 controlled measurement and bounded runtime change
Progression: Request mix → service targets → concurrency/queueing → tail latency → overload tradeoffs.
Weekend evidence: Bounded serving-load report separating queue and model time, plus changed-mix diagnosis.
Pass criterion: Report latency distribution, throughput, queue versus model time, and the cost of an intervention under a changed request mix.
Boundary: Label local simulation separately from runtime measurement. Detailed lessons and environment remain to be specified.
Week 10 · Nov 30–Dec 6 · planned
Single device to cluster
Before this: Week 9 serving limits and Week 2 memory arithmetic
Progression: Capacity → sharding → communication volume → topology → design comparison.
Reading: Wafer AI: distributed inference
Weekend evidence: Declared model/workload memory and communication calculation with topology-aware assumptions.
Pass criterion: Defend declared memory, communication, and topology assumptions; separate calculation from execution.
Boundary: Multi-GPU reasoning is required; a multi-GPU benchmark is not. Detailed lessons remain to be authored.
Week 11 · Dec 7–13 · planned
Capstone investigation
Before this: One bounded question supported by Weeks 1–10 evidence
Progression: One hypothesis → frozen workload → controls → bounded intervention → predicted failure.
Weekend evidence: Raw evidence, code, config, stop rule, and competing explanations; a supported negative result counts.
Pass criterion: Preserve raw evidence, code, config, controls, stop rule, and competing explanations, including a supported negative result.
Boundary: Use one investigation from prior work, not a new project. Detailed lessons depend on the selected hypothesis.
Week 12 · Dec 14–20 · planned
Capstone review
Before this: Week 11 frozen capstone workload and raw evidence
Progression: Reproduce → vary workload → find regressions → challenge mechanism → conclude.
Weekend evidence: Write-up with reproduction steps, limitations, uncertainty, changed-workload evidence, and ship/reject/revise decision.
Pass criterion: Reproduce and vary workload, identify regressions and uncertainty, then defend ship/reject/revise from evidence.
Boundary: One favorable run is not a universal speed claim. Detailed lessons remain to be authored.
Week 13 · Dec 21–27 · planned
Unfamiliar-problem assessment
Before this: Weeks 1–12 traces, measurements, and reviewed limitations
Progression: Tensor/correctness → cache/memory → measurement trap → serving diagnosis → review.
Weekend evidence: Predict, diagnose, interpret evidence, and defend the next test on unfamiliar inputs.
Pass criterion: On unfamiliar inputs, predict, choose a diagnostic, interpret evidence, and defend the next test without claiming schedule-based mastery.
Boundary: Finishing the schedule does not itself establish professional seniority. Detailed assessment prompts remain to be authored.