01 / In the garden
Find your ingredients.
A model’s ingredients are text. Walk through the available sources and choose what belongs in your basket.
You’re collecting material.
No weights have changed.
You’re making a three-ingredient dinner without dairy: rice, beans, and tomatoes. You don’t begin with the stove. You begin by finding ingredients that can make the dish you want.
Model-building starts the same way. You gather explanations, conversations, code, or other material from relevant sources. What you pick shapes what the model will have a chance to learn. These small, handwritten samples are our teaching harvest.
Your first decision
What goes in your basket?
Your basket: empty. Nothing has been cleaned or trained yet.
Read the samples before deciding. The sink removes exact repeats and malformed text. A recipe with cheese or an unrelated forecast can still be structurally clean; choosing relevant material is your decision.
02 / At the cutting board
Turn text into workable pieces.
A model processes token IDs. See how a sentence becomes smaller units it can work with.
You’re changing the representation.
The model stays unchanged.
Put a sentence on the board. In our tiny trainer, cutting it into model inputs means encoding its UTF-8 bytes. Larger LLMs often use subword pieces instead. The knife is a way to understand tokenization, not how every tokenizer literally works.
Try the knife / real byte tokenization
A preview of the tool, not a prepared training batch. Try “café” to see one character occupy multiple bytes.
We’re inspecting tokenization before preparing the full harvest. In the actual data pipeline, clean the source text first, then encode the finished batch. That is what we’ll do at the next stop.
The pieces still need checkingNext: wash and sort →03 / At the sink
Wash and sort your batch.
Inspect the harvest. Remove exact repeats and deliberately broken samples, then tokenize what remains.
The batch is becoming usable.
Only the data changes.
Now your earlier choices matter. Did you pick the same sample twice? Did a broken one slip into the basket? Clean the text before the final encoding and inspect what survives.
At the sink / real data preparation
Wash, sort, and check your batch.
What changed? Nothing yet. Model weights unchanged.
Under the hood: prepared text & actual tokens
Choose your text, then prepare a batch.
No batch prepared yet.
This trainer uses one token per UTF-8 byte, with a vocabulary of 256. “Rice” → [82, 105, 99, 101]. Non-ASCII characters can take several bytes. Larger LLMs commonly use subword tokenizers.
Handwritten teaching samples, not a production dataset. The trainer will offer an import; it won’t start training automatically. A tiny batch shows the mechanism, not general assistant capability.
You have ingredients ready for cooking. You still don’t have a trained model.
Carry the prepared batch to the stoveNext: start cooking →04 / At the stove
Start cooking the base recipe.
Predict the next token. Compare with the text’s actual next token. Use the error to guide a weight update.
Cook trains on your batch.
Your model’s weights change here.
Chopping wasn’t cooking. Washing wasn’t cooking. Pre-training is the cooking: repeated practice develops the foundation. A generative model commonly learns by predicting the next token in a sequence.
Your batch → your weights
Turn on the heat.
Train on your batch, right here in your browser.
Prepare your ingredients at the sink first. Nothing trains automatically.
What is in the pot?
No prepared batch yet.
One transformer layer, byte tokens, existing WASM trainer. Up to 256 optimizer steps per click, with a 30-second stop. This small recipe teaches training; it does not create a general assistant.
A real optimizer repeats this cycle across many batches. Lower training loss is a signal about the objective, not a guarantee of usefulness. Set aside evaluation data to test whether the learning carries beyond the training examples.
What the processes share: repeated trials can change the recipe. The difference is that a model learns numerical weights through an objective; it isn’t a cook understanding the dish.
The foundation is cooked. Make it fit the brief.Next: taste, refine, and serve →05 / At the tasting and serving table
Refine it. Then serve it.
Worked examples and tasting feedback guide further training. A new order then uses the resulting model.
Refining can change weights.
Plating and serving do not.
You’ve developed a base recipe, but the brief is specific: three ingredients, no dairy, a short answer. Learn from a worked example, taste two versions, and revise toward that brief. That is the post-training part of our journey.
First, learn from a worked recipe.
Supervised fine-tuning continues training on demonstrations of the desired response. Like practicing a particular recipe, it teaches the result you want to serve. With LoRA, that update can be carried in adapter weights while the base stays frozen.
Illustrative checkpoints
Suggest a three-ingredient dinner without dairy. Keep the answer brief.
Authored teaching outputs. No model was trained here.
Develop the base recipe
Pre-training
Inspect the example or training objective
Raw text → predict the next token → update weights. A base model may continue text rather than follow an instruction.
Real run: Weights during pre-training. Here: illustrative output only.
Then taste two versions.
The tasting table gives you a comparison. Choosing an answer records a preference; a separate optimization step uses it to change trainable weights. Both SFT and preference optimization belong within post-training.
Your turn / become the taste tester
Which meets the brief?
Suggest a three-ingredient dinner without dairy. Keep the answer brief.
No preference collected. Model weights unchanged.
Your preference tray / collected data
Your tray is empty. Choose a response to collect a preference pair.
Illustration only / no model optimized
Pair collected Optimization step Candidate weights Held-out evaluation
DPO can optimize on preference pairs directly. RLHF can first fit a reward model and use it during reinforcement learning. This is a conceptual illustration of the separation, not a numerical simulation of either algorithm.
Finally, plate, garnish, and serve.
A finished recipe can fulfil another order without being rewritten. That is ordinary inference: fixed weights, a new prompt, a new response. Garnish is presentation—it can make an answer easier to read, but presentation alone isn’t post-training.
A last distinction / same answer, different presentation
Rice, beans, tomatoes: cook, warm, and combine. Dairy-free.
Illustrative answer, not live generation. Model weights unchanged.
Gather. Prepare. Cook. Refine. Serve. Now you know which steps change data, weights, or the answer’s presentation.
Your model is ready at the tableTake an order →Take your harvest into the real lab
Your next dish can be a real run.
Return to your prepared batch and offer it to the existing browser trainer. It demonstrates tiny next-token training from scratch, not the SFT or preference scenes above.
06 / Dinner is served
Meet the model you cooked.
Start a sentence. Your model predicts what comes next using what it learned from your batch.
Same model. Same weights.
No pretrained assistant behind the curtain.
This is a tiny byte-level text generator. Short training may produce fragments or nonsense; more practice can lower training loss without proving useful generalization. Its weights stay in this tab until you download them or leave the page.
Cook your batch first, then try the model here.
Your own model will be served here.
What is running?
The same isolated WASM worker trained at the stove generates 48 new byte tokens. No external model downloads, no system prompt, and no canned replies. Download exports the trainer’s .tinygpt format for the Web Lab. Your existing saved Web Lab model is untouched. Changing ingredients does not silently retrain or replace this model.
Read the technical sources.
The kitchen is a metaphor. These sources explain the mechanisms.