FDEInterviews logo

Train an LLM From Scratch: What We Learned Running train-llm-from-scratch on a CPU in Two Minutes

You can train a small language model from raw text on an ordinary CPU in about two minutes. We ran the open-source train-llm-from-scratch repo, recorded the numbers, and explain each stage (tokenizer, pretraining, SFT, reward model, DPO, PPO, GRPO), the gotchas, and why an FDE should do it once.

BY HANNAH BRYANT · FDEINTERVIEWS EDITORIAL · UPDATED OCTOBER 7, 2026 · 6 MIN READ

PRACTICE THIS:LLM and GenAI questions ·Fine-tuning vs RAG vs prompting, the concept ·PPO and GRPO, the concept ·The must-know FDE questions

Yes, you can train a language model from scratch on an ordinary computer: on 7 October 2026 we trained a 369,024-parameter model on an 8-core desktop CPU in 111 seconds, from raw text to generated stories, using the open-source train-llm-from-scratch repository. It reached a dev loss of 2.78 and wrote sentences like "Once upon a time, there was a little girl named Sue. Sue loved to look on." That is the right size for learning, not for using: the point is to see every stage of how a model is made, in plain PyTorch you can read. The repository (MIT licensed, about 12,000 GitHub stars as of 7 October 2026, by Fareed Khan) now goes well past pretraining, through supervised fine-tuning, a reward model, DPO, PPO and GRPO. This post covers what we ran, what each stage teaches, the gotchas, and why a forward deployed engineer should do it once.

What the repository contains

It implements a Transformer by hand, based on "Attention Is All You Need", and then a whole training pipeline on top of it, with no trl, peft or transformers libraries. Every algorithm is a short PyTorch file you can open. The author's framing is the right one: turn text into numbers, predict the next token, then keep changing the data and the loss until the model does what you want.

Pretraining Raw textTinyStories, Pile TokenizerBPE, ids Transformerattention + MLP Next-token losscross-entropy Base modelcontinues text Post-training, same backbone SFTloss on assistanttokens only Reward modelBradley-Terry onpreference pairs PPO or DPORL loop, or learnfrom pairs GRPOgroup-relative,verifiable reward Evaluategreedy GSM8Kaccuracy Only two things change between stages: the data, and the loss. The model class never does.

There are two model architectures. The classic one is the 2017 design in its GPT-2 form, best for learning. The modern one swaps in what newer open models use: rotary position embeddings (RoPE), RMSNorm, SwiGLU feed-forward layers, grouped-query attention, a KV cache for generation, and optional mixture of experts. Any script takes --arch modern. A design choice worth copying in your own code is what the author calls "wrap, do not rewrite": the base Transformer gains one method that returns hidden states before the output layer, and the reward head, value head and log-probability math all compose around it.

What we ran, and the numbers

We followed the laptop track at commit bdd480a, forcing CPU only (no GPU visible), on an AMD Ryzen 7 9700X with 8 cores and 16 threads.

StepWhat happenedOur measurement
Prepare dataDownloads TinyStories, trains a 4,096-token BPE tokenizer from scratch, encodes7.6 s; 6,095,088 training tokens, 991,095 validation tokens; 4.02 characters per token
Train, modern architecture--preset tiny --arch modern369,024 parameters; 111 s; dev loss 2.78; about 65,000 tokens per second
Train, classic architecture--preset tiny --arch classic636,288 parameters; 116 s; dev loss 3.38
Dev loss by training step, tiny preset, CPU only (our run, 7 Oct 2026) 2 4 6 8 0 300 600 900 1200 uniform guess: ln(4096) = 8.32 classic, 636K params: 3.40 modern, 369K params: 2.81 training step (last logged eval; final dev loss after the run: modern 2.78, classic 3.38)

Three things stand out. First, the modern architecture has 42 percent fewer parameters and still reaches a lower loss (2.78 against 3.38). The README reports 2.78 against 3.30, so the modern number matched ours exactly and the classic one came out slightly worse on our run. Second, both runs start at a loss between 8.2 and 8.4, close to ln(4,096) = 8.32, the loss of a model that guesses every token uniformly. That is the most useful sanity check in machine learning: if your starting loss is far from ln(vocabulary size), something is wrong before training even begins. Third, most of the learning happens early. The modern model's dev loss halves in the first 150 steps and spends the rest of the run grinding out the last 1.3.

The output reads like a model that has learned the statistics of children's stories and nothing else. Ours wrote "Sue saw a big, red ball and asked", with grammar, recurring characters and a vague plot, and no world knowledge at all.

What each stage teaches

Tokenization. Training your own 4,096-token BPE tokenizer shows why tokens, not words, are the unit of cost and context. A small vocabulary keeps the embedding table small, which is what makes the tiny model fast. The bigger presets switch to the 50,304-token GPT-2 style tokenizer, where uniform-guess loss is 10.83. The README's 77-million-parameter run on two L40 GPUs starts at 11.14 and reaches about 3.73 train and 3.76 dev after 2,000 steps.

Supervised fine-tuning and the loss mask. SFT is still next-token prediction, with one change: the loss counts only the assistant's tokens. The repository builds a 0 or 1 mask alongside the token ids, so the model learns to write answers rather than to repeat prompts. After SFT, the model answers inside the <think> and <answer> structure it was trained on, which is the shape the later stages optimise.

The reward model. A scalar head on the SFT backbone, trained with the Bradley-Terry loss so the chosen answer scores higher than the rejected one. On 7,974 real preference pairs it reached 0.574 held-out preference accuracy in the author's run, above the 0.5 chance line and well short of useful. That small margin is the honest picture of a small reward model, and the reason reward models are where RLHF quality is won or lost.

DPO, PPO and GRPO. DPO skips the reward model and learns directly from preference pairs against a frozen reference copy. PPO is the classic RLHF loop with a value network and a clipped update. GRPO, the method popularised by DeepSeek-R1, removes the value network: it samples a group of answers per prompt and scores each against its own group. The core of it fits in five lines from the repository:

def group_advantages(rewards, group_size, eps=1e-4):
    r = rewards.view(-1, group_size)
    mean = r.mean(dim=1, keepdim=True)
    std = r.std(dim=1, keepdim=True)
    adv = (r - mean) / (std + eps)     # how much better than the rest of my group
    return adv.reshape(-1)

The repository also implements the 2025 corrections to GRPO (Dr. GRPO, DAPO, GSPO) as flags on the same trainer. It runs a short arithmetic curriculum first, because a tiny model on full GSM8K math problems gets zero reward and so learns nothing.

Gotchas we hit, and ones to expect

  • Install it the documented way. We ran the scripts against an existing PyTorch environment instead of pip install -e . and hit missing imports one at a time (tiktoken, then jaxtyping). The editable install declares them; use it, or uv sync.
  • There are two config systems. config/config.py drives the original pretraining script, while JSON files under configs/ drive everything else. Editing the wrong one changes nothing, without an error.
  • The GPU table is a rough guide. The README lists what each GPU can train (a free T4 handles the 13-million-parameter model, an 8 GB card tops out around 1 billion parameters, and the 2-billion-parameter configuration needs 24 GB or more). Memory savers such as --amp, --grad-checkpointing and --grad-accum are opt-in, not on by default.
  • Small models do not reason. The pipeline runs GRPO end to end, but a few-million-parameter base will score near zero on GSM8K. The stages are there to learn the mechanics, and the README says scores rise with base model size and pretraining compute.
  • Mind the data license for the large path. The pretraining path uses an uncopyrighted subset of The Pile. If you swap in a customer's or a scraped corpus, the licensing question becomes yours.

The FDE lens: why do this if you deploy models, not train them

Forward deployed engineers rarely pretrain anything. Customers do ask, often in the first meeting, whether they should train their own model. The answer is almost never from scratch, and the order to try is prompting, then retrieval, then fine-tuning an open model, as laid out in the fine-tuning vs RAG vs prompting concept. An engineer who has trained a toy model says that with authority instead of from a slide, and can explain why: the cost scales with tokens and parameters, and the customer's data rarely has the volume a base model needs.

The hands-on run also pays off in debugging. When a customer's fine-tune "forgot how to follow instructions", you will think of the loss mask and the chat template before you blame the model. When a reward model looks fine on paper and the RL run degrades, you will remember what 0.574 preference accuracy feels like. And when a bill arrives per token, you will have watched 24 million characters become 6 million tokens on your own machine.

It also covers a block of interview questions directly. What loss should training start at, and why? Why mask the prompt during SFT? What does GRPO remove from PPO, and what does DPO remove from both? How does a KV cache change generation cost? The transformer architecture and PPO and GRPO concepts go deeper, and the LLM and GenAI bank has the questions to practise. If you have an afternoon, run the laptop track first; it takes less time than reading this post twice.

THE ONE-PAGE VERSION
Infographic on training a small language model from scratch: tokenize text into ids because tokens are the unit of cost and context, pretrain by predicting the next token with a starting loss near the log of the vocabulary size, post-train with a loss on the answer only and then preference methods, and evaluate every stage on one fixed benchmark.
↧ DownloadShare on X ↗Share on LinkedIn ↗
PRACTICE THIS

Turn it into offers. Work the real questions and concepts this maps to:

FAQ

Can you train an LLM from scratch on a laptop?▲

A small one, yes. With the train-llm-from-scratch repository's tiny preset, we trained a 369,024-parameter model on TinyStories using only an 8-core desktop CPU in 111 seconds, reaching a dev loss of 2.78, and it wrote simple, mostly grammatical children's stories. That teaches the whole pipeline. It does not produce a model you would use for anything; useful assistants take billions of parameters and far more compute.

What is train-llm-from-scratch?▼
What loss should a language model start at?▼
Should a company train its own LLM from scratch?▼
What is the difference between PPO, DPO and GRPO?▼

Discussion (5)

Mei LinEditor

The step-zero loss check is the one I teach every new engineer. If your fine-tune starts at a loss nowhere near what it should, stop and look at the data before burning a GPU day. It caught a shifted-label bug for me once in under a minute.

Hannah BryantEditor

Same habit, and the loss mask is the second check. If SFT loss counts the prompt tokens, the model learns to repeat questions. Printing one encoded example with its mask aligned underneath is the cheapest debugging step in the whole pipeline.

Cole SullivanContributor

Reader question: I am not an ML engineer. Is this worth a weekend if my job is deploying agents for customers?

Adam ReyesEditor

Yes, the laptop track at least. After you have watched a loss curve flatten and a 369K model invent a character called Sue, you stop treating the model as magic. That shows when a customer asks why their fine-tune forgot how to follow instructions, or why the bill is per token.

Lukas HoffmannEditor

Worth adding: the reward model in their run reached 0.574 preference accuracy against 0.5 for chance. That small margin is what a small model on 8,000 pairs looks like, and it is why RL on top of a weak reward model mostly learns the reward model's quirks.