nanochat depth-8 — Agentic LLMs Assignment 1
Checkpoints from Assignment 1 of Agentic LLMs (LIACS, Leiden University), trained with nanochat.
A depth-8 model pre-trained from scratch on 503M tokens of karpathy/climbmix-400b-shuffle, then
fine-tuned in two stages, plus two control runs that make the Task 3 result interpretable.
The checkpoints
| folder | what it is | steps | val bpb | ARC-E | ARC-C | GSM-8K |
|---|---|---|---|---|---|---|
base/ |
end of pre-training | 1920 | 0.939 | 25.1 | 23.0 | 0.0 |
stage1_mid/ |
+ MMLU ×3 and GSM-8K ×4, from base | 464 | 0.865 | 32.6 | 31.3 | 3.3 |
stage2_sft/ |
+ SmolTalk only, from stage 1 | 1435 | 0.429 | 24.5 | 23.2 | 0.0 |
control_A_full/ |
one mixed stage, from base | 1868 | 0.421 | 31.2 | 29.2 | 0.2 |
control_B_replay/ |
SmolTalk + MMLU + GSM-8K, from stage 1 | 1559 | 0.429 | 32.3 | 27.5 | 0.0 |
tokenizer/ |
32,768-token byte-level BPE |
Scores are percentages on the full ARC-Easy and ARC-Challenge test sets and all 1,319 GSM-8K problems. Chance is 25 / 25 / 0. The base model never emits the end-of-turn token, so its GSM-8K run was stopped at 300 problems.
val bpb is bits per byte on the validation split of whatever that stage trained on, so the numbers
are comparable within a stage but not across stages — stage 1 and stage 2 use different mixtures.
What these show
Mid-training lifts ARC-Easy from chance to 32.6%, and the SmolTalk stage that follows it puts the model back at chance. The two controls separate the causes: replaying a small amount of stage-1 data during stage 2 keeps the gain (32.3%), and a single mixed stage straight from base reaches the same place in one run instead of two (31.2%). At this scale the two-stage split buys nothing.
Architecture
Depth 8, d_model 512, 4 heads of width 128, 125,829,354 total parameters (41.9M excluding
embeddings), 503,316,480 training tokens, vocabulary 32,768. Pre-training took 41.6 minutes on one
RTX 5070 Ti at 80.5% MFU, using about 0.2 kWh.
Each folder holds the weights (model_*.pt) and the resolved configuration for that run
(meta_*.json). Optimizer state is not included, so these can be loaded for inference and
evaluation but not resumed for further training.
Loading
import torch
sd = torch.load("base/model_001920.pt", map_location="cpu")
Then load into nanochat's GPT using the hyperparameters recorded in the matching meta_*.json.
Caveats
Single seed throughout, so every number is n = 1. The ± ranges quoted in the report are binomial standard errors from test-set size and do not include run-to-run training variance. This is a 125M-parameter model trained on half a billion tokens; it is a teaching artefact, not a useful assistant.