English
nanochat
coursework

nanochat depth-8 — Agentic LLMs Assignment 1

Checkpoints from Assignment 1 of Agentic LLMs (LIACS, Leiden University), trained with nanochat.

A depth-8 model pre-trained from scratch on 503M tokens of karpathy/climbmix-400b-shuffle, then fine-tuned in two stages, plus two control runs that make the Task 3 result interpretable.

The checkpoints

folder what it is steps val bpb ARC-E ARC-C GSM-8K
base/ end of pre-training 1920 0.939 25.1 23.0 0.0
stage1_mid/ + MMLU ×3 and GSM-8K ×4, from base 464 0.865 32.6 31.3 3.3
stage2_sft/ + SmolTalk only, from stage 1 1435 0.429 24.5 23.2 0.0
control_A_full/ one mixed stage, from base 1868 0.421 31.2 29.2 0.2
control_B_replay/ SmolTalk + MMLU + GSM-8K, from stage 1 1559 0.429 32.3 27.5 0.0
tokenizer/ 32,768-token byte-level BPE

Scores are percentages on the full ARC-Easy and ARC-Challenge test sets and all 1,319 GSM-8K problems. Chance is 25 / 25 / 0. The base model never emits the end-of-turn token, so its GSM-8K run was stopped at 300 problems.

val bpb is bits per byte on the validation split of whatever that stage trained on, so the numbers are comparable within a stage but not across stages — stage 1 and stage 2 use different mixtures.

What these show

Mid-training lifts ARC-Easy from chance to 32.6%, and the SmolTalk stage that follows it puts the model back at chance. The two controls separate the causes: replaying a small amount of stage-1 data during stage 2 keeps the gain (32.3%), and a single mixed stage straight from base reaches the same place in one run instead of two (31.2%). At this scale the two-stage split buys nothing.

Architecture

Depth 8, d_model 512, 4 heads of width 128, 125,829,354 total parameters (41.9M excluding embeddings), 503,316,480 training tokens, vocabulary 32,768. Pre-training took 41.6 minutes on one RTX 5070 Ti at 80.5% MFU, using about 0.2 kWh.

Each folder holds the weights (model_*.pt) and the resolved configuration for that run (meta_*.json). Optimizer state is not included, so these can be loaded for inference and evaluation but not resumed for further training.

Loading

import torch
sd = torch.load("base/model_001920.pt", map_location="cpu")

Then load into nanochat's GPT using the hyperparameters recorded in the matching meta_*.json.

Caveats

Single seed throughout, so every number is n = 1. The ± ranges quoted in the report are binomial standard errors from test-set size and do not include run-to-run training variance. This is a 125M-parameter model trained on half a billion tokens; it is a teaching artefact, not a useful assistant.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train chicoballer/nanochat-d8-assignment1