Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
raincandy-uย 
posted an update 1 day ago
Post
2670
20K parameters can tell a story. ๐Ÿš€

๐Ÿค— We trained a ~20k-parameter Transformer that can actually write stories!

raincandy-u/MacroStories

โ†’ ~50ร— smaller than the 1M-parameter TinyStories model
โ†’ ~3,000ร— smaller than AlexNet
โ†’ 81 KB in FP32

yayyy the whole model. เซฎ หถแต” แต• แต”หถ แƒ

She has a 32-dimensional hidden state, a 378-token vocabulary, and just one decoder block โ€” recurrently applied 4 times with shared weights.

Despite having only 19,969 parameters, she can maintain a simple narrative across 100โ€“300 words: establish a goal, encounter a problem, take relevant actions, and reach an outcome.

She runs extremely fast on CPU โ€” no GPU required. The entire model is tiny enough to load almost instantly! โ˜บ๏ธ

This is a really neat result. A few things that stood out to me reading the card:

The 60.6% embedding dominance is striking โ€” at 12,096 of 19,969 params, the shared input/output matrix does most of the "work" in parameter-count terms. That's a very different ratio from typical GPT-2-style models where the embedding is maybe 20โ€“40% of the total. The word-level 378-token vocab is clearly what makes this ratio possible; a subword vocab would blow up that number fast.

The 4-pass recurrent depth with the value-residual mixing (passes 2โ€“4 blending current with pass-1 values) is a clever way to buy extra compute without extra parameters. Curious whether you ablated the pass count โ€” does 2 or 3 passes degrade the "all four criteria" rate noticeably, or is 4 just where it plateaued?

65/100 meeting all four criteria at 20K params is a solid data point for the "how small can you go" question. The Ouro lineage makes sense as a base.

Congrats on shipping it โ€” the card is well-structured and the quick-start is clean.

Cool!๐Ÿ˜Ž