Mnemonic
Small chat models where every part (data pipeline, tokenizer, GPU training, quantization and inference) is hand-written assembly: x86-64 on the CPU, PTX on the GPU, WebAssembly in the browser.
Code: https://github.com/Supergoatscriptguy/Mnemonic
| file | model | weights | size |
|---|---|---|---|
mnemonic-300m-q8.mnm |
300M | int8, one scale per row | 291 MB |
mnemonic-300m-q4.mnm |
300M | int4, one scale per group of 32 | 181 MB |
mnemonic-q8.mnm |
126M | int8, one scale per row | 121 MB |
mnemonic-q4.mnm |
126M | int4, one scale per group of 32 | 75 MB |
Models. Both are Llama-style decoders (RoPE, RMSNorm, SwiGLU, grouped-query attention, tied embeddings, no biases), with a byte-level BPE vocab of 32768.
- 300M: 24 layers, d_model 1024, 16 query / 4 key-value heads, ffn 2816, context 2048.
- 126M: 16 layers, d_model 768, 12 query / 4 key-value heads, ffn 2048, context 1024.
Training. Both were trained on one RTX 5070 Ti at home.
- 300M: pretrained on 10B tokens of FineWeb-Edu (about 90 hours, validation loss 2.84). Then fine-tuned for chat in one pass of smol-smoltalk plus an identity set, at 2048 context (3.4 hours).
- 126M: pretrained on 5B tokens (19 hours, validation loss 3.10), then fine-tuned for chat on smol-smoltalk plus a smaller identity set.
Format. .mnm is the project's own format (a 4 KB header, then the tensors), read by the engines in the repo (chat/engine.asm, site/engine.wat). It is not a transformers checkpoint. For llama.cpp and Ollama there are GGUF builds: ollama run supergoatscriptguy/mnemonic.
Limitations. These are small models: they make confident mistakes (especially with dates and numbers), are weak at math, and lose the thread in long conversations. The 300M is noticeably better than the 126M at all three.
Made by Supergoatscriptguy.