Mnemonic

Small chat models where every part (data pipeline, tokenizer, GPU training, quantization and inference) is hand-written assembly: x86-64 on the CPU, PTX on the GPU, WebAssembly in the browser.

Code: https://github.com/Supergoatscriptguy/Mnemonic

file model weights size
mnemonic-300m-q8.mnm 300M int8, one scale per row 291 MB
mnemonic-300m-q4.mnm 300M int4, one scale per group of 32 181 MB
mnemonic-q8.mnm 126M int8, one scale per row 121 MB
mnemonic-q4.mnm 126M int4, one scale per group of 32 75 MB

Models. Both are Llama-style decoders (RoPE, RMSNorm, SwiGLU, grouped-query attention, tied embeddings, no biases), with a byte-level BPE vocab of 32768.

  • 300M: 24 layers, d_model 1024, 16 query / 4 key-value heads, ffn 2816, context 2048.
  • 126M: 16 layers, d_model 768, 12 query / 4 key-value heads, ffn 2048, context 1024.

Training. Both were trained on one RTX 5070 Ti at home.

  • 300M: pretrained on 10B tokens of FineWeb-Edu (about 90 hours, validation loss 2.84). Then fine-tuned for chat in one pass of smol-smoltalk plus an identity set, at 2048 context (3.4 hours).
  • 126M: pretrained on 5B tokens (19 hours, validation loss 3.10), then fine-tuned for chat on smol-smoltalk plus a smaller identity set.

Format. .mnm is the project's own format (a 4 KB header, then the tensors), read by the engines in the repo (chat/engine.asm, site/engine.wat). It is not a transformers checkpoint. For llama.cpp and Ollama there are GGUF builds: ollama run supergoatscriptguy/mnemonic.

Limitations. These are small models: they make confident mistakes (especially with dates and numbers), are weak at math, and lose the thread in long conversations. The 300M is noticeably better than the 126M at all three.

Made by Supergoatscriptguy.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support