Training a 430M Swedish LM
I trained a 430M-parameter Swedish language model on 10B tokens costing about 70$, which includes failed runs.
The reason for building this model is to learn a bit about the scaling of language models. Previously I have only built models up to around 100M params for the BabyLM competition. It gets interesting when multiple GPUs are utilized, as these GPUs need to synchronize at some point. In our training code, we utilize Distributed Data Parallel (DDP), where each GPU gets a copy of the model, and then each GPU trains for one step in parallell and then synchronizes through a collective operation allreduce which takes the mean of the gradients. We effectively increase the batch size through this method, which drastically decreases pretraining wall time.
The corpus chosen for this model was FineWeb-2. The dataset was built by taking lots of Common Crawl snapshots and employing many different deduplication and curation techniques. The dataset separates entries by language, we pick Swedish as our language. In our case we have about 35B words in the swedish dataset. Compared to FineWeb which has 18.5T tokens of Enlish data (using gpt2 tokenizer). If we were to scale our model even more we would quickly run out of tokens to train on.
Evaluation
I evaluated the 430M param model on four Swedish multiple-choice benchmarks with lm-evaluation-harness zero-shot. Each answer is scored by its log-likelihood. The model is architecturally very similar to the dense Qwen3 models. The model was exported as a Hugging Face model and run through the harness without custom code.
As a reference I used Qwen2.5-0.5B, a base model of similar size. The main difference is that Qwen2.5 was pretrained on around 18T multilingual tokens while ours was trained on about 10.5B swedish tokens.
| Task | Metric | ours (430M) | Qwen2.5-0.5B |
|---|---|---|---|
| HellaSwag-sv | acc_norm | 38.2 ± 0.5 | 30.4 ± 0.5 |
| ARC-sv | acc_norm | 28.2 ± 1.3 | 25.8 ± 1.3 |
| Belebele (swe_Latn) | acc | 23.2 ± 1.4 | 30.6 ± 1.5 |
| Global-MMLU-sv | acc | 22.9 ± 0.4 | 33.1 ± 0.4 |
For HellaSwag and ARC we report acc_norm, which divides the log-likelihood by its length in characters. This takes into account the quality degrading with longer context.
HellaSwag-sv measures how the model understands the Swedish language. My model is clearly ahead despite seeing roughly 1700 times fewer tokens. Although not very shocking given how Qwen2.5 is trained on many more languages.
The Belebele and Global-MMLU focus one-shot question answering and zero-shot knowledge QAs. It is certainly not suprising Qwen2.5 beats ours in these tasks.
Architecture
The architecture is based on Qwen3 with Deepseek v3 Multi-Token-Prediction heads. The MTP heads are only used for a denser training signal and not used during inference.
It is a decoder-only transformer with pre-norm residual blocks. Every block is
\[x \leftarrow x + \text{Attn}(\text{RMSNorm}(x)), \qquad x \leftarrow x + \text{FFN}(\text{RMSNorm}(x)),\]followed by one last RMSNorm before the output head.
| Layers | 20 |
| Model dimension $d$ | 1280 |
| Query heads / KV heads | 20 / 4 |
| Head dimension | 64 |
| FFN hidden dimension | 3456 |
| Vocabulary | 65,536 |
| Context length | 2048 |
| Positional encoding | RoPE, $\theta = 10^4$ |
| Normalization | RMSNorm, $\epsilon = 10^{-5}$ |
| Parameters | 428M (84M as tied embeddings) |
| MTP module (training only) | 20.5M |
Tokenizer
The tokenizer is a byte-level BPE tokenizer with a vocabulary of $2^{16} = 65{,}536$ on
one million Swedish FineWeb-2 documents. The size is chosen so a token id fits in a uint16
which makes the tokenized corpus 21GB instead of 42GB.
The reason for a custom tokenizer is fertility, the average number of tokens per word. A tokenizer trained specifically on Swedish data should in theory outperform one trained on a diverse multi-lingual corpus. On 3000 held-out FineWeb-2 documents:
| Tokenizer | Vocabulary | Tokens per word | Relative |
|---|---|---|---|
| This model | 65k | 1.34 | 1.00 |
| Viking-7B (Nordic) | 131k | 1.41 | 1.05 |
| EuroLLM-1.7B | 128k | 1.76 | 1.31 |
| Mistral-Nemo | 131k | 1.89 | 1.41 |
| Qwen2.5 / Qwen3 | 151k | 2.02 | 1.50 |
| GPT-2 | 50k | 2.54 | 1.89 |
The comparison is slightly favouring our tokenizer because the test documents come from the same distribution tokenizer was trained on, while the other tokenizers were trained for many langauges.
Additionally, 64 ids were reserved for special tokens, which can be used later for finetuning for instructions. These tokens represent 0.005% of the token ids and are probably negligible for fertility. These reserved tokens can be used later for fine-tuning tasks.
Training Schedule
| Tokens per step | 128 sequences $\times$ 2048 = 262,144 |
| Steps | 40,000 (10.5B tokens, 1 epoch) |
| Precision | bfloat16 autocast |
| Hidden matrices | Muon, lr 0.02, Nesterov momentum 0.95, 5 Newton-Schulz steps |
| Embedding and norms | AdamW, lr $10^{-3}$, $\beta = (0.9, 0.95)$, $\epsilon = 10^{-10}$ |
| Weight decay | 0.01 for both |
| Gradient clipping | 1.0 |
The optimizer is split in two. Muon takes all 2D matrices inside the blocksi, using five Newton-Schulz iterations. The embedding and all norm gains are trained with Adam.
We employ a warmup-stable-decay (WSD) schedule, applied as one multiplier on both optimizers. It warms up linearly for 500 steps, stays constant until step 32,000, and the last 20% decays to zero as $1 - \sqrt{p}$ where $p$ is the progress through the decay phase.
Hugging Face
The model is released on Hugging Face