I trained a 430M-parameter Swedish language model on 10B tokens costing about 70$, which includes failed runs. The reason for building this model is to learn a bit about the scaling of language models. Previously I have only built models up to around 100M params for the BabyLM competition. It gets interesting when multiple GPUs are utilized, as these GPUs need to synchronize at some point. In our training code, we utilize Distributed Data Parallel (DDP), where each GPU gets a copy of the model, and then each GPU trains for one step in parallell and then synchronizes through a collective operation allreduce which takes the mean of the gradients. We effectively increase the batch size through this method, which drastically decreases pretraining wall time.

The corpus chosen for this model was FineWeb-2. The dataset was built by taking lots of Common Crawl snapshots and employing many different deduplication and curation techniques. The dataset separates entries by language, we pick Swedish as our language. In our case we have about 35B words in the swedish dataset. Compared to FineWeb which has 18.5T tokens of Enlish data (using gpt2 tokenizer). If we were to scale our model even more we would quickly run out of tokens to train on.

Evaluation

I evaluated the 430M param model on four Swedish multiple-choice benchmarks with lm-evaluation-harness zero-shot. Each answer is scored by its log-likelihood. The model is architecturally very similar to the dense Qwen3 models. The model was exported as a Hugging Face model and run through the harness without custom code.

As a reference I used Qwen2.5-0.5B, a base model of similar size. The main difference is that Qwen2.5 was pretrained on around 18T multilingual tokens while ours was trained on about 10.5B swedish tokens.

Task Metric ours (430M) Qwen2.5-0.5B
HellaSwag-sv acc_norm 38.2 ± 0.5 30.4 ± 0.5
ARC-sv acc_norm 28.2 ± 1.3 25.8 ± 1.3
Belebele (swe_Latn) acc 23.2 ± 1.4 30.6 ± 1.5
Global-MMLU-sv acc 22.9 ± 0.4 33.1 ± 0.4

For HellaSwag and ARC we report acc_norm, which divides the log-likelihood by its length in characters. This takes into account the quality degrading with longer context.

HellaSwag-sv measures how the model understands the Swedish language. My model is clearly ahead despite seeing roughly 1700 times fewer tokens. Although not very shocking given how Qwen2.5 is trained on many more languages.

The Belebele and Global-MMLU focus one-shot question answering and zero-shot knowledge QAs. It is certainly not suprising Qwen2.5 beats ours in these tasks.

Architecture

The architecture is based on Qwen3 with Deepseek v3 Multi-Token-Prediction heads. The MTP heads are only used for a denser training signal and not used during inference.

It is a decoder-only transformer with pre-norm residual blocks. Every block is

\[x \leftarrow x + \text{Attn}(\text{RMSNorm}(x)), \qquad x \leftarrow x + \text{FFN}(\text{RMSNorm}(x)),\]

followed by one last RMSNorm before the output head.

   
Layers 20
Model dimension $d$ 1280
Query heads / KV heads 20 / 4
Head dimension 64
FFN hidden dimension 3456
Vocabulary 65,536
Context length 2048
Positional encoding RoPE, $\theta = 10^4$
Normalization RMSNorm, $\epsilon = 10^{-5}$
Parameters 428M (84M as tied embeddings)
MTP module (training only) 20.5M

Tokenizer

The tokenizer is a byte-level BPE tokenizer with a vocabulary of $2^{16} = 65{,}536$ on one million Swedish FineWeb-2 documents. The size is chosen so a token id fits in a uint16 which makes the tokenized corpus 21GB instead of 42GB.

The reason for a custom tokenizer is fertility, the average number of tokens per word. A tokenizer trained specifically on Swedish data should in theory outperform one trained on a diverse multi-lingual corpus. On 3000 held-out FineWeb-2 documents:

Tokenizer Vocabulary Tokens per word Relative
This model 65k 1.34 1.00
Viking-7B (Nordic) 131k 1.41 1.05
EuroLLM-1.7B 128k 1.76 1.31
Mistral-Nemo 131k 1.89 1.41
Qwen2.5 / Qwen3 151k 2.02 1.50
GPT-2 50k 2.54 1.89

The comparison is slightly favouring our tokenizer because the test documents come from the same distribution tokenizer was trained on, while the other tokenizers were trained for many langauges.

Additionally, 64 ids were reserved for special tokens, which can be used later for finetuning for instructions. These tokens represent 0.005% of the token ids and are probably negligible for fertility. These reserved tokens can be used later for fine-tuning tasks.

Training Schedule

   
Tokens per step 128 sequences $\times$ 2048 = 262,144
Steps 40,000 (10.5B tokens, 1 epoch)
Precision bfloat16 autocast
Hidden matrices Muon, lr 0.02, Nesterov momentum 0.95, 5 Newton-Schulz steps
Embedding and norms AdamW, lr $10^{-3}$, $\beta = (0.9, 0.95)$, $\epsilon = 10^{-10}$
Weight decay 0.01 for both
Gradient clipping 1.0

The optimizer is split in two. Muon takes all 2D matrices inside the blocksi, using five Newton-Schulz iterations. The embedding and all norm gains are trained with Adam.

We employ a warmup-stable-decay (WSD) schedule, applied as one multiplier on both optimizers. It warms up linearly for 500 steps, stays constant until step 32,000, and the last 20% decays to zero as $1 - \sqrt{p}$ where $p$ is the progress through the decay phase.

Hugging Face

The model is released on Hugging Face