Kiswahili language models / Open weights
A language at the weights.
Open-weight Kiswahili language models in 109M and 149M sizes, trained from scratch. Compare base and instruction-tuned releases, inspect training methods and benchmark results, and review the findings that guide subsequent work.
Explore the modelsA language at the weights
KW5 learns directly from Kiswahili text. The 109M and 149M generations each have a base checkpoint and instruction-tuned variants, released under Apache 2.0.
The release names identify the generations. The 149M generation stores 173.6M parameters: a 123.9M transformer body, separate 24.6M input and output embeddings, and the Canon layers. Use the stored count when comparing memory and model size.
Built for a closer look
The training stack supports document-boundary attention and label masking, checkpoint resume, and a cooldown transition within the run. A frozen Kiswahili corpus retains its provenance across eight language registers. These choices make the training process inspectable alongside the resulting weights.
The v2 history also documents an embedding-tie error at load time. The base weights were recovered and verified on reload. The subsequent instruction diagnosis showed why falling training loss alone could miss repetitive generation: the samples had already entered a loop while the loss continued to improve.
Small models make these failures easier to reproduce and investigate. Compression, extraction, instruction following and factual recall need separate measurements; success in one does not establish success in the others.
Inside the 149M generation
A decoder trained for Kiswahili, with a 32,000-token SentencePiece vocabulary, byte fallback and NFC text normalization. The context window is 2,048 tokens.
| Depth and width | 20 layers · 768 hidden dimensions · 2,048 feed-forward dimensions |
|---|---|
| Attention | 12 query heads · 3 key/value heads · grouped-query attention |
| Normalization | Pre-norm RMSNorm, with normalization on queries and keys |
| Activation and position | SwiGLU · rotary position encoding |
| Canon | Causal depthwise convolutions at four points in each block |
| Released precision | FP32 weights; training used bf16 autocast on TPU v5e-8 |
The model card includes the inference implementation and prompt conventions. KW5 v2 is a custom architecture; loading it as an unrelated transformer can change behaviour or fail outright.
Open releases
| Model | Use and context |
|---|---|
| 149M Base ↗ | Pretrained on 1.97B tokens. A foundation for language research and further training. |
| 149M Instruct r2 ↗ | The revised instruction release. More abstention, with documented limits in follow-up turns and factual reliability. |
| 149M Instruct ↗ | The first v2 fine-tune. Retained for comparison; repetition behaviour is documented in the audit. |
| 109M Base ↗ | The first-generation base model, pretrained on 1.41B tokens. |
| 109M Instruct ↗ | The first-generation instruction model and a baseline for later evaluations. |
What we measured
The 14 September comparison measured language compression and likelihood-based classification. Bits per byte supports comparison across tokenizers; perplexity does not. These results do not measure the quality of a conversation.
Scores below describe that run, rather than a general ranking. The manifest pins model revisions and the result files include confidence intervals.
Bits per byte
149M base on 120 held-out Kiswahili documents. The lowest value among completed models in this run.
MasakhaNEWS
Topic classification. The interval overlaps XGLM-564M, so the small point-score difference does not establish a lead.
AfriXNLI
Natural language inference, against a 33.3% chance rate. This run does not establish reliable inference.
The September audit
The later audit examined generated behaviour, prompt formatting and the evaluation itself. It found positional bias in grounded questions and training-template overlap in parts of the behavioural test. An instruction checkpoint scored on unformatted completion prompts is not a test of instruction following.
On a small out-of-template set, r2 produced 11 genuine refusals out of 20 unanswerable questions after manual review. Follow-up questions still triggered excessive refusal. All evaluated models failed the 30-item arithmetic probe. These are small diagnostic sets, with their method and outputs retained.
Use with the limits in view
These are research models. Fluent text can contain invented facts, repetitive output or an inappropriate refusal. They are not a reliable basis for consequential decisions.
For the 149M instruction releases, use the exact chat format in the model card: an explicit beginning-of-sequence token and the correct end-of-turn stop token. The base and instruction models use different prompt conventions.
V3 is in progress. The published generations remain available as baselines; this page does not announce a new checkpoint.
Reproduce the comparison
149M comparison · 14 September
Seed 20260914, Tesla T4, float16. The run scores 900 Belebele items, 600 AfriXNLI items and 476 MasakhaNEWS items. Compression uses 320,701 bytes across 120 held-out documents. Two checkpoints could not load; their failures remain in the table.
| Model | Bits / byte ↓ | Belebele | AfriXNLI | MasakhaNEWS | Run status |
|---|---|---|---|---|---|
| KW5-149M | 0.9262 | 29.3% | 33.7% | 73.1% | Completed |
| KW5-149M-instruct | 0.9329 | 29.1% | 34.8% | 72.5% | Completed |
| XGLM-564M | 1.1216 | 28.8% | 35.3% | 72.5% | Completed |
| BLOOM-560M | 1.6994 | 26.2% | 32.2% | 18.3% | Completed |
| GPT-2 124M | 2.1256 | 26.4% | 33.3% | 24.4% | Completed |
| Pythia-410M | 2.2855 | 27.2% | 33.7% | 26.1% | Completed |
| Qwen2.5-0.5B | 2.3930 | 27.0% | 33.3% | 25.0% | Completed |
| Goldfish swa 100MB | Unmeasured | Unmeasured | Unmeasured | Unmeasured | Failed to load |
| Gemma 3 270M | Unmeasured | Unmeasured | Unmeasured | Unmeasured | Failed to load |
109M comparison · 4 September
The earlier run uses seed 20260904 and its own pinned model revisions. Classification uses PMI-normalized suffix likelihood. Entries flagged for single-label collapse remain unscored. This is a separate experiment, so its values should not be treated as a matched comparison with the later 149M run.
| Model | Bits / byte ↓ | Belebele | AfriXNLI | MasakhaNEWS | Run status |
|---|---|---|---|---|---|
| KW5-Lite (base) | 1.2009 | 26.1% | 32.5% | 23.1% | Completed |
| KW5-Lite (instruct) | 1.3016 | 25.0% | 33.5% | 29.0% | Completed |
| swa-gpt2-100mb | 1.2851 | 26.2% | 36.2% | 28.1% | Completed |
| XGLM-564M | 1.3360 | 29.2% | 34.3% | 63.4% | Completed |
| BLOOM-560m | 2.1054 | 28.1% | 33.8% | 41.2% | Completed |
| Qwen2.5-0.5B | 2.9629 | 25.2% | 32.7% | 37.4% | Completed |
| Pythia-410m | 2.7877 | 28.1% | Unscored | 30.3% | Completed |
| GPT-2-124m | 3.4553 | 27.7% | Unscored | 21.6% | Completed |
| Gemma-3-270m | Unmeasured | Unscored | Unscored | Unscored | Failed to load |











