Regnant
Request access
← Research

Kiswahili language models / Open weights

A language at the weights.

Open-weight Kiswahili language models in 109M and 149M sizes, trained from scratch. Compare base and instruction-tuned releases, inspect training methods and benchmark results, and review the findings that guide subsequent work.

Explore the models
109M / 149MBASE / INSTRUCTAPACHE 2.0

A language at the weights

KW5 learns directly from Kiswahili text. The 109M and 149M generations each have a base checkpoint and instruction-tuned variants, released under Apache 2.0.

The release names identify the generations. The 149M generation stores 173.6M parameters: a 123.9M transformer body, separate 24.6M input and output embeddings, and the Canon layers. Use the stored count when comparing memory and model size.

Built for a closer look

The training stack supports document-boundary attention and label masking, checkpoint resume, and a cooldown transition within the run. A frozen Kiswahili corpus retains its provenance across eight language registers. These choices make the training process inspectable alongside the resulting weights.

The v2 history also documents an embedding-tie error at load time. The base weights were recovered and verified on reload. The subsequent instruction diagnosis showed why falling training loss alone could miss repetitive generation: the samples had already entered a loop while the loss continued to improve.

Small models make these failures easier to reproduce and investigate. Compression, extraction, instruction following and factual recall need separate measurements; success in one does not establish success in the others.

Inside the 149M generation

A decoder trained for Kiswahili, with a 32,000-token SentencePiece vocabulary, byte fallback and NFC text normalization. The context window is 2,048 tokens.

Depth and width20 layers · 768 hidden dimensions · 2,048 feed-forward dimensions
Attention12 query heads · 3 key/value heads · grouped-query attention
NormalizationPre-norm RMSNorm, with normalization on queries and keys
Activation and positionSwiGLU · rotary position encoding
CanonCausal depthwise convolutions at four points in each block
Released precisionFP32 weights; training used bf16 autocast on TPU v5e-8

The model card includes the inference implementation and prompt conventions. KW5 v2 is a custom architecture; loading it as an unrelated transformer can change behaviour or fail outright.

Open releases

ModelUse and context
149M Base ↗Pretrained on 1.97B tokens. A foundation for language research and further training.
149M Instruct r2 ↗The revised instruction release. More abstention, with documented limits in follow-up turns and factual reliability.
149M Instruct ↗The first v2 fine-tune. Retained for comparison; repetition behaviour is documented in the audit.
109M Base ↗The first-generation base model, pretrained on 1.41B tokens.
109M Instruct ↗The first-generation instruction model and a baseline for later evaluations.

What we measured

The 14 September comparison measured language compression and likelihood-based classification. Bits per byte supports comparison across tokenizers; perplexity does not. These results do not measure the quality of a conversation.

Scores below describe that run, rather than a general ranking. The manifest pins model revisions and the result files include confidence intervals.

0.926

Bits per byte

149M base on 120 held-out Kiswahili documents. The lowest value among completed models in this run.

73.1%

MasakhaNEWS

Topic classification. The interval overlaps XGLM-564M, so the small point-score difference does not establish a lead.

33.7%

AfriXNLI

Natural language inference, against a 33.3% chance rate. This run does not establish reliable inference.

The September audit

The later audit examined generated behaviour, prompt formatting and the evaluation itself. It found positional bias in grounded questions and training-template overlap in parts of the behavioural test. An instruction checkpoint scored on unformatted completion prompts is not a test of instruction following.

On a small out-of-template set, r2 produced 11 genuine refusals out of 20 unanswerable questions after manual review. Follow-up questions still triggered excessive refusal. All evaluated models failed the 30-item arithmetic probe. These are small diagnostic sets, with their method and outputs retained.

Read the audit ↗

Use with the limits in view

These are research models. Fluent text can contain invented facts, repetitive output or an inappropriate refusal. They are not a reliable basis for consequential decisions.

For the 149M instruction releases, use the exact chat format in the model card: an explicit beginning-of-sequence token and the correct end-of-turn stop token. The base and instruction models use different prompt conventions.

V3 is in progress. The published generations remain available as baselines; this page does not announce a new checkpoint.

Reproduce the comparison

Results CSV ↗ · Significance CSV ↗ · Run manifest ↗

Training and evaluation code ↗

149M comparison · 14 September

Seed 20260914, Tesla T4, float16. The run scores 900 Belebele items, 600 AfriXNLI items and 476 MasakhaNEWS items. Compression uses 320,701 bytes across 120 held-out documents. Two checkpoints could not load; their failures remain in the table.

ModelBits / byte ↓BelebeleAfriXNLIMasakhaNEWSRun status
KW5-149M0.926229.3%33.7%73.1%Completed
KW5-149M-instruct0.932929.1%34.8%72.5%Completed
XGLM-564M1.121628.8%35.3%72.5%Completed
BLOOM-560M1.699426.2%32.2%18.3%Completed
GPT-2 124M2.125626.4%33.3%24.4%Completed
Pythia-410M2.285527.2%33.7%26.1%Completed
Qwen2.5-0.5B2.393027.0%33.3%25.0%Completed
Goldfish swa 100MBUnmeasuredUnmeasuredUnmeasuredUnmeasuredFailed to load
Gemma 3 270MUnmeasuredUnmeasuredUnmeasuredUnmeasuredFailed to load

149M benchmark figures

Grouped bar chart of accuracy for every model on Belebele-sw, AfriXNLI-sw and MasakhaNEWS-sw, with each task's chance line marked
Accuracy by task, against chance
Bar chart of bits per byte on held-out Swahili text; KW5-149M is lowest in the field
Bits per byte on held-out Swahili
Quality plotted against parameter count for every model in the run
Quality against parameter count
Generation throughput in tokens per second against peak GPU memory for each model
Throughput against peak memory
Repetition and lexical-diversity measures over the generation probe, per model
Generation quality probe
Overview panel combining accuracy, compression and efficiency for the whole run
The run, in one panel

109M comparison · 4 September

The earlier run uses seed 20260904 and its own pinned model revisions. Classification uses PMI-normalized suffix likelihood. Entries flagged for single-label collapse remain unscored. This is a separate experiment, so its values should not be treated as a matched comparison with the later 149M run.

ModelBits / byte ↓BelebeleAfriXNLIMasakhaNEWSRun status
KW5-Lite (base)1.200926.1%32.5%23.1%Completed
KW5-Lite (instruct)1.301625.0%33.5%29.0%Completed
swa-gpt2-100mb1.285126.2%36.2%28.1%Completed
XGLM-564M1.336029.2%34.3%63.4%Completed
BLOOM-560m2.105428.1%33.8%41.2%Completed
Qwen2.5-0.5B2.962925.2%32.7%37.4%Completed
Pythia-410m2.787728.1%Unscored30.3%Completed
GPT-2-124m3.455327.7%Unscored21.6%Completed
Gemma-3-270mUnmeasuredUnscoredUnscoredUnscoredFailed to load

109M benchmark figures

Grouped bar chart of PMI-normalized accuracy for every model on Belebele, AfriXNLI and MasakhaNEWS, with the random-chance line marked on each panel
Accuracy by task, against chance
Forest plot of KW5-Lite base accuracy minus each baseline, with 95% bootstrap confidence intervals; Belebele and AfriXNLI intervals all cross zero, MasakhaNEWS intervals do not
Paired bootstrap against every baseline
Bar chart of bits per byte on held-out Swahili text; KW5-Lite base is lowest in the field
Bits per byte on held-out Swahili
Accuracy plotted against parameter count, grouping models into size classes
Accuracy against parameter count
Generation throughput in tokens per second against peak GPU memory for each model
Throughput against peak memory
Composite score across all normalized metrics, ranked by model
Composite score

Run artifacts