The September evaluation of KW5 compares nine models on Kiswahili. Seven completed. Their vocabularies run from 32,000 tokens (KW5) to 256,008 (XGLM-564M), and that one fact decides which numbers can be put side by side.
What perplexity measures, and what it does not#
Perplexity is the exponentiated average loss per token. A model with a large vocabulary cuts the same sentence into fewer, longer tokens, so each prediction carries more of the text and the per-token loss is not the same quantity as a small-vocabulary model's. Comparing the two is comparing distances measured in different units.1
A byte is the same unit for every model. Sum the loss over a document in bits, divide by the document's length in UTF-8 bytes, and the result no longer depends on how the tokenizer cut it up:
where is the number of tokens the model's own tokenizer produced, is its mean loss per token, and is the byte count of the text, which is identical for every model.
The September numbers, worked#
For the 149M base model the run records a mean loss of 3.1705 nats over 64,937 tokens of held-out text, and the text is 320,701 bytes:
import math
mean_loss_nats = 3.1704845 # results.csv: lm_loss
tokens = 64_937 # results.csv: ppl_tokens
text_bytes = 320_701 # manifest.json: bpb_bytes
bits = mean_loss_nats * tokens / math.log(2)
print(round(bits / text_bytes, 4)) # 0.9262The same computation over the same 120 documents gives the column below. Vocabulary size is shown so the reader can see what perplexity would have been confounded with.
| Model | Vocabulary | Bits per byte |
|---|---|---|
| KW5-149M | 32,000 | 0.926 |
| KW5-149M-instruct | 32,000 | 0.933 |
| XGLM-564M | 256,008 | 1.122 |
| BLOOM-560M | 250,880 | 1.699 |
| GPT-2 124M | 50,257 | 2.126 |
| Pythia-410M | 50,304 | 2.285 |
| Qwen2.5-0.5B | 151,936 | 2.393 |

Where accuracy still belongs#
Compression is not the whole evaluation. The run also scores three multiple-choice tasks: reading comprehension on Belebele [1], news topic classification on MasakhaNEWS [2] and natural language inference on AfriXNLI [3]. The harness records both raw and PMI-normalised suffix likelihood, the latter to correct for answers that are likely simply because they are short or common [4]; the dossier reports raw accuracy, consistently.
Each score comes with a 95% interval and a one-sided test against that task's chance rate, because a small model near chance can look like a result when it is not one. On AfriXNLI no model in the run clears chance, and the dossier says so rather than ranking them.
Failures stay in the record#
Two of the nine models did not load: one on a tokenizer error, one because its repository is gated. They are listed as failures. Dropping them would make the table look complete, and scoring them zero would report something nobody measured.
Notes
-
The 149M release stores 173.6M parameters because its input embeddings and output head are untied. Parameter counts have the same problem in miniature: say which one you mean. ↩
References
- (2023). The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants.
- (2023). MasakhaNEWS: News Topic Classification for African languages.
- (2024). IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models.
- (2021). Surface Form Competition: Why the Highest Probability Answer Isn't Always Right. EMNLP.
