Note 07 August 12, 2026
We showed 593 community-validated Somali submissions to the newest models from OpenAI, Google, and Anthropic and asked the question our volunteers answer every day: is this good Somali, or not? The models placed the texts in the right domain two-thirds of the time.
Evaluation / Frontier models / Data quality · Research note
Note 06 August 5, 2026
We gave three open models the easiest version of our community's validation job: tell a real Somali sentence from the same words shuffled.
Evaluation / Multilingual / Somali NLP · Research note
Note 05 July 30, 2026
We released the first version of an open Somali corpus.
Datasets / Low-resource languages / Provenance · Research note
Note 04 July 26, 2026
We trained a lying model against two detectors blind in different ways and asked whether it rebuilt its lie to suit whichever one it faced.
Safety / Interpretability / Evaluation · Research note
Note 03 July 24, 2026
We gave three safety filters the same 100 harmful requests in English and in Somali.
Safety / Multilingual safety / Evaluation · Research note
Note 02 July 22, 2026
We gave a model a passage it could not read and asked it to judge claims about that passage.
Safety / Scalable oversight / Evaluation · Research note
Note 01 July 20, 2026
Every major tokenizer cuts Somali into roughly twice as many pieces as English.
Safety / Evaluation / Tokenization · Research note
Lab essay July 19, 2026
Why we started a lab to measure whether AI systems are safe in Somali, what the name means, and why safety that only works in English is not safety.
The lab / Safety · Essay