Articles

Research notes with their evidence attached, and the occasional piece about building the lab itself. Failed predictions stay published.

Note 07

The frontier can read Somali. It cannot judge it.

We showed 593 community-validated Somali submissions to the newest models from OpenAI, Google, and Anthropic and asked the question our volunteers answer every day: is this good Somali, or not? The models placed the texts in the right domain two-thirds of the time.

Evaluation / Frontier models / Data quality · Research note

Note 05

Every sentence has a name

We released the first version of an open Somali corpus.

Datasets / Low-resource languages / Provenance · Research note

Note 04

The lie moved. The liar did not.

We trained a lying model against two detectors blind in different ways and asked whether it rebuilt its lie to suit whichever one it faced.

Safety / Interpretability / Evaluation · Research note

Note 03

The guard that waves Somali through

We gave three safety filters the same 100 harmful requests in English and in Somali.

Safety / Multilingual safety / Evaluation · Research note

Note 02

The overseer that cannot read

We gave a model a passage it could not read and asked it to judge claims about that passage.

Safety / Scalable oversight / Evaluation · Research note

Note 01

The Somali token tax

Every major tokenizer cuts Somali into roughly twice as many pieces as English.

Safety / Evaluation / Tokenization · Research note

Lab essay

Unkad: creation from nothing

Why we started a lab to measure whether AI systems are safe in Somali, what the name means, and why safety that only works in English is not safety.

The lab / Safety · Essay

Research notes also live on the research page, with BibTeX and artifacts.