About Unkad Labs
Unkad Labs is an independent AI research lab working on one question: do the safety properties of AI systems survive a change of language? The name comes from Somali — unkad, creation from nothing, shares a root with unug, the cell — and describes the method: large things get built by assembling small verified units until something holds together.
What we study
When a lab reports that a model refuses harmful requests reliably, that finding was almost certainly established in English. The model is then deployed to people who will not address it in English, and the safety claim travels with the deployment as though language had nothing to do with it.
We measured what actually happens. Putting identical harmful requests to open-weight models in English and in Somali, Llama 3.1 refuses 97 percent of the time in English and 7 percent of the time in Somali. Aya falls from 80 percent to 5 percent. For most of the world’s languages, nobody has run this test at all.
The scope is deliberate: safety evaluation, oversight under language limits, and the data infrastructure both require. We do not train large models, and we do not make capability claims — we would rather know whether a system is safe in Somali than celebrate that it speaks it.
Why Somali
It is our language, which matters both for the work and for who gets to do it. It is also a genuinely useful hard case: more than twenty million speakers, an enormous oral tradition, a writing system standardised only in 1972, and real dialect variation — Maay is spoken by millions and nearly absent from digital text. Somali is thin enough in training data that the failures we care about are visible rather than subtle. Methods that hold up under those conditions have a reasonable chance of holding up for the hundreds of languages in similar positions.
People and structure
- LeadershipFounder-led, with a volunteer community of Somali-speaking contributors and reviewers, at home and in the diaspora.
- CommunityVolunteer contributors write, translate, and peer-validate the corpus; the live count is on the Qor page, and contribution is open to any Somali speaker at qor.unkad.com.
- Review structureTrusted linguist reviewers give final sign-off on every released item and settle validation disagreements. Reviewers are appointed from proven contributors; nobody reviews their own work.
- AuthorshipResearch notes carry citable authorship; BibTeX for every note is on the research page.
Legal status and funding
- Legal statusUnkad Labs operates as an independent research lab and is not yet a registered legal entity.
- FundingSelf-funded to date; no external funders. For grant and funding conversations, write to research@unkad.com.
- LicensingDatasets under CC BY-SA 4.0; code under open-source licenses on GitHub.
Timeline
- Lab founded; SomaliBench measures the English–Somali refusal gap
- Qor Af-Soomaali opens; the first corpus sentence is written
- First four research notes: the token tax, the overseer, the guard collapse, the lie that moved
- First open dataset release (v0.1.0) on Hugging Face
- Grammaticality note: one model reads Somali, another approves word salad
- Corpus v0.2.1: 2,282 verified sentences
- Frontier evaluation: GPT-5.6, Gemini 3.1 Pro, and Claude Sonnet 5 against the community’s quality judgments
Working with us
We are interested in conversations with people working on multilingual evaluation, AI safety institutes and their capacity-building work, and researchers building language resources for under-served languages. If your work touches whether models behave safely outside English, we would like to compare notes. Get in touch.