Research
Our research agenda is alignment and evaluation infrastructure for Somali. We build safety and toxicity test sets, refusal evaluations, translation and comprehension benchmarks, and red-teaming methods designed for low-resource languages. The goal is simple: frontier labs and deployers should be able to measure how their systems behave in Somali, and improve them. Today they cannot. We intend to make that measurement routine.
The work runs in both directions. Somali is a hard case for alignment — little data, rich dialect variation, a vast oral tradition and a young written one — and methods that hold up here transfer to the hundreds of languages in the same position. We contribute what we learn back to the broader alignment community: benchmarks, methodology, and honest negative results included.
PublicationsOur first benchmark paper is in preparation (2026).
Research principles- Open release by default. Datasets, benchmarks, and papers are published under open licenses unless there is a specific safety reason not to.
- Human data dignity. Contributors are credited and consented. Nobody’s words enter a dataset without their knowledge.
- Dialect inclusion. Somali is not one uniform language. Our benchmarks and datasets cover Maxaa tiri and Maay from the start, not as an afterthought.
- Safety evaluation before capability hype. We measure whether systems are safe and accurate in Somali before we celebrate that they speak it at all.
- Alignment is for every language. Findings, methods, and failures are shared with the broader alignment community, so that work on one low-resource language moves them all.