Refusal and safety evaluation
Question. Do models refuse the same harmful requests in Somali that they refuse in English?
Current experiment. SomaliBench puts identical harmful requests to models in both languages, verified by native speakers, and publishes the refusal gap per model on a live leaderboard.
Latest result. Measured gaps of 0.97 → 0.07 (Llama 3.1) and 0.80 → 0.05 (Aya 23).
Open uncertainty. A low refusal rate does not always mean fluent harmful compliance; some Somali failures surface as incoherent output. Separating genuine compliance from low-quality generation is the central task for the next version.