The frontier can read Somali. It cannot judge it.
GPT-5.6, Gemini 3.1 Pro, and Claude Sonnet 5 place community-labeled Somali texts in the right domain 64–70% of the time, but reproduce almost none of the community’s quality rejections (0/23, 0/23, 3/23) while still scoring ~96% agreement — an illusion made of base rates.