MaisMedia Lab · Annual study — first edition results
How well do AI models speak Galician?
Galician is a Romance language spoken by around 2.5 million people — and, like most
minoritized languages, nobody was measuring how well mainstream AI actually handles it. So we
did: five leading models, all on free tiers, took our public 10-test battery
(orthography, verb morphology, translation, a real customer-service email, language
persistence, terminology). These are the first-edition results.
Final leaderboard
| # | Model (free tier) | Score /100 |
| 1 | DeepSeek | 100 |
| 2 | Claude Sonnet | 99 |
| 2 | ChatGPT | 99 |
| 4 | Gemini 3.6 Flash | 98 |
| 5 | Copilot (Smart) | 95 |
Tested on September 12, 2026. Deductions, one by one: the Spanish-style progressive gerund
(«estamos xestionando» instead of normative «estar a + infinitive») cost ChatGPT
and Copilot one point each; a gender-agreement slip («coa mellor
leite galega» — “leite” is masculine) cost Gemini and Copilot; a
gender-agreement slip («praza monumentais») cost Claude one; and Copilot lost three points
for an invented figure — see below.
Four findings
1. The risk isn't where you'd expect
In the language-persistence test we asked factual questions in Galician without saying which language we wanted. All five answered in Galician. The failure mode of AI + minoritized languages in 2026 isn't switching to the dominant language — it's losing points in the fine grain: syntactic calques and agreement slips.
2. A hallucination with a citation
Copilot claimed the province has 302,250 inhabitants “according to the latest official INE registry (2026)”. That figure doesn't exist and that registry hasn't been declared: the official figure for Jan 1, 2025 is 326,013 (Royal Decree 1117/2025). We verified it against Spain's statistics institute (INE), the Galician statistics institute (IGE) and the official gazette (BOE). Its five municipalities, though, were digit-for-digit exact.
3. Errors repeat across models
The progressive-gerund calque fell to ChatGPT and Copilot. The gender-agreement slip («leite galega») fell to Gemini and Copilot. These aren't anecdotes; they're patterns — exactly what a benchmark should surface.
4. AI models are not deterministic
Claude ran the battery twice in a row with identical prompts. The first run was flawless; the second introduced a brand-new agreement slip. Repeating a test doesn't guarantee repeating quality — which is why we publish the full transcripts as evidence.
An accidental experiment: same model, two speeds
Claude took the battery twice: once on its slowest, most powerful mode, once on medium. The
slow run wrote flawless Galician (100) but kept our reviewer waiting over five minutes. The
medium run answered instantly and made a single agreement slip: 99. Practical takeaway for
any business: more thinking time does not guarantee more quality — the slip
showed up precisely on the fast run.
Why this matters
For the 2.5 million people who speak Galician — and for any organization
serving them with chatbots, automation or content: mainstream free-tier AI now handles
normative Galician at a level that was unthinkable a few years ago. The differentiators are
price, latency and integration, not the language. But always demand samples in
Galician before you pay: what failed here (an invented statistic, a calqued verb)
is invisible until it happens to you.
For researchers: this is, to our knowledge, the first public benchmark of Galician
capability in mainstream commercial models. The battery, full transcripts and scoring are
open — replicate it, extend it to more models, challenge our scoring.
Methodology & downloads
Two fresh conversations per model (the persistence test alone and first; the rest batched),
literal prompts, model version and tier logged, screenshots of every answer, scoring against
a public rubric (v1.4). No sponsorship, no manufacturer involvement: we measure the
machines, not companies.
The battery (.md, in Galician) Full transcripts (.md) Scores & deductions (.md)
The battery scored 98.4 on average — too easy. The 2027 edition will add tests that spread
the field: complex syntax, colloquial vs. normative register, and speech.