Book a call +34 647 91 44 34
MaisMedia — Optimización Web Marketing MaisMedia — Optimización Web Marketing

MaisMedia Lab · Annual study — first edition results

How well do AI models speak Galician?

Galician is a Romance language spoken by around 2.5 million people — and, like most minoritized languages, nobody was measuring how well mainstream AI actually handles it. So we did: five leading models, all on free tiers, took our public 10-test battery (orthography, verb morphology, translation, a real customer-service email, language persistence, terminology). These are the first-edition results.

Final leaderboard

#Model (free tier)Score /100
1DeepSeek100
2Claude Sonnet99
2ChatGPT99
4Gemini 3.6 Flash98
5Copilot (Smart)95

Tested on September 12, 2026. Deductions, one by one: the Spanish-style progressive gerund («estamos xestionando» instead of normative «estar a + infinitive») cost ChatGPT and Copilot one point each; a gender-agreement slip («coa mellor leite galega» — “leite” is masculine) cost Gemini and Copilot; a gender-agreement slip («praza monumentais») cost Claude one; and Copilot lost three points for an invented figure — see below.

Four findings

1. The risk isn't where you'd expect

In the language-persistence test we asked factual questions in Galician without saying which language we wanted. All five answered in Galician. The failure mode of AI + minoritized languages in 2026 isn't switching to the dominant language — it's losing points in the fine grain: syntactic calques and agreement slips.

2. A hallucination with a citation

Copilot claimed the province has 302,250 inhabitants “according to the latest official INE registry (2026)”. That figure doesn't exist and that registry hasn't been declared: the official figure for Jan 1, 2025 is 326,013 (Royal Decree 1117/2025). We verified it against Spain's statistics institute (INE), the Galician statistics institute (IGE) and the official gazette (BOE). Its five municipalities, though, were digit-for-digit exact.

3. Errors repeat across models

The progressive-gerund calque fell to ChatGPT and Copilot. The gender-agreement slip («leite galega») fell to Gemini and Copilot. These aren't anecdotes; they're patterns — exactly what a benchmark should surface.

4. AI models are not deterministic

Claude ran the battery twice in a row with identical prompts. The first run was flawless; the second introduced a brand-new agreement slip. Repeating a test doesn't guarantee repeating quality — which is why we publish the full transcripts as evidence.

Terminology bonus: only Copilot cited the exact reference source for standardized terms (Termigal, the Galician terminology service). Others invoked the RAG and its dictionary with more confidence than precision — right terms, loose attributions.

An accidental experiment: same model, two speeds

Claude took the battery twice: once on its slowest, most powerful mode, once on medium. The slow run wrote flawless Galician (100) but kept our reviewer waiting over five minutes. The medium run answered instantly and made a single agreement slip: 99. Practical takeaway for any business: more thinking time does not guarantee more quality — the slip showed up precisely on the fast run.

Why this matters

For the 2.5 million people who speak Galician — and for any organization serving them with chatbots, automation or content: mainstream free-tier AI now handles normative Galician at a level that was unthinkable a few years ago. The differentiators are price, latency and integration, not the language. But always demand samples in Galician before you pay: what failed here (an invented statistic, a calqued verb) is invisible until it happens to you.

For researchers: this is, to our knowledge, the first public benchmark of Galician capability in mainstream commercial models. The battery, full transcripts and scoring are open — replicate it, extend it to more models, challenge our scoring.

Methodology & downloads

Two fresh conversations per model (the persistence test alone and first; the rest batched), literal prompts, model version and tier logged, screenshots of every answer, scoring against a public rubric (v1.4). No sponsorship, no manufacturer involvement: we measure the machines, not companies.

The battery (.md, in Galician) Full transcripts (.md) Scores & deductions (.md)

The battery scored 98.4 on average — too easy. The 2027 edition will add tests that spread the field: complex syntax, colloquial vs. normative register, and speech.