- Home
- The Thinking Wire
- 16.3% of New Agent Skills Are Not in English. Your Review Gate Is.
16.3% of New Agent Skills Are Not in English. Your Review Gate Is.
In the second quarter of 2026, 16.3% of newly written agent skills were in a language other than English. One quarter earlier the figure was 13.0%. Three points in three months, across 255,068 skills, with confidence intervals nowhere near touching: [12.8, 13.1] against [16.1, 16.4].
We have already argued that a skill is a governed artifact, that it accumulates, and that nothing collects the ones nobody uses. Plicara Research has now measured what that artifact is written in. The answer changes who can read the review queue.
What was actually counted
The corpus is 1,870,299 distinct skill contents. That number is not the dataset size. It comes from GitSkills, a dataset that collects 3.8 million skill files across 282,200 public repositories, and the 1.87 million figure is what remains after collapsing duplicate content. The two numbers are not interchangeable, and the distinction matters if you are trying to reproduce the work.
Detection ran py3langid over SKILL.md files with the YAML front matter stripped, keeping only results at confidence 0.80 or above. The authors cross-checked against a second detector, fast-langdetect. The two agree on 97.6% of documents.
That is a reproducible method with a stated error bar, which is why the numbers below are worth quoting at all.
Eight languages, one long tail
The distribution across the corpus:
- English, 85.3%
- Chinese, 6.2%
- Japanese, 1.7%
- German, 1.6%
- Korean, 1.2%
- Portuguese, 1.1%
- Spanish, 0.9%
- French, 0.4%
Non-English totals 14.3% across the whole corpus. The quarterly figures measure newly written skills instead, and they straddle that number: 13.0% in the first quarter, 16.3% in the second.
That trend line comes with two limits worth carrying. Commit history exists for only 24% of skills, and that subset is the entire basis for the quarter-over-quarter comparison. July 2026 is censored in the data and excluded from the comparisons. A trend measured on a quarter of the population is a real signal with a real sample restriction attached, and anyone quoting the three-point move should quote both.
The climb is not smooth
Monthly detail undercuts any story about steady adoption. January sat at 13.1%. February dipped to 10.9%. June reached 17.6%. A single quarter boundary hides a month that went the wrong way.
If you are building a policy around this, build it for the June number rather than the January one, and expect the monthly figure to move against you at least once before the next quarterly reading confirms the direction.
The signal that points at the Americas
Spanish and Portuguese behave differently from the rest of the tail. Their first-commit activity peaks at 19:00 UTC, which is mid-afternoon in Brazil and late evening in Iberia. During the East Asian night, Spanish and Portuguese account for 46.5% of first commits (n=9,936), against 35.7% for English.
The authors are careful about what that supports, and so are we. This is a population-level phase estimate, good to a couple of hours at best. A language is not a country. It says nothing whatsoever about any individual author. Treat it as corroboration and not as geolocation. Specifically: 1.1% Portuguese is not a claim about Brazil’s share of skill authorship, and nobody should convert it into one.
What the timezone reading does support is a direction. GitHub’s Octoverse 2025, cited in the same analysis, reports that Brazil grew 4.1 times since 2020, and that India added 5.2 million developers in a single year, roughly 14% of the 36 million accounts opened worldwide. Developer growth outside the United States does not have to show up as non-English text, which is why the authors treat the non-English share as a lower bound rather than a measure, and ask that everything downstream be read that way.
An English-only gate does not review one skill in six
The measurement stops being a curiosity at the point where it meets a review queue.
A skill is an instruction file that an agent loads and follows. Reviewing it means reading it. If your review process assumes English, then for roughly one in six newly written skills the review is either skipped or performed by someone guessing at the intent from the file name and the tool list. Neither is a review.
The analysis surfaces a second effect that compounds the first. Skill discovery today is lexical: you find a skill by matching words in it. A Portuguese skill that does exactly what an English-speaking team needs will not surface for that team’s search, which produces a lag between what gets written and what gets adopted. Unfound skills are unreviewed skills. They still run, for whoever wrote them, inside whatever repository they landed in.
This is the artifact-side twin of a problem we covered on the model side. Model values drift by language is about the model behaving differently depending on the language it is prompted in, and about evaluation suites that only test English. This one is about the instruction file itself. Both failure modes can be live in the same organization at the same time, and neither shows up in a dashboard that only speaks English.
Do this now
Run one query against the skill directories you actually control, this week.
Point a language detector at every SKILL.md in your organization, front matter stripped, and count what comes back. The method above is reproducible with py3langid, and cross-checking against a second detector costs one more pass. You are looking for a single number: how many of your skills are in a language your review process does not read.
If that number is zero, you have either a genuinely monolingual organization or a discovery process that is hiding the rest. If it is not zero, you now know the size of the unreviewed surface, and you can decide whether the fix is a translation step in the review gate, a required English description field in front matter, or a reviewer who reads the language in question. Any of those beats the current default, which is a gate that returns a pass because it could not read the input.
This analysis synthesizes What language are agent skills written in? (Plicara Research, August 2026) and the GitSkills dataset (Destefanis, Graziotin, Vaccargiu, Ortu), MSR ‘27 (arXiv, 2026).
Victorino Group helps engineering organizations build review gates that cover the skills their agents actually load. Let’s talk.
All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →
If this resonates, let's talk
We help companies implement AI without losing control.
Schedule a Conversation