Who Authorizes Your AI? Increasingly, It's HR.

TV
Thiago Victorino
8 min read
Who Authorizes Your AI? Increasingly, It's HR.

Rippling now decides which AI models an employee can call by reading that employee’s attributes out of the HR system. Not a group in Okta. Not a project tag in the cloud console. The employment record: role, department, level, manager. If the org chart says you are a senior backend engineer in the platform group, the gateway lets your request through to the expensive model. If it says you moved to support last month, it does not.

That shipped in August 2026, in the same fortnight that Coinbase published a year of rebuilding its engineering interview around AI supervision, and a pre-registered field experiment across more than 70,000 job applicants reported an AI interviewer beating human interviewers on offers, starts, and one-month retention. Three artifacts, three companies, no coordination. All three moved a piece of AI authority into the people system.

The Hiring Filter Now Tests Supervision

Coinbase reports that AI-generated code went from 5.7% of everything merged in Q1 2025, past 50% in Q4 2025, to what the company calls “roughly 100%” today. Humans still review all of it. The interview process had to be rebuilt to match, and the reasoning in the post is unusually blunt: “We want the bar we interview against to reflect the bar we operate against. We can’t hire engineers to work alongside AI if we’re still selecting for the ability to work without it.”

The resulting rubric scores AI Fluency on three dimensions: Usage, Application, and Understanding Limits. The first two are what most companies would guess. The third is the governance one. It asks whether the candidate identifies privacy and security implications in what the model produced, and whether they apply human judgment as a guardrail rather than accepting output. That is a control competency being tested at the door, before the person has any system access at all.

Coinbase also did something most hiring teams never do: it measured its own process. An internal analysis found an 84% correlation between two interview rounds, which is a polite way of saying one of them was redundant, and the redundant one was cut. New hires get pulse surveys at 45 and 90 days, with a stated goal that 100% of them maintain or improve their AI Fluency rating after joining. The rollout ran in three phases, starting with a frontend pilot in the second half of 2025, extending to backend in January 2026, and going company-wide in March 2026.

An engineering leader quoted in the post frames the economics: “When the cost of building goes to zero, the cost of identifying what to build, verifying it’s correct, and getting it out safely becomes the limiting factor.” If that is true, the hiring filter is the first place the verification capacity of the organization gets set. Everything downstream inherits it.

The Interviewer Itself Became a Model

The Philippines experiment is the piece that will make people uncomfortable. In a pre-registered 2025 field study (SSRN abstract 5395709), more than 70,000 applicants for entry-level customer-service roles were randomized: 20% to a human interviewer, 60% to an AI interviewer, and 20% given the choice. The AI-interviewed candidates were 12% more likely to receive an offer, 18% more likely to start the job, and 17% more likely to still be employed a month later. Among those given a choice, 78% picked the AI.

Two qualifiers carry real weight here. First, humans made the hiring decision in every arm. The model conducted the conversation; it did not sign. Second, MeasuringU’s write-up notes that the AI interview questions were “mostly closed, factual, verification-style questions, things like commute time, salary fit, availability, and contact info.” That is screening, and screening is exactly the task structured enough for a model to do well.

The boundary is visible elsewhere in the same research thread. NN/Group tested ten experienced researchers against two AI moderators and found the AI adequate for structured interviews and not yet adequate for semi-structured ones. And 419 research professionals have signed letters opposing the use of AI for qualitative research, which is a strong signal about where practitioners believe the competence ends.

We have argued before that correlated hiring models create algorithmic monoculture, and that risk has not gone away. What the randomized design adds is inconvenient: on retention, at scale, in a structured screening task, the model outperformed. A governance position that rests on “the AI interviewer is worse” now needs a different foundation. The defensible position is about scope and about who signs, not about capability.

Spend Policy Keyed to the Org Chart

Rippling’s own numbers, self-reported and without published methodology, describe the problem that produced the product. AI token spend was growing 80% month over month, on a trajectory toward 40% of the R&D headcount budget in year one and 90% the year after. Between 10% and 15% of employees drove roughly 60% of the spend. One engineer was at $50,000 per month.

The AI Spend Console is what they built, and the design matters more than the savings figure. Their gateway enforces three things at request time: per-tool monthly dollar caps, automatic routing to approved LLMs, and model-access policy keyed to employee attributes. Spend came down to 10-15% of headcount budget. They also nominated 20 “AI Captains” across the organization, and their summary of the whole exercise is worth keeping: “The constraint created the innovation.”

That third enforcement point is the one to sit with. Model-access policy keyed to employee attributes means the entitlement decision is a function of HR data. When someone changes teams, their model access changes, because the employment record changed. When a contractor’s classification is wrong in the HRIS, their entitlements are wrong in the gateway. The org chart became a policy input evaluated at request time.

One thing Rippling explicitly does not do: their mapping of spend to outcomes, run against performance ratings and PR volume, only reports. Nothing gates on it. That restraint is the right call today, and it is also the exact place where the next line will be crossed by someone less careful.

The Direction of Travel

We wrote earlier that credit governance is the working template for AI spend policy. In that model, finance grants the authority: a limit, an approver, a statement. What these three artifacts show is a different grantor. Who gets hired, who evaluates them, and which models they may spend on are now all decided by data that lives in the people system.

Almost nobody governs that system as a control plane. HRIS data was designed for payroll accuracy, headcount reporting, benefits eligibility, and compliance filings. It was built with the tolerances those uses allow. A stale department field for three weeks after a transfer is a rounding error in a headcount report. It is an authorization defect in a model gateway. Job-title free text that HR business partners edit by hand is fine for an org chart and dangerous as a policy predicate. Contractor and vendor records that were never meant to carry entitlements now carry them.

The failure mode is not dramatic. It is a promotion that silently grants access to a frontier model with no spend ceiling, or a transfer that silently keeps access it should have dropped. Neither generates an alert, because the HR system does not know it is now an authorization system. The same logic we applied to agents as headcount with OKRs and budgets runs in reverse here: the people system started making machine decisions before anyone gave it machine-grade controls.

What to Do This Month

Pull the list of attributes your AI gateway, model router, or agent platform reads from the HRIS. If nobody can produce that list in ten minutes, that is the finding, and the rest of this exercise is theoretical until it exists.

Then, for each attribute on it, answer four questions. Who can change this field, and does that person know it grants model access? How stale can it get, measured in days between the real-world event and the record update? What happens on the null or default value, and does that fail open? Is a change to this field logged where a security team would look, or only in the HR audit trail?

Take the same list to whoever owns your interview loop. If your hiring process now screens for AI supervision competence, as Coinbase’s does, that rubric is a control. It deserves version history, a named owner, and periodic review, the way an access policy does. If an AI conducts any part of that screening, write down which questions it is allowed to ask and who signs the decision, and treat the boundary between structured and semi-structured work as the line the SSRN results actually support.

Organizations that spent the last two years building AI governance for engineering will find the HR conversation harder, because the owners are different, the vocabulary is different, and the data quality assumptions were set by a very different set of requirements. Start anyway. The entitlement source already moved.


This analysis synthesizes Interviewing Engineers in the AI Era: Lessons from a Year of Rebuilding (Coinbase, July 2026), Can We Trust AI to Moderate UX Interviews? (MeasuringU, Jeff Sauro, Lucas Plabst and Jim Lewis, August 2026), and Introducing AI Spend Console (Rippling, Whitney Zack and Catalina Zhao, August 2026).

Victorino Group helps organizations audit the HR attributes their AI platforms now treat as authorization inputs. Let’s talk.

All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →

If this resonates, let's talk

We help companies implement AI without losing control.

Schedule a Conversation