Washington Will Make Them Publish How Often the Guardrail Failed

TV
Thiago Victorino
7 min read
Washington Will Make Them Publish How Often the Guardrail Failed

Starting January 1, 2027, a company operating an AI companion chatbot in Washington State has to publish how many crisis-referral notifications it sent in the previous year. What the law demands is a count. A description of the safety protocol does not satisfy it.

That single reporting line, buried in HB 2225, is the most operationally interesting thing any US state has done to AI governance so far. Venita Subramanian, writing for Tech Policy Press, frames the design principle behind it in one sentence: “Regulatory regimes should be built to expose imperfections, not describe perfect guardrails.”

Every AI governance program we have reviewed does the opposite. It documents controls. It maps them to a framework. It produces a policy document that describes what should happen when a model behaves badly. What almost none of them produce is a number telling you how often that actually happened.

What the Statute Actually Requires

HB 2225 applies to AI companion chatbots, which the law defines as systems “that provide adaptive, human-like responses and sustain relationships across multiple interactions.” The scope is narrow on purpose. The obligations are specific enough to audit.

The disclosure rule is one of them. The chatbot must state that it is not human at the beginning of an interaction and at least every three hours during continued use. For minors, that disclosure repeats at least every hour. A fixed cadence is a testable property. You can instrument it, sample it, and prove the interval held.

The design rule is another. Companies must prevent manipulative engagement techniques involving minors, which the statute enumerates: prompting a child to return for companionship, creating emotional attachment, promoting isolation, or encouraging secrecy from trusted adults. Those are product mechanics, described in product language. An engagement loop optimized for retention is now a compliance surface.

Then the crisis rule. Companies must run protocols for detecting suicidal ideation and self-harm, refer users to crisis resources, and prevent the chatbot from encouraging self-harm. They must publicly describe those protocols. And they must report how many crisis-referral notifications they issued in the previous year.

The first three obligations are the kind of thing a policy document handles well. The last one is not. A count cannot be written aspirationally.

Duty of Care Is the Spine

Subramanian anchors the whole argument in a legal concept older than any AI framework: duty of care, defined as “the legal obligation to behave in a reasonably safe manner and not a reckless one.” What it demands, in her reading, is transparency, guardrails against harmful validation, human crisis-response resources, meaningful escalation pathways, and accountability when safeguards fail.

Read that list against how most AI governance is assembled. Four of the five items describe things a company builds. The fifth describes what happens after the build fails. That fifth item is where nearly every program we see stops short, because it is the only one that requires admitting a failure occurred, with a number attached, on a schedule.

Duty of care is a useful spine precisely because courts already know how to reason about it. It does not require a regulator to define what a safe model is. It requires the operator to behave reasonably given what it knew, and it makes the record of what it knew discoverable.

Why the Counter Is Enforceable and the Policy Is Not

Consider what a regulator can do with a published control description. It can read it. It can compare it to a framework. It can ask whether the described control exists. Every one of those checks is satisfied by writing better prose.

Now consider what a regulator can do with a published crisis-referral count. It can compare this year to last year. It can compare one operator to another operator of similar scale. It can ask why a platform with millions of minor users referred almost no one. It can ask why the count collapsed in the quarter after a model upgrade. Each of those questions has a factual answer that prose cannot supply.

The count also changes internal behavior before any regulator arrives. A number that will be published creates an owner, a measurement pipeline, a definition dispute, and an argument about edge cases. Those arguments are the actual governance work. We made a version of this point about verification you cannot see: a control that produces no observable artifact is indistinguishable from a control that does not run.

This is also why the disclosure obligations, on their own, do not carry the weight. Character.AI’s disclosure read: “Remember: Everything Characters say is made up!” The case that put the company in front of a court involves the February 2024 death of 14-year-old Sewell Setzer, with a lawsuit filed by his mother. A disclosure line was present. What was absent was any published record of how often the system detected and escalated a user in crisis.

Washington Is More Prescriptive Than What Came Before

New York and California enacted comparable laws in 2025. Oregon followed in March 2026. Washington goes further in prescriptiveness, and the direction of travel across four states is hard to read as an isolated event.

For anyone building agentic systems outside the companion-chatbot category, the specific scope is less important than the mechanism. Once a legislature learns that a reportable failure count is enforceable and a control description is not, that instrument gets reused. Fraud detection, medical triage, hiring screens, customer escalation: every one of them has a guardrail whose firing rate is currently known only to the operator, if it is known at all.

The proposal Subramanian builds on top of the statute extends the logic to evaluation itself: publish evaluation criteria and failure rates. That is a materially harder ask than publishing a referral count, and no jurisdiction has legislated it. As a design principle for an internal program, it costs nothing to adopt today.

Do This Now: Pick One Guardrail and Count It

Choose the single guardrail in your system that exists to prevent the worst outcome your product can cause. One, out of the entire control catalog.

Then answer four questions in writing this week:

  1. How many times did it fire in the last 12 months? If the answer is “we do not instrument that,” you have found the real state of the control.
  2. What is the definition of a firing? Two engineers will disagree. The disagreement is the finding.
  3. How many times should it have fired and did not? This number is harder and usually requires sampling. An honest estimate with a stated method beats a precise number with no method.
  4. Who would own the count if it had to be published on a fixed schedule? If no name fits, the control has no owner today either.

Publish the answers internally, on a page with a date on it, and repeat in a quarter. The comparison between the two dates is worth more than either number alone. It is also the artifact you will need if a regulator, a customer’s procurement team, or a plaintiff’s lawyer asks what you knew and when.

Governance programs that describe perfect guardrails will keep passing internal review and keep failing the first external test. The ones that publish their own failure counts get to argue from evidence. As we argued in trust is the UX and in the confidence convention, the credibility of an AI system is built from what it exposes about its own limits. Washington just wrote a narrow version of that into law, with a deadline.


This analysis synthesizes A UX Design Perspective on Improving AI Safety by Venita Subramanian, Aspen Policy Academy Science and Technology Policy Fellow (Tech Policy Press, August 2026).

Victorino Group helps engineering organizations turn AI control descriptions into measured, reportable failure rates. Let’s talk.

All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →

If this resonates, let's talk

We help companies implement AI without losing control.

Schedule a Conversation