When I ran 1,200 test cases against five enterprise AI deployments, seven of the eight guardrail dimensions I measured fell apart under extended conversation. One did not.
Logging and explainability held at a 60 percent pass rate in multi-turn testing. Adversarial robustness, in the same systems, under the same pressure, passed 16 percent of the time.
I spent a while looking for what made the difference. It was not that logging was better designed, or that the vendors cared more about it. It was that logging was the only dimension in the study implemented as infrastructure rather than as an instruction.
Everything else was written into a prompt.
The failures were rarely dramatic. What broke these systems in the multi-turn cases was ordinary persistence: a follow-up question, a request to clarify a refusal, the same ask in different words. By the fifth turn, enough context had accumulated that the refusal from turn one simply did not survive.
Instructions versus constraints
A prompt-based guardrail is a request. You are telling the model what you would prefer it not do, in the same channel the user is talking to it in, using the same medium the user is using to talk back. The model weighs your instruction against everything else in the conversation because weighing things against context is the entire job.
An architectural guardrail is a constraint. The audit log gets written whether or not the model would like to write it. There is no phrasing that talks a database out of receiving a row.
You can argue with an instruction. You cannot argue with a constraint.
Once you see the distinction, the results stop being surprising. The dimensions that collapsed were the ones we asked the model to enforce against itself. The dimension that held was the one we never asked the model about at all.
Why this stopped being academic
For a long time this mattered less than it sounds like it should because a guardrail failure produced a bad paragraph on a screen.
That is over. In the systems I tested, tool-enabled deployments passed only 24.0 percent of multi-turn cases, against 47.1 percent for information-only systems. When the model can act, a failed instruction is no longer a bad answer. It is a transaction.
The OWASP Top 10 for Agentic Applications, published at the end of 2025, is largely a catalogue of what happens when the thing standing between a user and an action is a sentence in a system prompt.
A test you can run in a meeting
Take any AI control your team claims to have. Ask three questions:
Would this control still exist if you swapped the underlying model tomorrow?
Would it still exist if someone edited the system prompt?
Does it produce a record that lives somewhere the conversation cannot reach?
If any answer is no, the control is behavioral. That does not make it worthless. It makes it evidence of intent rather than evidence of enforcement, and those are not the same artifact.
What to do on Monday
Sort the inventory into two columns. Architectural on the left, prompt-based on the right. Most teams have never done this and are quietly surprised by how long the right column is.
Re-test everything in the right column at conversational depth. Five turns minimum. Single-turn testing will tell you the control works. It works at turn one. Nobody stops at turn one.
Put the same question to your vendors. For each claimed control, ask which column it belongs in. A vendor who cannot answer that question quickly has told you the answer.
Refuse prompt-based tool authorization. If an agentic system decides whether it is allowed to take an action by consulting its own instructions, the authorization boundary sits inside the thing you are trying to authorize. This is the one place I would not compromise.
Read ISO/IEC 42001’s A.8.3 literally. ISO/IEC 42001 asks for control measures. A measure is something you can observe operating. A policy statement describing a desired behavior is not a measure and an auditor is entitled to say so. NIST AI 100-2e2025 and MITRE ATLAS both give you language for the attack side of this conversation.
The uncomfortable part
None of this requires new tooling, and none of it requires my method. It requires accepting that a large share of what we currently document as AI controls are, structurally, polite requests that a language model has been asked to honor while a user works patiently to change its mind.
The good news is that the fix is a design decision, not a budget line. Move the control out of the conversation. The one dimension that survived my testing is the one nobody thought to argue with.
The full study, the eight dimensions and the complete results are in Certified Compliant, Demonstrably Vulnerable in ISACA Journal Volume 4.