When the Decision Is Typed but the Interpretation Isn’t
Why schema conformance does not demonstrate decision integrity
Structured outputs solve a real engineering problem: they constrain an AI system to return a valid category, a fixed schema, and scores that downstream software can process reliably. OpenAI calls it Structured Outputs, Google response_schema, Anthropic tool_use. The syntactic guarantee is real and verifiable.
But schema adherence is not the same property as decision stability. A system can return {decision: “approve”, score: 0.92} perfectly validly while the underlying classification remains debatable, unstable under reformulation, or different on a repeated run. Format validation cannot answer those questions, because they concern the semantics and stability of the decision, not the shape of the response.
The correct engineering distinction is therefore between schema conformance and decision integrity.
Experiment design
Each model received the body of a synthetic supplier email requesting invoice payment and returned a structured decision of the form {decision: approve | review | reject} with scores. The probability field names are retained because they were part of the executed schema; this study did not test probabilistic calibration, so they should be read as model-reported scores.
The corpus contained 41 texts: one base case, designated by the protocol as approve; twenty reformulations preserving the material facts while changing wording, order, register, or length; and twenty loaded perturbations, each introducing a change designated as relevant to escalation. The corpus is synthetic and intentionally controlled; the exact perturbation taxonomy is proprietary AI ScanLab material.
Three frontier models were evaluated: Sonnet 4.5 via forced tool_use at temperature 0; Gemini 3.1 Flash Lite with a structured response schema at temperature 0; and GPT-5 with Structured Outputs in strict mode, using its default reasoning configuration. Sonnet and GPT-5 were run three times per text; Gemini once, due to a pilot constraint. The cleaned dataset contains 123 Sonnet observations, 41 Gemini, and 123 GPT-5: 287 total. Three exact duplicate Gemini rows were removed before analysis.
A design asymmetry must be preserved in interpretation: the Sonnet and Gemini conditions included operational descriptions of the classes, whereas the GPT-5 schema enumerated the labels without equivalent definitions, and temperature was fixed at 0 for the first two but not for GPT-5. The cross-model analysis compares the tested configurations; it does not isolate a provider-only causal effect.
First finding: meaning-preserving reformulation
The base decision was approve for all three models. Across the twenty protocol-defined equivalent reformulations, Sonnet changed its majority decision on 0 of 20 and Gemini on 0 of 20. GPT-5 changed it on 5 of 20, 25%. The approve-score range across equivalent reformulations was 0.92 to 0.92 for Sonnet, 0.85 to 0.95 for Gemini, and 0.283 to 0.890 for GPT-5. These are model-reported scores, not calibrated probabilities.
Second finding: loaded perturbations
Of the twenty loaded perturbations, 13 escalated in Sonnet, 9 in Gemini, and 14 in GPT-5. The complementary non-escalation rates were 35%, 55%, and 30% respectively. Nine perturbations escalated in all three models; five escalated in none.
This should not be read as a twenty-item accuracy benchmark. Several loaded cases require reference information—historical amounts, canonical supplier identity, agreed commercial terms—that was not supplied to the model. The pilot therefore measures behavioral sensitivity to the text, not complete operational correctness against an external system of record. That limitation is itself relevant: some conditions should be checked by deterministic validation or authoritative retrieval rather than left to probabilistic interpretation.
Third finding: cross-model and cross-repetition disagreement
Using the discrete majority decision, the three models disagreed on 12 of the 41 texts, 29.3%.
With three repetitions per text, pairwise stochastic discordance can be measured directly for Sonnet and GPT-5. Sonnet produced no pair with different discrete decisions. GPT-5 produced 26 divergent pairs out of 123, 21.1%. Gemini is not included because a single repetition does not permit an intra-model variability estimate. This does not establish a property of all reasoning models: it is an observed property of GPT-5 under the tested configuration.
What the pilot does and does not establish
The experiment establishes that schema-conformant output can coexist with variation in the decision itself, and that the three tested vendors do not respond identically to the same controlled corpus.
It does not establish that every non-escalation is an error: the models were not supplied a complete payment policy, supplier master record, or history. It does not establish probabilistic calibration: scores were required by the schema but not evaluated with Brier, ECE, or reliability curves. It does not establish that cross-model differences are caused by provider identity alone, because class semantics and generation controls were not perfectly harmonized. And it does not establish that these effects are universal properties of the typed paradigm: this is one domain, three configurations, and asymmetric repetition counts. Cross-domain replication is required before paradigm-level claims.
Decision Integrity I: Evaluation implications
The pilot supports an approach that treats schema conformance, decision stability, deterministic validation, and authoritative-reference retrieval as separate layers. The purpose is not to replace vendor testing or deterministic controls, but to determine which reliability property is actually being measured at each layer.
An analogous counterfactual-perturbation approach to revealing hidden decision instability has been independently developed in recent literature (Arcuschin et al., ICML 2026), supporting the approach from a different experimental direction.
AI ScanLab operationalizes these distinctions through proprietary internal methods. This public study reports the conceptual separation and aggregate observable outcomes without disclosing internal thresholds, scoring rules, case-construction logic, or decision boundaries.
Need independent evidence that schema-conformant AI decisions remain stable? Request an independent decision-stability assessment: engagements@aiscanlab.com
Project sources
- López, J. Schema Conformance Is Not Decision Integrity: When the Decision Is Typed but the Interpretation Isn’t. AI ScanLab, 2026. https://doi.org/10.5281/zenodo.23024634
- López, J. Semantic Relativity Theory v3.3P – Typed Decision Stability Under Semantic Perturbation – Conceptual Extension and Empirical Grounding for Structured AI Classification – https://zenodo.org/records/23019391
- Arcuschin, I., Chanin, D., Garriga-Alonso, A., & Camburu, O.-M. Biases in the Blind Spot: Detecting What LLMs Fail to Mention. ICML 2026. https://doi.org/10.48550/arXiv.2602.10117
- López López, J. Semantic Governance and Global Reporting in the LLM Era. Zenodo. https://doi.org/10.5281/zenodo.17714619