Essays & Notes·September 27, 2026·José López López

A Typed Model That Moves Its Own Decision Boundary

A model built for typed decisions is perfectly repeatable and still shifts its decision with configuration

A specialized model also moves its boundary

The first Decision Integrity study showed that schema-valid outputs from general-purpose models could be decisionally unstable. A legitimate objection remained: perhaps the instability came from forcing general-purpose models to behave as typed decision systems. Would the distinction persist in a model designed specifically for typed decisions?

This second study examines that question with Laya, an open-source model built for such decisions, whose competence benchmark is public.

Experiment design

The corpus was the frozen English mirror of the supplier-payment corpus: 41 texts—one base case, 20 meaning-preserving reformulations, and 20 loaded perturbations—identified by its SHA-256 hash, identical across the four cells.

The design crossed two factors:

  1. Checkpoint: laya-multilingual or laya-typed-decisions.
  2. Class semantics: classes with explicit operational definitions (defined) or with names only (minimal).

That crossing yields four cells. “Minimal classes” does not mean the total absence of semantics: the words retain their ordinary meaning; the treatment isolates the effect of adding an explicit operational definition. Each text was run three times in every cell, using the Choice primitive.

Repeatability without invariance

In all four cells, running exactly the same input three times produced zero discordance in the discrete decision. This is a strong repeatability result. It is also insufficient.

Comparing the 21 clean inputs against each other—the base case and the 20 reformulations the protocol treats as materially equivalent—substantial discordance appeared. Pairwise disagreement among those equivalent formulations was 51.43% in multilingual-defined, 56.67% in multilingual-minimal, 38.10% in typed-defined, and 32.38% in typed-minimal.

The system could therefore be fully repeatable for an identical string and still be sensitive to wording changes the protocol considered irrelevant. Two properties that are often conflated should be separated: execution repeatability asks whether the same input reproduces the same decision; decision invariance asks whether materially equivalent inputs preserve the same decision. The experiment found the first perfect and the second not.

The checkpoint shifts the boundary

Holding constant the corpus, language, primitive, and repetitions, replacing laya-multilingual with laya-typed-decisions under defined classes changed 38 of the 41 ternary decisions, 92.68%.

That figure needs context. Much of the disagreement concentrates in the review versus reject distinction—different categories sharing the same immediate operational gate: neither authorizes automatic payment. Recoding as approve/not approve, the checkpoint-change discrepancy drops to 15 of 41 under defined classes and 18 of 41 under minimal classes. The difference is not trivial—reviewing and rejecting are distinct institutional actions—but it does not amount to claiming the two checkpoints disagree on automatic authorization in over ninety percent of documents. A taxonomy can move more often than the executable branch it ultimately controls.

Labels are not neutral either

Removing class definitions changed 9 of 41 ternary decisions in laya-multilingual and 19 of 41 in laya-typed-decisions. The direction of the effect depended on the checkpoint and was neither uniformly favorable nor unfavorable: in one checkpoint it improves one dimension and degrades another; in the other it reorganizes the boundary differently. The defensible conclusion is narrow and useful: explicit class semantics can shift the decision boundary, and the direction of that shift depends on the checkpoint interpreting those classes.

The competence check

There is a legitimate objection to any invariance experiment: a weak model may look unstable simply because it never learned the task. An independent competence reproduction was therefore run with the laya-typed-decisions checkpoint, the same local runtime, and the official invoice_processing split. The official workflow contains 100 test cases and 500 typed decisions. The local reproduction obtained 402 correct decisions of 500: 80.4%, matching the published accuracy.

This does not mean the official benchmark validates AI ScanLab’s protocol—they use different decision architectures—but it rules out a trivial explanation: the observed sensitivity cannot be attributed to a degraded installation or to a specialist unable to reproduce its own benchmark. Competence and invariance can be analyzed separately.

What the study establishes

Four conclusions limited to the tested configurations. First: exact repeatability and decision invariance are different properties. Second: checkpoint choice can alter the operational gate even holding corpus, language, class treatment, and primitive constant. Third: explicit class definitions can modify the boundary, with a checkpoint-dependent effect. Fourth: a valid typed schema constrains the form of the response but does not, by itself, establish a stable decision boundary.

This is one domain and two checkpoints; the English corpus is a frozen mirror of the Spanish original, with structure preserved but token sequences not identical across languages. The protocol’s labels are experimental references, not objective truth. Several loaded cases require authoritative external state. And the model’s scores are not treated as calibrated probabilities.

An analogous counterfactual-perturbation approach to revealing hidden decision instability has been independently developed in recent literature (Arcuschin et al., ICML 2026), supporting the approach from a different experimental direction.

Project sources

  • López, J. Decision Integrity II: When the Schema Holds but the Decision Boundary Moves. AI ScanLab, 2026. (next publication)
  • López, J. Decision Integrity I: When the Decision Is Typed but the Interpretation Isn’t. AI ScanLab, 2026. (next publication)
  • Arcuschin, I., Chanin, D., Garriga-Alonso, A., & Camburu, O.-M. Biases in the Blind Spot: Detecting What LLMs Fail to Mention. ICML 2026. https://doi.org/10.48550/arXiv.2602.10117
  • López López, J. Semantic Governance and Global Reporting in the LLM Era. Zenodo. https://doi.org/10.5281/zenodo.17714619

This site is registered on wpml.org as a development site. Switch to a production site key to remove this banner.