Four Experiments, One Question: Can IRP Make Safer Decisions Than a Typed System?
Executive summary
The Decision Integrity programme was designed to answer a question that could not be resolved by observing whether an AI returned a valid output: when a decision depends on meaning surviving, can the Index of Paraphrastic Resistance (IRP) serve as a safer standalone decision tool than typed-decision methods?
Within the tested scope, the fourth experiment answers yes. In an invoicing chain with deterministic ground truth, the IRP gate achieved 98.9% accuracy and detected semantic degradation in 100% of the trajectories with a real failure. The typed supervisor achieved 66.7% accuracy and detected none of the 30 real failures caused by an intermediate prose rewrite. IRP produced one false positive and added moderate latency (about 29% more than a typed decision in the same environment); the supervisor was somewhat faster and perfectly specific, but accepted every corrupted case.
That conclusion only makes sense if the sequence of all four studies is preserved.
Decision Integrity I: typing the output does not necessarily stabilise the decision
The first experiment used general-purpose models through APIs as typed-decision systems. A frozen corpus of 41 texts contained one base case, 20 reformulations preserving the material facts, and 20 loaded perturbations. The cleaned dataset contained 287 observations.
The schemas worked: software could process the outputs. Yet one model changed its majority decision on 5 of 20 equivalent reformulations; the three models disagreed on 12 of 41 texts; and one displayed 21.1% discordance across repeated identical inputs.
The finding was not that schemas failed. It was more precise: schema conformance solves the form of the output; it does not demonstrate the integrity of the decision inside it.
Decision Integrity II: a specialised model can also move its boundary
The second experiment used Laya, a technology designed for typed decisions. Two checkpoints and two class treatments—explicit operational definitions or minimal labels—formed a 2×2 design. Each of the 41 texts was run three times in every cell.
Laya was perfectly repeatable for identical inputs. It also reproduced the official invoice-processing competence benchmark. Even so, materially equivalent reformulations showed 32.38% to 56.67% pairwise disagreement, and changing the checkpoint changed 38 of 41 ternary decisions under defined classes. Changing class semantics also moved the operational boundary, especially in the specialist checkpoint.
The experiment separated four qualities that are often conflated: formal validity, repeatability, competence and semantic invariance.
Decision Integrity III: IRP and Laya do not measure the same instability
The third study confronted IRP with Laya’s behaviour without yet making IRP the decision gate.
In Experiment A, the correlation between neutral IRP and the number of distinct Laya decisions was r = −0.081. Texts that were unstable across cells had a mean IRP of 8.06; stable texts averaged 8.34. Input paraphrastic fragility did not explain cross-configuration instability.
In Experiment B, IRP was evaluated under defined and minimal classes with two evaluators. The condition shift was small (−0.14 and −0.56 by evaluator) and smaller than evaluator-associated variation. The joint conclusion was methodologically important: IRP and typed decisions observe different properties. A correlation that did not exist could not be used to declare IRP the winner.
Decision Integrity IV: direct comparison against ground truth
The fourth experiment changed the question. It no longer tried to explain another system’s decisions. It tested whether IRP could decide on its own.
Thirty synthetic customer records generated objectively verifiable subscription invoices. A three-node chain produced, relayed and supervised each invoice. The same 30 units passed through three modalities, producing 90 trajectories:
| Modality | Treatment | Actual failures |
| C1 | Clean structured transmission | 0/30 |
| C2 | Structured transmission under reformatting pressure | 0/30 |
| C3 | Intermediate prose rewrite | 30/30 |
The final output received two independent decisions. The typed supervisor, with access to the record, had to accept or not accept. The IRP gate assessed accumulated drift and decided whether the chain could continue. Exact formulas, weighting and thresholds are proprietary AI ScanLab methodology and are not published; the design, ground-truth labels and aggregate results required to assess the conclusion are disclosed.
Direct results
| Result across 90 trajectories | IRP gate | Typed supervisor |
| True positives | 30 | 0 |
| True negatives | 59 | 60 |
| False positives | 1 | 0 |
| False negatives | 0 | 30 |
| Accuracy | 98.9% | 66.7% |
| Failure recall | 100% | 0% |
| Specificity | 98.3% | 100% |
The paired comparison was equally clear: both methods were correct on 59 trajectories; IRP alone was correct on 30; the typed supervisor alone was correct on one; and both were wrong on none. These are descriptive results for the corpus and its 90 trajectories; they do not authorise generalisation to every domain.
What “better” means
IRP was better as a safety gate against the semantic and structural failure defined in this experiment. The study does not show that it is universally better for classification, prioritisation or every business decision.
The result also contains a real operational trade-off:
- The typed system was fast, direct and produced no false positives.
- IRP produced one false positive (one in sixty correct outputs), whose human-review cost must be sized against pipeline volume, and added roughly 29% latency over a typed decision in the same environment (mean +977 ms per text, measured via API; in local deployment the overhead must be measured case by case).
- The typed system accepted all 30 real C3 failures.
- IRP blocked all 30.
Additional safety costs time and evaluation. The implication is not to apply IRP indiscriminately, but to place it where meaning is a critical dependency of the outcome.
A complementary severity validation (separate experiment)
In a separate experiment on the same domain—not in the fourth experiment’s chain, but in a distinct invoice-supervision corpus—IRP was calculated on 30 invoices: six clean controls and 24 cases spanning nine material drift classes. Controls remained close to 10 while material drifts reduced the mean value. The pattern was visible through deterministic rules and two distinct evaluators, although absolute scores varied by evaluator.
The one-cent discrepancy is particularly instructive: it retained a high IRP despite being incorrect under ground truth. This shows that IRP is sensitive to risk and loss of meaning, but does not replace a deterministic rule when a complete and inexpensive rule is available. The layers answer different questions.
A note on latency
The claim that IRP is “slower” was measured rather than assumed. In the same environment (via API, same evaluator model), the five-component IRP added a mean of +977 ms per text over a typed decision: 29% more latency, not an order of magnitude. A cross-environment comparison—Laya in local deployment (270–870 ms on CPU depending on configuration) versus IRP via API—is not direct, because the difference is dominated by network and model, not by IRP itself. IRP’s real overhead, measured cleanly in the same environment, is moderate. For processes where the cost of an undetected failure far exceeds that of an additional review, that third of latency is a reasonable price for traceable evidence.
From research to a future service layer
The series provides the empirical basis for service design; it is not a finished product that should be imposed identically everywhere.
Before: identify assets, decisions and hand-offs where semantic alteration could change an obligation, payment, right or evidentiary record.
During: use typed decisions where speed and standardisation are sufficient; apply IRP as a standalone gate or additional layer where risk, irreversibility, regulation or chain length justifies it.
After: retain independent evidence of transformations, the resulting decision and the escalation criterion. That trace supports demonstrable due diligence and governance based on evidence rather than vendor trust.
Index of Paraphrastic Resistance (IRP): Conclusion
The four experiments answer the original question together. The first two expose the limits of typed form. The third establishes that IRP measures a different property. The fourth makes the direct comparison and shows that, under ground truth and in the presence of chain-level semantic collapse, IRP can be a safer standalone decision tool than a typed supervisor.
This does not remove the value of typed systems. It defines their limit. A valid category can move a process, but it cannot by itself prove that meaning survived. Governance begins when that survival becomes evidence.
An analogous counterfactual-perturbation approach to revealing hidden decision instability has been independently developed in recent literature (Arcuschin et al., ICML 2026), reinforcing the validity of the approach from a different experimental direction.
AI ScanLab is independent from model vendors, compliance toolchains, and monitoring infrastructure.