Essays & Notes·September 27, 2026·José López López

Paraphrastic Resistance: Text Fragility is not Decision Instability

Why Semantic Robustness and Decision Stability Need Separate Evidence

A system can preserve the format while moving the decision

Many evaluations of AI decision systems stop too early. If the response complies with the schema, the class belongs to the permitted set, and repeated runs produce similar outputs, the system is treated as stable.

Those tests describe only one configuration. They do not show what happens when the checkpoint changes, when labels receive different operational definitions, or when another evaluator interprets the same textual transformation.

AI ScanLab examined that separation in three stages. Decision Integrity I studied general-purpose models acting as typed decision systems. Decision Integrity II studied Laya through a 2×2 factorial design. Two later cross-experiments, A and B, connected that material to the Index of Paraphrastic Resistance (IRP).

The purpose was not to put a semantic metric and a decision model into a contest. It was to determine whether two phenomena that are often conflated—text fragility and decision instability—could be separated through evidence.

What the Laya experiment measured

The shared corpus contained 41 English texts: one base text, 20 semantically equivalent variants, and 20 loaded variants. The Laya study combined two factors:

  1. Checkpoint: laya-multilingual or laya-typed-decisions.
  2. Class semantics: operationally defined classes or minimal labels.

Their combination produced four experimental cells. These were not four native modes offered by Laya, but four conditions designed to isolate each factor.

Decision Integrity II had already shown that formally valid output does not guarantee a stable decision boundary. The new experiments added the missing question: do decisions change because some texts are less able to survive semantic transformation?

Experiment A: neutral Index of Paraphrastic Resistance (IRP)does not explain decision diversity

IRP predates this series. It estimates a text’s resistance to paraphrase: how much of its intent, constraints, and architecture of meaning survives reformulation. It does not produce approve, review, or reject.

Experiment A assigned a neutral IRP score to each of the 41 texts and crossed it with the four Laya decisions.

The central result was r = −0.081 between IRP and the number of distinct categories assigned to a text. The 38 texts that did not retain one decision across all four cells had a mean IRP of 8.06; the three stable texts had a mean of 8.34.

The difference is small and the linear correlation is close to zero. In this dataset, cross-cell instability therefore cannot be explained as a consequence of low paraphrastic resistance.

Breaking the result down by factor shows where attention is required:

Isolated changeExact decisions changed
Defined → minimal classes, multilingual checkpoint9/41
Defined → minimal classes, typed-decisions checkpoint19/41
Multilingual → typed checkpoint, defined classes38/41
Multilingual → typed checkpoint, minimal classes32/41

The checkpoint was the dominant factor in this corpus. Class definitions also changed the output, however, and their effect depended on the checkpoint. Referring only to “architecture” would hide part of the result; referring only to “text semantics” would do the same. What the evidence shows is configuration sensitivity.

Experiment B: semantic measurement also has a context

Experiment B brought the class factor into the Index of Paraphrastic Resistance (IRP) procedure itself. The same texts were assessed under two frames:

  • IRP-defined: the evaluator received operational definitions for approve, review, and reject.
  • IRP-minimal: the evaluator received only the class names.

Two Gemini-family evaluators were used. The aim was to measure both the effect of class context and the effect of replacing the evaluator.

All 41 texts were scored. The treatment effect was estimated on the 40 variants; the base case was retained as a reference and excluded from these means:

EvaluatorDefined IRPMinimal IRPDefined − minimal
E17.797.93−0.14
E26.637.19−0.56

The class treatment produced a modest but evaluator-dependent shift. The evaluators also preserved a broadly similar ordering of the texts: r≈0.68 under defined classes and r≈0.67 under minimal classes.

The global average, however, hides a by-family interaction. In the equivalent variants the class definition barely moved IRP (Δ≈+0.03), whereas in the loaded variants minimal classes raised it by an average of 0.73 points over defined classes (Δ≈−0.73; between-family difference at p≈0.006). That is: class semantics does not alter the evaluation of equivalent reformulations, but it does change how texts containing a material modification are interpreted.

The rigorous interpretation is not that context “does not matter.” Evaluator 2 responded more strongly than Evaluator 1. Nor can the study claim independence across model architectures, because both evaluators belong to the same family. The useful result is that IRP retains meaningful ordinal agreement while its absolute level depends on the evaluator and, to a lesser extent, the decision frame.

The distinction organisations need

Together, the experiments provide an operational separation.

1. Semantic robustness of the input

This asks whether material preserves meaning when it is summarised, paraphrased, translated, retrieved, or passed between components. A resistant text may preserve its semantic structure even when its linguistic surface changes.

2. Decision integrity

This asks whether the system converts that meaning into a stable decision under defensible configurations. An output can be repeatable within one cell and change when the checkpoint or class semantics changes.

One property does not substitute for the other. A high IRP does not guarantee the same decision. A repeatable decision does not prove that the input preserved all of its meaning.

From observation to a future evidence service

These findings provide the basis, not the launch, of a future AI ScanLab evidence layer. Its design must distinguish three moments:

  • Before the decision: evidence that the input preserved intent, constraints, and meaning.
  • During the decision: evidence of sensitivity to checkpoint, taxonomy, class definitions, and evaluator.
  • After the decision: evidence that the output retained the category, application conditions, and traceability required for justification.

The goal is not to replace the model that decides or turn IRP into a decision engine. It is to let an organisation document which property it tested, under which configuration, and within which limits. That documentation becomes especially important when the output controls access, approval, prioritisation, compliance, hiring, or another material outcome.

The evidence must also be independent. AI ScanLab is independent from model vendors, compliance toolchains, and monitoring infrastructure. The system producing a decision should not be the only source validating it.

What these results do not say

The study does not show that IRP “decides better” than Laya. It lacks external ground truth for ranking the two methods by decision efficacy. Nor can it generalise to every domain, checkpoint, or model. The corpus contains 41 texts from one domain, only three were stable across all four cells, and both evaluators in Experiment B belong to Gemini.

The time and economic cost of each layer also remains to be benchmarked. IRP generates richer semantic evidence, but it should not be assumed to have the same latency as direct classification.

What has been established is enough to change an evaluation protocol: semantic stability of the text and stability of the decision are different properties. If an organisation measures only one, it cannot infer the other.

AI ScanLab’s next phase is to convert this distinction into scoping, sampling, and escalation criteria for infrastructures where the cost of an interpretively unstable decision justifies an additional evidence layer.

Project sources

  • López, J. Decision Integrity I: When the Decision Is Typed but the Interpretation Isn’t. AI ScanLab, 2026. (next publication)
  • López, J. Decision Integrity II: When the Schema Holds but the Decision Boundary Moves. AI ScanLab, 2026. (next publication)
  • AI ScanLab. Experiment A: Cross-analysis of neutral IRP and Laya instability. 2026.
  • AI ScanLab. Experiment B: IRP conditioned by class semantics and evaluator. 2026.
  • López López, J. Semantic Governance and Global Reporting in the LLM Era. Zenodo. https://doi.org/10.5281/zenodo.17714619
  • Arcuschin, I., Chanin, D., Garriga-Alonso, A., & Camburu, O.-M. Biases in the Blind Spot: Detecting What LLMs Fail to Mention. ICML 2026. https://doi.org/10.48550/arXiv.2602.10117

This site is registered on wpml.org as a development site. Switch to a production site key to remove this banner.