Directly supported by a defined artifact, test, build, execution record or comparison.
Evidence
What can actually be claimed.
Repeatedly seen in actual use but not presented as a controlled benchmark.
Derived from comparative analysis, correction cycles, conventional-equivalent effort or modeled operating impact.
A future deployment effect or scale assumption that has not yet occurred.
Example / same underlying model
Same underlying model. Noticeably different result.
Across our comparative work so far, we estimate that governance produces a noticeable difference in roughly 99% of cases. The size and value of that difference vary by task. In some categories, the gain appears substantial.
Current operating estimate based on comparative runs and analysis, not a controlled benchmark.Model answers from the context and assumptions available to it.
Source, context, state, evidence and authority conditions are actively challenged around the same underlying model.
Where percentages or gain ranges are eventually published, they should carry their evidence class, comparison basis and limitations with them.
Open the work register →