Directly supported by a defined artifact, test, build, execution record or comparison.
Evidence
What can actually be claimed.
Repeatedly seen in actual use but not presented as a controlled benchmark.
Derived from comparative analysis, correction cycles, conventional-equivalent effort or modeled operating impact.
A future deployment effect or scale assumption that hasn’t yet occurred.
Example / same underlying model
Same model. Smarter system.
The underlying model is the same. What changes is the environment it works inside. Better organization of memory, state, authority, scope and evidence can reduce how much of the model's attention is spent reconstructing the situation before it can deal with the actual question.
Across our comparative work so far, we estimate that governance produces a noticeable difference in roughly 99% of cases. The size and value of that difference vary by task. In some categories, the gain appears substantial.
Current operating estimate based on comparative runs and analysis, not a controlled benchmark.Model answers from the context and assumptions available to it.
Source, context, state, evidence and authority conditions are actively challenged around the same underlying model.
Where percentages or gain ranges are eventually published, they should carry their evidence class, comparison basis and limitations with them.
Open the work register →