How this thing is allowed to be wrong

Method

Academy is built in two lanes: one that makes the material real, and one whose entire job is to find out whether the product should exist in its current form. The second does not wait for the first.

The falsification run

One outsider. One real repository task. The docs and whatever assistant they already use. No Academy, no new code.

What is being tested

Not "do you like this". One question: can a competent outsider finish a real task here using only the documentation and an assistant they already have? And if not, what specifically was missing at the moment they stopped?

The version where they succeed easily is the one that most changes the architecture — and it is a good outcome. It would mean the scarce thing was never a workspace, and that Practice should observe, constrain, record and verify work done elsewhere rather than being somewhere you work.

Why n=1 is right

This is a falsification test, not a study. No population rate is being estimated. If one competent outsider sails through, the premise that people need a workspace is already in trouble, and no larger sample is needed to justify re-examining it.

Designing a bigger, better experiment is the most attractive way to avoid running the small one.

The protocol

  1. Give them the task and the docs. Nothing else — no walkthrough, no orientation.
  2. Before they touch anything, write down verbatim where they expect to look first. Then say nothing. This is the baseline the run is measured against. If their guess was wrong and reading the docs corrected it, Read is working. If the assistant corrected it while the docs went unread, the documentation did none of the explanatory work.
  3. State that using an AI assistant is expected, not cheating. The comparison is against how they would really work.
  4. Record where they searched, what they asked the assistant verbatim, every hesitation over ~30 seconds, and what evidence changed their next action. The trace matters more than the outcome. They will not be able to articulate what was missing; their first question will show it.
  5. Record every time they ask you something. Each one is a deficit the system should have covered, and each one contaminates the run.
  6. Stop at 90 minutes regardless. An unfinished run is a result, not a failed session.
  7. Debrief with two questions: what did you need to know, and where did you expect to find it? The second locates the fix. Somewhere the docs do not go is a Read problem; somewhere that would have to watch the running system is a Practice problem.

Never ask whether they liked it, and never ask what features they want.

The doctrine

Not designed up front. Noticed after the same shape kept appearing in subsystems that have nothing to do with each other.

Situation
What the system does
Insufficient claims to ground a page
refuses, and states the reason
The semantic extractor was not run
reports no disagreement rate — not zero
No human has adjudicated a sample
reports no winning arm
A declaration restates the default
flags it as ceremony
A repository rule overrules an author
reports the override by name
A document is private
removed from the model, not hidden from one output
A panel is not implemented
carries a mock or blocked edge
One run cannot establish something
labelled unmeasured
Absence of evidence must remain distinguishable from evidence of absence, and every consequential transformation must expose why it occurred.

Both halves carry weight. The first is why an unrun measurement reports nothing rather than zero, and why refusals are published rather than swallowed — a zero and a blank look identical in a table and mean opposite things. The second is why every withheld page and every overridden declaration is named at build time: a transformation that changes what the world sees, without saying why, is indistinguishable from a bug.

What is unproven

Stated here so it cannot be quietly dropped.

No engine probes

Nothing here says "cited in 2 of 3 runs where it used to be 0 of 3". Until it does, the citation thesis is unproven. Crawlable words are an infrastructure result, not a citation result.

No user validation

No B1 session has been run. Every claim about what Practice should be is an assumption with an argument attached, not an observation.

No semantic extraction

The disagreement harness is built and tested against a known perturbation. The semantic arm has never been run, so no disagreement rate exists.