How this thing is allowed to be wrong
Academy is built in two lanes: one that makes the material real, and one whose entire job is to find out whether the product should exist in its current form. The second does not wait for the first.
One outsider. One real repository task. The docs and whatever assistant they already use. No Academy, no new code.
Not "do you like this". One question: can a competent outsider finish a real task here using only the documentation and an assistant they already have? And if not, what specifically was missing at the moment they stopped?
The version where they succeed easily is the one that most changes the architecture — and it is a good outcome. It would mean the scarce thing was never a workspace, and that Practice should observe, constrain, record and verify work done elsewhere rather than being somewhere you work.
This is a falsification test, not a study. No population rate is being estimated. If one competent outsider sails through, the premise that people need a workspace is already in trouble, and no larger sample is needed to justify re-examining it.
Designing a bigger, better experiment is the most attractive way to avoid running the small one.
Never ask whether they liked it, and never ask what features they want.
Not designed up front. Noticed after the same shape kept appearing in subsystems that have nothing to do with each other.
Absence of evidence must remain distinguishable from evidence of absence, and every consequential transformation must expose why it occurred.
Both halves carry weight. The first is why an unrun measurement reports nothing rather than zero, and why refusals are published rather than swallowed — a zero and a blank look identical in a table and mean opposite things. The second is why every withheld page and every overridden declaration is named at build time: a transformation that changes what the world sees, without saying why, is indistinguishable from a bug.
Stated here so it cannot be quietly dropped.
Nothing here says "cited in 2 of 3 runs where it used to be 0 of 3". Until it does, the citation thesis is unproven. Crawlable words are an infrastructure result, not a citation result.
No B1 session has been run. Every claim about what Practice should be is an assumption with an argument attached, not an observation.
The disagreement harness is built and tested against a known perturbation. The semantic arm has never been run, so no disagreement rate exists.