There is an old joke about the marksman who fires at a barn wall and then paints the target around the bullet hole. Nobody in research does anything that crude — and almost everybody does a mild version of it, by deciding what counts as a good result only after seeing the result. This page is about one small laboratory's attempt to make that impossible for itself, the exact machinery it uses, and an honest bill for what the machinery has cost: five campaigns, and every one of them came back with an answer we did not want.
How to read this page
It starts in plain language and gets steadily more technical. Sections 01–02 assume nothing. Section 03 onward describes the actual mechanism and uses the vocabulary of reproducible computing; sections 05–08 quote the internal records directly. The checklist at the end is deliberately written to be usable without reading the rest.
01 · Painting the target after the shot
#Imagine you want to know whether a coin is fair. You flip it a hundred times, look at the results, notice a run of heads in the middle, and announce that the coin is biased in its middle stretch. You have not lied about anything. Every number you reported is real. But the pattern you tested was chosen after you saw where the pattern was, and in a hundred flips of a perfectly fair coin there is always something somewhere. The test proved nothing, and it feels like it proved something, which is the dangerous part.
Almost nobody in research does this deliberately. What people do instead is a thousand tiny versions of it. The cut-off was going to be five, but four is more natural given what the data look like. The comparison was going to be against that baseline, but this other baseline is fairer. The analysis window was two years, but the last six months are clearly contaminated. Each decision is defensible on its own. Made after seeing the outcome, all of them together quietly guarantee that the answer comes out interesting.
The fix is not cleverness and not more statistics. It is ordering. Decide what would count as success, and what would count as failure, and write it down, and lock it — before you can see which way the answer falls. After that, the rules are allowed to disappoint you, and when they don't, the result means something.
That is the whole idea. Everything else on this page is machinery for making the locking real rather than a good intention: how you prove to a stranger — and, more importantly, to yourself six months later — that the rule genuinely was written first. And then the uncomfortable part, which is what the rules actually said when they were applied.
02 · The failure this is built against
#In an adjacent project, a modelling stack was pointed at a national lottery — a domain where the fairness of the equipment is about as strongly warranted as any prior in applied statistics, so that a reported edge is far more likely to be a defect than a discovery. Over roughly sixty versions it repeatedly reported an edge. Every one of those edges turned out to be a defect in the method: a headline that averaged two windows out of five, a calibration trained on its own data, a portfolio compared against a single random ticket instead of a matched set, a random seed frozen in code so that a dead pipeline looked impressively stable.
None of it was dishonest. All of it was invisible until someone reran the comparison. And the only reason it was ever caught is that the prior against a real effect was so strong that every positive was worth treating as a bug until proved otherwise. On real data — cosmology, gravitational waves, anything where the truth is unknown — the same defects produce publishable results instead of caught ones.
That is the problem this protocol exists for. It does not make analysis correct. It makes analysis unable to be quietly reshaped after the answer is visible, which is a smaller claim and a much more achievable one.
03 · What “frozen” means, exactly
#Freezing is not a mood. Before any campaign runs, four things are written, hashed, and recorded together in a manifest:
The manifest then records the one field that makes the whole thing checkable by a stranger: output_existed_at_freeze = false. That single boolean, next to the hashes, is what separates pre-registration from a description written afterwards.
What a hash does and does not prove
This is the weakest joint in the protocol and it should be stated plainly rather than glossed. A hash proves that a file has not changed relative to that hash. It proves nothing whatever about when the file existed. A private manifest asserting output_existed_at_freeze = false is an assertion, not evidence — anyone willing to lie can compute the hashes in any order they like.
What actually supplies ordering here is weaker than it should be: the freeze manifests are committed to version control before the run, so the commit chain records a sequence, and each campaign's freeze commit precedes its result commit. That is genuine evidence against accidental drift, which is the failure this protocol is built for. It is not a third-party attestation: the repository is local, commit dates are author-controlled, and nothing here is signed or externally timestamped.
An honest statement of the guarantee, therefore: the protocol makes it impossible to reshape a rule without leaving a trace in the working record, and it makes self-deception hard. It does not make fraud detectable by a stranger. Closing that gap needs an external timestamp — a public registry entry, a signed tag, or an immutable log — and this programme does not yet have one. It is on the list.
04 · Determinism is a design requirement, not a nicety
#Every campaign is run twice and required to produce byte-identical output. That requirement propagates backwards into design decisions that would otherwise never be made:
- No wall-clock fields in any hashed artifact. A timestamp inside the output makes two runs differ, which destroys the check. Timing information lives in the execution report, never in the object being compared.
- No unseeded randomness, and no seeded randomness that hides staleness. A frozen seed can make a broken pipeline look stable across data updates — that exact bug cost the adjacent project weeks. Where randomness is used, the resampling count is in the spec and the result is checked for sensitivity to it.
- Serialisation is pinned. Key order, line endings, float formatting: identical or the comparison is meaningless. Two renders of the same object must be the same bytes, which sounds pedantic until the first time a dictionary reordering makes two identical runs look different.
- Every artifact is checksummed into a registry. Scripts, outputs, specifications, and the manifest itself, with a repository-wide hash recomputed and required to match on every close-out. The registry is the difference between “the result is in the repository somewhere” and “this exact result came from that exact code”.
05 · Rules written to be able to lose
#The test of a protocol is not whether it exists; it is what happens the first time it returns something unwelcome. Here is the complete record of frozen verdicts across the programme's campaigns, positive and negative alike.
| Campaign | Question | Frozen verdict |
|---|---|---|
| Historical reconstruction | Does the archived flagship value survive a re-fixed construction? | Non-adjudicable at every level, both scalings. Supports none of the three candidate values. |
| Construction sweep | Does any documented reading of the construction reproduce it? | All six cells eliminated. Also exposed that a key ratio was forced by its own input. |
| Fresh campaign | Does a construction built to succeed produce the effect? | No refinement-stable effect. Smooth alternative preferred in all twelve cells; zero of 256 bootstrap resamples supported the hypothesis. |
| Perturbation capsule | Does the injected effect localise, and does the code find the predicted boundary? | Rule 1 negative. Rule 2 negative on operational grounds, with the underlying theorem explicitly not falsified. Rule 3 not adjudicable. |
| Dynamics scan | Which candidate evolutions are admissible? | Two admissible, three eliminated, three non-adjudicable on a declared numerical-validity ground rather than on physics. |
| Ringdown regression | What does the stacked fit say? | Template only, no event measurements. The harness exists; the analysis has not been authorised and the output says so in its own status field. |
06 · The rule that referenced a tolerance that did not exist
#One decision rule required a measured response to exceed three times a matched control, while a global drift stayed below a tolerance carried over from an earlier campaign. The raw diagnostics came back fine: the response cleared its control in 23 of 24 cells. Then the second clause was checked, and the referenced tolerance turned out never to have been defined anywhere.
There were two available moves. Pick a defensible tolerance now — the numbers were already visible, and almost any reasonable choice would have let the rule pass. Or record that the rule cannot be evaluated.
What was actually recorded
The classification field was written as null, with an explicit status: not adjudicable, missing referenced tolerance. The note states in its own text that this null records a missing adjudication and is not the registered “no effect” class — the two mean different things and conflating them would be a quiet upgrade of a void result into a negative one. No interpretation beyond that status is licensed.
That is the entire discipline compressed into one decision. No tolerance is invented after viewing the output. Everything else in this protocol — the hashes, the double renders, the registry — exists to make that one refusal enforceable rather than merely intended. A rule you may repair after seeing its result is not a rule; it is a description.
07 · Acquisition is a separate stage, and it is blind
#The sharpest application of the protocol is not analysis at all. It is data acquisition — the stage where, in most projects, someone downloads a catalogue, opens it, has a look, and only then decides what the study will be. That look is unrecoverable: from then on every subsequent choice is informed by data that was supposed to be tested against.
So acquisition was made its own contract, with a single governing rule: read no scientific value. Not one posterior sample, not one strain point, not one calibration number. The stage's output carries a counter, scientific_array_values_read, and its required value is zero.
| Element | How it was handled |
|---|---|
| Event list | Frozen in a previous stage from catalogue metadata only — twelve primaries and one ordered reserve — with the substitution rule for a deficient event sealed in advance, including the count below which the study stops as a successful non-adjudicating outcome. |
| Inventory | Two metadata requests, no file transfers. Every filename, size, provider checksum and download URL recorded and hashed before a single byte of payload was fetched. Near-namesake events on the same date forced matching on the full event token, which was specified rather than discovered. |
| Transfer | 49 files, 30.9 GB. Every parameter-estimation file verified against the provider's declared size and checksum. The strain archive publishes no checksums, so files above 100 MB were size-verified before and after transfer with a local hash recorded, and smaller files were fetched twice and required to be byte-identical. |
| Structure check | Files opened, but only dataset and attribute names read — enough to confirm the required quantities exist and to record the waveform family verbatim, and not enough to learn anything about any event. Reading a value was defined as a contract violation. |
| Payload handling | Multi-gigabyte payloads deliberately excluded from version control; only the freeze records, receipts and code committed. The receipts are what make the payloads re-derivable by anyone else — the bytes themselves are not the artifact. |
The result is a dataset that is fully specified, fully checksummed, and completely unexamined. Whatever the analysis stage concludes, it cannot be accused of having chosen its events after seeing them, because the event list is hashed and dated earlier than the first download.
08 · How to change a frozen plan without destroying it
#Frozen plans meet reality. The useful question is not whether to amend them but what an amendment has to look like to be legitimate. Three were issued during the acquisition campaign, and their shape is the transferable part.
09 · When role separation breaks, say so in the artifact
#The protocol assumes three separated roles: a planner who writes the contract, an executor who runs it, and a verifier who checks the result against the contract. During the acquisition campaign that separation broke. The contracted executor twice ended its sessions without completing the first step, and the author directed the planner to execute the remainder directly. For that stage, planner, executor and verifier were the same agent.
What was written into the record
The deviation is disclosed inside the frozen contract, not in a private note — with the reason, the date, the authorising instruction, and two compensating controls: the stage produces no scientific values at all, so the conflict of interest has nothing to act on; and the downstream analysis contract is required to include an independent re-verification of the entire freeze chain before any value is decoded.
This is the part of the protocol most worth copying and least likely to be copied. A governance deviation that is disclosed, scoped, and compensated is a manageable weakness. The same deviation, undisclosed, is indistinguishable from the result being unreliable — and the reader has no way to tell which one they are looking at unless the record says.
10 · What this is not
#- It is not peer review. A frozen manifest says a rule was fixed in advance. It says nothing about whether the rule was a good one, whether the model is appropriate, or whether the whole question is worth asking. Those require a reader, and no amount of hashing substitutes for one.
- It is not a correctness guarantee. A deterministic pipeline reproduces its own bugs perfectly. Several campaigns above found real defects despite the protocol, and the protocol's contribution was that the defects surfaced as failed checks instead of as results.
- It is not clinical pre-registration. There is no external registry, no third-party timestamp, no institutional oversight. It is self-imposed and self-enforced, which means its guarantees are only as strong as the willingness to publish the losses. That willingness is the actual mechanism; the hashes just make it checkable.
- It does not prove chronology to an outsider. Hashes fix content, not time. The ordering evidence here is a local commit chain with author-controlled dates and no signatures — good against drift, useless against a determined bad actor. Anyone treating these records as third-party-verified pre-registration is reading more into them than they carry.
- It is not free. Everything below.
11 · What it costs
#An honest method note has to price its own method.
- Speed. Writing the specification, the extraction rule and the decision rules before running anything typically takes longer than the run. For exploratory work that is simply the wrong trade, and the protocol is not applied to exploration — it is applied at the point where an exploration is about to become a claim.
- Findings. Five campaigns, no positive headline. Under looser rules at least three of them would have produced quotable results, and every one of those results would have been wrong. That is the trade, stated plainly: the protocol converts probable false positives into certain nulls.
- Comfort. Publishing a null that contradicts your own flagship claim is unpleasant, and doing it while the corresponding paper is under review is worse. The compensation arrived unexpectedly — the reason for the null turned out to be a theorem, which converted a retraction into a prediction. That is not guaranteed and should not be counted on.
- Flexibility. Sometimes the frozen rule is genuinely the wrong rule, and you have to record a bad verdict from a rule you now know to be poorly designed. The correct response is a new campaign with a better rule, not a repair of the old one — which means paying for the same computation twice.
Against that: every number in the adjacent research pages can be traced to an executable and a hash, every negative verdict on those pages was recorded before anyone knew whether it would be negative, and the one deviation from the protocol is written into the artifact it affected. Whether that is worth the cost depends entirely on whether you would rather be fast or be checkable.
12 · The checklist
#Reduced to something another small group could adopt on a Monday, without infrastructure and without permission from anyone.
The last line is the one that does the work. Everything above it is bookkeeping that becomes theatre the moment an inconvenient result goes unpublished.
- Where the failure mode was measured — the lottery audit: sixty model versions against a domain where a reported edge is almost certainly a defect, and a complete catalogue of the ways a pipeline manufactures one.
- The programme this governs — the open programme: what the campaigns above were campaigns about, and what remains gated.
- Protocol under maximum stress — the correlation budget: a proven theorem, certified arithmetic, and three of four rules still returning negative or void.
- Data terms — the acquisition campaign uses public releases under their published licences, with provider acknowledgements and release identifiers retained in the sealed metadata records. Reuse terms are recorded at freeze time and re-checked before any use.
| Claim | Status | Artifact and how to check it |
|---|---|---|
| Historical reconstruction returned non-adjudicable at every level. | Executed — frozen | Freeze manifest outputs/t1_freeze_manifest.json (records ladder_output_existed_at_freeze = false); spec, extraction rule and script hashed; ladder output outputs/t1_step_height_adjudication.json, SHA-256 a09e04bc…2398d751. |
| Fresh campaign found no refinement-stable step; smooth preferred in all twelve cells, 0/256 bootstrap support. | Executed — frozen | outputs/t2_freeze_manifest.json, script t2_horizon_profile_campaign.py; output outputs/t2_horizon_profile_campaign.json, SHA-256 6022816d…1edd0c8e32c, byte-identical across two post-freeze runs. |
| Perturbation capsule: one rule as intended, two negative, one void. | Executed — frozen | outputs/fw18_correlation_handle.json, 2,399,706 bytes, SHA-256 155354db…4157d7d8; repository manifest of 108 entries, SHA-256 6ffeabea…451df364. |
| Dynamics scan: two admissible, three eliminated, three non-adjudicable on a declared numerical ground. | Executed — frozen | Spec 5a1277a9… and script cf52eca1… hashed into outputs/p9_freeze_manifest.json while the scan output did not yet exist; result outputs/p9_pisr_pilot.json, byte-identical on rerun. |
| Value-blind acquisition of 49 files and 30,918,549,479 bytes with no scientific value read. | Executed — frozen | Pre-download freeze outputs/gwtc5_np1_stageC_freeze.json, SHA-256 94f24bf5…264de036; receipt outputs/gwtc5_np1_stageC_acquisition_receipt.json carrying scientific_array_values_read: 0, per-file provider md5 where supplied, and local SHA-256 for every file. |
| The rules were fixed before the outputs existed. | Asserted, with local ordering evidence only | Each manifest carries the output_existed_at_freeze = false flag, and freeze commits precede result commits in version control. There is no third-party timestamp, no signature and no external registry entry — see section 03. Treat this as evidence against drift, not as proof against fraud. |
| Public availability of the artifacts named above. | Not public | These are internal working artifacts. The public release is the frozen framework-article snapshot, which does not contain them. They are available on request, and the hashes quoted here are what an independent check would have to reproduce. |
Research and AI disclosure
This page describes an internal working protocol, not a standard, and not a peer-reviewed methodology. It is self-imposed and self-enforced; its guarantees extend only to what the artifacts can demonstrate. Nothing here should be read as an endorsement of the underlying scientific claims, several of which are explicitly unresolved.
AI agents performed most of the planning, execution and verification described here, in nominally separated roles, and one breakdown of that separation is documented above. AI-generated output can contain errors; the protocol exists partly because it does.
Research and correction enquiries: contact@kort-x.com.
- Author
- Ciprian Stoichici, KORT-X Research, Bucharest, Romania
- Version
- 1.0
- Published
- 2026-08-09
- Updated
- 2026-08-09
- Licence
- CC BY 4.0
- Cite as
- Stoichici, C. (2026). "Freeze First: pre-registration for computational research." KORT-X Research. https://kort-x.com/indexfiles/research-freeze-first.html
This is a laboratory write-up, not a refereed publication. Where a claim rests on an executed artifact, the evidence record above names the artifact and its status; internal working artifacts are not part of the public release and are available on request. Corrections are welcome and are applied in place with the update date changed.