LABORATORIES
Knowledge Orchestration for Radical Transformation - X
GitHub

FREEZE FIRST

Research Governance · Pre-registration · Reproducibility

There is an old joke about the marksman who fires at a barn wall and then paints the target around the bullet hole. Nobody in research does anything that crude — and almost everybody does a mild version of it, by deciding what counts as a good result only after seeing the result. This page is about one small laboratory's attempt to make that impossible for itself, the exact machinery it uses, and an honest bill for what the machinery has cost: five campaigns, and every one of them came back with an answer we did not want.

How to read this page

It starts in plain language and gets steadily more technical. Sections 01–02 assume nothing. Section 03 onward describes the actual mechanism and uses the vocabulary of reproducible computing; sections 05–08 quote the internal records directly. The checklist at the end is deliberately written to be usable without reading the rest.

4
Artifacts hashed before any result exists
2×
Every run repeated, requiring byte-identical output
0
Tolerances invented after seeing an output
30.9
GB acquired with zero scientific values read
1
Governance deviation, disclosed inside the artifact

01 · Painting the target after the shot

#

Imagine you want to know whether a coin is fair. You flip it a hundred times, look at the results, notice a run of heads in the middle, and announce that the coin is biased in its middle stretch. You have not lied about anything. Every number you reported is real. But the pattern you tested was chosen after you saw where the pattern was, and in a hundred flips of a perfectly fair coin there is always something somewhere. The test proved nothing, and it feels like it proved something, which is the dangerous part.

Almost nobody in research does this deliberately. What people do instead is a thousand tiny versions of it. The cut-off was going to be five, but four is more natural given what the data look like. The comparison was going to be against that baseline, but this other baseline is fairer. The analysis window was two years, but the last six months are clearly contaminated. Each decision is defensible on its own. Made after seeing the outcome, all of them together quietly guarantee that the answer comes out interesting.

The fix is not cleverness and not more statistics. It is ordering. Decide what would count as success, and what would count as failure, and write it down, and lock it — before you can see which way the answer falls. After that, the rules are allowed to disappoint you, and when they don't, the result means something.

That is the whole idea. Everything else on this page is machinery for making the locking real rather than a good intention: how you prove to a stranger — and, more importantly, to yourself six months later — that the rule genuinely was written first. And then the uncomfortable part, which is what the rules actually said when they were applied.

02 · The failure this is built against

#

In an adjacent project, a modelling stack was pointed at a national lottery — a domain where the fairness of the equipment is about as strongly warranted as any prior in applied statistics, so that a reported edge is far more likely to be a defect than a discovery. Over roughly sixty versions it repeatedly reported an edge. Every one of those edges turned out to be a defect in the method: a headline that averaged two windows out of five, a calibration trained on its own data, a portfolio compared against a single random ticket instead of a matched set, a random seed frozen in code so that a dead pipeline looked impressively stable.

None of it was dishonest. All of it was invisible until someone reran the comparison. And the only reason it was ever caught is that the prior against a real effect was so strong that every positive was worth treating as a bug until proved otherwise. On real data — cosmology, gravitational waves, anything where the truth is unknown — the same defects produce publishable results instead of caught ones.

That is the problem this protocol exists for. It does not make analysis correct. It makes analysis unable to be quietly reshaped after the answer is visible, which is a smaller claim and a much more achievable one.

03 · What “frozen” means, exactly

#

Freezing is not a mood. Before any campaign runs, four things are written, hashed, and recorded together in a manifest:

1 · The specification
What is being computed, on what domain, at what resolutions, with which parameters. Every nuisance value fixed. If a range is being swept, the range is in the spec, not chosen while looking at the sweep.
2 · The extraction rule
How a number is pulled out of the output — which window, which estimator, which model comparison, which margin counts as decisive. This is the single most abused degree of freedom in computational science and the one most worth nailing down, because it is where an ambiguous result becomes a clear one.
3 · The executable
The script itself, hashed. Not the repository, not the branch — the exact bytes that will run. A change to the script after the freeze is a new campaign with a new manifest, not a fix.
4 · The decision rules
Named, numbered, and stated as comparisons with thresholds: what observation would count as the effect existing, what would count as it not existing, and what would count as the test being unable to decide. All three branches must be written, because a rule with no losing branch is not a rule.

The manifest then records the one field that makes the whole thing checkable by a stranger: output_existed_at_freeze = false. That single boolean, next to the hashes, is what separates pre-registration from a description written afterwards.

What a hash does and does not prove

This is the weakest joint in the protocol and it should be stated plainly rather than glossed. A hash proves that a file has not changed relative to that hash. It proves nothing whatever about when the file existed. A private manifest asserting output_existed_at_freeze = false is an assertion, not evidence — anyone willing to lie can compute the hashes in any order they like.

What actually supplies ordering here is weaker than it should be: the freeze manifests are committed to version control before the run, so the commit chain records a sequence, and each campaign's freeze commit precedes its result commit. That is genuine evidence against accidental drift, which is the failure this protocol is built for. It is not a third-party attestation: the repository is local, commit dates are author-controlled, and nothing here is signed or externally timestamped.

An honest statement of the guarantee, therefore: the protocol makes it impossible to reshape a rule without leaving a trace in the working record, and it makes self-deception hard. It does not make fraud detectable by a stranger. Closing that gap needs an external timestamp — a public registry entry, a signed tag, or an immutable log — and this programme does not yet have one. It is on the list.

04 · Determinism is a design requirement, not a nicety

#

Every campaign is run twice and required to produce byte-identical output. That requirement propagates backwards into design decisions that would otherwise never be made:

  • No wall-clock fields in any hashed artifact. A timestamp inside the output makes two runs differ, which destroys the check. Timing information lives in the execution report, never in the object being compared.
  • No unseeded randomness, and no seeded randomness that hides staleness. A frozen seed can make a broken pipeline look stable across data updates — that exact bug cost the adjacent project weeks. Where randomness is used, the resampling count is in the spec and the result is checked for sensitivity to it.
  • Serialisation is pinned. Key order, line endings, float formatting: identical or the comparison is meaningless. Two renders of the same object must be the same bytes, which sounds pedantic until the first time a dictionary reordering makes two identical runs look different.
  • Every artifact is checksummed into a registry. Scripts, outputs, specifications, and the manifest itself, with a repository-wide hash recomputed and required to match on every close-out. The registry is the difference between “the result is in the repository somewhere” and “this exact result came from that exact code”.

05 · Rules written to be able to lose

#

The test of a protocol is not whether it exists; it is what happens the first time it returns something unwelcome. Here is the complete record of frozen verdicts across the programme's campaigns, positive and negative alike.

Every frozen verdict, as recorded
Campaign Question Frozen verdict
Historical reconstruction Does the archived flagship value survive a re-fixed construction? Non-adjudicable at every level, both scalings. Supports none of the three candidate values.
Construction sweep Does any documented reading of the construction reproduce it? All six cells eliminated. Also exposed that a key ratio was forced by its own input.
Fresh campaign Does a construction built to succeed produce the effect? No refinement-stable effect. Smooth alternative preferred in all twelve cells; zero of 256 bootstrap resamples supported the hypothesis.
Perturbation capsule Does the injected effect localise, and does the code find the predicted boundary? Rule 1 negative. Rule 2 negative on operational grounds, with the underlying theorem explicitly not falsified. Rule 3 not adjudicable.
Dynamics scan Which candidate evolutions are admissible? Two admissible, three eliminated, three non-adjudicable on a declared numerical-validity ground rather than on physics.
Ringdown regression What does the stacked fit say? Template only, no event measurements. The harness exists; the analysis has not been authorised and the output says so in its own status field.
The pattern
Not one campaign returned the answer that motivated it.
That is not a sign the protocol is too strict. It is a measurement of how often a pre-registered rule disagrees with a well-motivated expectation — and the honest reading is that without the rules, several of these would have been reported as findings. The most instructive is the perturbation capsule, where the mathematics was a proven theorem, the arithmetic was certified to thirty significant figures, and three of the four rules around it still came back negative or void.

06 · The rule that referenced a tolerance that did not exist

#

One decision rule required a measured response to exceed three times a matched control, while a global drift stayed below a tolerance carried over from an earlier campaign. The raw diagnostics came back fine: the response cleared its control in 23 of 24 cells. Then the second clause was checked, and the referenced tolerance turned out never to have been defined anywhere.

There were two available moves. Pick a defensible tolerance now — the numbers were already visible, and almost any reasonable choice would have let the rule pass. Or record that the rule cannot be evaluated.

What was actually recorded

The classification field was written as null, with an explicit status: not adjudicable, missing referenced tolerance. The note states in its own text that this null records a missing adjudication and is not the registered “no effect” class — the two mean different things and conflating them would be a quiet upgrade of a void result into a negative one. No interpretation beyond that status is licensed.

That is the entire discipline compressed into one decision. No tolerance is invented after viewing the output. Everything else in this protocol — the hashes, the double renders, the registry — exists to make that one refusal enforceable rather than merely intended. A rule you may repair after seeing its result is not a rule; it is a description.

07 · Acquisition is a separate stage, and it is blind

#

The sharpest application of the protocol is not analysis at all. It is data acquisition — the stage where, in most projects, someone downloads a catalogue, opens it, has a look, and only then decides what the study will be. That look is unrecoverable: from then on every subsequent choice is informed by data that was supposed to be tested against.

So acquisition was made its own contract, with a single governing rule: read no scientific value. Not one posterior sample, not one strain point, not one calibration number. The stage's output carries a counter, scientific_array_values_read, and its required value is zero.

One value-blind acquisition, as executed
Element How it was handled
Event list Frozen in a previous stage from catalogue metadata only — twelve primaries and one ordered reserve — with the substitution rule for a deficient event sealed in advance, including the count below which the study stops as a successful non-adjudicating outcome.
Inventory Two metadata requests, no file transfers. Every filename, size, provider checksum and download URL recorded and hashed before a single byte of payload was fetched. Near-namesake events on the same date forced matching on the full event token, which was specified rather than discovered.
Transfer 49 files, 30.9 GB. Every parameter-estimation file verified against the provider's declared size and checksum. The strain archive publishes no checksums, so files above 100 MB were size-verified before and after transfer with a local hash recorded, and smaller files were fetched twice and required to be byte-identical.
Structure check Files opened, but only dataset and attribute names read — enough to confirm the required quantities exist and to record the waveform family verbatim, and not enough to learn anything about any event. Reading a value was defined as a contract violation.
Payload handling Multi-gigabyte payloads deliberately excluded from version control; only the freeze records, receipts and code committed. The receipts are what make the payloads re-derivable by anyone else — the bytes themselves are not the artifact.

The result is a dataset that is fully specified, fully checksummed, and completely unexamined. Whatever the analysis stage concludes, it cannot be accused of having chosen its events after seeing them, because the event list is hashed and dated earlier than the first download.

08 · How to change a frozen plan without destroying it

#

Frozen plans meet reality. The useful question is not whether to amend them but what an amendment has to look like to be legitimate. Three were issued during the acquisition campaign, and their shape is the transferable part.

Amendment 1 — pacing only, explicitly
The remaining steps were declared a single indivisible unit of work, so that a step boundary stopped being a permitted place to stop. The amendment ends with an unambiguous sentence: this changes execution pacing only; no scientific rule, gate, cap, or prohibition changes. Naming the scope of an amendment is what keeps it from becoming a rewrite.
Amendment 2 — corrected from sealed provider metadata
The original transfer caps were a planner's estimate written before file sizes were known, and one file turned out to be more than twice the per-file cap. The caps were raised — but the justification is derived exclusively from metadata sealed before any download, involves no scientific value whatsoever, and the freeze is required to record both the original and corrected numbers with the reason. An amendment that cannot be traced to information available before the freeze is not an amendment.
Amendment 3 — a fallback the contract already contained
The specified short-duration data products did not exist in the sealed records for any event. Rather than improvising, the contract's own pre-written fallback applied, mechanically. The integrity rule for the checksum-less archive was also changed on a stated engineering argument — transport-layer integrity plus verified sizes, instead of doubling twenty gigabytes of transfer for no additional assurance — and the argument is in the amendment where it can be disputed.

09 · When role separation breaks, say so in the artifact

#

The protocol assumes three separated roles: a planner who writes the contract, an executor who runs it, and a verifier who checks the result against the contract. During the acquisition campaign that separation broke. The contracted executor twice ended its sessions without completing the first step, and the author directed the planner to execute the remainder directly. For that stage, planner, executor and verifier were the same agent.

What was written into the record

The deviation is disclosed inside the frozen contract, not in a private note — with the reason, the date, the authorising instruction, and two compensating controls: the stage produces no scientific values at all, so the conflict of interest has nothing to act on; and the downstream analysis contract is required to include an independent re-verification of the entire freeze chain before any value is decoded.

This is the part of the protocol most worth copying and least likely to be copied. A governance deviation that is disclosed, scoped, and compensated is a manageable weakness. The same deviation, undisclosed, is indistinguishable from the result being unreliable — and the reader has no way to tell which one they are looking at unless the record says.

10 · What this is not

#
  • It is not peer review. A frozen manifest says a rule was fixed in advance. It says nothing about whether the rule was a good one, whether the model is appropriate, or whether the whole question is worth asking. Those require a reader, and no amount of hashing substitutes for one.
  • It is not a correctness guarantee. A deterministic pipeline reproduces its own bugs perfectly. Several campaigns above found real defects despite the protocol, and the protocol's contribution was that the defects surfaced as failed checks instead of as results.
  • It is not clinical pre-registration. There is no external registry, no third-party timestamp, no institutional oversight. It is self-imposed and self-enforced, which means its guarantees are only as strong as the willingness to publish the losses. That willingness is the actual mechanism; the hashes just make it checkable.
  • It does not prove chronology to an outsider. Hashes fix content, not time. The ordering evidence here is a local commit chain with author-controlled dates and no signatures — good against drift, useless against a determined bad actor. Anyone treating these records as third-party-verified pre-registration is reading more into them than they carry.
  • It is not free. Everything below.

11 · What it costs

#

An honest method note has to price its own method.

  • Speed. Writing the specification, the extraction rule and the decision rules before running anything typically takes longer than the run. For exploratory work that is simply the wrong trade, and the protocol is not applied to exploration — it is applied at the point where an exploration is about to become a claim.
  • Findings. Five campaigns, no positive headline. Under looser rules at least three of them would have produced quotable results, and every one of those results would have been wrong. That is the trade, stated plainly: the protocol converts probable false positives into certain nulls.
  • Comfort. Publishing a null that contradicts your own flagship claim is unpleasant, and doing it while the corresponding paper is under review is worse. The compensation arrived unexpectedly — the reason for the null turned out to be a theorem, which converted a retraction into a prediction. That is not guaranteed and should not be counted on.
  • Flexibility. Sometimes the frozen rule is genuinely the wrong rule, and you have to record a bad verdict from a rule you now know to be poorly designed. The correct response is a new campaign with a better rule, not a repair of the old one — which means paying for the same computation twice.

Against that: every number in the adjacent research pages can be traced to an executable and a hash, every negative verdict on those pages was recorded before anyone knew whether it would be negative, and the one deviation from the protocol is written into the artifact it affected. Whether that is worth the cost depends entirely on whether you would rather be fast or be checkable.

12 · The checklist

#

Reduced to something another small group could adopt on a Monday, without infrastructure and without permission from anyone.

BEFORE THE RUN [ ] specification written and hashed [ ] extraction rule written and hashed (which window, which estimator, which margin) [ ] script hashed (exact bytes, not the branch) [ ] decision rules named, with a pass branch, a fail branch and a cannot-decide branch [ ] manifest records: output_existed_at_freeze = false DURING THE RUN [ ] no wall-clock fields in any hashed artifact [ ] run twice; require byte-identical output [ ] every artifact checksummed into a registry [ ] amendments: scope stated, justified only from pre-freeze information AFTER THE RUN [ ] apply the rules exactly as written [ ] a rule that cannot be evaluated is recorded as not adjudicable, never repaired [ ] a failed control makes a data point unusable, never supportive [ ] role deviations disclosed inside the artifact, with compensating controls [ ] publish the losses

The last line is the one that does the work. Everything above it is bookkeeping that becomes theatre the moment an inconvenient result goes unpublished.

  • Where the failure mode was measured — the lottery audit: sixty model versions against a domain where a reported edge is almost certainly a defect, and a complete catalogue of the ways a pipeline manufactures one.
  • The programme this governs — the open programme: what the campaigns above were campaigns about, and what remains gated.
  • Protocol under maximum stress — the correlation budget: a proven theorem, certified arithmetic, and three of four rules still returning negative or void.
  • Data terms — the acquisition campaign uses public releases under their published licences, with provider acknowledgements and release identifiers retained in the sealed metadata records. Reuse terms are recorded at freeze time and re-checked before any use.
Evidence record — the frozen artifacts behind this page
Claim Status Artifact and how to check it
Historical reconstruction returned non-adjudicable at every level. Executed — frozen Freeze manifest outputs/t1_freeze_manifest.json (records ladder_output_existed_at_freeze = false); spec, extraction rule and script hashed; ladder output outputs/t1_step_height_adjudication.json, SHA-256 a09e04bc…2398d751.
Fresh campaign found no refinement-stable step; smooth preferred in all twelve cells, 0/256 bootstrap support. Executed — frozen outputs/t2_freeze_manifest.json, script t2_horizon_profile_campaign.py; output outputs/t2_horizon_profile_campaign.json, SHA-256 6022816d…1edd0c8e32c, byte-identical across two post-freeze runs.
Perturbation capsule: one rule as intended, two negative, one void. Executed — frozen outputs/fw18_correlation_handle.json, 2,399,706 bytes, SHA-256 155354db…4157d7d8; repository manifest of 108 entries, SHA-256 6ffeabea…451df364.
Dynamics scan: two admissible, three eliminated, three non-adjudicable on a declared numerical ground. Executed — frozen Spec 5a1277a9… and script cf52eca1… hashed into outputs/p9_freeze_manifest.json while the scan output did not yet exist; result outputs/p9_pisr_pilot.json, byte-identical on rerun.
Value-blind acquisition of 49 files and 30,918,549,479 bytes with no scientific value read. Executed — frozen Pre-download freeze outputs/gwtc5_np1_stageC_freeze.json, SHA-256 94f24bf5…264de036; receipt outputs/gwtc5_np1_stageC_acquisition_receipt.json carrying scientific_array_values_read: 0, per-file provider md5 where supplied, and local SHA-256 for every file.
The rules were fixed before the outputs existed. Asserted, with local ordering evidence only Each manifest carries the output_existed_at_freeze = false flag, and freeze commits precede result commits in version control. There is no third-party timestamp, no signature and no external registry entry — see section 03. Treat this as evidence against drift, not as proof against fraud.
Public availability of the artifacts named above. Not public These are internal working artifacts. The public release is the frozen framework-article snapshot, which does not contain them. They are available on request, and the hashes quoted here are what an independent check would have to reproduce.

Research and AI disclosure

This page describes an internal working protocol, not a standard, and not a peer-reviewed methodology. It is self-imposed and self-enforced; its guarantees extend only to what the artifacts can demonstrate. Nothing here should be read as an endorsement of the underlying scientific claims, several of which are explicitly unresolved.

AI agents performed most of the planning, execution and verification described here, in nominally separated roles, and one breakdown of that separation is documented above. AI-generated output can contain errors; the protocol exists partly because it does.

Research and correction enquiries: contact@kort-x.com.

Continue
The campaigns this protocol governed, and the case study where it was calibrated.
A rule you are allowed to repair after seeing its result is not a rule. It is a description of what you were going to say anyway.
Article record
Author
Ciprian Stoichici, KORT-X Research, Bucharest, Romania
Version
1.0
Published
2026-08-09
Updated
2026-08-09
Licence
CC BY 4.0
Cite as
Stoichici, C. (2026). "Freeze First: pre-registration for computational research." KORT-X Research. https://kort-x.com/indexfiles/research-freeze-first.html

This is a laboratory write-up, not a refereed publication. Where a claim rests on an executed artifact, the evidence record above names the artifact and its status; internal working artifacts are not part of the public release and are available on request. Corrections are welcome and are applied in place with the update date changed.

↑ Top
SINGULARITY_V2