LABORATORIES
Knowledge Orchestration for Radical Transformation - X
GitHub

THE CORRELATION BUDGET

Linear Algebra · Numerical Analysis · Pre-registered Computation

Some tables of numbers describe a world that cannot exist. Tell me Alice always copies Bob, and Bob always copies Carol, and that Alice and Carol have nothing to do with each other, and you have said something impossible without using a single hard word. This page is about the exact point at which such a table tips over into impossibility — a question with a closed-form answer that fits on one line — and about what happened when a computer was asked to confirm the answer. Proving it took an afternoon. Confirming it took 120-digit arithmetic, and mostly failed.

How to read this page

It starts in plain language and gets steadily more technical. Section 01 needs no mathematics. Sections 02–03 state the result properly and sketch the proof — undergraduate linear algebra is enough. From section 05 the page turns to numerical analysis and pre-registered testing, and section 07 is close to the raw record.

1
Closed-form bound, no approximation
1076
Condition number of the worst declared matrix
120/160
Decimal digits required to evaluate it
24
Pre-registered cells, all completed twice
1 / 4
Decision rules that returned a positive verdict

01 · Why some tables of numbers are impossible

#

Take the Alice–Bob–Carol example above. Each of the three statements is perfectly reasonable on its own. Together they cannot be true, and you can feel it without doing any arithmetic: if Alice tracks Bob and Bob tracks Carol, Alice cannot be unrelated to Carol. Relationships constrain each other, the way the three sides of a triangle do — you are free to choose two of them, and then the third is boxed in.

A correlation matrix is exactly that kind of table, written out for every pair at once. And the trap scales up badly. With three quantities you can spot the contradiction by eye. With fifty, a table can be full of individually sensible numbers and still describe a world that cannot exist, and nothing about looking at it will tell you. The technical name for a table that can exist is positive semi-definite. The plain version: no combination of the quantities is allowed to come out with a negative amount of variability. Negative variance is not a small error to be tolerated — it is nonsense, like a negative length.

This matters well outside mathematics. Portfolio risk models, weather and mining models, sensor networks, machine-learning kernels and Gaussian-process models are all built on such tables, and all of them break when somebody adjusts one entry by hand to reflect a judgement. Entire software libraries exist for nothing but taking a table that a human has quietly ruined and pushing it back to the nearest valid one.

So here is a natural question which, pleasingly, has an exact answer. Start with a table that is valid. Pick two of the quantities and start adding extra correlation between them — turning a knob up from zero, changing nothing else. At zero the table is fine by assumption. Turn the knob far enough and it is certainly nonsense. Where precisely is the cliff edge?

It is worth saying where the question came from, because the origin is unusual and the answer is not. It arose in a physics programme where the table is a correlation structure and distance is defined from it, so that adding correlation between two far-apart regions is a statement about geometry. None of that motivation is needed below. What follows is finite-dimensional linear algebra and applies to anyone's covariance matrix.

02 · Stating the question precisely

#

Fix a valid base structure K₀. Pick two localised profiles — think of them as two regions, two sensors, two clusters of assets — and add extra correlation strictly between them, in the smallest possible way: a symmetric rank-two term, scaled by the knob λ. At λ = 0 the structure is valid by assumption; at large λ it is not. The question is where the transition happens, and the answer should be exact rather than found by trial and error.

Two words carry weight in what follows. Positive semi-definite is the validity condition from section 01, written as a condition on eigenvalues: every one of them must be at least zero. And rank-two means the perturbation is the simplest object that can couple two profiles symmetrically and nothing else — the minimal edit, so that whatever we find is a property of the base structure rather than of an elaborate perturbation.

03 · The exact answer

#

Write the perturbed structure as K(λ) = K₀ + λ(uAuBᵀ + uBuAᵀ) with unit-norm profiles and λ ≥ 0. Assume K₀ is symmetric and strictly positive definite. Define three numbers — the two self-terms and the cross-term, all measured in the metric induced by the inverse of the base:

a = uᴬᵀ K₀⁻¹ uᴬ b = uᴮᵀ K₀⁻¹ uᴮ c = uᴬᵀ K₀⁻¹ uᴮ 1 λ_crit = ───────────────── K(λ) ⪰ 0 ⟺ 0 ≤ λ ≤ λ_crit √(ab) − c

That is the whole result. Below the boundary the perturbed structure is strictly positive definite; at the boundary it is semi-definite with exactly one new null direction; above it, indefinite — no longer a covariance structure. The endpoint is included, which matters, because pre-registered tests tend to be written at the endpoint.

The proof is short enough to give in full outline. Whiten: set x = K₀−1/2uA and y = K₀−1/2uB. Congruence by K₀−1/2 preserves inertia — it cannot change how many eigenvalues sit on each side of zero — and turns the problem into I + λ(xyᵀ + yxᵀ), which is the identity everywhere except on the two-dimensional span of x and y. On that span the matrix is 2×2, and its determinant factors:

1 + 2λc − λ²(ab − c²) = [1 + λ(c + √(ab))] · [1 + λ(c − √(ab))]

For λ ≥ 0 the first factor is at least one. The second is non-negative exactly up to 1/(√(ab) − c). Cauchy–Schwarz guarantees |c| ≤ √(ab), so the denominator is non-negative and the bound is meaningful. At the boundary the null direction is √b·K₀⁻¹uA − √a·K₀⁻¹uB, written out explicitly because a test that wants to detect the boundary numerically needs to know what it is looking at.

04 · The degenerate cases, which are where the intuition is

#

The two excluded cases are the interesting ones, because they say what the budget is really measuring.

Perfectly aligned profiles — unlimited budget
If the cross-term saturates Cauchy–Schwarz from above, the second profile is a positive multiple of the first, the perturbation is itself positive semi-definite, and you may add as much as you like forever. Correlation you are “injecting” between two regions that the base already treats as the same region costs nothing, because it is not new information — it is a rescaling.
Perfectly anti-aligned profiles — halved budget
If the cross-term saturates from below, the only non-unit eigenvalue is linear in λ and the boundary sits at 1/(2√(ab)). Two profiles the base structure already regards as opposed are the cheapest way to break positivity, and they break it twice as fast.
Rank-deficient bases still work, with one caveat
If the base is only semi-definite and both profiles lie in its range, the same proof runs on that range with the Moore–Penrose pseudo-inverse in place of the inverse; the old null space survives untouched. If either profile has a component outside the range, the inverse-form statement is simply not licensed and a separate null-space analysis is required. Stating that limit is not pedantry — it is the case a practitioner is most likely to hit, because empirical covariance matrices are routinely rank-deficient.

One structural remark worth extracting: spatial separation never enters the algebra. The bound is identical whether the two profiles sit at opposite ends of a domain or overlap completely. Distance only enters through the base matrix, via the three inner products. The budget is a statement about consistency of the correlation structure, not about geometry — and any claim that it constrains something geometric has to earn that separately.

05 · Why the numerics were far harder than the theorem

#

The executed test used a smooth base on a periodic chain: correlation falling off as a Gaussian in the circular distance between sites, evaluated on grids of 256 and 512 points at several correlation lengths. This is about as benign as a covariance matrix gets. It is also, at the declared parameters, numerically catastrophic.

Because the base is circulant, its eigenvalues are available in closed form as a cosine sum, and they can be bounded against a theta series. The result at the worst declared cell: the smallest eigenvalue is positive, but positive at the scale of 10−76, while the largest is about 15. The condition number exceeds 7 × 1076. Every one of those matrices is mathematically, provably strictly positive definite — and completely unresolved in ordinary double precision, where the smallest eigenvalues are indistinguishable from zero and forming the inverse would produce numerical noise dressed as an answer.

Two decisions that made the evaluation honest

Never form the inverse. Because the base is circulant and the two profiles are translates of one another, the whole bound collapses to a single spectral sum in which the dangerous subtraction is done analytically rather than numerically: the difference of the self- and cross-terms becomes a sum of non-negative contributions weighted by 1 − cos(2πrΔ/N), divided by the eigenvalues. Two nearly equal huge numbers are never subtracted, because the subtraction was carried out symbolically first.

Certify the arithmetic, do not trust it. Every prediction was computed at a fixed 120 decimal digits, then recomputed at 160, with certification requiring the two to agree to a relative 10−30. Failure to certify would have been reported as unresolved — explicitly without adjusting precision until it passed, which is the standard way this kind of check silently becomes circular. All 24 cells certified, with observed agreement around 10−46.

06 · What was fixed before anything ran

#

The bound is a theorem, so it does not need an experiment. What needed testing was everything around it: whether the boundary is where the code says it is, whether the perturbation does anything geometrically visible, and whether the effect is localised to the pair it was injected between. Those are empirical questions about an executable, and they were pre-registered — profiles, grids, tolerances, contrasts and decision rules all hashed before the scan existed.

The three registered decision rules, as written before execution
Rule What it required Why it was written that way
DR1 Every distance curve monotone, and the smallest total decrease across the injected pair must exceed three times the largest drift seen at any control pair, anywhere in the battery. A localisation test. A perturbation that moves the injected pair and everything else equally has not localised anything, however impressive the raw movement looks.
DR2 In every cell, the first eigenvalue observed below −10⁻¹⁰ on a fixed λ grid must fall within 2% of the certified analytic prediction. An implementation test, not a test of the theorem. It asks whether the code, the grid and the tolerance actually bracket the boundary the mathematics predicts.
DR3 Response at the boundary must exceed three times the matched control response, while the global curve stays below a tolerance carried over from an earlier campaign. A specificity test with a guard against a global artefact. Its second clause referenced a number that, it turned out, had never been defined.

07 · What came back

#

The capsule completed twice with byte-identical output — a 2.4 MB result file with a recorded SHA-256, inside a repository manifest of 108 checksummed entries. All 24 cells completed; all 24 precision certifications passed; all 24 distance-validity and spectral-trace checks passed. Then the decision rules were applied exactly as written.

DR1 — the injected correlation did not localise
All 24 distance curves were valid and monotone, which is the behaviour you would want. But the smallest total decrease at the injected pair was 1.973, against a required threshold of 152.105 — three times the largest control drift observed anywhere in the battery, which came from an entirely different cell. Verdict recorded: handle_forms = false. The curves move; they do not move specifically.
DR2 — the rule failed without the theorem failing
In 20 of 24 cells the first crossing landed at 1.01× the predicted boundary — inside the 2% window, exactly as intended. In the four most ill-conditioned cells, the minimum eigenvalue at 1.01× the boundary sat around −4 × 10−15: real, negative, and correct in sign, but never reaching the registered −10−10 threshold, because at those conditioning levels double-precision residuals never get that large. The registered rule tests a crossing, not the boundary itself. So the battery-wide flag stands at lemma_verified = false, and the note states plainly that the lemma is not falsified — the broken hypothesis is threshold resolution, and it is named as such.
DR3 — not adjudicable, and left that way
The raw diagnostics were fine: endpoint response exceeded three times its matched control in 23 of 24 cells, the single exception being a ratio of 1.235 where both quantities sat at machine scale. But DR3's second clause referred to a global tolerance from an earlier campaign that does not exist. There is no defensible way to finish the comparison, and inventing a tolerance after seeing the output would have made the rule meaningless. The result was recorded as not adjudicable, missing referenced tolerance — a status, not a class — and no interpretation beyond that is licensed.
Summary
An exact theorem, a certified evaluation, and three registered rules of which one passed in the sense intended.
That distribution is worth sitting with. The mathematics is not in doubt, the arithmetic is certified to 30 significant figures, and the empirical wrapper still returned one localisation failure, one resolution failure, and one rule that could not be evaluated at all. This is the normal condition of computational research. It is usually invisible because the wrapper gets adjusted until it agrees with the theorem.

08 · One remark that had to be made explicitly

#

In the parent programme there is a proposition stating that the observable correlator is real-analytic regardless of how singular the underlying structure is — smoothing screens out sharp features. It is tempting to reach for it here and conclude that the whole question is closed by fiat.

It is not, and the note says so. The perturbation used here is built from analytic bump profiles, so the perturbation is itself jointly analytic — a finite sum of products of analytic functions. The smoothing proposition screens non-analytic sharpness; it has nothing to say about a smooth rank-two off-diagonal term. Analyticity neither excludes nor establishes the effect being tested. Recording that is a small thing, but it is exactly the kind of small thing that otherwise gets used later as a proof of something it never proved.

09 · Honest limits

#
  • The bound is finite-dimensional. Everything above is about matrices. The infinite-dimensional statement, on operators, is not proved here and is not implied by these numerics.
  • The battery is one base family. Twenty-four cells across two grid sizes, two correlation lengths, two profile widths and two separations — all with a Gaussian base on a periodic chain. The theorem is general; the executed evidence is not.
  • Two of the three empirical verdicts are negative and one is void. Nothing on this page should be read as a demonstration that the effect being tested exists. The strongest supported statement is that the boundary is exactly where the algebra says it is, wherever the numerics can resolve it.
  • Ratios at machine scale are descriptive only. Several endpoint-to-control ratios in the battery are enormous purely because the control response is at floating-point noise level. They are reported, and they are explicitly not evidence.
  • The related question about causal structure is untouched. The programme's own formulation of it is gated behind an explicitly declared continuation assumption and stops at a question — including the observation that stationarity alone does not exclude closed timelike curves, and that the missing object for any exclusion argument is a derived global time function. No theorem was attempted, and none is claimed.

10 · Where this sits

#

As a piece of applied linear algebra the bound is reusable well outside its origin: any setting where a covariance matrix is edited, stress-tested, or synthesised with an imposed dependence between two blocks has the same question and the same closed-form answer, including the pseudo-inverse variant for rank-deficient estimates. It is short-note grade mathematics, offered as such.

As part of a physics programme it is one executed capsule among several, and it is a good example of why that programme publishes its negative verdicts: the theorem is the cheap part.

  • Programme context — the open programme, which explains why a correlation kernel is the object of study in the first place, and what remains gated behind it.
  • Companion result — why the Gaussian keeps winning: the same positivity requirement, used the other way round, eliminates an entire family of candidate dynamics.
  • The protocol — freeze first. Why the decision rules above could not be adjusted once the output existed, and what that cost.
  • Classical background — the result is elementary given Schur complements, congruence-invariance of inertia, Cauchy–Schwarz, Bochner's theorem for the circulant spectrum, and the Moore–Penrose inverse for the semi-definite extension. No modern machinery is involved.
Evidence record — what each claim rests on
Claim Status Artifact and how to check it
λ_crit = 1/(√(ab) − c) is the exact positive-semi-definiteness boundary for a rank-two symmetric perturbation. Derived Half-page proof, given in outline in section 03 and in full in the internal note FW18_NOTE.md §1.2. Hypotheses: finite dimension; base symmetric and strictly positive definite; profiles not positively proportional. Verifiable by hand without any artifact.
Every base matrix in the declared battery is strictly positive definite, despite a condition number above 7 × 10⁷⁶. Derived Circulant-eigenvalue bound against a theta series, worst cell N = 256, ξ = 6: positive term 1.056 × 10⁻⁷⁶ against a remainder below 1.579 × 10⁻⁹⁹. A statement about the declared cells only, not about every Gaussian kernel on every grid.
The analytic prediction is certified in all 24 cells at 120 and 160 decimal digits. Executed outputs/fw18_correlation_handle.json — 2,399,706 bytes, SHA-256 155354db…4157d7d8, produced twice byte-identically, inside a repository manifest of 108 entries (SHA-256 6ffeabea…451df364). Observed relative agreement ≈ 10⁻⁴⁶.
DR1: the injected correlation did not localise (1.973 against a required 152.105). Executed — negative Same JSON. Rule text and threshold fixed before execution; verdict recorded as handle_forms = false.
DR2: the registered crossing test failed in 4 of 24 cells; the lemma itself is not falsified. Executed — negative on operational grounds Same JSON, non-bracketing audit table. The broken hypothesis is threshold resolution at double precision, named as such in the note; the battery flag stays lemma_verified = false.
DR3: could not be evaluated, because the tolerance it referenced was never defined. Not adjudicable Same JSON: classification null, status not_adjudicable_missing_referenced_t2_global_shift_tolerance. No tolerance was chosen after the output existed.
Peer-review status of everything above. None Not submitted, not refereed, not preprinted. The parent framework article is in journal review; this note is not part of that submission.

Research and AI disclosure

This page describes a short mathematical result and a pre-registered numerical study. The result is elementary and stated with its hypotheses; the empirical verdicts around it are mostly negative and are reported as such. Nothing here has been peer reviewed.

AI assistance was used in deriving, executing and writing up this work, and AI-generated output can contain errors. The lemma should be checked by anyone intending to rely on it — it is short enough to verify by hand in under an hour, which is the main reason it is worth publishing.

Research and correction enquiries: contact@kort-x.com.

Continue
Related records from the same programme.
The theorem took an afternoon. Establishing that the code agreed with it took 120 digits, and it still only agreed in twenty cells out of twenty-four.
Article record
Author
Ciprian Stoichici, KORT-X Research, Bucharest, Romania
Version
1.0
Published
2026-08-09
Updated
2026-08-09
Licence
CC BY 4.0
Cite as
Stoichici, C. (2026). "The Correlation Budget: an exact bound on perturbing a covariance structure." KORT-X Research. https://kort-x.com/indexfiles/research-correlation-budget.html

This is a laboratory write-up, not a refereed publication. Where a claim rests on an executed artifact, the evidence record above names the artifact and its status; internal working artifacts are not part of the public release and are available on request. Corrections are welcome and are applied in place with the update date changed.

↑ Top
SINGULARITY_V2