LABORATORIES
Knowledge Orchestration for Radical Transformation - X
GitHub

NULL SIGNAL

Applied Statistics · Randomness Auditing · Revealed Preference

Roughly sixty model versions were built to predict a national lottery, across three games and two and a half years. Each generation looked better than the last. None of them worked. This is the complete record — every architecture, every bug, every honest win, every illusion that had to be dismantled — and the one measurable signal that survived, which was never in the draws at all.

~60
Model versions built and tested
2,565
Draws under audit
70
Statistical tests, one Holm family
0
Anomalies surviving correction
845
Draws of prize data — where signal was found

01 · A laboratory where you know the answer in advance

A mechanical lottery draw is about the closest you get to a process whose null hypothesis you can assume before you start: the machine is built and audited to be fair, and the balls have no memory. That is an assumption about the equipment, not something the data can prove — but it is the strongest prior available anywhere, which makes this the rare case where you roughly know the right answer in advance. Normally that is a reason to leave the problem alone. Here it became the reason to take it up.

A pipeline that reports an edge on a domain with this strong a fairness prior has not discovered anything about lotteries. It has discovered something about itself. Every false positive it produces is a defect in the method — one that would have passed unnoticed on messier data where the truth is unknown. So the project was run to its natural end: build a genuinely serious modelling stack, point it at a process that should contain nothing, and record honestly what it claims.

The subject was the Romanian national lottery — Loto 6/49, Loto 5/40 and Joker (5 of 45 plus a Joker ball of 20). The 6/49 history runs to 2,565 validated draws going back to 1993; Joker to 398; 5/40, a newer game, to 61. Every record was eventually reconciled against the official operator archive and cross-checked against independent secondary sources — but, as section 06 records, several integrity failures surfaced only months in, after they had already shaped results.

What follows is written the way a laboratory notebook should be: the failures at the same resolution as the results.

02 · The complete version history

It began in January 2024 with a set of TensorFlow specifications and a spreadsheet of 102 draws. It ended on 8 August 2026 with a seventy-test randomness audit and a decision to stop. Note that the file timestamps do not show this: archiving into DEPLETED reset them, so nothing on disk appears older than March 2026. The working record kept in conversations/ covers the final five months in detail; the two years before that are the neural-network era, which survives as designs and results rather than as dated files. Seven distinct eras, each ending because the era's core assumption broke.

Era 0 · January 2024 — the networks that could not be run
Eleven TensorFlow architectures, several literally filenamed “nu am resurse” — I don't have the resources. Up to 1,000 hidden layers of 2,048 neurons; 3D-interconnected designs with 13,884 units; LSTM stacks with hyperparameter search. None had ever been executed. All eleven were ported to scikit-learn MLPClassifier on CPU so that they could at least be measured. The port itself was the first result: eight of the ten converged on a prediction that simply echoed the previous draw. With 102 training rows in a row-N → row-N+1 setup, the networks were memorising, not learning.
Era 1 · ALG12-PRO V1 – V3 — the post-mortem trap
A “proximity calibration” pass was fitted knowing the target draw, and reported 4 of 5 hits. It was tuned on the answer. The one durable idea from this era was a reframing: stop asking “which number goes in position 3?” and start asking, for each of the 45 numbers independently, “will this number appear?” That binary per-number framing survived every later rewrite. The weight search also revealed that the neural network contributed 6.4% of the final signal — co-occurrence and proximity carried 79%.
Era 2 · V4 – V12 — signal blending, and the bug museum
Frequency, recency, overdue/gap, transition matrices, Markov chains, Monte Carlo, statistical-constraint generators, Fibonacci gap analysis, 9×5 spatial grids, pair binomials, exponential decay. Backtest averages crawled 0.7 → 1.0 → 1.2 → 2.0 hits per five-number ticket. Almost all of that movement turned out to be bugs: edge numbers 1 and 45 receiving artificial boosts, hard-coded random seeds freezing predictions across data updates, six “sticky” numbers appearing in the top-15 at every single state, and an anti-repeat penalty so aggressive it excluded numbers that genuinely do repeat.
Era 3 · V13 – V19.2 — machine learning versus domain knowledge
V17 replaced hand-tuned weights with logistic regression and a clean SQLite core. Its candidate pool was measurably better — 3.36 of 5 true numbers versus 3.08 — and its tickets were far worse: selection efficiency collapsed from 23.2% to 14.5%, and p rose from 0.026 to 0.719. The ML ranked individual numbers well and had no idea what a plausible combination looks like. V18 hybridised the two: ML pool, heuristic selector. That worked, and set the pattern for everything after.
Era 4 · V19 — the matched-cost baseline
Every portfolio result up to this point had been compared against one random ticket. Comparing three model tickets against three random tickets ended the illusion in a single line: random best-of-three scored 1.104 hits, the optimised portfolio 1.025, p = 0.881. Below random. Four further evaluation defects were fixed in V19.1 — a Platt-calibration leak, a Brier score divided by five, a Joker distribution normalised to sum 5 instead of 1, and an alpha/beta mismatch between backtest and live. None of them changed the prediction. All of them changed what the numbers meant.
Era 5 · V20 – V27 — novelty mathematics and genetic search
Ten exotic modules were implemented in one pass: persistent-homology void pressure, quantum-walk graph diffusion, Fisher-information velocity, symbolic-dynamics grammar, Kolmogorov/LZ complexity weighting, harmonic wavelet momentum, Tsallis non-extensive weighting, combinatorial design deficit, Gibbs/Potts field energy, and Joker–main transfer entropy. A genetic algorithm then took joint control of 33 genes. It drove four of the ten module weights to exactly zero. V26 improved robustness genuinely; V27 added a two-ticket portfolio objective. And the headline “V23 average 0.825” that had justified the whole line was found to be the mean of the last two walk-forward windows, not all five. The honest figure was 0.610.
Era 6 · V28 – V31 and the parallel games
Chaos-theory diagnostics — recurrence quantification, detrended fluctuation analysis, entropy regimes — folded into five-ticket portfolio search over the full pool. In parallel, the 6/49 track built a matrix-binomial family whose database was rebuilt from 2,565 draws, and the 5/40 track began with a database that did not exist: 38 draws were read directly out of a screenshot and written to a workbook. V31 reported p = 0.0000 on a 30-draw best-of-five backtest. That is the multi-ticket illusion again, now with more tickets.
Era 7 · July – August 2026 — audit, gate, stop
An independent audit found systematic model-selection leakage. A purpose-built evidence-gated selector declared NO MODEL PASSED and refused to ship. A seventy-test randomness audit then answered the prior question directly and closed number prediction as an avenue. The remaining question — what a win pays — turned out to have a real, validated, exactly computable answer.

03 · The compute wall that turned out not to matter

The project's original ambition was limited by hardware, and several source files carried that limitation in their names. The specifications were sized for machines that were never available. Estimated requirements, as assessed at the time of the port:

Original TensorFlow specifications and their hardware requirements
Design Original specification Est. VRAM Realistic hardware
ALG 71,000 layers × 2,048 neurons~16 TBImpossible as designed — over 4 billion parameters
ALG 310 layers, ~85M neurons/layer32 GB+A100 80 GB
ALG 10–123D interconnected, 13,884 units~3–4 GBRTX 3060+
ALG 22LSTM + hyperparameter search2–4 GBRTX 3060+
ALG 650 layers × 1,024 neurons~2 GBGTX 1660+
ALG 54 layers × 10,000 neurons~1.5 GBAny modern GPU
ALG 92 layers × 13,878 neurons~1.5 GBGTX 1660+
ALG 20Conv1D + 2,056-unit ensemble~1 GBAny modern GPU

Every one of these was ported down to a CPU-sized scikit-learn equivalent, and the ported versions ran in about two minutes on a laptop. Then came the finding that retired the entire complaint: an architecture search across the ported family selected the smallest network in the set — two hidden layers of 32 units — as the best performer. With roughly 100 training rows, the small network generalised and the large ones memorised.

The binding constraint was never compute. It was data, and the data was intrinsically uninformative. An A100 would have reached the same conclusion faster and at greater expense. This is worth stating plainly because “if only I had more GPU” is one of the most common and most comfortable explanations for a null result, and here it was measurably false.

04 · What actually worked

Very little, and none of it in the way it was expected to. These are the ideas that survived contact with an honest backtest — several of them survived by being useful engineering rather than useful prediction.

Binary per-number framing
Asking “will number n appear?” for each of the 45 numbers independently, instead of “what number sits in position 3?”. This single reframing ended the memorisation failure of the neural era and became the input layer of every subsequent version.
Structural combination scoring
Sum band, spread, odd/even balance, decade coverage and pair co-occurrence, applied as soft penalties rather than hard filters. Replacing hard pass/fail gates with continuous penalties doubled the ≥2-hit rate from 8% to 16% in a matched comparison. It is the difference between a plausible ticket and an arbitrary one — and it is the one place where domain knowledge consistently beat machine learning.
Hybrid architecture: ML pool, heuristic selector
Logistic regression ranked individual numbers better than hand-tuned weights (3.36 vs 3.08 true numbers in pool) but selected terrible tickets. Using ML for the shortlist and structural scoring for the final combination kept both advantages. Selection efficiency stayed at 21.2% against ML's own 14.5%.
Robustness-weighted fitness
Optimising mean ÷ (1 + σ) across walk-forward windows instead of raw mean. This genuinely stabilised the model: window variance fell from σ = 0.187 to σ = 0.053 while the mean also rose. It did not create an edge, but it stopped the search from chasing lucky peaks.
Conditional anti-repeat
Penalising a number from the previous draw only when its support comes from raw frequency noise, not when transition and context signals independently back it. The flat penalty had been costing real hits — repeats happen in roughly half of all draws.
The evaluation framework itself
Confidence intervals, exact hypergeometric baselines, p-values, Brier scores, matched-cost baselines, and a frozen pre-draw ledger. This produced no better tickets whatsoever and was by a wide margin the most valuable thing built. Every real finding in this project came from the evaluation layer, not the model layer.

05 · What did not work

Each of these was implemented, measured, and rejected on its own evidence. Several had been reported as successes first.

Overdue / gap analysis
The most intuitively appealing signal in the whole project, and the most thoroughly dead. Numbers at 3.9× and 4.5× their expected gap were tracked repeatedly and did not arrive. Tested again on 700+ draws: no predictive value. Every attempt to weight it upward degraded the backtest.
The ±1 near-miss pattern
For months, nearly every prediction landed within one or two of a drawn number, and it felt like a systematic bias waiting to be corrected. It was measured directly: the neighbour transition rate is 11.03% against a random expectation of 11.11%. A lift of 0.99×. The pattern was pure apophenia — with five numbers drawn from 45, something is always nearby.
Multi-ticket portfolios
Reported as the strongest result in the project for several versions running, at p = 0.000. That comparison was against a single random ticket. Against three random tickets, three optimised tickets scored 1.025 versus 1.104 — worse than random, p = 0.881.
Consensus across models
Combining twelve generated sets and taking the most frequent numbers produced an average ticket that resembled no real draw — five odd numbers against a typical distribution of two or three. A weighted-vote ensemble across six 6/49 models likewise scored below its own best single component. Averaging smooths variance and destroys whatever unique signal any one model had.
Window-length optimisation
Ten history windows from 100 to 701 draws were compared and 350–400 declared the “sweet spot”, later revised to 200. The comparison rested on two draws. Ten hypotheses ranked on two outcomes is not an optimisation, it is a coin flip with extra steps.
Sum filters
Audited properly with paired bootstrap intervals over 220 walk-forward predictions. Every interval crossed zero. The filter mainly forces a historically ordinary-looking ticket while discarding highly scored numbers. It was also discovered that the published benchmark used a 30–70% band while the live configuration used 20–80% — the headline never validated the shipped setting.
Filter-first architecture
Converting the scoring modules into consensus gates — a weighted bitmask with dynamic entropy thresholds, forbidden-pair rejection and anti-saturation logic. Architecturally the most elegant idea attempted. It collapsed the shortlist too early: 0.525 average against 0.825 for the version it was meant to replace.
Long-term historical priors
Frequency charts back to 1993 show numbers 10–13 persistently under-represented and 33, 29, 5 over-represented across 3,000+ draws. Injecting that as a prior degraded the backtest at every weight tested. The apparent thirty-year bias is within sampling variance, and the model was already capturing anything usable.
Fibonacci gap structure
48.4% of gaps between drawn numbers were Fibonacci against an 18.2% expectation — briefly the most exciting number in the project, used as a tie-breaker in V9. The Fibonacci set {1, 2, 3, 5, 8, 13, 21} covers most of the plausible gap range for 5-from-45; the “finding” was a property of the definition, not of the draws.
Everything else
Monte Carlo generators and Markov chains (both fixated on edge numbers 1 and 45), positional CDFs, hot–cold blends, neighbour boosts, decade-balance guarantees, spatial-grid heatmaps, transfer-entropy Joker coupling, quantum-walk diffusion, Tsallis weighting, symbolic dynamics, and a deliberately non-statistical SHA-256 “world-state hash oracle” seeded on moon phase and asteroid-mission headlines. All measured. None beat chance.

06 · The bug museum

These are not incidental defects. Each one produced a number that was believed, acted on, and in several cases carried forward across multiple sessions before it was caught. The pattern is worth more than the individual entries: every single one inflated the result rather than deflating it.

  • The headline that was two windows. “V23 average 0.825” was quoted as the project's best result for weeks. It was the mean of the last two walk-forward windows out of five. The honest five-window figure was 0.610 — below the version it supposedly beat.
  • The 4-of-5 blind hit that was a stale row. A blind prediction scored 4 of 5 plus the Joker. The database had been supplied with an off-by-one final row, so the “future” draw being predicted was already in the training data. Corrected, the same predictions scored 0–1 of 5.
  • Frozen random seeds. Seeds 42, 123 and 77 were hard-coded. Every database update produced identical predictions, which read as impressive stability and was actually a dead pipeline. Caught only because the user noticed the same numbers returning.
  • Edge-number bias. Numbers 1 and 45 appeared in nearly every backtest prediction. They sit at the ends of the range, so they have neighbours on one side only and escaped a dispersion penalty that applied to everything else.
  • Sticky numbers. Six numbers — 5, 37, 29, 10, 12, 28 — occupied the top fifteen at every one of twenty tested states, regardless of what had been drawn. Frequency and recency together were overwhelming every conditional signal.
  • The self-flattering backtest. One version tuned and reported on the same short recent slice, compared itself against a remembered figure for the previous version rather than re-running it, blended the Joker into the average as +0.5, and described 19 predictions as 20.
  • Silent exception swallowing. The 5/40 walk-forward caught every predictor exception and recorded it as zero hits. A completely broken model would have scored as a merely bad one, and model selection would never have noticed.
  • An evaluation loop that never ran. A variable named t in the rolling-origin evaluator shadowed the loop index, so the function returned empty every time. The metrics printed anyway.
  • Two versions that were one version. V13 and V14 emitted identical tickets — Platt calibration is a monotone transform and cannot change a ranking. Separately, V15 turned out to be V12 under a new filename.
  • Statistical hygiene defects. A Platt calibrator fitted on its own training fold; a Brier score divided by five; a Joker distribution normalised to sum 5 instead of 1; blend weights of 0.3/0.7 in backtest against 0.1/0.9 in production.
  • Data-integrity incidents. A database believed to span three years actually covered one. Pre-2024 Joker balls ranged to 45, not 20, silently corrupting 478 rows. A draw was dropped when its successor was appended. An Easter supplementary draw was recorded as a regular one. A to_excel round-trip destroyed a multi-sheet workbook. A search summary reported a drawn number that two official sources contradicted.
  • The sweep that was 1,712 hypotheses. A binomial parameter sweep ranked 1,712 configurations on exactly the outcomes it then reported as its test result — and every top configuration emitted the same live ticket, which was a property of the recent window rather than a tunable edge.
  • The bin mismatch, found by an external reviewer after publication. In the draw-sum test, observed counts were bucketed with [lo, hi) while the expected probabilities used (lo, hi]. On integer sums, boundary draws landed in different bins on the two sides of the comparison. Correcting it moved that test to p = 0.0312. Entry thirteen in a list written to argue that this kind of thing is never finished.

07 · The wins, honestly counted

There were real hits. Every ticket below was fixed before its draw. They are recorded here in full, best and worst together, because a results table that only contains the good draws is exactly the failure mode this project exists to document.

Selected pre-draw tickets and their outcomes
Draw Game Ticket Actual Hits Source
2026-06-036/498-18-19-39-40-462-6-18-32-39-403 / 6matrix_binomes_ge3
2026-06-036/497-14-18-20-39-462-6-18-32-39-402 / 6consensus_selector
2026-04-10Joker5-14-18-28-385-25-34-38-402 / 5V19.1
2026-03-08Joker14-18-23-38-41 + J1316-18-20-24-41 + J102 / 5StatCons v2 (W350 S3)
2026-03-08Joker7-13-29-37-44 + J1016-18-20-24-41 + J10Joker hitV4 rebalanced
2026-08-066/496-18-19-25-39-465-6-16-25-41-422 / 6external chat
2026-03-21Joker5-15-18-24-37 + J158-9-22-33-40 + J150 / 5, Joker hitV9
2026-06-066/4922-30-39-43-45-465-14-21-34-40-431 / 6matrix_binomes_ge3
2026-05-31Joker7-10-18-28-37 + J151-9-26-36-43 + J140 / 5V31 A + V19.2 joker
2026-05-27all three3 tickets—0, 0, 0best-by-backtest each
2026-07-126/4929-32-35-46-47-499-20-23-24-36-430 / 6evidence-gated fallback
2026-08-026/493-33-36-37-43-4612-13-30-32-48-490 / 6evidence-gated fallback

The 3-of-6 is the best live result the project ever produced, and it came from the simplest model in the repository — strict pair-binomials over three-draw matrices, no frequency term, no machine learning, no tuning. At the time it was taken as validation. Its own backtest reported a ≥3 rate of 4.11%, and the exact probability of 3 or more matches on a random 6/49 ticket is 1.86%. One event at roughly 1-in-50 odds, drawn from dozens of tickets across three games, is precisely what chance looks like.

Two hits happen 13.24% of the time by pure chance. Across the whole ledger the average sits where the mathematics says it must: around 0.735 hits per 6/49 ticket, 0.735 being exactly 36/49.

08 · How the edge was manufactured

An independent audit of the 6/49 stack found that the chronology was clean — no target draw was ever fed to a predictor meant to forecast it. The leakage sat one level up, in model selection, and it was systematic. This is the failure that matters, because it requires no carelessness about time ordering, which is the one everybody is trained to look for.

01 Sweep-and-report 1,712 configurations ranked on the very outcomes then quoted as the test result.
02 Design-on-test Thresholds and weights tuned against the same last-100 window used to report them.
03 Post-miss refitting Every post-mortem mechanism was, by construction, guaranteed to improve the window it was built from.
04 False replication “Recent” and “full-history” comparisons shared their final 100 targets — never independent confirmations.
05 Mismatched baseline Multi-ticket portfolios scored against a single random ticket, turning arithmetic into significance.
06 Convergent artefact All 1,712 configurations emitted the identical ticket — one property of the window, not a tunable edge.

None of this needed bad intent or sloppy code. It only needed a large model space, a short evaluation window, and the entirely natural habit of reporting the winner without paying for the search.

09 · The honest comparison

Every model was re-run strictly causally over the full history — 2,561 predictions — against the exact hypergeometric expectation of 36/49 = 0.734694 matches per ticket.

Full-history causal evaluation · 2,561 predictions · Loto 6/49
Model Avg hits ≥3 rate Observed ≥3 Expected p
product0.75281.76%4547.70.682
markov0.74191.99%5147.70.343
decade0.73761.87%4847.70.513
binom0.73291.68%4347.70.778
random0.71611.64%4247.70.819

The binomial model — the one with the most favourable recent window, the most development effort behind it, and the project's single best live result — scores below chance across the full history. The spread of the whole table is consistent with noise around 0.735.

An evidence-gated selector was then written to make the standard explicit rather than optional: declare a small model family in advance; evaluate strictly causally over 2,365 targets; compare against the exact hypergeometric distribution; Holm-correct across the entire family; require significance on both the full run and the most recent 400 draws, plus improvement in four of five chronological blocks. Its verdict on all nine declared models was Holm p = 1.0000 — NO MODEL PASSED — and it fell back automatically rather than shipping a number it could not defend.

10 · Advanced mathematics, and what it returned

Before accepting a null, it was worth bringing the strongest available structure-detection tools to bear. If instruments this powerful find nothing, the finding is close to definitive. Each was run against a shuffled control of the same data.

Structure-detection diagnostics against shuffled controls
Method What it would detect Real Shuffled
Takens embeddingA low-dimensional dynamical attractor2.292.35
Persistent homologyTopological features above noise2.5252.496
Spectral gapSlow mixing, i.e. memory of prior state0.911≈ max
Lag-1 recurrenceOverlap beyond the hypergeometric law0.6110.556

The correlation dimension of the real series is indistinguishable from that of its own shuffle — there is no attractor, the draws fill the space uniformly. Topological feature lifetimes give a real-to-shuffled ratio of 1.011. The transition matrix has eigenvalues λ₁ = 1.0 and λ₂ = 0.089, a spectral gap of 0.911: the system forgets its previous state almost immediately, where a structured process would sit below 0.3. An FFT did surface a periodic component near 2.4 draws with a signal-to-noise ratio of 6.67, which on inspection is the natural oscillation of draw sums — high sum, low sum, high sum — and carries no information about which numbers appear.

Two research programmes were investigated at the suggestion that deep modern mathematics might apply. The Langlands programme — automorphic forms, Galois representations, Shimura varieties — concerns hidden symmetry in arithmetic structures. A lottery draw has no number field, no Galois extension and no L-function to analyse; the connection is philosophical, not technical. Floer homology and gauge theory operate on continuous 3- and 4-manifolds, where a lottery is a finite discrete space with no non-trivial topology. The one genuinely transferable descendant of that world was topological data analysis, which was implemented, run, and returned the null above.

The instructive part

A genetic algorithm given free control over the ten exotic modules drove four of their weights to exactly zero — persistent homology, symbolic dynamics, Tsallis weighting and analogue matching. The optimiser was reporting the same null the diagnostics found, months earlier, in a form nobody read. When a search assigns zero weight to a feature family, that is a measurement, not a tuning artefact.

11 · Asking the prior question

Backtesting a predictor asks a weak and multiplicity-prone question. The prior question is stronger and has one answer: is this draw process distinguishable from a fair 6-of-49 sampler at all?

Seventy heterogeneous tests were run and corrected as a single Holm family — marginal frequency χ² on the full history and on the 2024+ era, per-position uniformity for all six drawn positions, lag-1/2/3 overlap against the hypergeometric law, 49 Wald–Wolfowitz runs tests, gap and waiting-time against the geometric law, draw-sum against an exact enumeration of all C(49,6) subsets, odd-count and low-count against hypergeometric, pair co-occurrence dispersion, first-half against second-half frequencies, and pooled autocorrelation at lags 1, 2, 3 and 7.

raw p < 0.05 : 3 of 70 (expected by chance ~3.5) Holm-adjusted p < 0.05 : 0 global KS uniformity : D = 0.1577, p = 0.0548

Fewer raw hits than chance alone would produce. The borderline KS statistic was checked for direction, because a two-sided p near 0.05 is exactly the kind of result this project had learned to over-read. The deviation runs toward large p-values (one-sided p = 0.027), the signature of conservatism from discreteness and cell pooling. The direction that would indicate real structure gives p = 0.77.

Verdict
No departure from fairness detected — in Loto 6/49.
Across 2,565 Loto 6/49 draws, seventy tests found nothing that survives correction. That is a failure to detect, not proof of absence — it bounds any surviving structure to effects too small for this much data to see. For the practical question it is still decisive: it explains why every model in the repository, across sixty versions and three games, lands on the same 0.735 average. Joker and 5/40 were modelled but never given this audit, so the formal result covers 6/49 alone.

Calibrating the audit against simulated fair histories

External review made the point that the p-values above are analytic, and that several of these statistics cannot have the null distribution the formula assumes — gap observations, pair counts and the 49 per-number indicator series are all mutually dependent inside a draw, and the global KS test inherits that dependence. So the battery was calibrated empirically: 400 complete fair 6-of-49 histories of the same length, each pushed through the identical code. Same review also found a real defect in the draw-sum test, where observed counts were bucketed left-inclusive while the expected probabilities were right-inclusive; that is fixed, and the corrected test moved to p = 0.0312, taking the raw count below 0.05 from two to three of seventy.

The calibration rewrites the interpretation, and not in the direction that flatters the original write-up. On genuinely fair data this battery fires far more often than its nominal level: the median fair history produces nine raw p < 0.05 out of seventy, not the 3.5 the nominal level implies, and more than half of all fair histories contain at least one test that survives Holm correction outright. Measured against that null, the observed history is not merely unremarkable — it is quieter than chance. The borderline KS statistic that needed a directional argument to explain away is simply ordinary: empirical p = 0.177. The verdict stands and stands harder. What does not stand is the precision the original analytic p-values implied.

statistic observed null median null 5% null 95% emp. p ------------------------------------------------------------------ min raw p 0.0312 0.0000 0.0000 0.0000 1.000 count p < 0.05 3 9 6.95 13.00 1.000 global KS D 0.1577 0.1182 0.0857 0.1975 0.177 min Holm p 1.0000 0.0000 0.0000 0.0000 1.000 400 simulated fair histories · 2,565 draws each · L649_audit_calibration.py

12 · The signal was never in the draws

The balls are random. The players are not — and in a pari-mutuel game the players are half of the outcome. Categories I, II and III split a fixed pool among everyone holding the winning combination, so what a win pays depends on how many other people chose the same numbers.

That makes the official winner counts a revealed-preference measurement of human number choice, published draw after draw for years, entirely free of the randomness that defeats prediction. The count of category IV winners is high precisely when the drawn numbers happen to be ones players favour.

A harvester was written against the operator's public results archive and recovered 845 draws, 2018-01-04 to 2026-08-06, with per-category winner counts and prize values. Ticket sales were proxied by the category III pool and the proxy was verified against the data: cat3_winners × cat3_value equals exactly one third of the category I contribution on every draw checked, so both are fixed shares of sales and neither is disturbed by a rolled-over jackpot.

catIV_winners ~ Poisson( sales · exp( a + Σ βᵢ·xᵢₜ ) )

Fitted by ridge-penalised Poisson maximum likelihood, coefficients centred for identifiability, sales entering as a log offset. Each βᵢ is the log-scale popularity of number i.

13 · Three independent validations

After sixty versions of being fooled by a flexible model on a flat surface, the burden of proof here was set deliberately high. Three checks, each answering a different objection, each capable of killing the result on its own.

Validation of the player-preference model
Check Objection it answers Result
Rolling-origin CV, held-out Poisson devianceDoes it generalise out of sample?1,080,301 → 659,282  (−38.97%)
Permutation test, numbers shuffled against countsIs it signal, or just model flexibility?p ≤ 0.002  (500 shuffles; null mean −15.4%)
Temporal split, 2018–2022 vs 2022–2026Is it stable enough to act on?r = 0.890  · 7/10 overlap in least-popular top ten
Bootstrap, 400 resamplesHow much of this is one lucky sample?34 of 49 coefficients exclude zero

The permutation result matters most. When the association between numbers and counts is destroyed, the fit collapses to well below the null — the improvement is not the model bending to accommodate noise, which is exactly what every earlier “win” in this project turned out to be. The temporal result is what makes the effect usable: an eight-year-stable human habit rather than a moving target.

Estimated player preference · log-scale β
Over-pickedβ Under-pickedβ
13+0.25046−0.183
27+0.24242−0.179
7+0.21849−0.172
9+0.17743−0.166
10+0.15832−0.164
5+0.14639−0.158

A textbook birthday effect: everything above 40 avoided, low numbers and lucky seven crowded. It independently reproduces published findings on the Israeli and UK lotteries — with one local inversion. Here 13 is the single most over-picked number, the opposite of the Western superstition, and a useful reminder that a preference model estimated in another country does not transfer without being re-measured.

What changed when these were made executable

Every figure in the table above is new. The original three validations were reported from a working session; only the cross-validation existed in code, and it used random folds despite being named for blocked CV. Written properly as L649_ev_validation.py — rolling-origin folds that only ever predict forward, 500 permutations, an explicit temporal split, 400 bootstrap resamples, and the ridge penalty selected by CV rather than fixed at 25 by hand — the effect survives, and three of the four published numbers move. Held-out improvement falls from 48.68% to 38.97%: random folds were letting the model peek at its own future, worth about ten points. Temporal correlation falls from 0.924 to 0.890 and the least-popular overlap from 8/10 to 7/10. The permutation test, previously three loose percentages, is now a proper empirical p at the floor of what 500 shuffles can report: the observed 38.97% sits far outside a null whose mean is −15.4% and whose best case is −6.9%.

The bootstrap adds something the original never checked, and it is the most important line here. The direction is solid — 34 of 49 coefficients have intervals excluding zero, and 13, 27 and 7 are unambiguously over-picked. The specific ticket is not: the exact six-number anti-crowd core reappeared in 1 of 400 resamples, with 237 distinct cores across the run. So the strategy replicates and the recommendation does not. Any single ticket presented as optimal is one draw from a very wide distribution of near-equivalent tickets.

14 · Exact combinatorics, not simulation

Under the fitted model a combination is chosen with probability proportional to the product of its numbers' weights. Normalising that over all C(49,6) combinations is the elementary symmetric polynomial e₆(w), computed exactly by the standard dynamic-programming recurrence rather than sampled. Expected rival winners in each category then follow:

λ_k = n_tickets · e_k(w_D) · e_{6−k}(w_¬D) / e₆(w_all) E[ P / (1+X) ] = P · (1 − e^(−λ)) / λ for X ~ Poisson(λ)

Categories I, II and III are evaluated by enumerating every draw consistent with 6, 5 and 4 matches — 1, 258 and 13,545 draws respectively. The inner polynomials over the 43 non-drawn numbers reduce to power sums, so each draw costs constant time. Two unit checks anchor the implementation: the dynamic program matches brute-force enumeration on a reduced pool to thirteen significant figures, and e₆(1,…,1) returns exactly 13,983,816. Category IV pays a fixed amount and is immune to sharing.

15 · What the measurable effect is worth

Exact expected return per ticket · rollover jackpot 50,062,357 · ~801,058 tickets in play
Ticket cat I cat II cat III cat IV Total EV vs birthday
32-39-41-43-46-493.5440.8510.6420.8835.919+20.2%
10-32-38-39-42-433.5270.6590.5530.8835.621+14.2%
03-07-11-17-22-283.4020.2940.3440.8834.9220.0%
The result in full
A 20.2% improvement on a −15.4% bet.
The optimised ticket returns 5.92 against a 7-unit stake, and even that figure is flattered by an unusually large rollover. The gain is real, exactly computed and reproducible — and it changes nothing about the probability of winning, which remains 1 in 13,983,816 for every combination. The honest summary of the whole thing: the draws carry no information, the players do, and even perfect knowledge of the players does not make this a good bet.

Stated limits of the preference model:

  • It does not improve the probability of winning. Only the conditional value of a win.
  • The product form assumes players choose numbers independently. The combination-level penalty for consecutive runs, arithmetic ladders and calendar clustering is literature-derived, not fitted.
  • β is estimated from 3-of-6 matches, so its transfer to exact six-number popularity is an extrapolation.
  • An all-high ticket is itself a known anti-crowd strategy. The model observes average behaviour and cannot see sophisticated players crowding the same idea.
  • The effect is measured on one operator's game in one country. The inverted position of 13 is direct evidence that these coefficients do not travel.
  • Ticket volume is estimated from a single draw’s category-IV winner count divided by P(exactly 3 of 6). But the whole preference model says that count depends on which numbers were drawn — so the sales estimate is confounded with the very effect being measured, and rests on one draw. It should come from published sales or the statutory pool split.
  • The enumeration is exact; the expected value is not. It inherits estimated popularity coefficients with no reported intervals, an independence assumption between numbers, an unfitted pattern penalty, a Poisson approximation for rival winners, and the ticket-volume estimate above.
  • The recommended ticket is not stable. Under 400 bootstrap resamples the exact six-number core reappeared once. Treat any single published combination as an illustration of the strategy, not as the output of it.
  • Ticket volume swings between roughly 615,000 and 1,480,000 across the last ten draws when estimated from the category-III pool, so any expected value quoted for one draw inherits that spread.

16 · Why the negative result was the point

The transferable output of this project is not a ticket. It is a calibrated set of practices, each one earned by a specific, documented failure observed on a surface where the right answer was known in advance:

Test the null first. Ask whether the domain contains signal before ranking models that assume it does. Pay for the search. A sweep of 1,712 configurations is 1,712 hypotheses, not one result. Declare the family in advance. Correction is meaningless if the model list grows after seeing outcomes. Match the baseline to the cost. Three tickets must be compared against three random tickets. Gate the decision, not the paper. A selector that refuses to ship is worth more than one that always answers. Freeze predictions before the event. An append-only ledger that rejects conflicting entries removes hindsight entirely. Check the direction of a borderline p. Two-sided significance near 0.05 is where over-reading lives. Never trust a remembered number. Re-run the baseline; do not quote the previous version from memory. Audit the data before the model. Half the apparent findings here were database defects. Read the zeros. When an optimiser assigns zero weight to a feature family, it is reporting a null. Look for signal in the observers. When the process is random, the humans around it usually are not.

This matters directly to the rest of the KORT-X programme. A framework such as Projective Correlation Theory lives or dies on whether its confirmations survive the same scrutiny, and cosmological data does not come with a known answer at the back of the book. A pipeline that cannot resist manufacturing an edge on a lottery cannot be trusted to report an honest null on anything harder. Every practice above now applies to that work.

17 · Algorithm sources

Every number on this page is derived from the archived code and the validated draw databases, but strict independent reproduction is not yet packaged — see the end of this section for exactly what is missing, with the exceptions named in section 13. The working tree holds forty-one Python modules across the three games, plus superseded versions kept for provenance. Grouped by track:

Code inventory by track
Track Modules Role
Randomness audit & EV L649_randomness_audit.py · L649_harvest_prize_data.py · L649_ev_optimizer.py · L649_evidence_guarded_selector.py The terminal stage. Seventy-test Holm-corrected audit, the 845-draw prize harvester, the exact expected-value optimiser, and the gate that refuses to ship an unsupported model.
6/49 binomial family L649_matrix_binomes_ge3_single.py · binom_ge3_sweep.py · L649_matrix_binomes_sumfilter.py · L649_decade_balanced_binom.py · L649_shrunk_binom_v2.py · L649_binom_recalibrated.py · L649_anchor_binom_ticket.py · build_l649_binom_databases.py Pair binomials over three-draw matrices. Produced the project's best live result and, on full history, scores below chance.
6/49 general & ensemble l649_model_compare_predict.py · L649_predictor_v2.py · L649_ge3_hunter_v3.py · consensus_selector_649.py · L649_markov_conditional.py · L649_binom_markov_product.py · L649_ensemble_debiased.py · L649_adaptive_companion.py · L649_postmiss_rebalanced.py · compare_all_backtest.py · v6_hybrid 649.py · run_v8_correct.py · run_v8_clean.py · v8_prediction.py Model comparison, Markov conditioning, weighted-vote consensus and the honest full-history head-to-head that ended the line.
Joker — production JokerPredictor_V19.2.py · V28 · V29 · V29_2 · V30 · V31 · JokerPredictor_SingleTicket.py · JokerPredictor_PairMemory.py · JokerPredictor_DecadeBalanced.py · JokerPredictor_Recalibrated.py The surviving Joker line: hybrid ML/heuristic scoring, chaos diagnostics, pair memory, and multi-ticket portfolio search.
Joker — archive V16_1 · V18_hybrid · V19 · V19.1 · V20 · V21 · V22 · V23 · V24 · V25 · V27  (Depleted_code/) Retained deliberately. The filter-first V24 that failed, the ten novelty modules of V23, and the genetic-search checkpoints — kept so the failures stay auditable.
Loto 5/40 FiveDin40Predictor.py · _V2.py · _V3_experiments.py · _V4_calibrated.py Built on a database that did not exist: 38 draws read directly out of a screenshot. Strict walk-forward with exceptions re-raised after the silent-swallow defect was found.
Non-statistical controls L649_world_state_hash_oracle.py · L649_lunar_world_experiment.py Deterministic SHA-256 tickets seeded on moon phase, asteroid-mission headlines and world news. Built as a deliberate absurdity; retained because a reproducible arbitrary ticket is a legitimate anti-crowd device and an honest control.

Alongside the code sit twelve dated session records under conversations/ — the primary source for everything above, including the post-mortems that were wrong at the time. The draw databases are validated on every write: row counts reconciled across CSV and workbook, every row six distinct in-range values, dates ascending, Joker balls in 1–20, and a canonical SHA-256 fingerprint recorded in the prediction ledger. What is still missing before anyone should call this reproducible in the strict sense: a README, a dependency lock, a licence, a tagged release with a DOI, and one command that regenerates every table on this page.

18 · Sources and bibliography

Published literature

Primary data sources

  • Official results archive, Loto 6/49 & Noroc — loto.ro
  • Official results archive, Joker & Noroc Plus — loto.ro
  • Official results archive, Loto 5/40 & Super Noroc — loto.ro
  • Official per-draw results PDFs and the published draw schedule — loto.ro — used to resolve a dropped draw and to identify a supplementary Easter draw that had been recorded as a regular one.
  • Independent secondary cross-checks — alba24.ro, oficiuldestiri.ro. Every disputed draw was resolved by two independent sources against the search summary, not by the summary.

Methods reviewed and not adopted

  • LottoExpert — MCMC over transition matrices. Same family as the Markov models already implemented here; no reported edge over baseline.
  • Callam7 / LottoPipeline — github.com/Callam7/LottoPipeline — clustering, decay weighting and Monte Carlo, combined as an ensemble. Structurally identical to the stack already built.
  • “Bayes always wins the lottery in Monte Carlo” — theoretical, on Monte Carlo estimation rather than lottery draws; reviewed and found not applicable.
  • Pure mathematics consulted and ruled out — the Langlands programme (automorphic forms, Galois representations, Shimura varieties) and Floer homology / gauge theory. Neither has an object to act on in a finite discrete uniform draw; the transferable descendant, topological data analysis, was implemented and returned a null.

A broader literature survey conducted during the project found no published method that statistically beats the random baseline on real lottery draws. The most-cited recent frameworks combine decay-weighted frequency, transition matrices, clustering and Monte Carlo into a single ensemble — precisely the architecture this project had already built and measured. That absence of a positive result in the literature is itself corroborating evidence, and it was available from the start.

19 · Acknowledgements

This was not solo work, and the most useful contributions were almost always the ones that removed a result rather than added one.

The critics
Several rounds of unsparing external review — on calibration leakage, on the matched-cost baseline, on the difference between internal consensus and independent validation. Almost every genuine correction in this project began as someone saying the backtest was flattering itself. That criticism was right every time.
The AI collaborators
Claude (Opus 4.7, Fable 5, Opus 5) and GPT-5 / Codex, working across roughly thirty sessions. They built the models, and — more valuably — they were the ones that found the frozen seeds, the shadowed loop variable, the two-window headline and the portfolio illusion. They also produced several of the errors on this page.
Loteria Română
For publishing a complete, free, machine-readable results archive going back to 1993, together with per-category winner counts and prize values. The randomness audit and the entire revealed-preference result exist only because that data is public.
The open-source stack
NumPy, pandas, SciPy, scikit-learn, statsmodels, openpyxl. The entire final analysis — 2,565 draws, seventy tests, exact enumeration over 13,983,816 combinations — runs on a laptop in minutes.
Ana Caraiani and Ciprian Manolescu
Whose work prompted the deepest mathematical detour in the project. That the Langlands programme and Floer homology turned out to have no purchase here is not a criticism of either — it is the clearest possible demonstration that a problem with no structure offers nothing for structure-detecting mathematics to find.
Everyone who asked for one more version
The instruction “do not tell me it is statistically impossible — find a way” produced ten more versions, ten more measurements, and eventually a far stronger negative result than any early concession would have. The null had to be earned.

Research, AI and responsible-play disclosure

This page documents a statistical case study whose primary result is negative. It is not a betting system, a prediction service, or encouragement to gamble. Lottery draws are random; nothing described here increases the probability of winning, and the expected return on every ticket analysed remains negative. Play only with money you are comfortable losing.

AI assistance was used throughout the analysis and in preparing this write-up, and AI-generated output can contain errors — several are documented above. Figures should be verified against the archived scripts and source databases. Research and correction enquiries: contact@kort-x.com.

Continue
This case study sits alongside the laboratory's primary research programme.
The pipeline that finds nothing where there is nothing is the only one you can believe when it finds something.
SINGULARITY_V2