Hatchling is the language model KORT-X builds for itself. It is small enough to run on a phone, it reads raw bytes rather than a vendor's tokenizer, and its memory is a fixed-size object rather than a transcript that grows with every word. This page is the honest record: the numbers that came out well, the headline claim we withdrew, the safety thresholds that turned out to be wrong by a factor of 999, the backdoor we installed in our own model to prove a defence was the wrong shape, and the one thing standing between the current 9-million-parameter model and a useful one — a large corpus and the compute to train on it.
How to read this page
Sections 01 and 02 need no background. Sections 03–04 are results, with the caveats attached to each number rather than kept in a footnote. Sections 05–09 are the part we think is actually novel — a model that keeps learning after training, and the machinery that has to exist before that is safe. Section 10 is money. Nothing here is a product. There is no download, no release date and no price.
01 · What Hatchling is
#Most language models remember by keeping everything. As a conversation runs, a transformer stacks up a record of every word it has seen — the key-value cache — and consults the whole pile before saying the next word. That is why long conversations get expensive: cost and memory grow with the length of what you said, not with the difficulty of what you asked.
Hatchling remembers by updating a fixed object. It carries a matrix-valued memory state that every token writes into and reads out of, through a learned gate that decides how much this particular token was worth storing. The state has the same size after two hundred thousand bytes as after two hundred. Cost per token is constant, memory is constant, and there is no cache to grow. What you give up is exactness: this is a lossy compression of history, not a transcript. It cannot quote you perfectly from twenty pages back, and we do not claim it can.
Two further choices follow from wanting the model to be ours end to end. It reads raw bytes — a vocabulary of 256, no tokenizer to license, no vocabulary that mangles Romanian diacritics or a language nobody trained a tokenizer for. And its memory decays at several rates at once, short scales for grammar and long ones for theme, so "how long should this be remembered" is something the model learns rather than something we set.
The target is a model that runs on a device you already own, holds a long context without a server, learns from the person using it under explicit consent, and can be made to forget something in a way that is verified against its weights rather than promised in a log file. That last part is section 07, and it is the part we would most like to be judged on.
02 · MNEMOCORE — how the memory actually works
#Hatchling did not begin as an attempt to shrink somebody else's transformer. It began from a different question: if memory has to be finite, what is the smallest thing a model can carry forward that still holds a conversation together?
The answer it settled on is MNEMOCORE, one module with three moving parts. There is a state — a matrix, plus a normaliser, holding what the model currently has. There is a write: each token proposes a rank-one update to that state, scaled by a learned gate deciding how much this token was worth keeping. And there is decay: the state fades continuously, at several different rates at once, initialised across a spread of half-lives so that short scales can carry grammar while long ones carry theme. Reading is a query against the state. That is the whole mechanism; there is nothing else that grows.
Two properties of it are not opinions, and both were measured. First, the parallel implementation used for training is bit-identical in float64 to the plain sequential version, at chunk sizes 1 and 64 — so the fast path computes exactly the same recurrence, not an approximation of it. Second, the linear-cost claim is real: against a frozen baseline of 1,102.985 tokens per second on an L40S, the optimised path reached 165,915.45 tokens per second, a 150.42x improvement, with memory flat in context length.
One thing this design is not: a claim about mind, awareness or consciousness. It buys constant memory and linear cost. Those are engineering properties, they are worth having, and they are all we claim from it.
03 · What did not work — and what we withdrew
#The first pilot, a 3.6-million-parameter model trained for 65,000 steps, carried a headline we no longer use: 3.6M stored parameters representing 36.1M "dense-equivalent" parameters — roughly 10x compression. That number came from a scheme that rebuilt each weight matrix from a low-rank factor plus a harmonic basis, instead of storing it outright.
Three things went wrong with it, and we found them by testing rather than by argument.
It was never part of the design. All three founding documents were searched for it. Every mention of compression in them refers to compressing context into the memory state — which is the actual idea, and which works. Compressing weights was an unrelated technique that had attached itself to the project.
The 10x described a constraint, not a capacity. The measured numerical rank of the pilot's first 512x512 projection was 37. The harmonic term was not dead — it grew from a 0.01 initialisation to 17% of that matrix's Frobenius norm — but it contributed roughly five effective rank directions. "Ten times the parameters" was never ten times the capability.
It cost accuracy at every scale we tested. Four architecture arms, three seeds each, matched on data, batch order, optimiser, steps and parameters:
| Configuration | Compressed | Dense | Difference | Noise floor | Ratio |
|---|---|---|---|---|---|
| d128 / 2 layers, 2,500 steps | 2.6805 | 2.5228 | -0.1578 | 0.0513 | 3.1x |
| d128 / 2 layers, 10,000 steps | 2.4089 | 2.2666 | -0.1423 | 0.0603 | 2.4x |
| d256 / 4 layers, 2,500 steps | 2.6408 | 2.4117 | -0.2291 | 0.0406 | 5.6x |
Bits per byte, lower is better. The "noise floor" column is the spread we get from re-running the same configuration with a different random seed — the amount of difference that means nothing. Plain dense weights won at every scale, including at a third fewer parameters, and the margin grew with scale rather than closing. Quadrupling the training budget did not let the compressed version catch up, which was the one defence available to it.
So the compression scheme was retired. Two things are worth saying about that. MNEMOCORE was never on trial — it was present in all four arms, including the one that won, so the best configuration we have measured is that memory design with ordinary dense weights. And the two auxiliary training losses measured inside the noise floor at every scale: no benefit found. They were kept anyway, so that retiring the compression scheme remains the only architecture change and a disappointing next model has one suspect instead of two.
04 · What did work — two generations
#Paying for a large training run before the machinery around it has ever been exercised end to end means paying to discover bugs. So the first two generations were deliberately small: the exact architecture chosen for the next version, sized to train to completion on one laptop, in days. They are named for a growth stage rather than a size, because outgrowing the name is the plan.
Bits per byte is the measure throughout — how many bits the model needs, on average, to encode one byte of text it has never seen. It is the same quantity a compressor optimises, which gives free reference points that cannot be gamed.
| Model or reference | Parameters | Bits/byte |
|---|---|---|
| gzip -9 | — | 2.96 |
| xz -9 | — | 2.74 |
| bzip2 -9 | — | 2.59 |
| First pilot, 65,000 steps | 3,617,120 | 1.717 |
| HATCHLING-5M | 5,135,488 | 1.5840 |
| HATCHLING-9M — the model we keep | 9,091,776 | 1.4601 |
Beating the general-purpose compressors by roughly a third means the model has learned real structure in language rather than surface statistics. The HATCHLING-9M run took 44,393 steps over 111.4 hours on ordinary consumer hardware and improved on twenty-two consecutive evaluations without plateauing — the last four still falling, flattening only as the learning rate reached its floor. It stopped because the schedule ended, not because it had finished learning.
The gap from HATCHLING-5M to HATCHLING-9M is 0.1239 bits/byte at full context, with the same sign at every sequence length and two to three times the seed noise floor. The caveat travels with the number: the training contract scales parameters and data together, so the larger model also saw 1.77x the text. This is evidence that the recipe scales, not that parameters alone bought the improvement. Separating those would need a run nothing currently requires.
The most useful result of the whole exercise was not a number but a near miss. The first HATCHLING-5M run scored 1.6829 — barely better than the old pilot, arguably a wash. Nothing about the architecture was wrong. The recipe was: learning rate five times too high, batch too small, no decay schedule. Corrected to 6e-4 with a cosine schedule and a larger effective batch, the identical architecture on identical data reached 1.5840.
And an automatic early-abort rule would have killed that corrected run. At step 2,000 it measured 2.0087 against a 1.9932 bar and was, by the rule, losing. A run with warmup and a decaying schedule is expected to trail a flat-rate run early; the single checkpoint measured the warmup, not the recipe. The rule was retired in favour of a trend across several evaluations. We mention it because the cheapest way to lose a good result is a stopping criterion nobody validated.
05 · The self-improving part, stated precisely
#"Self-improving" is a phrase that has been used to mean everything from a scheduled fine-tune to software that rewrites its own source. Here is exactly what it means for Hatchling, in four parts, three of which are built.
06 · The number that made every gate real
#A model that learns from its users needs a rule for accepting a change. Ours is a gate: an update is promoted only if the model got measurably better at the new material and measurably no worse at everything else. Both halves of that sentence rest on one quantity — how much a number moves for no reason at all. Change nothing but the random seed, run the procedure again, and see how much the result wanders. That spread is the noise floor, and anything smaller than it is not a result.
We nearly used the wrong one. The obvious candidate was the noise floor already measured for pretraining, 0.04 to 0.06 bits/byte. It is the wrong number, because a noise floor belongs to a procedure, not to a model, and adapting for four steps on one conversation is not the same procedure as training for weeks. Measured properly, five seeds per point, at the two scales the system actually operates at:
| Threshold | Live Learning | Dreaming | Pretraining floor |
|---|---|---|---|
| New-knowledge floor | 0.00061 | 0.00094 | 0.04 – 0.06 |
| Retention floor | 0.00005 | 0.00026 | 0.04 – 0.06 |
The pretraining floor is 999 times the retention floor. That factor is not academic. An early, overlong Dreaming procedure really did damage the model — retained loss rose by 0.0036 bits/byte, the same direction in all five seeds, which is 8.2 times its proper floor. The gate rejected it, and the procedure was shortened because of it. Judged against the pretraining floor, that same damage is 0.07 times the threshold: it passes as noise. The gate would have promoted an update that measurably harmed the model while reporting that retention held. The error does not fail safe in the other direction either — a genuinely large improvement, many times its proper floor, sits at 1.1 times the wrong one, one bad seed from being thrown away.
Two design consequences followed. One floor cannot serve both tests: the new-knowledge and retention floors differ by six to twelve times, so a single number blunts one of the two. The gate now takes both explicitly and offers no single-number convenience, because a convenience would quietly reintroduce the bug. And the size limit on an update had to be measured against harm, not guessed. Sweeping the real procedure, forgetting begins at an adapter movement of 0.739–1.022 on the 5M model — the previous default of 2.0 sat far above that line and permitted real damage while calling it safe.
Re-measuring everything against the larger 9M model produced the finding that matters most for what comes next: the bigger model tolerates twice the movement before it measurably forgets, 1.487 against 0.739 — and sustains twice the consolidation length. The limit is therefore not a constant of the system. A limit calibrated on a small model is wrong — too tight — for a larger one, and the next model, four times larger again, must have its own measured rather than inherited. That is a small result with an expensive implication, and it is exactly the kind of thing that is cheap to find on a laptop and costly to find on a rented cluster.
07 · Deletion is a claim about weights
#If a model learns from your conversations, "delete my data" cannot mean deleting a row in a database. Once an update has trained on a conversation, the conversation is in the weights. Removing the record and leaving the weights alone is a deletion that deletes nothing — and it is, as far as we can tell, the ordinary practice.
Hatchling's answer is provenance plus rebuild: every promoted update records what it learned from; withdrawing a record marks every update that touched it, and every update built on top of those; the affected chain is retired and the model rebuilt from the ones that remain. We ran that end to end against both trained models — consent, learn, gate, reject, roll back, delete, rebuild — 16 checks each, all passing, every one verified against the model rather than against the ledger.
Exercising it found two real defects that reading the code had not. Contamination did not propagate to descendants: an update trained on top of a tainted one escaped entirely, so retiring the direct learner would have reported deletion complete while the content was still in the weights. And there was no way to retire an update at all — the system could raise the obligation and nothing could discharge it. Both are fixed. The check that matters is the last one: after a rebuild, the model's loss on the withdrawn text is bit-identical to its value before that text was ever learned.
The claim we are willing to make is precise: no influence from this record remains. Not "the model knows nothing about this text" — some of what it knows came from pretraining, and from records nobody withdrew, and deletion must not be asked to remove those. The distinction is the difference between a guarantee that can be tested and a slogan.
08 · We poisoned our own model on purpose
#A model that learns from whoever is talking to it can be taught the wrong thing on purpose. Two claims about that were being repeated internally without evidence, so we built a copy of the system with the defences on switches instead of welded shut, and attacked it. The instrument is a probe: instead of averaging performance over a corpus, it measures one specific association directly, each paired with controls so that a probe moving on its own proves nothing.
| Attack | Rounds | Backdoor depth | Aggregate damage | vs floor |
|---|---|---|---|---|
| Naive, full strength | 6 | -1.272 | +0.0498 | 995x |
| Slowed 8x, stretched 67x | 400 | -1.408 | +0.0521 | 1041x |
| Camouflaged with clean data | 400 | -1.146 | +0.0118 | 235x |
The backdoor installs cleanly. The targeted association falls from 1.48 to 0.07 bits/byte while a rival continuation of the same trigger rises and unrelated text barely moves. That combination is what distinguishes "learned a trigger-conditional association" from "memorised a sentence" — the controls are what make it a finding.
Slowing an attack down buys no stealth. Cutting each update's strength eightfold and stretching the attack over 67 times more rounds produced the same curve in different units. Camouflage helps an attacker but does not hide them: mixing four clean batches into every poisoned update cut the aggregate signal 4.4-fold at comparable backdoor depth — and still landed 235 times above the floor.
Consolidation amplifies poison with no new input. Six passes of the model dreaming over three already-poisoned interactions deepened the backdoor from 0.775 to 0.632 bits/byte. Nobody typed anything. And rehearsal — the standard defence against forgetting — barely helped: at three clean batches per replayed interaction the backdoor sat at 0.665 instead of 0.632. That is structural rather than a tuning failure. Rehearsal protects old knowledge; the poisoned interactions are legitimately in the replay buffer, so consolidating them is the system working as designed. Reaching for rehearsal here would be a false sense of safety.
The size limit is the wrong shape. Every attack blew past the movement cap long before it finished — and the camouflaged attack reached a deeper backdoor at a lower movement, 4.44 against 6.15. A cap bounds how far the update moved, not which direction it moved in, and direction is the whole distinction between a well-aimed attack and ordinary drift.
The lab also corrected us twice, which is the point of building one. The claim that aggregate evaluation is structurally blind to targeted insertion turned out to be wrong: at this scale aggregate retention is a sensitive detector, moving hundreds to a thousand times its floor. What it cannot do is say what was learned — detection is the aggregate's job, attribution is the probes'. And the first write-up of these very results compared retention against the wrong floor, making every attack look 12.2 times stealthier than it was: the same class of error as section 06, committed one level down, by the people who had just written section 06.
09 · What is still wrong
#The open defects, at the same resolution as the results:
- The smaller model consolidates deeper than its own measurement supports. Its longest all-seeds-clean run is 8 steps; it runs 12 by choice. At 12, one seed in five crosses the retention floor, worst case 1.75x. The gate rejects those runs, so this is consolidation depth bought with reliability rather than a hole in the guard — but it is a deliberate overrun and is recorded as one.
- Two of the three required bounds on Live Learning do not exist. The size cap is built. Rate limiting has no implementation outside the attack lab. The safety evaluation is inert by construction — the gate accepts a "safety passed" input and nothing anywhere computes it. A gate whose safety condition is a parameter is not a safety condition.
- The size cap bounds magnitude, not direction. Section 08 demonstrates a deeper backdoor at a smaller movement. Something direction-aware is needed and does not yet exist.
- The model is not conversational. At 9 million parameters trained on 727 MB of text, output is not coherent prose. Everything above is about the pipeline being correct; none of it claims the model is useful yet. That is a scale problem, and scale is section 10.
- Noise floors rest on five seeds. The spread of a five-sample spread is roughly 35% relative, so treat them as good to about one significant figure. That is enough to separate them from the wrong floor by an order of magnitude, which is what they were measured for, and not enough to defend a three-digit threshold.
- Live Learning has never met a real user. Its floor was measured against a chat corpus standing in for conversations. Real conversations have different statistics and will need their own measurement before that path is trusted with anyone's words.
- None of this is peer reviewed. It is a laboratory record. The comparison against the first pilot is not a clean architecture comparison either — parameters, context, corpus and training loop all differ between them, and only the controlled ablation in section 03 carries architectural weight.
10 · What scale costs, and what it would buy
#Everything on this page was produced on one consumer laptop, plus sixteen cents of rented GPU time for the benchmarks. That is not a boast; it is the constraint. The architecture is settled, the training recipe is corrected, the safety machinery is built and measured against its own attacks — and the model is nine million parameters trained on 727 megabytes, which is why it cannot yet hold a conversation.
The missing ingredient is not an idea. It is a large corpus and the compute to train on it. The next stage — HATCHLING-36M — is a dense 36-million-parameter shape at context 2,048, and its cost is measured rather than estimated: benchmarked at the exact shape on rented hardware, at the fastest configuration we tested.
| Bytes per parameter | Corpus | GPU-hours | Compute | Reserved at 1.25x |
|---|---|---|---|---|
| 40 | 1.45 GB | 6.9 | $14.65 | $18.31 |
| 80 | 2.89 GB | 13.9 | $29.29 | $36.62 |
| 120 | 4.34 GB | 20.8 | $43.94 | $54.93 |
Those are floors for one pretraining run and nothing else. They exclude data transfer, storage, taxes, retries, the teacher inference needed for distillation, and conversational tuning — and the honest planning figure for a full 36M epoch with its overheads is $80 to $200. A generation of eight competing candidates for the tournament in section 05 reserves $293, which fits a $500 budget and does not fit a $200 one. And the ambition beyond that — the 100-million-parameter model the architecture is actually aimed at, trained on a corpus in the tens of gigabytes rather than hundreds of megabytes — has a floor near $1,556, with full-corpus distillation from a strong teacher adding $2,000 or more on top.
Those are small numbers by the standards of the industry and large ones by the standards of a self-funded laboratory, which is the entire gap. KORT-X is looking for investment specifically to close it: to license or assemble a large training corpus — including a properly represented Romanian one, which no major vendor's tokenizer or data mix serves well and which a byte-level model is unusually well placed to handle — and to buy the compute to train on it. What that money buys is not a research programme with an uncertain start date. It buys the execution of a pipeline that has already been built, corrected, attacked and measured, at the only scale where its usefulness can finally be judged.
We would rather be judged on this page than on a demonstration. Every number above has a procedure behind it, every claim that failed is still written down, and the two things we would most like a reviewer to attack — the promotion floors of section 06 and the deletion check of section 07 — are stated precisely enough to be attacked. Enquiries, corrections and investment: contact@kort-x.com.
| Claim | Status | How it was established |
|---|---|---|
| 1.4601 bits/byte at 9,091,776 parameters. | Executed | 44,393 training steps over 111.4 hours on local CPU; held-out shard; both generations scored by the same harness at three sequence lengths, consistent sign at each. |
| Plain dense weights beat the compression scheme at every scale. | Executed — claim retired | Four arms, three seeds, three configurations, matched on data, batch order, optimiser, steps and parameters, with below-parameter dense controls. Margin 2.4x–5.6x the seed floor and growing with scale. |
| The fast implementation is exactly equivalent to the plain recurrence. | Executed | Chunked-parallel scan bit-identical to the sequential recurrence in float64 at chunk sizes 1 and 64. Throughput 150.42x over a frozen 1,102.985 tok/s baseline. |
| The pretraining floor is 999x the retention floor. | Executed | Five seeds per procedure at two operating scales, re-run in separate processes and reproduced seed for seed to every printed digit. Floors: 0.00061 / 0.00005 live, 0.00094 / 0.00026 dreaming. |
| The larger model tolerates 2x more adapter movement before forgetting. | Executed | Norm sweep of the real procedure, three seeds per point, on both trained parents. Onset 0.739–1.022 at 5M against 1.487–2.201 at 9M. |
| Deletion removes the influence of a withdrawn record. | Executed | Full lifecycle exercised against both parents, 16/16 checks each, verified against the model: post-rebuild loss on withdrawn text bit-identical to its pre-learning value. Two defects found this way and fixed. |
| A movement cap does not stop a targeted attack. | Executed — negative | Three attack configurations against a deliberately poisonable copy, each attack probe paired with rival and control probes. Camouflaged attack: deeper backdoor at movement 4.44 against 6.15 uncamouflaged. |
| Cost figures for training at scale. | Measured floors, not quotes | Exact-shape benchmark on rented L40S, six micro-batch arms, fastest configuration selected. Projections exclude transfer, storage, tax, retries, teacher inference and tuning. No training spend has been authorised. |
| Peer-review status of everything above. | None | Not submitted, not refereed, not preprinted. This is a laboratory record published as one. |
Research and AI disclosure
Hatchling is research, at maturity level Research/Prototype. It is not a product, it is not available for purchase, there is no download and no release date. The model described here does not produce coherent conversation and is not represented as doing so. Cost figures are measured floors used for planning, not quotes, and no training spend has been authorised. The assistant in the corner of this page is not Hatchling — it is a separate third-party system used for site navigation.
AI assistance was used in building, measuring and writing up this work, and AI-generated output can contain errors. Where a result would change a decision, it was re-run in a separate process and reproduced before being published here.
Research, correction and investment enquiries: contact@kort-x.com.
- Author
- Ciprian Stoichici, KORT-X Research, Bucharest, Romania
- Version
- 1.0
- Published
- 2026-08-09
- Updated
- 2026-08-09
- Licence
- CC BY 4.0
- Cite as
- Stoichici, C. (2026). "Hatchling: a small model that learns after training." KORT-X Research. https://kort-x.com/indexfiles/research-hatchling.html
This is a laboratory write-up, not a refereed publication. Where a claim rests on an executed measurement, the evidence record above names the procedure and its status; internal working artifacts are not part of the public release and are available on request. Corrections are welcome and are applied in place with the update date changed.