What a decision model is allowed to see: Tetris, Laya and a native model
"System One" models have made a lot of noise recently. In September 2026, TypeSafe AI introduced Jev, which it describes as the first model of a new class. Jev does not write text. It reads a state, evaluates every question about it in parallel, and returns typed decisions with calibrated probabilities. TypeSafe reports it is 40 to 200 times faster than frontier language models on its own workflow evaluations. The pitch is that most production AI is not writing but deciding: routing a ticket, flagging a message, picking an option. Shortly after, Convai Innovations released Laya, an open counterpart under Apache 2.0. Laya is trained with reinforcement learning against strictly proper scoring rules, and once fine-tuned, its authors report that it outperforms Jev on their typed-decisions benchmark.
Laya is the one we can open up. Give it a state and a set of typed questions, and it returns calibrated probabilities over the answers in a single forward pass, without generating text. It is built on a 395M-parameter ModernBERT encoder plus a 26M decision head, and it is good at what it was trained for: routing, triage, guardrails, reranking.
Behind this post is a question we meet on every project. When a task is well
defined and will run a million times, do you take a large general model and
phrase the task in its terms, or do you build a small model around the task's
own structure? Tetris is a convenient place to test it. Every move is a choice
among a variable set of options (where to drop the piece), which is exactly the
shape of a Laya choice question, and the game gives an unambiguous score. So:
can Laya play Tetris? And if it plateaus, is it the model or the way the
problem is put to it?
Laya can play, and it learns, but the way we had to describe the game to it caps it at the level of a 14-parameter lookup table. A model that reads the board itself, with 0.53M trained parameters against 26M for the part of Laya we trained, reaches expert play in about half an hour on a laptop GPU, and within 90 minutes on each of three seeds. The same design, applied to chess, gives a 6.4M-parameter model playing at roughly 1,900–2,050 Elo against a strength-limited Stockfish.
The setup
The action is the final placement of the piece (rotation and column), the standard formulation in the Tetris literature. The board is 10 × 20 with the seven tetrominoes drawn from a 7-bag. Every evaluation in this post uses the same protocol: four fixed seeds, greedy play, no randomness, games played until the stack tops out or a piece cap is reached, and the score reported as lines cleared per game. Whenever two systems are compared, they share the seeds, the cap and the engine.
Two hand-written references anchor the scale. The first is a four-feature heuristic (aggregate height, lines, holes, bumpiness). The second is El-Tetris, a linear evaluation over the six features defined by Pierre Dellacherie, with optimised weights, a classic expert controller.
Figure 1 sums up the whole post: the same move, put to two models in two different ways. The rest of the post measures what each framing costs.
Laya can play Tetris
Handed the raw board as text, Laya plays at the level of chance. A control explains why. In a position where one placement clears two lines and the others clear none, the winning option's probability follows its position in the list, not its content: 0.163 when it is listed first, 0.073 when it is listed eighth. Options that differ only by numbers ("2 lines, +0 holes") are outside what the model was trained to separate.
Show the data
| Position of the right option | Its probability |
|---|---|
| Listed 1st | 0.163 |
| Listed 8th | 0.073 |
| Listed 9th | 0.074 |
So we gave the model words, which is what it was trained on. The arithmetic moves into code: for each placement, a short sentence states its consequence in words ("the piece leaves one hole under it and makes a small bump on top"), and Laya answers one two-option question per sentence. Asking what the stack looks like worked; asking what to do did not.
With that formulation Laya plays from the start, at 142–150 lines per game on the first evaluation. We then trained its decision head with PPO (a GRPO-style group baseline, the 395M encoder frozen and cached), for 1,200 updates and 23 hours. Over the last 200 evaluations it averaged 244.8 lines per game (standard deviation 66), with a best single evaluation of 395.
The best single evaluation, 395, is worth a word of caution. With only four games per evaluation, two evaluations of the same weights can differ by more than 100 lines. The best of 200 noisy draws is a lucky draw, not a level, which is why we report the average.
Why it stops
The plateau is not an optimisation failure. Every stability signal had converged: KL around 0.001, flat entropy, near-zero gradient norms. What limits the model is what it is shown.
The sentences combine five levels of holes, four levels of bumpiness and the number of lines cleared: a hundred possible descriptions. Since Laya scores each sentence in its own question, its policy is in the end a ranking of those hundred phrases. To see what such a ranking can do, we ranked the same sentences with no model at all:
| Policy reading the same sentences | Lines per game |
|---|---|
| Hand-written linear table | 64.8 |
| Linear table, 3 fitted weights | 127.0 |
| Laya, zero-shot | 151.2 |
| Additive table, 14 values (cross-entropy method) | 240.8 |
| Laya + 23 h of PPO | 244.8 |
| Best free ranking found, one score per sentence | 250.0 |
| El-Tetris, forced to choose through the sentences | 397.8 |
Four seeds, 1,000-piece cap, same engine.
Twenty-three hours of reinforcement learning on 26M parameters learned what a 14-number table learns in three minutes. The language model's prior was useful: zero-shot, it already ranked the sentences better than a naive table. But why does everything stop around 250?
Not because of ties. On average 22.5 placements share only 10.9 sentences, so we let the expert pick its move, then played any placement with the same sentence. It lost nothing: 397.8 lines against 398.2. The right move is always expressible through the sentences.
What fails is reading one coarse sentence at a time. "No hole, small bump, one line" can be a fine move or a poor one depending on how high the stack is and where the piece lands, and the sentence says neither. We searched for the best free score per sentence, a hundred numbers, and the best we found averages 250: it reaches the cap on one seed and collapses to 93 on another. Our search is a lower bound, not a proof, so a better ranking may exist. But every one we found, Laya's included, stops in the same place, and none of them is stable.
So we gave Laya exact numbers: holes, new holes, maximum and total height, bumpiness, even the ten column heights. Those sentences are enough on their own, and by construction: they contain the four numbers of our reference heuristic, which, reading only them, one sentence at a time, plays 397.8 lines. Zero-shot, Laya reads them at chance, 1.8 lines per game. It understands "one hole, small bump" but not "3 holes, height 7".
We then trained Laya's head on those exact sentences, with the same PPO settings as before. It was slow: exact sentences are almost all different, so the trick that made the first run affordable, caching the encoder's output once per sentence, stopped working, and each update cost ten times more. And it learned nothing: 1.8 lines per game at the first update, 3.5 at the fifteenth, four hours later, games still over within 45 pieces. On the coarse sentences the same head had started at 150 and gained 70 lines over its first 200 updates. The richer description is not just unreadable for Laya; it also takes away what made training it cheap.
Re-modelling the problem
We kept what is good in Laya's design, an encoder plus a decision that scores a variable set of options in one pass, and changed what the experiments had pointed at.
- Read the board, not a description of it. Each column is a token (its 20
cells as bits, projected to 128 dimensions), plus a
[CLS]token. The model sees the full geometry. - Score each option on its own. The network evaluates the board that a
placement leaves behind (the afterstate), and a placement is worth
Q = lines + 0.99 · V(board after). Options never attend to each other, so a positional prior is impossible by construction. - Make it small. Four transformer layers, width 128, no dropout: 534,401 parameters, against 26.2M trained in Laya's head (421M in total).
The figure below draws Laya and the native model the way architecture figures usually are: inputs at the bottom, a repeated block of attention and feed-forward layers in the middle, heads on top. The building blocks are the same on both sides. What changes is how the problem is cut into tokens, and where the decision is read.
Training has two stages. First, imitation with DAgger: the model plays, the El-Tetris expert scores every placement on the positions the model itself reaches, and the model learns that ranking. Labelling the learner's own states rather than the expert's is what keeps it from collapsing after its first mistake. Second, Q-learning on afterstates, with the next piece known: the target for a board is the best value reachable with the next piece. This stage optimises lines directly and could, in principle, overtake the teacher.
Here is every Tetris system of this post, measured the same way:
Show the data
| Policy | Size | Lines per game |
|---|---|---|
| Random play | — | 0.0 |
| Hand-written table | 3 numbers | 64.8 |
| Fitted linear table | 3 numbers | 127.0 |
| Additive table | 14 numbers | 240.8 |
| Laya + 23 h of PPO | 26M trained of 421M | 244.8 |
| Column transformer | 0.53M | 398.5 |
| El-Tetris expert | 6 weights | 398.2 |
The training budgets differ just as much: 23 hours for Laya and PPO, under three minutes for the 14-number table, and about half an hour for the column transformer on a laptop GPU. Size did not decide the outcome either:
Show the data
| System | Parameters | Lines per game |
|---|---|---|
| 14-number table | 14 | 240.8 |
| Laya + PPO | 26,200,000 | 244.8 |
| Column transformer | 534,401 | 398.5 |
| El-Tetris expert | 6 | 398.2 |
The cap saturates. With 5,000 pieces, the expert clears 1,997.8 lines per game and never tops out. The column transformer reached the same 1,998.8 after 8,826 training steps, and its best evaluation is 1,999.0. Its longest game during training ran 17,308 pieces and 6,906 lines. Two more training runs with different seeds reached expert level at steps 3,760 and 5,234, in 54 and 72 minutes: this is not a lucky run. Played without any cap, both the expert and the model were still alive after 100,000 pieces.
One thing did not work. The reinforcement stage did not measurably beat the expert. It ran for 1.45M steps. Its evaluations averaged 1,669 lines, and only 41% of them had all four games reaching the cap. Six times the automatic rollback had to restore the best weights. At this cap, the imitation stage did the work. Which also means the comparison does not isolate the training signal: the native model reached the expert by imitating it, and a student that distils a teacher at the cap is expected to reach the cap. What the experiment shows is that this cheap route was open to the model that reads the board and closed to the one that reads sentences. We did not train the native model without a teacher.
A second demonstration: chess
Chess is not a second comparison: there is no Laya baseline here, and small transformers learning chess by imitation is established ground (Maia, Ruoss et al.). It is a second demonstration of the same design: one token per square, a learned token for the side's castling rights and en passant, and a level token. The level token tells the network to play like Stockfish, or like a human of a given Elo band.
A bilinear head scores every (from-square, to-square) pair in one pass, and the probability is spread over legal moves only, the same masking Laya applies to its options. A second head predicts the probability of winning. The model has 6.4M parameters. Figure 8 draws it in the same convention as the Tetris models.
It learned from 10 million positions analysed by Stockfish at depth 40–60 (Lichess's open evaluation database) and 10 million positions from 300,000 human games, all trained on the same laptop. At the end of imitation, on positions it had never seen:
- it played Stockfish's best move 48.1% of the time;
- it predicted the human move 50.9% of the time, in the range reported for Maia, the reference human-move model;
- Stockfish judged its moves 87.6% accurate (Lichess's accuracy formula), with 6% blunders.
We then ran DAgger against Stockfish, the chess counterpart of the Tetris recipe. The model played 840 games, half against itself and half against a strength-limited Stockfish whose level tracked the model's results. Stockfish labelled the 60,577 positions where the model was to move, and the real game outcome was mixed into the value target, as in AlphaZero.
We measured strength in 120 games per setting against Stockfish limited to 1,400–3,000 Elo, with every setting using the same engine:
Show the data
| Model | Play | Elo |
|---|---|---|
| Imitation | No search | 1744 ± 107 |
| Imitation + DAgger | No search | 1889 ± 106 |
| Imitation | 64-simulation search | 2058 ± 106 |
| Imitation + DAgger | 64-simulation search | 2033 ± 106 |
DAgger added about 145 Elo to the network's instinctive play, cutting blunders from 6% to 5%. It added nothing once a search was layered on top: the search was already catching the mistakes DAgger removed.
Stockfish's Elo limiter is calibrated at a different time control, so these figures are an order of magnitude, not a rating.
What we take from it
Representation set the ceiling, not scale. The same game, the same compute, the same kind of architecture. The difference between 245 and 398 lines is what the model was allowed to see. The words Laya reads well are too coarse to rank moves one at a time, and the numbers that would be enough are words it does not read. Laya's structure is a strength for problems that are naturally phrased in language, like a ticket to route or a passage to rerank. On a problem whose structure is geometric, forcing it through words cost more than the model's scale could buy back.
A big generalist is not always the answer. This is where we started. When a task is well defined and comes back often, the reflex is to take the largest general model available and phrase the task in its terms. Here that reflex cost 23 hours and 421M parameters for 245 lines. A 0.53M-parameter model built around the game reached expert play in half an hour on the same laptop; it trains faster, runs faster and plays better. The chess model, 6.4M parameters, took a couple of days on that laptop. Neither needed reinforcement learning from scratch: imitating an existing expert, on the learner's own states, did most of the work.
Controls matter more than curves. Each conclusion above rests on a control rather than a training curve:
- the permutation test showed what Laya was really reading;
- the expert forced through the sentences, and the search for the best ranking of them, showed where its ceiling came from;
- the identical-engine Elo matches separated DAgger's effect from the search's.
Along the way, the controls also caught real bugs: a dropout that silently corrupted the PPO ratio, castling notation that mislabelled 1.6% of the chess targets, and a Tetris engine that let pieces pass through blocks.
Generalists and specialists answer different questions. Our two models are not System One models, and they are not trying to be. Jev and Laya are generalists: one model, any state, any question written at run time, which is exactly what makes them useful in front of an LLM stack. The Tetris model knows one game and the chess model another, and neither will route a support ticket. Our point is narrower: a task with a structure of its own, that you will run again and again, deserves its own model. The two are not rivals. A generalist decides what kind of problem it is looking at; a specialist solves the problems that come back often enough to deserve one.
The full methodology, protocols and references are in the accompanying paper.