Blog

What a decision model is allowed to see: Tetris, Laya and a native model

"System One" models have made a lot of noise recently. In September 2026, TypeSafe AI introduced Jev, which it describes as the first model of a new class. Jev does not write text. It reads a state, evaluates every question about it in parallel, and returns typed decisions with calibrated probabilities. TypeSafe reports it is 40 to 200 times faster than frontier language models on its own workflow evaluations. The pitch is that most production AI is not writing but deciding: routing a ticket, flagging a message, picking an option. Shortly after, Convai Innovations released Laya, an open counterpart under Apache 2.0. Laya is trained with reinforcement learning against strictly proper scoring rules, and once fine-tuned, its authors report that it outperforms Jev on their typed-decisions benchmark.

Laya is the one we can open up. Give it a state and a set of typed questions, and it returns calibrated probabilities over the answers in a single forward pass, without generating text. It is built on a 395M-parameter ModernBERT encoder plus a 26M decision head, and it is good at what it was trained for: routing, triage, guardrails, reranking.

Behind this post is a question we meet on every project. When a task is well defined and will run a million times, do you take a large general model and phrase the task in its terms, or do you build a small model around the task's own structure? Tetris is a convenient place to test it. Every move is a choice among a variable set of options (where to drop the piece), which is exactly the shape of a Laya choice question, and the game gives an unambiguous score. So: can Laya play Tetris? And if it plateaus, is it the model or the way the problem is put to it?

Laya can play, and it learns, but the way we had to describe the game to it caps it at the level of a 14-parameter lookup table. A model that reads the board itself, with 0.53M trained parameters against 26M for the part of Laya we trained, reaches expert play in about half an hour on a laptop GPU, and within 90 minutes on each of three seeds. The same design, applied to chess, gives a 6.4M-parameter model playing at roughly 1,900–2,050 Elo against a strength-limited Stockfish.

The setup

The action is the final placement of the piece (rotation and column), the standard formulation in the Tetris literature. The board is 10 × 20 with the seven tetrominoes drawn from a 7-bag. Every evaluation in this post uses the same protocol: four fixed seeds, greedy play, no randomness, games played until the stack tops out or a piece cap is reached, and the score reported as lines cleared per game. Whenever two systems are compared, they share the seeds, the cap and the engine.

Two hand-written references anchor the scale. The first is a four-feature heuristic (aggregate height, lines, holes, bumpiness). The second is El-Tetris, a linear evaluation over the six features defined by Pierre Dellacherie, with optimised weights, a classic expert controller.

Figure 1 sums up the whole post: the same move, put to two models in two different ways. The rest of the post measures what each framing costs.

Figure 1 · One Tetris move, put to a model two ways
LAYA · 421M parameters, 26M trainedcode: geometryholes, bumps, lines“leaves one hole, small bump”≈100 possible sentencesLaya: clean or messy?P(clean) per sentenceNATIVE MODEL · 0.53M parameters10 column tokens20 bits → 128-d eachtransformer, 4 layerscolumns attend to each otherQ = lines + 0.99·Vplay the highest Q
Top: Laya never sees the board. Code computes each placement's consequences and writes them as a sentence; Laya answers one two-option question per sentence. Bottom: the native model reads the board the placement leaves behind, one token per column, and returns a single value for it. Every placement is scored on its own.

Laya can play Tetris

Handed the raw board as text, Laya plays at the level of chance. A control explains why. In a position where one placement clears two lines and the others clear none, the winning option's probability follows its position in the list, not its content: 0.163 when it is listed first, 0.073 when it is listed eighth. Options that differ only by numbers ("2 lines, +0 holes") are outside what the model was trained to separate.

Figure 2 · Same position, same options, only the order changes
uniform 0.111Listed 1st0.163Listed 8th0.073Listed 9th0.074
The probability Laya gives to the only placement that clears two lines, as its position in the list of nine changes. A model that read the options would give it the same probability wherever it sits. The dashed line is a uniform guess, 1/9.
Show the data
Position of the right optionIts probability
Listed 1st0.163
Listed 8th0.073
Listed 9th0.074

So we gave the model words, which is what it was trained on. The arithmetic moves into code: for each placement, a short sentence states its consequence in words ("the piece leaves one hole under it and makes a small bump on top"), and Laya answers one two-option question per sentence. Asking what the stack looks like worked; asking what to do did not.

With that formulation Laya plays from the start, at 142–150 lines per game on the first evaluation. We then trained its decision head with PPO (a GRPO-style group baseline, the 395M encoder frozen and cached), for 1,200 updates and 23 hours. Over the last 200 evaluations it averaged 244.8 lines per game (standard deviation 66), with a best single evaluation of 395.

Figure 3 · Twenty-three hours of PPO, and a ceiling
01132253384502004496989461,195El-Tetris expert · 39814-number table · 241
Laya + PPO14-number tableexpert
Each dot is one evaluation of Laya after a PPO update (four fixed seeds, 1,000-piece cap); the line is a rolling mean of 10. The model settles around the level of a 14-number table that reads the same sentences, far below the expert, which reaches the cap on every seed. Hover to read values.

The best single evaluation, 395, is worth a word of caution. With only four games per evaluation, two evaluations of the same weights can differ by more than 100 lines. The best of 200 noisy draws is a lucky draw, not a level, which is why we report the average.

Why it stops

The plateau is not an optimisation failure. Every stability signal had converged: KL around 0.001, flat entropy, near-zero gradient norms. What limits the model is what it is shown.

The sentences combine five levels of holes, four levels of bumpiness and the number of lines cleared: a hundred possible descriptions. Since Laya scores each sentence in its own question, its policy is in the end a ranking of those hundred phrases. To see what such a ranking can do, we ranked the same sentences with no model at all:

Policy reading the same sentencesLines per game
Hand-written linear table64.8
Linear table, 3 fitted weights127.0
Laya, zero-shot151.2
Additive table, 14 values (cross-entropy method)240.8
Laya + 23 h of PPO244.8
Best free ranking found, one score per sentence250.0
El-Tetris, forced to choose through the sentences397.8

Four seeds, 1,000-piece cap, same engine.

Twenty-three hours of reinforcement learning on 26M parameters learned what a 14-number table learns in three minutes. The language model's prior was useful: zero-shot, it already ranked the sentences better than a naive table. But why does everything stop around 250?

Not because of ties. On average 22.5 placements share only 10.9 sentences, so we let the expert pick its move, then played any placement with the same sentence. It lost nothing: 397.8 lines against 398.2. The right move is always expressible through the sentences.

What fails is reading one coarse sentence at a time. "No hole, small bump, one line" can be a fine move or a poor one depending on how high the stack is and where the piece lands, and the sentence says neither. We searched for the best free score per sentence, a hundred numbers, and the best we found averages 250: it reaches the cap on one seed and collapses to 93 on another. Our search is a lower bound, not a proof, so a better ranking may exist. But every one we found, Laya's included, stops in the same place, and none of them is stable.

So we gave Laya exact numbers: holes, new holes, maximum and total height, bumpiness, even the ten column heights. Those sentences are enough on their own, and by construction: they contain the four numbers of our reference heuristic, which, reading only them, one sentence at a time, plays 397.8 lines. Zero-shot, Laya reads them at chance, 1.8 lines per game. It understands "one hole, small bump" but not "3 holes, height 7".

We then trained Laya's head on those exact sentences, with the same PPO settings as before. It was slow: exact sentences are almost all different, so the trick that made the first run affordable, caching the encoder's output once per sentence, stopped working, and each update cost ten times more. And it learned nothing: 1.8 lines per game at the first update, 3.5 at the fifteenth, four hours later, games still over within 45 pieces. On the coarse sentences the same head had started at 150 and gained 70 lines over its first 200 updates. The richer description is not just unreadable for Laya; it also takes away what made training it cheap.

Re-modelling the problem

We kept what is good in Laya's design, an encoder plus a decision that scores a variable set of options in one pass, and changed what the experiments had pointed at.

  • Read the board, not a description of it. Each column is a token (its 20 cells as bits, projected to 128 dimensions), plus a [CLS] token. The model sees the full geometry.
  • Score each option on its own. The network evaluates the board that a placement leaves behind (the afterstate), and a placement is worth Q = lines + 0.99 · V(board after). Options never attend to each other, so a positional prior is impossible by construction.
  • Make it small. Four transformer layers, width 128, no dropout: 534,401 parameters, against 26.2M trained in Laya's head (421M in total).

The figure below draws Laya and the native model the way architecture figures usually are: inputs at the bottom, a repeated block of attention and feed-forward layers in the middle, heads on top. The building blocks are the same on both sides. What changes is how the problem is cut into tokens, and where the decision is read.

Figure 4 · Two ways of modelling one Tetris move
LAYA · 421Msoftmax over optionsP(option) for each [MASK]Scorer at each [MASK]LayerNorm → MLP → 1 logitDecision head2 transformer layers, 26M (trained)28×Feed forwardLayerNormMulti-head attentionquestion, options and state togetherLayerNormToken embeddingsModernBERT-large, 395M (frozen)[CLS]question[SEP][MASK]opt 1[MASK]opt 2…[SEP]stateone sequence: the options share the context,so their order can leak into the scoresInput: text (a sentence per placement)TETRIS · 0.53Margmax over legal placementsQ(a) = lines(a) + 0.99 · V(s′a)Value V(s′)read on [CLS], linear → 14×Feed forwardLayerNormMulti-head attentioncolumns attend to each otherLayerNormColumn embeddingLinear 20 → 128, + column positionCLSc1c2c3c4c5c6c7c8c9c10one board per placement, scored on its own:no option ever sees anotherInput: the board after the move (afterstate)
embeddingsattentionfeed forwardlayer normheads
Read bottom to top; each repeated block is pre-norm, with a residual connection around both sub-layers. Same building blocks on both sides, two ways of cutting the problem into tokens. Laya puts question, options and state in one text sequence and scores each option where it sits, so the options share their context. The native model gives each placement its own pass over ten column tokens and returns one value per afterstate.

Training has two stages. First, imitation with DAgger: the model plays, the El-Tetris expert scores every placement on the positions the model itself reaches, and the model learns that ranking. Labelling the learner's own states rather than the expert's is what keeps it from collapsing after its first mistake. Second, Q-learning on afterstates, with the next piece known: the target for a board is the best value reachable with the next piece. This stage optimises lines directly and could, in principle, overtake the teacher.

Here is every Tetris system of this post, measured the same way:

Figure 5 · Lines per game, same seeds, same cap
Random play—0.0Hand-written table3 numbers64.8Fitted linear table3 numbers127.0Additive table14 numbers240.8Laya + 23 h of PPO26M trained of 421M244.8Column transformer0.53M398.5El-Tetris expert6 weights398.2
Four fixed seeds, greedy play, 1,000-piece cap (about 400 lines is the most a game can clear). Orange: tables that read Laya's sentences, no model at all. The native model and the expert both reach the cap on every seed.
Show the data
PolicySizeLines per game
Random play—0.0
Hand-written table3 numbers64.8
Fitted linear table3 numbers127.0
Additive table14 numbers240.8
Laya + 23 h of PPO26M trained of 421M244.8
Column transformer0.53M398.5
El-Tetris expert6 weights398.2

The training budgets differ just as much: 23 hours for Laya and PPO, under three minutes for the 14-number table, and about half an hour for the column transformer on a laptop GPU. Size did not decide the outcome either:

Figure 6 · Size did not buy lines
010020030040011k1M1Bparameters (log scale)14-number tableLaya + PPOColumn transformerEl-Tetris expert
Lines per game against the number of learned parameters, on a log scale. The two systems that read the board clear about 400 lines whatever their size; the two that read the sentences stop near 240, from 14 numbers to the 26 million parameters trained in Laya's head.
Show the data
SystemParametersLines per game
14-number table14240.8
Laya + PPO26,200,000244.8
Column transformer534,401398.5
El-Tetris expert6398.2

The cap saturates. With 5,000 pieces, the expert clears 1,997.8 lines per game and never tops out. The column transformer reached the same 1,998.8 after 8,826 training steps, and its best evaluation is 1,999.0. Its longest game during training ran 17,308 pieces and 6,906 lines. Two more training runs with different seeds reached expert level at steps 3,760 and 5,234, in 54 and 72 minutes: this is not a lucky run. Played without any cap, both the expert and the model were still alive after 100,000 pieces.

Figure 7 · The native model's first half hour
05631,1251,6882,25093,8977,78412k16kEl-Tetris expert · 1,9984-feature heuristic · 1,488
Evaluations during the imitation stage (four seeds, 5,000-piece cap, where the expert clears about 2,000 lines without topping out). Each evaluation plays only four games, hence the jumps; expert parity is first reached at step 8,826, about 30 minutes in.

One thing did not work. The reinforcement stage did not measurably beat the expert. It ran for 1.45M steps. Its evaluations averaged 1,669 lines, and only 41% of them had all four games reaching the cap. Six times the automatic rollback had to restore the best weights. At this cap, the imitation stage did the work. Which also means the comparison does not isolate the training signal: the native model reached the expert by imitating it, and a student that distils a teacher at the cap is expected to reach the cap. What the experiment shows is that this cheap route was open to the model that reads the board and closed to the one that reads sentences. We did not train the native model without a teacher.

A second demonstration: chess

Chess is not a second comparison: there is no Laya baseline here, and small transformers learning chess by imitation is established ground (Maia, Ruoss et al.). It is a second demonstration of the same design: one token per square, a learned token for the side's castling rights and en passant, and a level token. The level token tells the network to play like Stockfish, or like a human of a given Elo band.

A bilinear head scores every (from-square, to-square) pair in one pass, and the probability is spread over legal moves only, the same masking Laya applies to its options. A second head predicts the probability of winning. The model has 6.4M parameters. Figure 8 draws it in the same convention as the Tetris models.

Figure 8 · The chess model
CHESS · 6.4Msoftmaxlegal moves onlysigmoidP(win), side to moveBilinear move head⟨Wq hᵢ, Wk hⱼ⟩, 64×64Value headread on [VAL]8×Feed forwardLayerNormMulti-head attentionsquares attend to each otherLayerNormPiece + square embeddingsd = 256, + 50-move and repetition inputsa1b1…h8castlee.p.levelVAL64 square tokens; “level” = Stockfishor a human Elo bandInput: the position, from the side to move
embeddingsattentionfeed forwardlayer normheads
Same convention as Figure 4. One token per square, plus castling rights, en passant, a level token (Stockfish, or a human Elo band) and a value token. Eight pre-norm blocks feed two heads: a bilinear head that scores every (from, to) pair at once, masked to legal moves, and a value head that reads the probability of winning.

It learned from 10 million positions analysed by Stockfish at depth 40–60 (Lichess's open evaluation database) and 10 million positions from 300,000 human games, all trained on the same laptop. At the end of imitation, on positions it had never seen:

  • it played Stockfish's best move 48.1% of the time;
  • it predicted the human move 50.9% of the time, in the range reported for Maia, the reference human-move model;
  • Stockfish judged its moves 87.6% accurate (Lichess's accuracy formula), with 6% blunders.
Figure 9 · Learning chess by imitation
20%29%38%47%56%1,00074k146k219k291kcosine decayDAgger
Stockfish's best movehuman move
Share of held-out positions where the model's most probable move is Stockfish's best move, and where it is the move a human actually played. The constant learning rate plateaus; a cosine decay brings a clear gain; DAgger then trades a little agreement for fewer blunders on the model's own positions.

We then ran DAgger against Stockfish, the chess counterpart of the Tetris recipe. The model played 840 games, half against itself and half against a strength-limited Stockfish whose level tracked the model's results. Stockfish labelled the 60,577 positions where the model was to move, and the real game outcome was mixed into the value target, as in AlphaZero.

We measured strength in 120 games per setting against Stockfish limited to 1,400–3,000 Elo, with every setting using the same engine:

Figure 10 · Elo against a strength-limited Stockfish
1600180020002200No searchImitation1744Imitation + DAgger188964-simulation searchImitation2058Imitation + DAgger2033
ImitationImitation + DAgger
120 games per row, 95% intervals. DAgger moves the search-free model up by about 145 Elo; once a search is added, both checkpoints land in the same place, because the search was already catching the mistakes DAgger removed.
Show the data
ModelPlayElo
ImitationNo search1744 ± 107
Imitation + DAggerNo search1889 ± 106
Imitation64-simulation search2058 ± 106
Imitation + DAgger64-simulation search2033 ± 106

DAgger added about 145 Elo to the network's instinctive play, cutting blunders from 6% to 5%. It added nothing once a search was layered on top: the search was already catching the mistakes DAgger removed.

Stockfish's Elo limiter is calibrated at a different time control, so these figures are an order of magnitude, not a rating.

What we take from it

Representation set the ceiling, not scale. The same game, the same compute, the same kind of architecture. The difference between 245 and 398 lines is what the model was allowed to see. The words Laya reads well are too coarse to rank moves one at a time, and the numbers that would be enough are words it does not read. Laya's structure is a strength for problems that are naturally phrased in language, like a ticket to route or a passage to rerank. On a problem whose structure is geometric, forcing it through words cost more than the model's scale could buy back.

A big generalist is not always the answer. This is where we started. When a task is well defined and comes back often, the reflex is to take the largest general model available and phrase the task in its terms. Here that reflex cost 23 hours and 421M parameters for 245 lines. A 0.53M-parameter model built around the game reached expert play in half an hour on the same laptop; it trains faster, runs faster and plays better. The chess model, 6.4M parameters, took a couple of days on that laptop. Neither needed reinforcement learning from scratch: imitating an existing expert, on the learner's own states, did most of the work.

Controls matter more than curves. Each conclusion above rests on a control rather than a training curve:

  • the permutation test showed what Laya was really reading;
  • the expert forced through the sentences, and the search for the best ranking of them, showed where its ceiling came from;
  • the identical-engine Elo matches separated DAgger's effect from the search's.

Along the way, the controls also caught real bugs: a dropout that silently corrupted the PPO ratio, castling notation that mislabelled 1.6% of the chess targets, and a Tetris engine that let pieces pass through blocks.

Generalists and specialists answer different questions. Our two models are not System One models, and they are not trying to be. Jev and Laya are generalists: one model, any state, any question written at run time, which is exactly what makes them useful in front of an LLM stack. The Tetris model knows one game and the chess model another, and neither will route a support ticket. Our point is narrower: a task with a structure of its own, that you will run again and again, deserves its own model. The two are not rivals. A generalist decides what kind of problem it is looking at; a specialist solves the problems that come back often enough to deserve one.

The full methodology, protocols and references are in the accompanying paper.