♚Open in the app

Morozov Balance Index (MBI): methodology

The Morozov Balance Index (MBI) answers one question for each of the 960 Chess960 starting positions: how big is White's opening advantage? This page is the full method — every number on this site can be recomputed from it.

Author of the methodology: Roman Morozov. About the author · Version 2.0 · 2026-09-18

CC BY 4.0 Methodology and data: CC BY 4.0 — free to use with attribution to the author.

What the index measures

In chess, White moves first. That first move is worth something, and in Chess960 it is worth a different amount in each of the 960 starting setups. The index measures the SIZE of that opening edge — not who is better. In 956 of the 960 positions the engine puts the start at equality or in White's favour; the other four land 2 to 7 centipawns on Black's side, which is less than the search moves between its own iterations — noise around zero, not an advantage for Black. The measure is the distance from zero whichever way it points, so the honest phrasing is "how large is the head start here", and a balanced position is one where that head start is as small as it gets.

MBI is a 0–100 scale. 100 means the smallest measured White edge among all 960 positions — the most balanced start. 0 means the largest measured edge. The index is a rank-based transform of one measured quantity (see below), so it says how a position compares with the other 959, not how many pawns anybody is up.

How each position is evaluated

Every starting position is analysed once by Stockfish from the initial setup, with a fixed budget of search NODES rather than a fixed amount of time. This is the key decision of the whole method: a node budget is machine-independent. A time budget would give a fast computer a deeper search and therefore a different number, and nobody could check the result.

Single-threaded search with a fixed hash and a fixed node count is deterministic: the same binary on any machine returns the same centipawn value, bit for bit. That is what makes an independent check possible rather than merely plausible.

From centipawns to the index

Centipawns are not a linear measure of anything human, and the difference between +18 and +22 is not meaningful on its own. The index therefore uses the RANK of a position among all 960, not the raw number:

rank(sp) = position of eval_cp(sp) in the sorted list of all 960 values
(ascending: smallest White edge first, ties share the mean rank)
MBI(sp) = round( 100 * (960 - rank(sp)) / (960 - 1) )
Class White edge Positions
A — most balanced ≤ +19 cp 240
B — balanced +20…+26 cp 240
C — noticeable edge +27…+35 cp 240
D — largest edge > +35 cp 240

Across all 960 positions the engine value runs from -7 to +83 centipawns, median +26.0.

How the 960 values are distributed

Each bar is a five-centipawn bucket; the number on the right is how many of the 960 starting positions fall into it.

-10..-6 cp1
-5..-1 cp3
0..4 cp19
5..9 cp25
10..14 cp72
15..19 cp135
20..24 cp159
25..29 cp173
30..34 cp131
35..39 cp96
40..44 cp63
45..49 cp43
50..54 cp18
55..59 cp8
60..64 cp7
65..69 cp3
70..74 cp0
75..79 cp2
80..84 cp2

Validation

An engine evaluation is a claim about a position. To check that the claim predicts actual results, a subset of positions is played out: engine against engine, many games per position, with the seeds fixed and published. If positions the index calls balanced really do score closer to 50% for White, the index measures what it says it measures.

Seed base: 20260825. Every game can be replayed exactly.

Position Group Engine, cp Games W/D/L White score
42 random +23 50 6/44/0 56.0% (42.3–68.8%)
52 random +18 50 5/45/0 55.0% (41.3–67.9%)
80 most skewed +83 50 7/42/1 56.0% (42.3–68.8%)
86 random +34 50 1/48/1 50.0% (36.6–63.4%)
102 most balanced 0 50 3/47/0 53.0% (39.5–66.1%)
131 most balanced 0 50 0/50/0 50.0% (36.6–63.4%)
204 most balanced 0 50 0/46/4 46.0% (33.0–59.6%)
217 most balanced +3 50 0/50/0 50.0% (36.6–63.4%)
259 most balanced +3 50 2/43/5 47.0% (33.9–60.6%)
262 random +43 50 6/44/0 56.0% (42.3–68.8%)
266 random +26 50 3/47/0 53.0% (39.5–66.1%)
268 most skewed +58 50 8/42/0 58.0% (44.2–70.6%)
269 most balanced 0 50 1/46/3 48.0% (34.8–61.5%)
272 most skewed +58 50 13/36/1 62.0% (48.2–74.1%)
292 random +49 50 1/49/0 51.0% (37.6–64.3%)
300 most balanced 0 50 3/47/0 53.0% (39.5–66.1%)
332 most skewed +59 50 10/40/0 60.0% (46.2–72.4%)
348 most skewed +62 50 6/44/0 56.0% (42.3–68.8%)
399 most skewed +65 50 7/38/5 52.0% (38.5–65.2%)
401 most balanced -7 50 0/49/1 49.0% (35.7–62.4%)
418 most balanced -2 50 1/49/0 51.0% (37.6–64.3%)
424 random +23 50 6/44/0 56.0% (42.3–68.8%)
429 most balanced 0 50 0/50/0 50.0% (36.6–63.4%)
432 most balanced +3 50 5/44/1 54.0% (40.4–67.0%)
433 most balanced -4 50 4/45/1 53.0% (39.5–66.1%)
467 most balanced +2 50 2/47/1 51.0% (37.6–64.3%)
471 most balanced 0 50 0/50/0 50.0% (36.6–63.4%)
481 random +5 50 1/49/0 51.0% (37.6–64.3%)
485 most balanced +3 50 1/49/0 51.0% (37.6–64.3%)
497 most balanced 0 50 0/48/2 48.0% (34.8–61.5%)
499 random +18 50 4/46/0 54.0% (40.4–67.0%)
531 random +13 50 4/46/0 54.0% (40.4–67.0%)
556 random +21 50 2/48/0 52.0% (38.5–65.2%)
557 most skewed +63 50 15/35/0 65.0% (51.1–76.7%)
565 most skewed +60 50 4/44/2 52.0% (38.5–65.2%)
576 most balanced 0 50 1/48/1 50.0% (36.6–63.4%)
578 most skewed +56 50 12/38/0 62.0% (48.2–74.1%)
593 most balanced +2 50 0/49/1 49.0% (35.7–62.4%)
604 most skewed +81 50 13/37/0 63.0% (49.1–75.0%)
620 most skewed +63 50 12/38/0 62.0% (48.2–74.1%)
637 random +21 50 7/43/0 57.0% (43.3–69.7%)
654 random +34 50 4/45/1 53.0% (39.5–66.1%)
665 random +22 50 6/44/0 56.0% (42.3–68.8%)
745 random +13 50 3/46/1 52.0% (38.5–65.2%)
748 random +42 50 6/43/1 55.0% (41.3–67.9%)
760 most skewed +61 50 8/40/2 56.0% (42.3–68.8%)
784 most skewed +58 50 10/38/2 58.0% (44.2–70.6%)
786 most balanced +2 50 4/45/1 53.0% (39.5–66.1%)
794 most skewed +67 50 12/38/0 62.0% (48.2–74.1%)
801 random +11 50 1/49/0 51.0% (37.6–64.3%)
814 random +32 50 5/45/0 55.0% (41.3–67.9%)
823 most balanced -5 50 2/48/0 52.0% (38.5–65.2%)
868 most skewed +75 50 9/41/0 59.0% (45.2–71.5%)
870 random +44 50 5/45/0 55.0% (41.3–67.9%)
879 most skewed +66 50 11/38/1 60.0% (46.2–72.4%)
880 most skewed +63 50 7/42/1 56.0% (42.3–68.8%)
885 random +14 50 1/49/0 51.0% (37.6–64.3%)
886 most skewed +58 50 5/45/0 55.0% (41.3–67.9%)
902 most skewed +62 50 5/44/1 54.0% (40.4–67.0%)
935 most skewed +78 50 11/38/1 60.0% (46.2–72.4%)

Deep verification and classification v2

Version 2.0 adds a second layer on top of the index. Every position was played out along the engine's main line for 30 moves (35 for the priority set), the most balanced positions and every zero or negative one were re-searched with a 3–10× larger node budget, and the zero/negative ones also received a three-line MultiPV search and a second self-play series.

Balance band (by the base evaluation)

Absolute centipawn thresholds instead of quartiles: E — even, |cp| ≤ 10; S — slight, 11–25; M — moderate, 26–40; L — large, ≥ 41. Quartiles guaranteed 240 “balanced” positions whatever the data; a band holds only as many as there really are. The base run (400 000 000 nodes, the same depth for all 960) sets the band — the deeper search verifies it.

Band Positions
E — even 64
S — slight edge 386
M — moderate edge 377
L — large edge 133

Dynamics (along the play-out)

The same engine plays both sides with a fixed budget of 5 000 000 nodes per move and the evaluation is recorded before every half-move. volatile — the evaluation swung by ≥ 60 cp or changed sign at least 2 times (only values beyond ±15 cp count as a sign); equalizing — the final edge is at least 10 cp smaller than at the start; sharpening — at least 15 cp larger; stable — everything else.

Evaluations at this per-move budget overstate White's edge early in the game and converge towards zero as the position simplifies, so the raw trajectory of almost any position “equalizes”. Before labelling, the median trajectory of all 960 positions (per half-move) is subtracted from each play-out: the common drift cancels out and what remains is how this position differs from a typical one. The raw checkpoints at 10, 20 and 30 moves are published as is.

Dynamics Positions
stable 542
equalizing 285
sharpening 25
volatile 108

Bands × dynamics

Band stable equalizing sharpening volatile
E 28 33 1 2
S 234 120 6 26
M 240 81 14 42
L 40 51 4 38

Deeper re-search

Positions with |cp| ≤ 25 were re-searched with 1 200 000 000 nodes, the zero and negative ones with 4 000 000 000 nodes — one thread, fixed hash, so every number is bit-reproducible. A position counts as verified when the deeper value stays within 10 cp of the base run (twice the iteration-to-iteration noise of the base search): 123 of 147 verified. The deeper search shifts these positions systematically upwards — median +3 cp at ×3 and +9 cp at ×10 — and not one of them came out negative; the band by the deeper value is published next to the base band, so the reader sees both.

The contested thirteen

Thirteen positions came out at zero or slightly negative in the base run. The deeper search decides the sign; the three-line MultiPV search (800 000 000 nodes) and a second self-play series (100 games per position at 1 000 000 nodes per move) are independent checks, not a veto.

Position Engine, cp Deep, cp Verdict Self-play v2
102 0 +20 White edge — the minus was noise 54.5% (44.8–63.9%)
131 0 +5 even — within engine noise 50.5% (40.9–60.1%)
204 0 +9 even — within engine noise 47.5% (38.0–57.2%)
269 0 +15 White edge — the minus was noise 48.5% (38.9–58.2%)
300 0 +12 White edge — the minus was noise 51.5% (41.8–61.1%)
401 -7 0 even — within engine noise 49.0% (39.4–58.7%)
418 -2 +5 even — within engine noise 50.5% (40.9–60.1%)
429 0 +10 even — within engine noise 51.0% (41.3–60.6%)
433 -4 +3 even — within engine noise 51.5% (41.8–61.1%)
471 0 +15 White edge — the minus was noise 50.0% (40.4–59.6%)
497 0 +6 even — within engine noise 49.0% (39.4–58.7%)
576 0 +13 White edge — the minus was noise 50.0% (40.4–59.6%)
823 -5 0 even — within engine noise 51.5% (41.8–61.1%)

Bands by position family

Nomenclature E S M L
BastionGuard 14 69 58 27
BastionScout 14 61 70 23
RampartGuard 11 53 67 13
RampartScout 2 30 30 10
ThroneGuard 16 98 65 25
ThroneScout 7 75 87 35

Limitations — read this before quoting the numbers

Reproduce it yourself

Everything needed to recompute the dataset is public: the exact engine build, the budget, the scripts and the raw output. Recomputing one position takes a few minutes on a single core; the whole set is an overnight run.

  1. Download the raw data and the scripts from the open-data page. Open data
  2. Check one position against the published dataset — this recomputes it from scratch and compares:
    python tools/balance/verify_reproduce.py --sp 518
  3. Or recompute the entire set (resumable; it appends to the same file if interrupted):
    python tools/balance/eval_all.py --nodes 400000000 --workers 20

If your value differs, the difference is the finding — the engine build, the node budget or the options will not match. All three are printed above.

Citation and licence

The methodology and the dataset are published under CC BY 4.0: use them for anything, including commercially, with attribution.

Morozov R. Chess960 Balance Index (MBI). ch960.com/balance, 2026. CC BY 4.0.

Balance of all 960 positions · Dataset and raw engine output