The LLWS simulator was overestimating favorites and ignoring games that had
already been played. Two separate causes:
1. Championship futures were used directly as single-game strength
(p1 / (p1 + p2)). A future already compounds the ~6 wins needed to take
the title, so this made every individual game as lopsided as the whole
tournament and re-compounded that edge round after round. Against a
representative 20-team board the favorite priced at 21.8% simulated at
44.9%, and the longest shot fell to ~0%.
Futures are now decompressed to single-game Elo via convertFuturesToElo,
the same pipeline the other bracket simulators use, and games are played
with eloWinProbabilityWithParity. The parity factor was calibrated by
sweeping it until a randomized-draw simulation reproduces the board it
was fed: at 1000 the favorite simulates at 21.8% and field-wide RMSE
drops from 0.062 to 0.003. It is overridable per season via config.
2. The simulator never read playoff_matches, so it re-ran the tournament
from an empty bracket every time and shuffled the draw at random each
iteration. A recorded loss changed nothing.
It now loads the seeded llws_20 bracket, places teams in their real
slots, and replays completed games from their recorded result instead of
re-simulating them, so an eliminated team correctly drops to zero. When
no bracket exists (or it has no participants seeded) it falls back to the
previous randomized-draw behavior, and a seeded bracket is authoritative
about which side a team is on, so externalId is only required on the
pre-bracket path.
Guards: a recorded result is only honored when its two participants are the
ones the simulation routed into that game, so a corrupt or out-of-order row
cannot desynchronize the rest of the bracket; brackets seeding an unknown or
duplicated participant now fail loudly.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X5vQZMPeokzfMqHQjq1RDZ