Use the win-rate calculator when you have one binary outcome and need to see how uncertain it still is.
Evidence after the playtest
Board Game Balance Test Analysis
Turn recorded human playtests into review signals for win rates, seat order, faction matchups, player counts, and scoring paths—without pretending that a small sample can certify balance.
Your playtest files stay on your device. CSV analysis runs inside this browser tab; no rows are uploaded or stored by Tabletop Maker Lab.
Five independent questions
Choose the evidence that matches the design decision.
Each tool uses a distinct dataset and comparison. Start with one hypothesis instead of pouring unrelated sessions into a single balance score.
Win Rate Confidence Calculator
Put an observed win rate beside a 95% Wilson interval and your own target band.
CSV analyzerFirst-Player & Seat Advantage
Compare each winning seat with an equal-seat baseline by version and player count.
CSV analyzerFaction Matchup Analyzer
Normalize reversed faction order, preserve draws, and flag matchup intervals for review.
CSV analyzerPlayer-Count Balance Analyzer
Compare duration and score-spread changes against a creator-selected baseline count.
CSV analyzerScore Path Analyzer
Measure category share and winner association from long-form player scoring data.
A defensible analysis loop
Segment → quantify → inspect → retest
- Keep version, player count, matchup, and recruitment conditions visible.
- Use confidence intervals and user-entered review rules to expose uncertainty.
- Inspect the sessions behind a signal before changing the design.
- Record the next version separately and test whether the pattern repeats.
What these tools do not decide
A flag is not a verdict.
Observational playtest data can be distorted by player skill, teaching, version drift, matchup selection, and small samples. These tools organize evidence; they do not prove fairness, causation, fun, or publication readiness.
Use a CSV analyzer when version, player count, seat, faction, or scoring category changes the question.
Retest flagged and under-sampled groups with comparable players and rules before treating the pattern as stable.