Draft Science · Auctions

How Pikola is tested

Pikola's homepage says "three times the championships". This page is the test behind that line: simulated Yahoo auctions in which a team drafted by Pikola's engine and a team drafted off a rankings list sit in the same room, and each finished auction is played out over 32 simulated seasons.

In simulated Yahoo auctions, against a drafter who buys off a rankings list, Pikola's team won the title 2.7 to 3 times as often in the three rooms built to draft like people.

The test

Every run is one auction: 14 teams drafting 13-man rosters from the same pool of players and the same season projections. Two seats matter. Pikola's engine holds one. It prices every player against the roster it has and the money left in the room, and it recomputes after every sale. The benchmark holds the other. The remaining twelve seats are the room, described below.

The benchmark is a drafter who buys off a rankings list. Picture a manager with a good cheat sheet and last year's draft results. He ranks every player by the sum of his z-scores across the nine categories, which is the standard way projections become a rankings list. Then he pays for the k-th player on his list what this league paid for its k-th most expensive player last year: the top player gets last year's top price, the fortieth gets last year's fortieth price, and so on down the list. He has no favourites, no punts and no opinions, and he never deviates.

Both drafters sit in fixed seats in the same auction, so neither gets a better draft position or nomination order, and both face the same room, the same prices and the same luck. When the auction ends, its rosters are played through 32 simulated seasons on the simulator's model of player production, scored head-to-head by category under the league's each-category rule, with playoffs. Ties in the auction and in the standings are broken by seeded coin flips, never by seat. A tied playoff week goes to the higher seed, as Yahoo's does. Every seat must hold a legal 13-man roster after every buy.

The title margin is the engine's championship rate minus the benchmark's, paired by auction, with the standard error across the 40 auctions. The playoff margin and the win-rate margin are the same difference in the rate of making the playoffs and in the regular-season win rate (category results won, ties at a half). A championship carries a lot of lottery. Those two carry far less, and they are the figures to read when a title margin is inside its error.

The rooms

The other twelve seats set the prices, so they decide how hard a title is to win. We draft in four rooms. Three are built to draft like people; the fourth is not.

human Human-like

Strategists who pay the room's book: last year's prices by rank, with Yahoo's numbers where the room shows them. Each has category fixations, and a couple have declared punts.

mixed Human-like

Strategists beside drafters who read the two numbers Yahoo shows beside every player, his projected auction value and his average price in Yahoo auctions, and pay somewhere between them.

yahoo Human-like

A room of those Yahoo readers alone. In this row the engine also reads the two numbers Yahoo shows, at the 35% weight the app uses by default; in this room that weight measured level with none.

z Not human-like

Twelve drafters of the benchmark's own kind: z-score rankings, paying prices that rise in a straight line by rank, so nobody pays a star's price. It is the easy room, and it is not how people draft.

In an auction the price is set by the second-highest valuation, not the highest. Each room is tuned so that the bidder who sets a price, the second-keenest, pays last year's price for a player of that rank.

Results

One row per room. The two stress rows replay the mixed room's auctions with the season played on a truth the drafters did not have. Rates are shares of the 1,280 seasons; margins are in percentage points, ± one standard error across the 40 auctions.

Championship, playoff and win-rate margins by room

Championship, playoff and win-rate margins by room
Room Pikola champ Benchmark champ Title margin Playoff margin Win-rate margin
humanstrategists on the room's book 26.4%9.4%17.0 ± 2.019.3 ± 2.14.9 ± 0.4
mixedstrategists beside Yahoo readers 25.7%9.6%16.1 ± 1.621.6 ± 2.25.6 ± 0.4
yahooYahoo readers alone; the engine reading Yahoo's numbers at 35% 26.6%9.0%17.6 ± 2.222.4 ± 2.66.1 ± 0.5
ztwelve z-score drafters; not human-like 30.6%3.8%26.9 ± 1.649.1 ± 2.59.4 ± 0.4
mixed, 15% projection noisestress test: the mixed room's auctions, seasons played on projections wrong by 15% 21.6%8.4%13.1 ± 2.722.4 ± 4.95.4 ± 0.9
mixed, the room half right about pricesstress test: the mixed room's auctions, production pulled halfway toward the room's prices 3.4%0.4%3.0 ± 0.727.7 ± 3.84.1 ± 0.6

"Champ" is the share of the 1,280 seasons that team won the title. The engine made the playoffs in 93.0% of seasons in the human room, 94.0% in mixed, 93.0% in yahoo and 95.3% in z; 88.8% under projection noise and 42.8% in the room that was half right about prices.

A five-to-six point edge in weekly category wins became a sixteen-to-eighteen point edge in titles. The playoffs multiply a small weekly edge, which is also why the title rows are the noisy ones.

Par in a 14-team league is a title one season in 14 (7.1%) and the playoffs eight seasons in 14 (57%). In the clean rows both seats beat par by a distance: the benchmark makes the playoffs 71% to 74% of the time in the human-like rooms. Both draft from the projections the season is then drawn from, while those rooms price off last year's results and the numbers Yahoo shows. So the levels are not what to expect in your league. The gap between the two seats is the test, and the two stress rows are what happens when the room knows something too.

Championship rate by room

Share of 1,280 simulated seasons won, the team drafted by Pikola's engine against the rankings-list drafter in the same auctions. The whisker is ± one standard error on the gap between the two bars, paired by auction. The figure at the right is the ratio of the two rates.

Drawn to scale from the CSV below.

The ratios people ask about are the engine's title rate divided by the benchmark's, room by room: human 2.8×, mixed 2.7×, yahoo 3.0×, and z 8× (30.6 ÷ 3.8 = 8.1). "Three times the championships" is the three human-like rooms, 2.7 to 3. Each ratio carries about ±0.2 at one standard error (the margin's error over the benchmark's rate); 2.7 to 3 is the spread across rooms, not an error bar. The z room's 8× is real, but it is a room of the benchmark's own kind, flat at the top, and it is not the number to quote.

We don't pool the rooms. They disagree by more than their errors (Q = 28.9 on 3 degrees of freedom), so the edge is a property of the room: read the rows.

Two stress tests

15% projection noise. The seasons are played on projections that are wrong by 15%, so both drafters bought from numbers that were off. The engine's title rate falls from 25.7% to 21.6% and the benchmark's from 9.6% to 8.4%. The title margin is 13.1 ± 2.7 against the clean row's 16.1 ± 1.6, inside the error, and the playoff margin (22.4 ± 4.9) and win-rate margin (5.4 ± 0.9) barely move. Worse projections cost the engine some titles, not its edge over the rankings list.

The room half right about prices. Here the room's prices carry information the projections did not. Each player's production is pulled halfway toward what the room paid for him: a player the room paid more for than the projections said produces more, one it let go cheap produces less. That is a world in which the whole room knew something, and in it both of our drafters are punished for trusting the projections. The engine wins 3.4% of titles and the benchmark 0.4%; the engine's playoff rate falls from 94.0% to 42.8%, below par (57%). In a world where the room's prices know something the projections don't, a projection-driven drafter is a below-average team, and Pikola is one. It still beats the benchmark (playoff margin 27.7 ± 3.8, win-rate margin 4.1 ± 0.6), but the title ratio means nothing at those rates. Pikola cannot know what the projections do not.

Calibration

A second check asks whether the engine's confidence is right. At the end of each auction the engine predicts, for its finished roster, the probability of winning each category. We bin those predictions and compare them with what happened in the simulated seasons, in the z room: the run keeps calibration for its first row, and the engine's category profile differs by room, so the slopes below are that room's. The season is played on the simulator's own model of production, so this checks the engine's arithmetic against that model, that a roster it gives 0.6 wins 0.6, not against a real season. One prediction per auction is compared with each of its 32 seasons, so n counts seasons.

Predicted category win probability against the realised rate, z room

Predicted category win probability against the realised rate, z room
Predicted binMean predictedRealisedn
0 to 0.10.0810.10464
0.1 to 0.20.1630.180448
0.2 to 0.30.2580.275640
0.3 to 0.40.3550.385992
0.4 to 0.50.4570.4731,440
0.5 to 0.60.5530.5591,920
0.6 to 0.70.6580.6382,432
0.7 to 0.80.7430.7272,624
0.8 to 0.90.8440.830928
0.9 to 10.9310.89132

The two columns are close on every populated row. The engine is a little overconfident at the top: rosters it gives 0.6 to 0.7 win 0.64, and the few it gives more than 0.9 win 0.89. It is a little underconfident below 0.5. By category, the slope of realised on predicted runs from 0.66 (steals) to 0.95 (free-throw percentage), where 1 would be perfect; the engine is most overconfident in steals, blocks and assists (slopes 0.66 to 0.72).

Data

The six rows above, with the standard errors and the engine's playoff rates, as a CSV: how-pikola-is-tested.csv. No player names and no per-player numbers.

Last run: 2 October 2026.

Cite as: Pikola, "How Pikola is tested", pikola.app/research/how-pikola-is-tested, October 2026.

Free to quote with a link; the data is CC BY 4.0. Questions about the method: hello@pikola.app.

Pikola's other study: Should you punt from pick one? (snake drafts). Try the panel in a free Yahoo mock draft, or read what Pikola is and isn't.