Loading…
Fetching the problem index.
Pick a benchmark and a problem on the left. The problem page shows the statement, the reference proof, the judge's grading rubric and the prior-attempt evidence packs. The grids below show every rollout of every method: rows are methods, columns are seeds, and each cell is the external judge's score (0–7, strict success = 7). ✓ means the trajectory's own verifier accepted the final candidate, ↻ means it hit the 3-candidate limit, and ⚠ marks protocol or budget errors.
Click a cell to open that rollout. Each model call is shown in order, with its prompt (collapsed), private reasoning (collapsed) and response or tool call. Verifier calls show their verdict. ⟦…⟧ chips inside prompts stand for text shown elsewhere (the problem, a candidate, an evidence pack, an earlier turn); click one to expand it. Continuation turns show only their new messages. Pin a rollout to compare it with another side by side. Use j/k to step through problems.
Summary aggregates every rollout per method: 7/7 rate with 95% intervals, tokens, how runs end, GVR verifier calls and value-tool queries. Analysis asks what each method does on problems Direct gets wrong: fix and break rates against Direct, and what GVR's verifier does with its first candidate. Hover or focus any mark for its numbers; every chart has a table view.
Fetching the problem index.