Criteria and weights
Why the rubric, not the judges#
A study of Olympic breaking at the Paris 2024 games found expert judge agreement between 0.21 and 0.45 on loosely defined criteria, while artistic gymnastics — which enumerates every observable element — reaches 0.94 to 0.98. Same caliber of judge; the rubric is the difference. A pitch jury scoring "Team" and "Market" on a bare 1-to-10 column sits on the breaking side of that gap.
The related failure is a scale that collapses. In an AIBS grant-review case study, reviewers used only the 1.3-to-4 part of a 1-to-5 scale: without anchors, a ten-point scale becomes a three-point one.
The default pitch rubric#
| Dimension | Weight | A 3 looks like | A 7 looks like |
|---|---|---|---|
| P1 Problem significance | 0.15 | No real problem articulated, or the pain is vague and unsubstantiated | A specific, frequent, costly pain with a clearly identified user and a credible reason it matters now |
| P2 Solution differentiation | 0.15 | Solution unclear, disconnected from the problem, or a thin wrapper over an existing tool | A coherent solution with a clear mechanism and a genuine, defensible difference from alternatives |
| P3 Market attractiveness | 0.20 | No market reasoning, no segment defined, or an implausible market claim | A well-sized, reachable market with a clear segment, a credible entry motion and a believable path to first customers |
| P4 Business model / GTM | 0.15 | No monetization logic, or pricing that ignores the buyer | Clear monetization, sensible pricing for the buyer, and a credible go-to-market motion with a beachhead |
| P5 Team / founder fit | 0.20 | No meaningful information about the team, or an obvious mismatch with what the venture requires | A capable, reasonably complete team with relevant experience and good fit to the problem |
| P6 Feasibility / readiness | 0.15 | The plan is implausible, internally inconsistent, or absent | A credible, well-sequenced plan with resources that broadly match the ambition and risks acknowledged |
In the product each dimension carries anchors for four bands — 0–3, 4–6, 7–8 and 9–10 — plus red flags, with the top band reserved for what is demonstrated rather than asserted. The two columns above are the working core: if judges can tell a 3 from a 7 the same way, most of the disagreement problem is already gone.
Red flags per dimension, and the full anchor set: Dimensions P1–P6.
Editing the weights#
Weights are yours. They are meant to be edited before scoring starts, and the defaults simply reflect what early-stage juries most often argue about — market and team carrying 0.20 each.
Two adaptations worth knowing:
- Demo day. A cohort that just finished a program has had equal coaching on story, so many organizers shift weight toward Team / founder fit (P5) and Feasibility (P6) and away from pitch polish.
- Hackathon. The rubric changes shape rather than weight: Execution and Demo carries 0.30 and Technical Depth 0.20, both weight-protected, with Problem Impact and Innovation at 0.15 and UX Clarity and Delivery Readiness at 0.10. See Hackathons.
The lock#
Once scoring begins, weights freeze. This is the one rule to keep even if you rewrite every anchor. A field ranked partly under one weighting and partly under another is not a ranking, and the lock is what lets you say every submission was ranked on the same standard.
Because weights apply at the leaderboard rather than inside each judge's reading, the same evidence can be re-ranked under different weights without re-running the batch — which is the supported way to explore "what if market mattered more" after a run, instead of editing mid-round.
Multiple tracks#
Tracks, criteria and weights are configured per project, and each track scores against its own rubric. The leaderboard respects the weighting of the track it belongs to.
Using this rubric without EvalLens#
The dimensions, weights, anchors and the spread rule work on paper, in a spreadsheet, or in any scoring tool — they are published to be copied. EvalLens becomes useful when the field is bigger than your judges' hours: the panel runs the first read on this same rubric and your judges decide with the evidence in front of them.
Next steps#
- Brief your jury — putting the anchors on the scorecard rather than in an appendix.
- Disagreement and spread — the threshold that turns disagreement into an action.
- How the score is built — where weights enter the arithmetic.