Overview

Brief your jury

The evidence, in three results#

Result Finding Source
ICC ≈ 0 43 trained NIH reviewers scoring the same 25 funded applications in a mock study section agreed at essentially zero. Expertise alone did not produce agreement Pier et al., PNAS, 2018
p = .228 Telling 42 reviewers how their scores differed from others, after the fact, in a randomized trial at a Norwegian funder. Agreement did not move Hesselberg et al., Research Integrity and Peer Review, 2021
0.61 → 0.89 Agreement before and after an 11-minute training video, in a randomized trial with 75 professors Sattler et al., PLOS ONE, 2015

One conclusion: calibrate before scoring, and write criteria a judge can verify rather than feel.

The pack: six blocks, in reading order#

Send it as one document a week before the event, with the video linked at the top. Everything else about your event belongs in a separate logistics email, so the pack stays about one thing: how to score.

1 · The rubric, with anchors. Not a list of criterion names. Each dimension needs an anchor description per band, so a judge can tell a 3 from a 7 before the first pitch. The full pitch rubric with weights and anchors is published to be dropped straight in.

2 · The scoring procedure. Three sentences do most of the work: cite the evidence behind every score, name the band before the number, and when the score sits on a boundary with evidence missing, take the lower band. Judges who follow this produce records, not impressions.

3 · The disagreement rule. Define in the pack what happens when scores diverge. In EvalLens the rule is spread — highest minus lowest per dimension, where under 1.5 is consensus, 1.5 to 2.99 is a split, and 3.0 or more is a conflict that goes to discussion instead of an average. Judges who know disagreement is expected stop softening their scores toward the middle.

4 · Conflicts of interest, upfront. Ask each judge to declare investments, employment, mentorship or personal ties to any team, and state the recusal mechanics. Collect it before scores exist — a recusal after the leaderboard looks like damage control. A log of who scored what makes any recusal verifiable later.

5 · The load and the schedule. Tell each judge exactly how many submissions they will read, how long one takes, and when breaks land. Unplanned load is where scoring quality quietly dies.

6 · The calibration exercise. An 11-minute video and one practice submission, done before event day. The highest-return block in the pack.

Building the calibration#

  1. Record one 11-minute video. Walk through your scale one band at a time, in your own words, with your own anchors. In the Sattler trial this format alone moved agreement from ICC 0.61 to 0.89 — less time than watching one extra pitch.
  2. Score one sample submission on camera. Take a real deck from a past event, walk the rubric, cite the evidence, name the band, pick the number. Judges copy what they see better than what they read.
  3. Show the cost of a wrong band. Trained reviewers in the same trial picked the correct band 74% of the time against 35% untrained. One concrete example of a misplaced score changing a ranking, and the scale stops being decorative.
  4. Lock independent scores before any discussion. Every judge submits before deliberation, the chair speaks last, and discussion resolves flagged conflicts rather than manufacturing agreement.

The load arithmetic#

Promise your judges a number, not an evening.

Figure What it is Source
4 min Per project per judge in the MLH budget: 2 minutes of demo, 1 for questions and scoring, 1 to walk to the next table MLH organizer guide
J = ceil(P × n × t / T) The MLH judge-count formula — 13 judges for its worked 500-attendee, 120-minute example at three rounds per project, plus a 2–3 judge no-show buffer MLH organizer guide
30 min Technovation's estimate for reviewing one submission properly, with judges committing to at least 5 and around 3 hours including training technovation.org
1.5 h Per business plan when written feedback is part of the ask, at Project ECHO's competition — 3 plans, roughly 6 hours projectecho.org

When the formula asks for more judges than you can recruit, the honest options are fewer rounds, a longer window, or a first read that lands before your judges start. That last one is what EvalLens does: the panel scores every submission against your rubric and briefs each judge on what to verify, so panel hours go to decisions instead of triage.

Making the pack load-bearing#

A judge can skip a PDF, but not a form field that asks which slide the score came from. Put the anchors on the scorecard itself rather than in an appendix, require an evidence line next to every score, and open the event by scoring one practice submission together.

Next steps#

Updated

Was this page helpful?