Overview

EvalLens documentation

These pages document how that works in practice: what to configure before a run, what happens to a deck inside the pipeline, how to read what comes back, and what EvalLens deliberately does not do.

Start here
Run a round
The panel
Scoring
Trust
By program type

What makes an EvalLens score different#

Evidence comes before the number. A judge must cite the slide, state what supports the score and what lowers it, name the rubric band, and only then pick a number inside that band. On a boundary with evidence missing, the rule is the lower band.

Six lenses, not one opinion. Six judges read each deck independently and never see one another's scores. Where they disagree, the report shows the spread instead of averaging it away.

The arithmetic is deterministic. No model call runs during final aggregation: the same judge outputs and weights produce the same AI Total Score every time.

The ranking is human. The leaderboard is built only from submitted Jury Scores and your criterion weights. The AI Total Score sits beside them as a read-only reference.

Where this fits#

EvalLens does not replace your judges, your intake tool, or your rules. It runs the first read of the whole field so that judge hours go to decisions instead of triage — and leaves a record that explains, months later, why a submission placed where it did.

Run it on your own field
Book a partner call

EvalLens is currently available through a limited partner program: there is no public sign-up. Tell us what your program reviews and roughly how many decks are in the pile, and access is set up for your team.

Updated

Was this page helpful?