Overview

Read a report

Layer 1 · Project Summary#

The fast read: what the project does, how it scored, where it looks strong, and what to verify live. This is the layer that replaces the first pass through the raw deck — you start with the report, not the file.

Layer 2 · AI Score Report#

Six things to look at, roughly in this order.

Per-dimension breakdown. For each dimension: the score, the confidence signal, what supports it, what lowers it, and what would change it. A dimension with a strong score and a thin "what supports it" list is a dimension to check.

Judge contribution matrix. Who contributed to each dimension and how much. Not every judge influences every dimension — a Team score is driven by the Team Readiness lens, with others contributing secondary or advisory reads. The matrix is also where strong disagreements are flagged.

Score formation. Each dimension score, its weight, and its contribution to the total. This is the line to show anyone who asks how a number was arrived at.

Methodology. The scale and scoring rules, applied identically across every deck in the batch.

Initial criteria. Your weights, read-only, and identical for every team. They are shown inside the report so that a team reading it can see the standard it was judged against.

Judge conclusions. From each judge: a takeaway, a main concern, and one live question.

Layer 3 · Questions for live Q&A#

Ready-to-use questions ranked by priority, each linked to the dimension it tests. The point is to see where a deck is thin before the team is in the room, so the minutes on stage go to the gap rather than to a recap of the deck.

Evidence: how to check a claim#

Every finding cites the exact slide — number, title and note — so a claim reads as an observation you can open, not an opinion you have to accept. The scoring procedure that produces this is fixed:

  1. Cite the evidence — slide-grounded facts only.
  2. Weigh it both ways — what supports the score, what lowers it, what the deck leaves unproven.
  3. Name the band the evidence falls into, and why.
  4. Then the score, inside that band. On a boundary with evidence missing, the lower band.

A worked example from the published methodology, on P3 Market: evidence on slides 6 and 8, a strength of a clear target segment, a weakness of an unsourced TAM, missing buyer validation, confidence medium — which lands the score in the band rather than above it.

Deck completeness#

The report checks ten core sections — Problem, Solution, Market, Business Model, Traction, Team, Roadmap, Financials, Ask, Other — and marks each present, thin or missing, with a severity of info, warning or critical and the dimension it affects.

Completeness is not a fact-check. Missing means the deck did not cover it. It does not mean a claim is false, and it does not validate the claims that are present. That remains a human judgment — see What EvalLens does not do.

Reading confidence and spread together#

Two signals sit beside the number and mean different things:

  • Confidence describes how well the evidence supports this judge's read. Low confidence can apply a downward adjustment of at most 15%, so a thin read does not present as a firm number.
  • Spread describes how much the judges differed. It never changes the score. Under 1.5 is consensus, 1.5 to 2.99 is a split, 3.0 or more is a conflict flagged for human review.

A high score with high spread is not a bad score — it is a score that needs a conversation.

The same report, five times over#

One report serves the whole round: reviewer prep before review, a common basis during shortlisting, structured feedback to teams afterwards, an evidence-linked basis for the committee, and the archive that answers "why did this team place here" six months later.

Next steps#

Updated

Was this page helpful?