Hackathons
The judging table, honestly#
| Figure | What it is |
|---|---|
| 4 min | All the time the MLH organizer guide budgets per project per judge: 2 minutes of demo, 1 for questions and scoring, 1 to walk to the next table |
| 18 judges | What the MLH formula J = ceil(P × n × t / T) demands for 175 projects in a two-hour expo at three rounds each |
| ~5% | Share of projects the average judge actually saw at HackMIT, where 100 judges covered more than 200 projects |
| 70% | Share of a standard Devpost rubric riding on technical execution and innovation, assessed from a demo video nobody is required to watch to the end |
The hackathon panel#
Hackathon mode runs five reviewer roles — Innovation, Technical Execution, Business Value, Pitch Quality and Feasibility — reading every submission independently across six dimensions.
The rubric changes shape rather than just weight:
| Dimension | Weight |
|---|---|
| Execution and Demo | 0.30 (weight-protected) |
| Technical Depth | 0.20 (weight-protected) |
| Problem Impact | 0.15 |
| Innovation / Divergence | 0.15 |
| UX Clarity | 0.10 |
| Delivery Readiness | 0.10 |
Execution and technical depth are protected precisely so that a polished story cannot outrank a working build. Weights are yours to set before the run and lock when it starts.
The six steps#
1 · Your rubric and tracks, locked. Criteria, weights and tracks configured per event, plus a methodology line you can publish in the rules. You get a rulebook judges and sponsors can read before the doors open.
2 · Submissions land on your event page. A public link or QR with a deadline and live statuses, or a manual batch you upload yourself. Completeness is checked automatically, so staff chase exceptions rather than the pile.
3 · The panel does the first read. Every submission scored on execution, technical depth, problem impact, innovation, UX clarity and delivery readiness — the whole field pre-read in hours.
4 · Every judge walks in with a briefing. Per team: scores with the evidence behind them, quotes tagged to the slide they came from, what to verify at the table, and three questions worth the four minutes. Table visits test the build instead of the pitch.
5 · The expo runs exactly as designed. Same tables, same judges, same closing ceremony. Judges score as usual, and where reviewers disagreed the report says so, so deliberation starts at the real argument.
6 · Leaderboard, then feedback for every team. The ranking is built from human Jury Scores and your criteria weights. Structured feedback is drafted from the evidence and approved by your staff before it goes out.
What it reads today, stated plainly#
Today the panel reads the submission you already collect: the deck, the project description and the team's own notes. Nothing changes for participants and no judge loses a role.
Reading a repository and a running demo end to end is the next build on the roadmap, not a claim made today. Anyone quoting execution scores in a closing ceremony should know exactly what those scores were computed from.
The Monday Discord thread#
A team that shipped a working build lost to a team that demoed well, and the thread is public with the sponsor cc'd. Today the honest answer is a shrug, because four minutes at a table is genuinely not a review.
With a record, the reply is one message: the Execution and Demo score and its weight, the finding — two of three feature claims are demonstrated; the third is described, not shown — the quote and slide it came from, the flagged split on Technical Depth, and the organizer's own Jury Score logged next to the AI read.
The disclosure kit#
Hackers notice everything and post about all of it. The risk is never the tool; it is defending the tool without a script.
- Opening ceremony: "Every submission gets a full read under identical rules, and humans decide every placement."
- Rules page: a methodology statement — what the panel assists with, what judges decide, how a team can ask about its own record.
- Submission form: plain language, so nobody discovers AI involvement after results.
- Conflicts of interest: your policy. The record logs who scored what, which is what makes a recusal verifiable.
Tracks and sponsor judges#
Tracks, criteria and weights are configured per event, and each track scores against its own rubric. Judge count does not go down — sponsors and alumni keep the floor and arrive briefed. Weights stay editable right up until the run starts, then lock so the field is scored on one standard end to end.
Data and team IP#
Team submissions are processed only for your event's evaluation and never used to train models — contractual. The event owns the reports, scores and decision log; retention and deletion follow your policy, a DPA is available, and student-data handling is structured to support an institution's obligations. PO and invoice accepted, security questionnaires supported, public sub-processor list, education discount for university programs.
Next steps#
- Criteria and weights — the pitch rubric this one diverges from.
- Brief your jury — the load arithmetic and the calibration video.
- Disagreement and spread — what a flagged split means at the table.