Sit in enough hiring debriefs and you notice the fog: forty minutes of impressions, someone's 'gut feeling', a strong personality anchoring the room, and a decision that emerges from social dynamics rather than evidence. The fog isn't a people problem — it's a format problem. When there's no agreed structure for the decision, the decision defaults to whoever speaks most confidently. The fix fits on one page: the outcome the role must deliver, four to six weighted criteria, the evidence collected against each, and a decision rule agreed before anyone met the candidate. One page is not a simplification of the decision. It's the claim that if the decision can't be expressed on one page, the process hasn't actually produced a decision — it's produced material for an argument.
The one-pager, part by part
The scorecard is written when the role opens, before a single conversation, because writing it after meeting candidates lets a charming candidate rewrite the bar. Part one: the role outcome — one or two sentences naming what this person must have shipped or changed by a defined horizon. Not responsibilities; outcomes. 'Within two quarters, our LLM features have an evaluation harness the team trusts and a measurable quality trend' is an outcome. Part two: four to six criteria that predict that outcome, each with a weight. Part three: an evidence column — for each criterion, which vetting step supplies the signal, filled in with specifics as the process runs. Part four: the decision rule, agreed in advance: what weighted score means offer, what means no, and what narrow band means one targeted follow-up conversation rather than another full round.
| Criterion | Weight | Evidence source | Score (1-5) |
|---|---|---|---|
| Has shipped LLM features to production, end to end | 30% | Shipped-work review; walkthrough of a real system | — |
| Evaluation-first mindset: can define and defend 'good' | 25% | Past-decision interview: how they measured their last feature | — |
| Pragmatic model/cost/latency judgment | 20% | Interview trade-off probes; decisions visible in reviewed work | — |
| Works without heavy direction; clarifies scope proactively | 15% | References: 'what happened when the spec was vague' | — |
| Communicates trade-offs to non-experts | 10% | Team conversation; clarity of the work walkthrough itself | — |
Anchored scoring kills vibes-voting
A criterion with a 1-5 scale and no anchors is a vibe with decimal places. Two interviewers scoring 'strong technical judgment' without anchors are reporting how impressed they felt, and how impressed someone feels tracks confidence, fluency and similarity at least as much as competence. The anchor fixes this by defining the scale in behaviors before anyone is scored. For the evaluation-mindset criterion above: a 2 is 'talks about quality in adjectives; no concrete measurement in any shipped example', a 4 is 'described the actual eval they built for a past feature, including a case it caught that offline metrics missed'. Anchors do their best work at the boundary you care about — the 2-versus-4 line — because that's where offers are won and lost. They also make disagreement productive: two interviewers who diverge on an anchored score are disagreeing about which evidence they saw, which is a resolvable question, not about whose gut is better, which isn't.
- Write anchors for at least levels 2 and 4 of every criterion — the boundary levels where decisions actually live.
- Anchors describe observable evidence ('walked through the eval they built'), never adjectives ('impressive depth').
- Interviewers score independently, before hearing anyone else's numbers. Anchors plus independence is the whole anti-anchoring recipe.
- Reuse anchors across roles in a family; they get sharper with each miss review.
Why four to six criteria, and not more
The cap is doing real work. Past six criteria, three failure modes arrive together. Weights flatten — with ten criteria nothing can weigh more than trivially, so the scorecard stops expressing what actually matters for the role. Evidence thins — the vetting process has perhaps three or four real signal sources, and ten criteria means most get scored from impression rather than evidence, reimporting the vibes the scorecard was built to exclude. And the wishlist effect appears: every stakeholder's pet requirement gets a row, no candidate clears all ten, so every decision becomes an argument about which failures to forgive — committee fog with a spreadsheet aesthetic. Forcing the cut to four to six is forcing the conversation that matters: of everything we'd like, what does this role's outcome actually require? That argument, had once at role-opening between the hiring manager and their key stakeholders, is cheaper than having it fresh at every debrief with a candidate waiting.
The debrief in minutes, not meetings
With scored cards submitted independently before the debrief, the meeting changes shape entirely. The facilitator reads the weighted totals. Where scores agree, there is nothing to discuss — consensus already happened, on paper, without anyone performing it. Discussion goes only to divergences: 'you scored evaluation mindset a 4, you scored it a 2 — what did you each see?' That conversation is short and concrete because it's about evidence, and it usually surfaces something real: one interviewer probed a follow-up the other didn't, someone's anchor slipped. Then the pre-agreed rule is applied: above the offer line, offer; below the no line, no; in the narrow band between, one targeted follow-up on the specific weak criterion — not another full round. Fifteen minutes, most of the time. The forty-minute fog version wasn't collecting more wisdom; it was performing deliberation while the loudest prior won.
Overriding the scorecard — and logging why
A scorecard that can never be overridden is a different mistake: it pretends the instrument is the reality. Sometimes the totals say no and the hiring manager has a concrete reason to believe the process missed something — evidence outside the criteria, a criterion that this candidate reveals to be badly weighted, information that arrived after scoring. The discipline isn't forbidding the override; it's pricing it. Every override is logged on the scorecard itself: what the rule said, what was decided instead, the specific reason, and who owns the call. Two things follow. First, overrides become deliberate and rare, because 'I just have a feeling' looks exactly as thin in writing as it is. Second, the log becomes the scorecard's own improvement loop: reviewed against 90-day outcomes, it tells you whether your overrides are catching real signal the criteria miss — in which case the criteria should change — or whether they're the old vibes sneaking back in through the exception door, in which case the log says so, in your own handwriting.
- Overrides are allowed in both directions — hiring below the line and passing above it — but never silently.
- The log entry names the rule's verdict, the actual decision, the concrete reason, and the owner.
- Review override outcomes at 90 days alongside the regular miss review.
- A pattern in the log is a design instruction: recurring override reasons should become criteria; recurring override failures should end that class of override.
