The One-Page Hiring Scorecard: Decisions Without the Committee Fog

If the decision doesn't fit on one page, the process is hiding the signal. The scorecard that forces clarity.

Elena Voss·Head of AI Delivery, Aiporate··7 min read·Share on XLinkedIn

Key takeaways

  • The scorecard has four parts: role outcome, 4-6 weighted criteria, evidence per criterion, and a pre-agreed decision rule. All on one page, written before the first interview.
  • Anchored scoring — concrete descriptions of what each score level looks like — is what separates evidence from vibes-voting.
  • Criteria are capped at six because beyond that, weights flatten and the scorecard becomes a wishlist that no candidate can pass and any candidate can argue into.
  • A scored debrief takes minutes: compare numbers, discuss only the divergences, apply the rule. The fog was never necessary.
  • Overrides are allowed — the scorecard is an instrument, not an oracle — but every override is logged with its reason and audited later. The log is how the scorecard improves.

Sit in enough hiring debriefs and you notice the fog: forty minutes of impressions, someone's 'gut feeling', a strong personality anchoring the room, and a decision that emerges from social dynamics rather than evidence. The fog isn't a people problem — it's a format problem. When there's no agreed structure for the decision, the decision defaults to whoever speaks most confidently. The fix fits on one page: the outcome the role must deliver, four to six weighted criteria, the evidence collected against each, and a decision rule agreed before anyone met the candidate. One page is not a simplification of the decision. It's the claim that if the decision can't be expressed on one page, the process hasn't actually produced a decision — it's produced material for an argument.

The one-pager, part by part

The scorecard is written when the role opens, before a single conversation, because writing it after meeting candidates lets a charming candidate rewrite the bar. Part one: the role outcome — one or two sentences naming what this person must have shipped or changed by a defined horizon. Not responsibilities; outcomes. 'Within two quarters, our LLM features have an evaluation harness the team trusts and a measurable quality trend' is an outcome. Part two: four to six criteria that predict that outcome, each with a weight. Part three: an evidence column — for each criterion, which vetting step supplies the signal, filled in with specifics as the process runs. Part four: the decision rule, agreed in advance: what weighted score means offer, what means no, and what narrow band means one targeted follow-up conversation rather than another full round.

CriterionWeightEvidence sourceScore (1-5)
Has shipped LLM features to production, end to end30%Shipped-work review; walkthrough of a real system
Evaluation-first mindset: can define and defend 'good'25%Past-decision interview: how they measured their last feature
Pragmatic model/cost/latency judgment20%Interview trade-off probes; decisions visible in reviewed work
Works without heavy direction; clarifies scope proactively15%References: 'what happened when the spec was vague'
Communicates trade-offs to non-experts10%Team conversation; clarity of the work walkthrough itself
Example scorecard skeleton for a senior applied-AI engineer

Anchored scoring kills vibes-voting

A criterion with a 1-5 scale and no anchors is a vibe with decimal places. Two interviewers scoring 'strong technical judgment' without anchors are reporting how impressed they felt, and how impressed someone feels tracks confidence, fluency and similarity at least as much as competence. The anchor fixes this by defining the scale in behaviors before anyone is scored. For the evaluation-mindset criterion above: a 2 is 'talks about quality in adjectives; no concrete measurement in any shipped example', a 4 is 'described the actual eval they built for a past feature, including a case it caught that offline metrics missed'. Anchors do their best work at the boundary you care about — the 2-versus-4 line — because that's where offers are won and lost. They also make disagreement productive: two interviewers who diverge on an anchored score are disagreeing about which evidence they saw, which is a resolvable question, not about whose gut is better, which isn't.

  • Write anchors for at least levels 2 and 4 of every criterion — the boundary levels where decisions actually live.
  • Anchors describe observable evidence ('walked through the eval they built'), never adjectives ('impressive depth').
  • Interviewers score independently, before hearing anyone else's numbers. Anchors plus independence is the whole anti-anchoring recipe.
  • Reuse anchors across roles in a family; they get sharper with each miss review.

Why four to six criteria, and not more

The cap is doing real work. Past six criteria, three failure modes arrive together. Weights flatten — with ten criteria nothing can weigh more than trivially, so the scorecard stops expressing what actually matters for the role. Evidence thins — the vetting process has perhaps three or four real signal sources, and ten criteria means most get scored from impression rather than evidence, reimporting the vibes the scorecard was built to exclude. And the wishlist effect appears: every stakeholder's pet requirement gets a row, no candidate clears all ten, so every decision becomes an argument about which failures to forgive — committee fog with a spreadsheet aesthetic. Forcing the cut to four to six is forcing the conversation that matters: of everything we'd like, what does this role's outcome actually require? That argument, had once at role-opening between the hiring manager and their key stakeholders, is cheaper than having it fresh at every debrief with a candidate waiting.

The debrief in minutes, not meetings

With scored cards submitted independently before the debrief, the meeting changes shape entirely. The facilitator reads the weighted totals. Where scores agree, there is nothing to discuss — consensus already happened, on paper, without anyone performing it. Discussion goes only to divergences: 'you scored evaluation mindset a 4, you scored it a 2 — what did you each see?' That conversation is short and concrete because it's about evidence, and it usually surfaces something real: one interviewer probed a follow-up the other didn't, someone's anchor slipped. Then the pre-agreed rule is applied: above the offer line, offer; below the no line, no; in the narrow band between, one targeted follow-up on the specific weak criterion — not another full round. Fifteen minutes, most of the time. The forty-minute fog version wasn't collecting more wisdom; it was performing deliberation while the loudest prior won.

Overriding the scorecard — and logging why

A scorecard that can never be overridden is a different mistake: it pretends the instrument is the reality. Sometimes the totals say no and the hiring manager has a concrete reason to believe the process missed something — evidence outside the criteria, a criterion that this candidate reveals to be badly weighted, information that arrived after scoring. The discipline isn't forbidding the override; it's pricing it. Every override is logged on the scorecard itself: what the rule said, what was decided instead, the specific reason, and who owns the call. Two things follow. First, overrides become deliberate and rare, because 'I just have a feeling' looks exactly as thin in writing as it is. Second, the log becomes the scorecard's own improvement loop: reviewed against 90-day outcomes, it tells you whether your overrides are catching real signal the criteria miss — in which case the criteria should change — or whether they're the old vibes sneaking back in through the exception door, in which case the log says so, in your own handwriting.

  • Overrides are allowed in both directions — hiring below the line and passing above it — but never silently.
  • The log entry names the rule's verdict, the actual decision, the concrete reason, and the owner.
  • Review override outcomes at 90 days alongside the regular miss review.
  • A pattern in the log is a design instruction: recurring override reasons should become criteria; recurring override failures should end that class of override.

Frequently asked questions

Doesn't a scorecard reduce a rich human judgment to a number?

It reduces the aggregation to a number — the judgment stays human. Interviewers still probe, weigh and interpret; the scorecard just forces each judgment to attach to evidence and to be recorded before social dynamics can rewrite it. What it removes isn't nuance, it's the debrief theater where the most confident voice becomes the decision.

Who writes the scorecard, and when?

The hiring manager drafts it when the role opens — outcome first, then criteria and weights — and pressure-tests it with one or two people who know the work. It's locked before the first interview. Written after meeting candidates, it stops being a bar and becomes a rationalization of whoever charmed the room.

What if two strong candidates both clear the decision rule?

Then you have a good problem the scorecard makes easy to reason about: compare the weighted profiles against the role outcome, not overall likability. One candidate's strength usually aligns better with the highest-weighted criterion — and if the choice is genuinely a coin flip on evidence, speed wins: make the offer you can close fastest rather than inventing a new round to break the tie.

How is this different from the scoring rubrics we already use per interview?

Per-interview rubrics structure each conversation; the one-page scorecard structures the decision. It sits above the rounds: every criterion names which round supplies its evidence, and the weighted total plus the pre-agreed rule turns the rounds' outputs into a verdict. Teams often have decent rubrics and still fog the debrief — because nothing forced the rubric outputs into a single decision instrument.

Head of AI Delivery, Aiporate

Elena has spent 12 years building and embedding AI and data teams inside B2B SaaS companies, from first pilot to enterprise-wide platform. At Aiporate she leads how forward-deployed talent is matched, onboarded and shipped to production.

Need the team to make this real?

Describe your need in plain English, get the exact hire, forward-deployed talent or a fractional leader, vetted and matched in 72 hours.

Scope your need →

Keep reading

The Weekly Brief

Intelligence for building AI-native organizations.

One email a week: the sharpest thinking on AI hiring, infrastructure, teams and strategy, for the people building the future of work.

Join operators, founders and CTOs. No spam, unsubscribe anytime.