Everything else in provider evaluation, the deck, the references, the RFP response, measures how well a provider presents. A pilot measures how they work: how their engineers handle your actual codebase, your actual data quirks, your actual standup at 9:30. Two weeks of that evidence outweighs every written claim in the file. But pilots only produce this evidence when they're structured to, and most aren't. A vague pilot produces a polished demo, a warm feeling, and no decision, which is to say, it produces exactly what the provider's marketing already did, at higher cost. Here's the structure that produces a real go/no-go instead.
Scope selection: real, bounded, and touching your stack
The pilot task must sit at the intersection of three properties. Real: pulled from your actual backlog, so that success produces value and, more importantly, so the work exercises the judgment you're actually buying. Bounded: completable by one or two engineers in two weeks, with a crisp definition of done, if the scope is elastic, the pilot measures scoping instead of delivery. And stack-touching: it must run against your repositories, your CI, your (permissioned, scoped) data, because 'how fast do they get productive in our environment' is the single most predictive thing a pilot measures, and a sandbox measures it at zero. Good candidates: a well-understood feature from the backlog, an evaluation harness for an existing AI feature, a bounded integration. Bad candidates: greenfield demos, anything on the critical path of a launch (too much pressure to accept mediocre work), and anything requiring three weeks of access approvals before work can start, do the access paperwork before day one, not during the pilot.
Paid, not free — and why that protects you
A free pilot sounds like the provider taking the risk. Look closer at what it actually does. First, selection: strong providers with full benches decline free work, so 'free pilot' filters your field for providers whose engineers weren't billing anyway, adverse selection dressed as generosity. Second, staffing: unpaid work gets whoever is idle, not whoever is right, and you learn nothing about the team you'd actually get. Third, leverage: once you've accepted free work, a soft obligation enters the relationship, and a no-go decision now carries the flavor of ingratitude, which is precisely the pressure a clean evaluation can't afford. Paying market rate for two weeks keeps the transaction honest in both directions: you have full standing to demand their genuine A-setup and to walk away without debt, and the provider's willingness to stake real staffing on a two-week evaluation is itself evidence. Two weeks of pilot spend is the cheapest insurance available against a six-month mistake.
Evaluation criteria: fixed before day one
Decide in writing, before the pilot starts, what you'll measure and what score means go. This is the step most teams skip, and skipping it converts the pilot into theater: humans who've spent two weeks with likable engineers will construct criteria that justify liking them. Share the criteria with the provider, none of this is a gotcha, and a provider who performs better because they know code quality and communication are being scored is showing you their reachable best, which is exactly the information you want.
| Dimension | What to look at | Weight |
|---|---|---|
| Delivery against the agreed scope | Was the definition of done met, at what quality, with how much rework | 30% |
| Code and engineering quality | Review the actual diffs: tests, structure, fit with your conventions — not just 'it works' | 25% |
| Ramp speed in your environment | Days from access to first meaningful commit; quality of questions asked in week one | 15% |
| Communication and collaboration | Standup signal, proactive flagging of blockers and risks, async writing quality | 15% |
| Provider-side behavior | Staffing as promised, account manager involvement, how the inevitable friction got handled | 15% |
The go/no-go debrief
Within two or three days of pilot end, while evidence is fresh, hold a scheduled debrief with everyone who touched the pilot: the engineers who reviewed the code, the manager who ran the standups, whoever owned the access and onboarding friction. Walk the scorecard dimension by dimension, score independently before discussing (anchoring is real), and require specifics for every score, 'their PRs needed one round of review, here are three examples' beats 'they were great.' Then make the call in the meeting: go, no-go, or, at most once, extend-two-weeks because a specific named uncertainty remains. Guard against the two classic escapes: the drift (no decision, pilot just quietly continues month-to-month, and you've acquired a provider by inertia rather than evaluation) and the split verdict that nobody resolves. If the decision is no-go, say so to the provider with the scorecard as the reason, you paid for the work, the exit is clean, and a professional provider takes the feedback.
Converting a go into the engagement — without losing the people
A successful pilot creates one asset that doesn't appear on any invoice: one or two engineers who now know your codebase, your conventions, and your team. That asset evaporates if conversion is slow or careless. Providers reassign good engineers within days of a pilot ending, benches don't idle, so the moment the debrief says go, move: have the SOW skeleton drafted during the pilot, not after, and issue it within the week. Name the pilot engineers in the SOW with a substitution-consent clause; without named-people language, 'the team that did the pilot' and 'the team that staffs the engagement' can be entirely different people, and the pilot's predictive value drops to roughly that of a reference call. Keep continuity on your side too, same manager, same repo, same rituals, so the engagement inherits the pilot's working rhythm instead of restarting it. And carry the scorecard forward: the pilot's evaluation dimensions make a natural quarterly review template for the engagement itself, which means the accountability the pilot created never actually ends.
