The Pilot Project: De-Risking a Big Engagement in Two Weeks

No reference call tells you what a two-week paid pilot does. How to structure a pilot that produces a real go/no-go — not a demo.

Elena Voss·Head of AI Delivery, Aiporate··7 min read·Share on XLinkedIn

Key takeaways

  • Pilot scope must be real work from your actual backlog, touching your real stack and data, bounded to two weeks. A sandboxed toy project measures nothing you're buying.
  • Pay for the pilot. Free pilots feel like a win and quietly invert the power dynamic: they select for providers with idle benches, remove your standing to demand priority staffing, and make the exit awkward instead of clean.
  • Fix the evaluation criteria in writing before day one. Criteria written after the pilot always drift toward justifying the impression you already formed.
  • The go/no-go debrief is a scheduled meeting with a scorecard, held within days of pilot end, not a mood that congeals over the following month.
  • If it's a go, convert fast and name the pilot engineers in the SOW, the people who just learned your codebase are the asset the pilot created, and they're only yours if the contract says so.

Everything else in provider evaluation, the deck, the references, the RFP response, measures how well a provider presents. A pilot measures how they work: how their engineers handle your actual codebase, your actual data quirks, your actual standup at 9:30. Two weeks of that evidence outweighs every written claim in the file. But pilots only produce this evidence when they're structured to, and most aren't. A vague pilot produces a polished demo, a warm feeling, and no decision, which is to say, it produces exactly what the provider's marketing already did, at higher cost. Here's the structure that produces a real go/no-go instead.

Scope selection: real, bounded, and touching your stack

The pilot task must sit at the intersection of three properties. Real: pulled from your actual backlog, so that success produces value and, more importantly, so the work exercises the judgment you're actually buying. Bounded: completable by one or two engineers in two weeks, with a crisp definition of done, if the scope is elastic, the pilot measures scoping instead of delivery. And stack-touching: it must run against your repositories, your CI, your (permissioned, scoped) data, because 'how fast do they get productive in our environment' is the single most predictive thing a pilot measures, and a sandbox measures it at zero. Good candidates: a well-understood feature from the backlog, an evaluation harness for an existing AI feature, a bounded integration. Bad candidates: greenfield demos, anything on the critical path of a launch (too much pressure to accept mediocre work), and anything requiring three weeks of access approvals before work can start, do the access paperwork before day one, not during the pilot.

A free pilot sounds like the provider taking the risk. Look closer at what it actually does. First, selection: strong providers with full benches decline free work, so 'free pilot' filters your field for providers whose engineers weren't billing anyway, adverse selection dressed as generosity. Second, staffing: unpaid work gets whoever is idle, not whoever is right, and you learn nothing about the team you'd actually get. Third, leverage: once you've accepted free work, a soft obligation enters the relationship, and a no-go decision now carries the flavor of ingratitude, which is precisely the pressure a clean evaluation can't afford. Paying market rate for two weeks keeps the transaction honest in both directions: you have full standing to demand their genuine A-setup and to walk away without debt, and the provider's willingness to stake real staffing on a two-week evaluation is itself evidence. Two weeks of pilot spend is the cheapest insurance available against a six-month mistake.

Evaluation criteria: fixed before day one

Decide in writing, before the pilot starts, what you'll measure and what score means go. This is the step most teams skip, and skipping it converts the pilot into theater: humans who've spent two weeks with likable engineers will construct criteria that justify liking them. Share the criteria with the provider, none of this is a gotcha, and a provider who performs better because they know code quality and communication are being scored is showing you their reachable best, which is exactly the information you want.

DimensionWhat to look atWeight
Delivery against the agreed scopeWas the definition of done met, at what quality, with how much rework30%
Code and engineering qualityReview the actual diffs: tests, structure, fit with your conventions — not just 'it works'25%
Ramp speed in your environmentDays from access to first meaningful commit; quality of questions asked in week one15%
Communication and collaborationStandup signal, proactive flagging of blockers and risks, async writing quality15%
Provider-side behaviorStaffing as promised, account manager involvement, how the inevitable friction got handled15%
A pilot scorecard that predicts the full engagement

The go/no-go debrief

Within two or three days of pilot end, while evidence is fresh, hold a scheduled debrief with everyone who touched the pilot: the engineers who reviewed the code, the manager who ran the standups, whoever owned the access and onboarding friction. Walk the scorecard dimension by dimension, score independently before discussing (anchoring is real), and require specifics for every score, 'their PRs needed one round of review, here are three examples' beats 'they were great.' Then make the call in the meeting: go, no-go, or, at most once, extend-two-weeks because a specific named uncertainty remains. Guard against the two classic escapes: the drift (no decision, pilot just quietly continues month-to-month, and you've acquired a provider by inertia rather than evaluation) and the split verdict that nobody resolves. If the decision is no-go, say so to the provider with the scorecard as the reason, you paid for the work, the exit is clean, and a professional provider takes the feedback.

Converting a go into the engagement — without losing the people

A successful pilot creates one asset that doesn't appear on any invoice: one or two engineers who now know your codebase, your conventions, and your team. That asset evaporates if conversion is slow or careless. Providers reassign good engineers within days of a pilot ending, benches don't idle, so the moment the debrief says go, move: have the SOW skeleton drafted during the pilot, not after, and issue it within the week. Name the pilot engineers in the SOW with a substitution-consent clause; without named-people language, 'the team that did the pilot' and 'the team that staffs the engagement' can be entirely different people, and the pilot's predictive value drops to roughly that of a reference call. Keep continuity on your side too, same manager, same repo, same rituals, so the engagement inherits the pilot's working rhythm instead of restarting it. And carry the scorecard forward: the pilot's evaluation dimensions make a natural quarterly review template for the engagement itself, which means the accountability the pilot created never actually ends.

Frequently asked questions

Why pay for a pilot when providers offer free ones?

Because free inverts the dynamic against you: it selects for providers with idle benches, staffs the pilot with whoever wasn't billing, and creates a soft obligation that makes a clean no-go socially expensive. Paying market rate preserves your standing to demand real staffing and walk away without debt, that protection is worth far more than two weeks of fees.

How big should a pilot be?

One or two engineers, two weeks, one bounded task from your real backlog with a crisp definition of done. Bigger pilots don't add proportional signal, four weeks tells you little that two doesn't, at twice the cost and delay. What multiplies signal isn't size, it's realness: your stack, your data, your review process.

What if the pilot result is mixed rather than clearly go or no-go?

Diagnose which dimension scored low. Provider-side failures, staffing games, absent account management, weak engineering quality, are structural: no-go. Environment failures, your access took a week, your team was unavailable, are yours: fix them and extend two weeks, once, against a named uncertainty. A second 'mixed' is a no.

How do we keep the pilot engineers for the full engagement?

Speed and contract language. Draft the SOW during the pilot and issue it within a week of the go decision, providers reassign good engineers to other clients within days. Name the individuals in the SOW with a clause requiring your written consent for substitution. Without both, the pilot's predictive value largely evaporates.

Head of AI Delivery, Aiporate

Elena has spent 12 years building and embedding AI and data teams inside B2B SaaS companies, from first pilot to enterprise-wide platform. At Aiporate she leads how forward-deployed talent is matched, onboarded and shipped to production.

Need the team to make this real?

Describe your need in plain English, get the exact hire, forward-deployed talent or a fractional leader, vetted and matched in 72 hours.

Scope your need →

Keep reading

The Weekly Brief

Intelligence for building AI-native organizations.

One email a week: the sharpest thinking on AI hiring, infrastructure, teams and strategy, for the people building the future of work.

Join operators, founders and CTOs. No spam, unsubscribe anytime.