No provider hands you an ambivalent reference. The two names you get are the two happiest clients in the book, briefed that you might call, and often genuinely fond of the provider, that's why they agreed. Treating this as disqualifying misses the point: curation caps how negative a reference can be, but it doesn't cap how specific one can be, and specificity is where the truth lives. A happy client can't tell you the engagement was bad. They can absolutely tell you, if asked precisely, what went wrong in month three and how long the fix took. Your job on the call isn't to find a hidden critic; it's to extract enough concrete detail from a friendly witness that the pattern shows through anyway.
Accepting the curation, and mining it anyway
The reference call's structural weakness, a friendly, prepared witness, is also its opening. Prepared witnesses have real experience to draw on; they're just defaulting to summary mode ('great partner, really responsive'). Your technique is to refuse summaries and ask for scenes: specific weeks, specific people, specific incidents. Sentiment is curated; episodic memory mostly isn't, because recalling what actually happened in a specific month takes more effort to sanitize than a general impression does. When a reference says 'they were great about replacing an engineer who wasn't working out,' the follow-ups, how many weeks from your first complaint to the new person being productive? who raised it first, you or them?, get answered from memory, factually, and those facts are the product. A reference call that ends with adjectives was a wasted call; one that ends with three dated, specific stories gave you real evidence, whatever the sentiment wrapped around it.
The question set
Run the call in twenty to thirty minutes, promise confidentiality, and ask open questions in the past tense about specifics, past tense retrieves memories, present tense retrieves opinions.
- 1'What went wrong at some point in the engagement, and how did the provider handle it?' Every real engagement has a worst month. A reference who can't name anything hasn't worked with them enough to be useful, note that, too.
- 2'Did you ever have an engineer replaced, or come close? Walk me through it: who flagged it, how long it took, what you paid for during the transition.' This is the highest-value question on the list, replacement behavior under stress is the product you're buying.
- 3'Who exactly did the work, and were they the people you were shown during the sale?' The bait-and-switch question. Also ask whether the good engineers stayed, or rotated off to newer clients once the engagement was stable.
- 4'How did quality and attention change after the first three months?' Providers court new clients and coast on old ones; the reference is an old one, and their answer describes your future.
- 5'Did you expand, shrink, or end the engagement, and why?' Follow with: 'would you expand with them today, on what workstream?' A concrete answer ('yes, we're adding two seats in Q1') is strong signal; a diplomatic one ('we'd certainly consider it') is a soft no from a witness who can't say no.
- 6'What would you tell a friend to watch out for in the first month with them?' The framing licenses honesty, advice to a friend, that a direct 'any weaknesses?' never unlocks.
Listening for pattern vs. anecdote
A single reference's story, good or bad, is an anecdote: it might describe the provider, or it might describe that client's environment, that engineer, that year. What you're listening for across calls is repetition of operational behavior. If both references independently mention that the account manager surfaces problems before the client notices, that's a pattern, and patterns predict. If both mention that invoicing was chaotic, or that the second engineer staffed was noticeably weaker than the first, same. Structure your notes to make this visible: score each call on the same handful of dimensions (staffing honesty, escalation behavior, quality-over-time, replacement handling) rather than keeping free-form impressions, so that when the pilot produces its own evidence, you can lay all three sources side by side and see what repeats. One useful discipline: before the calls, write down the two or three specific worries the sales process left you with, and treat the calls as hypothesis tests on those worries, not as general vibe collection.
| What you hear | Anecdote or pattern? | What to do with it |
|---|---|---|
| One reference had a rough onboarding | Anecdote — could be either side's fault | Note it; probe onboarding explicitly in the pilot |
| Both references say problems were flagged by the provider first | Pattern | Strong positive weight — proactive escalation is rare and valuable |
| Both describe the same engineer or lead by name, fondly | Pattern — but about a person, not the provider | Ask whether that person would be on your engagement; the provider may be one great TL deep |
| 'Would you expand with them?' gets enthusiasm but no specifics | Diplomatic no | Treat as mild negative, not neutral |
| Same complaint appears in a reference call and your pilot | Confirmed pattern | Believe it — this is the strongest evidence a reference process can produce |
Triangulating beyond the given references
The provider chose the two names you got; the more informative references are the ones they didn't choose. You can often reach them anyway, legitimately. Look at the provider's public case studies and past client logos, and check whether anyone in your network, investors, peer CTOs, your team's own LinkedIn graphs, has worked with them; a fifteen-minute back-channel conversation with a client the provider didn't select is worth more than both curated calls combined. Engineering communities are the other underused source: engineers who've worked through a provider talk candidly about how they were treated, paid, and supported, and a provider that treats its engineers poorly delivers you tired, rotating talent no matter how good the client-side account management looks. If a back-channel check contradicts the curated references, weight the back-channel, it has no reason to perform.
Weighting references against pilot evidence
Keep the evidence hierarchy straight. Your own pilot, real work, your stack, your reviewers, is the strongest evidence you'll collect. Back-channel references you sourced yourself come second. The provider's curated references come third: useful, cheap, and worth doing, but structurally the weakest tier, and never a substitute for the tiers above. The practical sequencing follows from that: run reference calls early, when they're cheapest, and use them to generate specific hypotheses, 'watch whether the strong engineer gets rotated off,' 'watch invoicing,' 'watch whether problems get flagged before we notice', that the pilot then tests against reality. References that merely confirm a warm feeling changed nothing; references that told the pilot where to look earned their thirty minutes. And if reference evidence and pilot evidence conflict, the pilot wins, always: what a provider did in someone else's engagement two years ago never outranks what they did in yours last week.
