How CTOs Should Vet Cloud Computing Service Providers
Author : James Smith | Published On : 02 Sep 2026
Vendor evaluation has quietly become one of the better-run processes in enterprise technology. Ten years ago it was a beauty contest settled by whoever brought the most convincing slide about digital transformation. Today the teams I compare notes with run structured evaluations, test before they sign, and end up with partners who last. The improvement is real, and it is worth writing down what changed.
What follows is the process I use. It is not complicated, but it does require being willing to spend two weeks on diligence that a procurement calendar would rather compress into two days.
Start with the workload, not the vendor deck
The first mistake is running a general evaluation. General evaluations produce general answers, and every provider looks competent in the abstract.
Pick the three workloads that would hurt most if they went badly. The transaction path that touches revenue. The data pipeline that half the company depends on. The legacy application everyone has been avoiding since 2019. Write down, for each, the current latency, the current failure modes, and the current cost. Then run the entire evaluation against those three.
This changes the conversation immediately. Providers who work mostly from templates begin to struggle around the second follow-up question. Providers with real engineering depth get more interested, not less, because you have finally given them something specific to think about.
Five signals that separate an engineering partner from a reseller
Both types will call themselves partners. These are the tells I look for.
They push back on your architecture
A provider who agrees with everything in your current design has either not read it or is not planning to say anything uncomfortable later. I want at least one substantive disagreement during evaluation, ideally one I end up conceding. The willingness to argue in a sales cycle predicts the willingness to argue at 2am, which is when it matters.
They lead with a landing zone, not a migration plan
Landing zone is the term for the account structure, network topology, identity model, and guardrails that everything else gets built on. Providers who bring this up unprompted are thinking about the second year. Providers who lead with a lift-and-shift schedule are thinking about the invoice. Both conversations are necessary, but the order tells you something.
Their FinOps answer contains numbers
Ask how they will manage spend and listen for whether the answer is a process or a tool. The good version sounds like: here is what we would tag, here is the reserved capacity position we would take in month three, here is the anomaly threshold, here is who reviews it weekly. The weak version is a dashboard screenshot.
They raise Day 2 before you do
Migration is a project with an end date, which makes it easy to sell and easy to scope. Running the estate afterwards is neither. If nobody on the provider's side brings up on-call, patching, drift, or the runbook library before you ask, you are buying a migration and hoping operations comes free with it.
They name the individuals
Not the org chart, the people. Who is the lead engineer, what have they shipped, and will they still be on this account in month six. The answer to the last question is often no, which is fine if it is disclosed and staffed for, and corrosive if it is discovered.
Run a paid pilot before a long contract
The single highest-return step in this process is a short, scoped, paid pilot on one real workload. Not a proof of concept in a sandbox. A real workload with real data and a real deadline.
Six weeks and a modest budget will tell you more than any reference call. You learn how they handle a surprise, how quickly they escalate, whether their documentation is written for you or for them, and whether their engineers ask questions that make your engineers think. Providers who resist a paid pilot are telling you something useful for free.
Where AIOps maturity should show up in their answers
Any credible provider now has an AI story. The interesting question is where AI sits in their operating model rather than whether they mention it.
Ask how alerts get correlated today across the accounts they already run. Ask what percentage of incidents close without a human touching them, and how that number moved over the last year. Ask what happens to a novel failure the model has not seen. The answers separate teams that operate an AIOps practice from teams that have bought one.
Gartner expects 40% of organisations deploying AI to adopt dedicated AI observability tooling by 2028, which tells you where the market is heading. Gartner's 2026 Hype Cycle for AI in IT Operations adds a caution worth repeating to any provider making bold claims: in the near term, these tools often add consoles rather than remove them. A provider who acknowledges that tension is more credible than one who does not. Most of the AIOps benefits that matter reduce to two measurable things, the volume of alerts a human has to read and the time between detection and resolution. Ask for both numbers from an account they already run.
The scorecard I hand my team
|
Evaluation area |
Weak answer |
Strong answer |
|---|---|---|
|
Architecture |
Agrees with your current design |
Disagrees on something specific and explains why |
|
Landing zone |
Raised only after you ask |
Presented before the migration schedule |
|
Cost management |
Shows a dashboard |
Names tags, thresholds, and a weekly reviewer |
|
Day 2 operations |
Bundled vaguely into support |
Separate runbooks, on-call model, escalation path |
|
AIOps |
Describes the category |
Quotes alert volume and MTTR from a live account |
|
Staffing |
Presents an org chart |
Names engineers and commits to tenure |
|
Exit |
Not discussed |
Documented data export and handover plan |
Contract terms worth an extra two weeks
- Data export in a documented, non-proprietary format, with a tested procedure rather than a clause that says it is possible.
- Named key personnel with a notice requirement before substitution.
- Service credits that scale with business impact rather than a flat percentage of monthly fees.
- A defined knowledge transfer obligation, with artefacts your team owns at the end.
- Security incident notification measured in hours.
None of these are unusual. All of them are easier to agree before signature than after.
A closing note on the shortlist
Most shortlists of cloud computing service providers contain one obvious name, one cheap name, and one that somebody on the board suggested. That is a fine starting point and a poor finishing point. Run the three workloads through all of them, insist on the pilot, and let the engineering conversation decide.
The right partner is usually identifiable within the first hour of technical discussion, and the cloud computing platform they recommend is rarely the interesting part of the answer. How they intend to run it is.
