Skip to content
GuideAI Ops
Xither Staff11 min read

AI Ops · Practical guide

AI Pilots and Maturity: Picking the First Project, 90-Day Metrics, and Why Pilots Fail

Run an AI pilot as a purchase of information, not a demo. Pick a workflow with a logged per-unit metric, baseline it before day one, measure against pre-agreed 30/60/90-day gates, and end with an explicit scale, iterate, or kill decision. Most pilot failures are decision failures — no owner, no metric, no gate — not model failures.

In this guide · 10 steps
  1. 01By the numbers: adoption is broad but shallow
  2. 02What a pilot buys you: information for a go/no-go
  3. 03Picking the first project
  4. 04The 90-day measurement plan
  5. 05Why pilots fail: twelve causes, four clusters
  6. 06A working maturity model: ad hoc to transformative
  7. 07Who approves what: the stakeholder map
  8. 08Honest objections
  9. 09The read
  10. 10How to apply this

An AI pilot is a purchase of information: you are paying 90 days of attention to learn whether a system deserves production investment. So run it like one — pick a workflow with a measurable baseline, agree the metrics and the decision rule before day one, and end with an explicit scale, iterate, or kill call. Pilots rarely fail on model quality; they fail on decision discipline.

1. By the numbers: adoption is broad but shallow

18% → 32%

Share of US firms that used AI in a business function during November 2025–January 2026: 18% of firms, rising to 32% on an employment-weighted basis — larger employers adopt at much higher rates.[^census-ces-wp-26-25]

US Census Bureau, BTOS AI supplement working paper

37%

Share of firms with at least 250 employees reporting AI use in business operations, versus overall business AI use hovering between 17% and 20% from December 2025 to May 2026.[^census-btos-firms-2026]

US Census Bureau, Business Trends and Outlook Survey

57%

Share of AI-using firms that deploy AI in three or fewer business functions — even among adopters, use is narrow, which means most organizations are still in pilot-and-early-scale territory.[^census-ces-wp-26-25]

US Census Bureau, BTOS AI supplement working paper

15%

Average productivity gain, measured as issues resolved per hour, for customer-support agents given access to a generative AI assistant in a large field study — with the biggest gains going to less-experienced workers.[^arxiv-2304-11771]

Brynjolfsson, Li & Raymond, Generative AI at Work (arXiv)

Read those four numbers together and the strategic picture is clear. AI adoption in US business is real but shallow: most firms that use AI at all confine it to a handful of functions, and among very large firms in knowledge-intensive sectors — where use rates reach 50%–60% of firms in Information, Professional Services, and Finance — the differentiator is no longer whether you run pilots but whether pilots convert into scaled, instrumented functions.[1] If your organization is still choosing its first serious pilot, you are not behind the market. You are, however, about to make a decision that sets the pattern for every AI investment that follows.

2. What a pilot buys you: information for a go/no-go

The most useful framing available comes from NIST's AI Risk Management Framework. Its MAP function "establishes the context to frame risks related to an AI system," and NIST is explicit about what that context work is for: after completing MAP, "Framework users should have sufficient contextual knowledge about AI system impacts to inform an initial go/no-go decision about whether to design, develop, or deploy an AI system."[4] That is exactly what a pilot is — a structured way to acquire the contextual knowledge that a go/no-go decision needs. If your pilot plan cannot name the decision it will inform, it is not a pilot; it is a demo with a budget line.

This reframing changes what counts as success. A pilot that reveals your data cannot support the use case, or that unit costs at production volume would be prohibitive, has succeeded — it bought you a cheap no before you paid for an expensive one. The distinction between a demo-driven pilot and a decision-driven pilot is the central tension of this entire topic.

DimensionDemo-driven pilotDecision-driven pilot
ObjectiveShow that the technology worksDecide whether to invest in production
Success definitionStakeholders impressed at the readoutPre-agreed metric thresholds met or missed
BaselineNone — comparisons are anecdotalMeasured before day one, on the same metric the pilot reports
End stateExtension, or quiet abandonmentExplicit scale / iterate / kill memo with an owner
DataCurated sample chosen to flatter the modelProduction-representative data, including the ugly parts
Cost modelPilot-period spend onlyProjected unit economics at production volume
The framing choice that predicts pilot outcomes better than any technology choice

About those famous failure-rate statistics

Headline claims that some large share of AI pilots "fail" circulate constantly in vendor decks and press coverage. Most trace back to surveys or preprints with loose definitions of both "pilot" and "failure," and the numbers mutate as they are retold. Do not anchor your business case — for or against piloting — on a secondhand failure rate. Build your own base rate: if every pilot ends in a written scale/iterate/kill memo, twelve months from now you will know your organization's actual conversion rate, which is the only one that matters.

3. Picking the first project

Three filters do most of the work of selecting a first pilot. First, the workflow must already produce a per-unit metric — issues resolved per hour, invoices processed per day, days from claim to decision. Second, the data the system needs must be accessible today, not after a data-platform project. Third, the workflow must tolerate error gracefully: a human should review or absorb the model's mistakes without customer or regulatory harm. A candidate that fails any of the three is a second or third pilot, not a first.

The first filter is why customer support became the early proving ground for generative AI. In the Brynjolfsson, Li, and Raymond field study of 5,172 customer support agents, access to a generative AI assistant raised productivity by 15% on average, measured as issues resolved per hour — a number that exists only because support operations log per-unit throughput natively.[3] The same study found the gains concentrated among less-experienced agents, while the most experienced workers saw small speed gains with small quality declines — a distribution worth remembering when you choose which team pilots first and how you set expectations with your best performers.[3]

It is also worth knowing where your peers actually deploy. Among US firms using AI, the most common business functions are sales and marketing, strategy and business development, and IT.[1] That concentration is informative in both directions: these functions have abundant text-heavy tasks and measurable outputs, which makes them natural first pilots — and it also means vendor tooling and reference patterns there are the most mature, which lowers your integration effort.

Most common business functions for AI, among US firms using AI (Nov 2025–Jan 2026)

US Census Bureau, BTOS AI supplement working paper CES-WP-26-25[^census-ces-wp-26-25]

With candidates that pass the three filters, prioritize on two axes: business impact and implementation effort. The quadrant logic is familiar, but the discipline is in scoring honestly — implementation effort must include data preparation, integration, security review, and change management, not just model work, because those are the line items that pilots systematically underestimate.

QuadrantImpact / effortWhat to do with it
Quick winHigh impact, low effortSelect as the first pilot — fast value demonstration builds the coalition for harder projects
Strategic investmentHigh impact, high effortRoadmap it for after the first pilot proves the operating pattern
Incremental gainLow impact, low effortLet teams pursue as tool adoption; do not spend pilot governance on it
AvoidLow impact, high effortDecline explicitly, in writing, so it does not resurface each budget cycle
Impact/effort prioritization for pilot candidates

Disqualifiers — walk away from a first pilot when

No measurable baseline exists and none can be constructed in two weeks; data access requires approvals that have not started; the output feeds a regulated decision (credit, employment, health) that your review process is not ready to govern; or no business owner will put their name on the go/no-go memo. Any one of these converts a 90-day pilot into a 9-month stall.

4. The 90-day measurement plan

Ninety days is long enough to observe a metric trend and short enough to keep attention. But direct financial ROI rarely materializes inside the window — the honest posture is to measure leading indicators now and project financial returns from them, explicitly, with the assumptions written down. For the full treatment of how pilot indicators roll up into a defensible ROI model, see /guides/measuring-ai-roi-guide; this section covers what to instrument inside the pilot itself.

  • Operational metrics — the per-unit throughput or quality measure the workflow already logs: cycle time, error or rework rate, units handled per person-hour. These are your primary evidence because they move within 90 days.
  • Adoption metrics — share of eligible users active weekly, share of AI outputs accepted versus overridden, time to first productive use. A pilot with strong operational numbers and collapsing adoption is telling you the scale-up will fail.
  • Cost and effort metrics — actual spend per unit of work (inference, licenses, review labor) and the integration hours consumed. These feed the production unit-economics projection, which is where scale decisions are won or lost — the full framework lives at /insights/enterprise-ai-tco-guide.

Baseline before day one, or the pilot is unmeasurable. Capture at least four weeks of the primary metric under current-state operations, from the same logs the pilot will report against. Then run stage gates at 30, 60, and 90 days with pre-assigned questions — the point of interim gates is to kill or fix early, while sunk costs are small.

GateQuestion it answersKill / fix trigger
Day 30Is the system actually wired in — data flowing, users onboarded, outputs landing in the real workflow?Integration or data-access work still incomplete: stop the clock, fix, restart — do not let calendar time substitute for run time
Day 60Is the primary metric trending against baseline, and are users still using it?Flat metric with healthy adoption → iterate the workflow; healthy metric with collapsing adoption → treat as a change-management failure, not a model failure
Day 90Scale, iterate, or kill — written memo with projected unit economics and named ownerNo one willing to own the scale recommendation is itself a kill signal
Stage gates for a 90-day pilot

Attribution deserves more care than it usually gets. If the pilot runs while a process change, a seasonal peak, or a reorg is also in flight, the metric movement is contaminated. Where volume allows, hold out a comparison group — a queue, region, or team that keeps current-state operations — and compare trends rather than absolute levels. Where it does not, at minimum document the concurrent changes in the decision memo so the readout is honest about confidence. And police scope: every mid-pilot feature addition resets your ability to say what caused the result.

A pilot that cannot fail was never a pilot. It was a slow-motion purchase order.

5. Why pilots fail: twelve causes, four clusters

Catalogs of pilot failure causes run long — unclear objectives, weak sponsorship, dirty data, over-scoping, talent gaps, thin budgets, no change plan, undefined metrics, vendor dependence, ignored compliance, unscalable architecture, no iteration loop. Useful as a checklist, but the twelve collapse into four root clusters, and each cluster has one owner who can prevent it.

Decision debt

No named business owner, no pre-agreed metric, no go/no-go rule. The pilot ends in a shrug and an extension. Owner: the executive sponsor, who should refuse to fund any pilot without a one-page charter naming the decision, the metric, and the decider.

Data and integration reality

The demo ran on curated data; production data is fragmented, stale, or locked behind approvals, and the output never lands inside the actual workflow. Owner: the platform lead, via a two-week data-access and integration spike before the pilot clock starts.

The adoption gap

The system works and nobody uses it — training was an email, incentives punish the new workflow, and skeptical veterans opt out. Owner: the business-line leader, with a real change plan rather than a launch announcement.

Scale economics

The pilot succeeds at 50 users and dies in the business case at 5,000 — unit costs, security review, and infrastructure were never projected. Owner: whoever writes the day-90 memo, which must include production unit economics, not pilot-period spend.

The adoption gap deserves particular respect because the national data says augmentation is the dominant mode: 66% of AI-using firms rely on AI solely to augment tasks, and AI-related employment decreases occur in only 2% of firms.[1] Augmentation means your return arrives only through changed human behavior — the model can be flawless and the pilot still fails if the humans route around it. That is a change-management problem with a known playbook; we cover it in depth at /guides/ai-change-management-adoption, and your pilot plan should budget for it from day one rather than discovering it at day 60.

6. A working maturity model: ad hoc to transformative

Maturity models attract deserved skepticism, so let us be precise about what this one is: a sequencing tool of our own construction, not an industry standard or a certification ladder. Its only job is to answer one question — given where the organization actually is, what is the next right move? Used that way, four stages are enough.

StageWhat AI looks like hereCharacteristic failure modeThe next right move
1 · Ad hocIndividual teams experiment; tools adopted bottom-up; no shared metrics or governancePilot sprawl — many demos, no decisions, no institutional learningRun one decision-driven pilot with a charter, a baseline, and a day-90 memo
2 · RepeatableA pilot playbook exists; charters, stage gates, and a stakeholder map are standard; 2–3 pilots run as a portfolioPilots succeed but stall before production — no platform, no MLOps, no budget path to scaleBuild the scaling path: shared platform services, security review patterns, unit-cost models
3 · IntegratedAI is embedded in core processes with monitored KPIs, feedback loops, and clear ownershipPortfolio drifts — models degrade unmonitored, costs creep, governance lags new use casesInstrument the portfolio: model monitoring, cost observability, periodic re-validation
4 · TransformativeAI reshapes products and operating models; new capabilities are designed AI-first with governance built inConcentration risk — deep dependence on a vendor, platform, or a few key peopleManage AI as strategic infrastructure: diversification, resilience, succession
A four-stage maturity model built for sequencing decisions, not benchmarking vanity

Two honest observations about using it. First, the national picture suggests most adopters sit at stages one and two: when 57% of AI-using firms confine AI to three or fewer business functions and 65% limit worker use to three or fewer tasks, integrated enterprise-wide AI is the exception, not the norm.[1] Second, stages are per-domain, not per-company — your support organization can be at stage three while finance is at stage one, and pretending the enterprise has a single maturity score obscures exactly the sequencing information the model exists to provide.

7. Who approves what: the stakeholder map

More pilot calendar time is lost to unsequenced approvals than to model problems. AI initiatives cut across data, infrastructure, legal, risk, and the business, and each constituency holds a genuine veto over some part of the work. Map them before kickoff, not when the first blocker surfaces.

  • Executive sponsor — approves scope, funding, and the strategic fit; signs the go/no-go memo.
  • Data owners — approve data access and use; the single most common source of hidden schedule risk, so start this approval first.
  • IT and platform teams — approve infrastructure, security posture, and integration points into production systems.
  • Legal and compliance — approve regulatory exposure, vendor contract terms, and data-protection obligations.
  • Risk and model validation — approve model performance evidence, bias testing, and human-oversight design for consequential decisions.
  • Operational leads and end users — approve workflow fit; without their sign-off, expect the adoption gap to do its work.

Put the map in a RACI matrix in the pilot charter, run the slow approvals (data access, security review) in parallel with pilot setup, and keep the same matrix alive for the scale decision — the approvers who felt consulted during the pilot are the ones who move quickly when you ask for production sign-off. NIST's MAP guidance points the same direction: incorporating perspectives from a diverse internal team and from actors outside the deploying team is what makes the context-mapping — and therefore the go/no-go — trustworthy.[4]

8. Honest objections

"Pilots are theater — just deploy and iterate." For low-risk assistive tooling, this is substantially right: a coding assistant or a meeting summarizer does not need a 90-day charter, and over-governing cheap experiments is its own failure mode. The pilot discipline in this guide is for systems that embed in workflows, touch customer or regulated outcomes, or carry production-scale cost commitments — where the information a pilot buys is genuinely expensive to acquire any other way. Match the governance weight to the decision weight.

"Ninety days is too short to prove ROI." Correct — which is an argument for honest metric design, not longer pilots. Leading indicators move in 90 days; dollars follow over quarters. The failure is claiming dollar ROI a pilot cannot show, then losing credibility when finance audits it. Project financial returns from measured operational deltas, show the assumptions, and let the projection be challenged — that discipline is exactly what separates a fundable scale case from an optimistic slide.

"Maturity models are consultant-ware." Often true: a maturity model used to produce a benchmark score and a transformation program is decoration. The defensible use is narrower — it stops two specific, expensive errors: attempting a stage-three deployment on stage-one foundations (no governance, no data readiness), and buying stage-four platform commitments while the organization is still learning to run one pilot well. If the model is not changing your next investment decision, discard it.

9. The read

Treat pilots as decision instruments and maturity as sequencing, and the operating pattern follows. Run a small portfolio — two or three decision-driven pilots, not one bet-the-year project and not a dozen ungoverned demos. Pre-commit the gates and the decision rule so the day-90 memo writes itself. Fund scaling only where three tests pass together: the operational metric moved against a clean baseline, adoption held without executive pressure, and projected production unit economics survive contact with the TCO model. And bank the losses properly — a killed pilot with a written memo is reusable organizational knowledge; a quietly abandoned one is pure cost.

The prize for getting this right is not the first pilot's ROI, which is usually modest. It is the repeatable capability — charter, baseline, gates, memo — that turns every subsequent AI investment decision from a faith-based debate into an evidence-based one. That capability is what moves an organization from stage one to stage two, and it costs far less than the pilots it will save you from.

10. How to apply this

The 90-day pilot discipline

  • Name the decision the pilot informs, the metric it will be judged on, and the person who makes the go/no-go call — in a one-page charter, before any build work.
  • Filter candidates on three gates: a logged per-unit metric, data accessible today, and graceful tolerance for model error.
  • Score the surviving candidates on impact versus honest effort — including data prep, integration, security review, and change management.
  • Run a two-week data-access and integration spike before starting the pilot clock.
  • Capture at least four weeks of baseline on the primary metric from the same logs the pilot will report against.
  • Instrument all three metric families: operational, adoption, and cost per unit of work.
  • Hold a comparison group where volume allows; document concurrent business changes where it does not.
  • Run 30/60/90 stage gates with pre-assigned questions and real kill triggers.
  • Map stakeholders into a RACI before kickoff and start the slowest approvals (data, security) first.
  • Close with a written scale/iterate/kill memo that includes projected production unit economics — and file it where the next pilot team will find it.

Sources

Every quantitative or attributed claim above is linked to a primary source. Last verified at publication.

  1. [1]
    The Microstructure of AI Diffusion: Evidence from Firms, Business Functions, and Worker Tasks (CES-WP-26-25)
    US Census Bureau, Center for Economic Studies · · accessed
  2. [2]
    Large Firms With at Least 20 Employees Biggest AI Users
    US Census Bureau · · accessed
  3. [3]
    Generative AI at Work
    arXiv (Brynjolfsson, Li & Raymond) · accessed
  4. [4]
Steps10