Skip to content
GuideAI Ops
Xither Staff13 min read

AI Ops · Practical guide

Making the AI Business Case: Board Presentations, Portfolio Management, and Realistic Expectations

A defensible AI business case prices the investment from published rates, forecasts benefits from measured evidence rather than vendor decks, and commits to portfolio discipline before the first dollar moves. This guide covers how to build that case, present it to a board, manage a suite of AI investments, and set expectations you can still defend a year later.

In this guide · 7 steps
  1. 01The gap between the trials and the aggregate is your expectations problem
  2. 02Build the cost side first — it is the side you can actually know
  3. 03The structure of a case that survives scrutiny
  4. 04Presenting to the board: what directors actually need
  5. 05Manage a portfolio, not a parade of projects
  6. 06Honest objections
  7. 07The read

A credible AI business case does three things: it prices the investment from published rates rather than guesses, it forecasts benefits from measured evidence rather than vendor promises, and it commits — in writing, up front — to the portfolio discipline that will kill weak projects. The board does not need bigger numbers. It needs numbers you can still defend in twelve months.

That standard is harder than it sounds, because the evidence base for AI value is genuinely split. Controlled trials show large task-level gains. Adoption surveys show a fast-growing but still minority share of firms using AI at all. Economy-wide productivity statistics show, so far, nothing dramatic. A business case that ignores any one of those three tiers is either underselling or overpromising — and boards have seen enough of the latter to discount it on sight.

19.8%

of US businesses reported using AI to produce goods and services as of May 3, 2026, per the Census Bureau's Business Trends and Outlook Survey — with overall usage hovering between 17% and 20% from December 2025 through May 2026.[^census-btos-2026]

U.S. Census Bureau BTOS

15%

average productivity gain — issues resolved per hour — measured across 5,172 customer support agents in the largest published workplace study of a generative AI assistant.[^arxiv-2304-11771]

Brynjolfsson, Li & Raymond

55.8%

faster task completion for developers using GitHub Copilot in a controlled experiment — the high end of the trial evidence, on a narrow, self-contained coding task.[^arxiv-2302-06590]

Peng et al.

+1.4%

US nonfarm business labor productivity growth in the second quarter of 2026 — the economy-wide number that every task-level gain must eventually show up in.[^bls-productivity-2026]

U.S. Bureau of Labor Statistics

1. The gap between the trials and the aggregate is your expectations problem

Three randomized or controlled studies anchor almost every serious claim about generative AI productivity. Brynjolfsson, Li, and Raymond studied the staggered rollout of an AI assistant to 5,172 customer support agents and measured a 15% average increase in issues resolved per hour, with the largest gains going to less experienced agents — the most experienced saw small speed gains and small quality declines.[2] Noy and Zhang assigned incentivized writing tasks to 453 college-educated professionals and found ChatGPT cut average time taken by 40% while output quality rose 18%.[5] Peng and colleagues timed developers implementing an HTTP server and found the Copilot group finished 55.8% faster.[3]

Set those against the aggregate statistics. As of May 2026, 19.8% of US businesses report using AI to produce goods and services[1] — a minority, though a fast-growing one. And US nonfarm labor productivity grew 1.4% in the second quarter of 2026 (manufacturing, 1.9%)[4] — healthy, but nothing resembling a 15-to-55% step change showing up in national accounts.

Evidence tierWhat it measuresHeadline findingWhat it means for your forecast
Controlled trialsOne task, selected workers, weeks15% (support agents)[^arxiv-2304-11771], 40% less time on writing tasks[^pubmed-noy-zhang-2023], 55.8% faster coding[^arxiv-2302-06590]Upper bound on the affected task — not on the job, the team, or the P&L
Firm adoption surveysWhether firms use AI at all19.8% of US businesses; 37% of firms with 250+ employees[^census-btos-2026]Context for 'why now' — most peers are still early, larger firms are moving faster
Aggregate productivityOutput per hour, whole economy+1.4% nonfarm, Q2 2026[^bls-productivity-2026]Realized, diluted, economy-wide effect — the honest floor for skeptical directors
Three tiers of evidence on AI productivity. A defensible business case forecasts somewhere between the floor and the ceiling — and says which assumptions move it.

The gap between these tiers is not a contradiction; it is a composition effect, and understanding it is the core of realistic expectation-setting. A trial measures the treated task in isolation. Your P&L measures the whole job — and the affected task might be 20% of it. Trials study workers who complete a defined task under observation; production involves onboarding, exception handling, review overhead, and the coworkers who never adopt the tool. And aggregate statistics fold in the 80% of firms not using AI at all. Amdahl's-law arithmetic, not pessimism, is why a 40% task-level gain becomes a low-single-digit gain at the business-unit level.

The forecasting rule this implies

Quote the trial numbers as ceilings on the affected task, then discount explicitly: share of the job the task represents, expected adoption rate, review and rework overhead, and ramp time. Show the arithmetic on the slide. A board that watches you discount your own headline number will trust every other number in the deck more.

One more reason to respect the trial evidence rather than round it up: the distribution of gains is not uniform. In the support-agent study, the productivity lift concentrated among less experienced, lower-skilled workers, while the most experienced agents gained little speed and lost a little quality.[2] If your business case assumes your senior staff will get the average lift, it is already wrong. The better model — and the better pitch — is that AI compresses the gap between your newest hires and your best people, which changes hiring, training, and staffing math more than it changes any single productivity line.

2. Build the cost side first — it is the side you can actually know

Most AI business cases fail on the benefit side, but they lose credibility on the cost side, because the cost side is checkable. Model usage is priced publicly, per token, by every major vendor. As of August 2026, Anthropic lists Claude Sonnet 5 at $2 per million input tokens and $10 per million output tokens, Claude Opus 5 at $5 and $25, and Claude Haiku 4.5 at $1 and $5, with a 50% discount for batch processing and cache reads at one-tenth the input rate.[6] OpenAI lists GPT-5 at $1.25 per million input tokens and $10.00 per million output tokens.[7] These are list prices you can put in an appendix and defend.

Worked at realistic volumes, raw inference is often startlingly cheap: Anthropic's own worked example prices 10,000 support-ticket conversations, at roughly 3,700 tokens each on Claude Haiku 4.5, at about $37 total.[6] That number is a trap if you present it as the cost of the initiative — and a gift if you present it correctly: inference is rarely the dominant cost. Integration engineering, data preparation, evaluation harnesses, security review, vendor management, and change management routinely dwarf the token bill. The full cost model — including the costs that only appear at enterprise scale — is covered in depth in /insights/enterprise-ai-tco-guide.

Structure the cost estimate in three layers so the board can see where uncertainty lives. Layer one is metered usage: tokens, hosting, per-seat licenses — knowable to within a factor of two from published pricing and your own volume data. Layer two is delivery: engineering time, data work, evaluation, and integration — estimate it the way you estimate any software project, with ranges. Layer three is the ongoing operating cost people forget to model: monitoring, model migration when a vendor retires a version, prompt and evaluation maintenance, and the human review capacity your risk posture requires. Layer three is where AI differs most from conventional IT projects — it behaves like a product you operate, not a project you finish.

Do not anchor on the pilot's cost

Pilot economics flatter you twice: volumes are small enough that usage costs round to zero, and the pilot team absorbs integration work that production will have to fund explicitly. Price the production state, then show the pilot as a cheap step toward it — not the other way around.

3. The structure of a case that survives scrutiny

With the evidence tiers and the cost layers in hand, the case itself has a standard shape. Start from the business problem, stated in the operational metric the company already tracks — handle time, cycle time, cost per claim, time to first response. A case that has to invent a new metric to show value is a case that will be unmeasurable in production. How to choose and instrument those metrics is its own discipline, covered in /guides/measuring-ai-roi-guide.

Quantify value drivers conservatively and separately. Tangible drivers — labor hours redeployed, error rates reduced, revenue per rep — get a number, a source, and a discount factor. Intangible drivers — decision speed, employee experience, option value on future capability — get named and argued, but never multiplied into the ROI figure. Blending them is the single most common way business cases destroy their own credibility: one challenged assumption in the intangible column and the whole number is suspect.

Then run the standard financial machinery — a multi-year cash flow with net present value, internal rate of return, and payback period — but present it as a sensitivity range, not a point estimate. The honest version has three columns: a floor case where the affected task shrinks less than hoped and adoption is partial; a base case built on discounted trial evidence; and a ceiling case that assumes trial-level gains hold. If the floor case still clears your hurdle rate, you have a fundable project. If only the ceiling case does, you have a science experiment — fund it as one, small and time-boxed, not as a transformation program.

Finally, stage the ask. A request for the full three-year budget on day one forces the board to underwrite your most speculative assumptions. A staged request — pilot funding now, scale funding contingent on named metric thresholds at a named review date — lets the board underwrite only the next tranche of risk, and it commits you to the measurement discipline that makes the next ask easier. The pilot-to-production journey, and the maturity gates worth attaching money to, are laid out in /guides/ai-pilots-and-maturity-guide.

4. Presenting to the board: what directors actually need

A board presentation is not a compressed version of the working deck. Directors are underwriting a decision, not learning a technology, so the deck should be organized around the decision: what you want approved, what it costs, what it returns under which assumptions, what could go wrong, and how you will know — by a specific date — whether it is working. Lead with the ask. Everything else on every slide either supports the ask or does not belong.

Adoption context is the most useful framing data you can bring, because it answers the question every director is silently asking: are we early, late, or reckless? The Census data gives you a defensible answer. As of May 2026, 37% of firms with 250 or more employees report using AI, against a 19.8% national average — and expected use over the following six months sat between 20% and 23% of all businesses.[1] For a large enterprise, that supports a specific message: among firms your size, AI use is approaching a plurality behavior, so the strategic risk is no longer being early — it is scaling without discipline.

AI use by US businesses, May 2026 (share reporting current use)

U.S. Census Bureau, Business Trends and Outlook Survey, May 2026[^census-btos-2026]

On the benefit slide, cite the evidence tiers explicitly — trial ceiling, discounted base case, aggregate floor — rather than a single blended number. This is the move that separates a fundable case from a hyped one, and it also inoculates you: when a director has read that economy-wide productivity grew 1.4% last quarter[4], a slide promising 40% enterprise-wide gains ends the meeting. A slide that says 'the best trial evidence shows 15% on the affected task[2], we are modeling 5% at the business-unit level, here is why' starts a different conversation.

Treat risk as a first-class section, not a compliance appendix. Directors' liability instincts are trained on exactly the failure modes AI introduces: undisclosed model errors reaching customers, data leaving approved boundaries, regulatory exposure, and vendor concentration. For each, name the control and its owner — evaluation gates before release, data-boundary architecture, a designated accountable executive, a documented exit path from each critical vendor. A one-page risk register with owners does more for approval odds than any benefit projection, because it signals the initiative is being run by people who expect to be audited.

The board does not need bigger numbers. It needs numbers you can still defend in twelve months — and a named date on which you will report against them.

Close the presentation by scheduling your own accountability: a specific future meeting at which you will report the named metrics against the thresholds in the staged ask. Then keep the paper trail honest — a one-page summary for the record, and the same metric definitions in every subsequent report. Boards forgive missed forecasts far more readily than they forgive redefined ones.

5. Manage a portfolio, not a parade of projects

By the time an enterprise has more than a handful of AI initiatives, project-level ROI stops being the right unit of analysis. AI investments vary wildly in time-to-value and measurability: a support-deflection deployment can prove itself in a quarter, while a document-intelligence platform may be infrastructure for use cases that do not exist yet. Judged one at a time against a uniform hurdle rate, the portfolio degenerates into whatever demos well. Judged as a portfolio, it can hold workhorses, scaling bets, and options simultaneously — each with expectations appropriate to its lane.

Production workhorses

Deployed systems with measured baselines. Judge on realized ROI against the business case of record, plus operating cost trend. These fund the rest of the portfolio's credibility.

Scaling pilots

Validated in a pilot, now crossing into production. Judge on whether the pilot's measured effect survives contact with real volumes, real users, and real integration costs — not on the pilot's numbers.

Options bets

Small, time-boxed explorations of capabilities that may matter in 18 months. Judge on learning per dollar and kill on schedule, not on ROI — an option that must show ROI stops being cheap.

The not-built register

The explicit list of projects deferred or shelved to fund the AI portfolio. Reviewed alongside the portfolio, so the opportunity cost of AI spending stays a decision, not an accident.

The fourth card deserves emphasis, because it is the one almost no one keeps. Every dollar and every senior engineer assigned to AI is a dollar and an engineer not assigned to platform modernization, technical-debt reduction, or the product roadmap. That opportunity cost never appears in the AI initiative's own accounting — which is exactly why it should appear in the portfolio's. Maintain a register of what was deferred to fund the AI portfolio, and put it on the table at every portfolio review. Sometimes the honest conclusion is that the marginal AI project is worth less than the integration work it displaced; a governance process that cannot reach that conclusion is not governing.

Portfolio metrics need three layers to be useful. The first is conventional financials per initiative — realized savings, revenue impact, payback against plan. The second is AI-specific risk adjustment: model performance drift, retraining and migration costs, evaluation coverage, and dependence on a single vendor's pricing and deprecation schedule. The third is strategic weighting — alignment with where the enterprise is actually going, and the reuse value of shared assets like data pipelines, evaluation infrastructure, and governance tooling that one project builds and five projects use. Shared-asset value is systematically undercounted by project-level ROI, and it is often the strongest real argument for the portfolio's existence.

Then rebalance on a cadence, with teeth. Quarterly is fast enough for a market where model prices and capabilities shift within a fiscal year, and slow enough to see real signal. The discipline that matters is symmetric: scale the initiatives whose floor cases are being beaten, and actually stop the ones that miss their gates. A portfolio that has never killed anything is not a portfolio — it is a backlog with a budget, and boards eventually notice the difference.

6. Honest objections

'Conservative cases lose to whoever promises more.' Sometimes true in the short run — and self-correcting in the medium run, at your expense, because the overpromised case sets the baseline you will be measured against. The stronger play is to be conservative on the numbers and aggressive on the cadence: a modest, staged case that hits its first gate earns scale funding faster than an inflated case that has to explain a miss. Credibility compounds; hype amortizes.

'The trials may understate the long-run gain.' This is the steelman for optimism, and it is legitimate. The headline studies measured existing workflows with an assistant bolted on, over weeks — they could not capture what happens when processes are redesigned around the capability, and the support-agent study itself found the tool accelerated worker learning.[2] Aggregate productivity is also a lagging indicator; earlier general-purpose technologies took years of complementary investment to appear in national statistics. The honest response is not to inflate the base case but to hold the option open: fund workflow-redesign experiments in the options lane, where the ceiling scenario can prove itself without the base case depending on it.

'While you deliberate, adoption is compounding.' Also legitimate — US business adoption roughly doubled across recent survey windows, and larger firms are adopting at nearly twice the national rate.[1] But the conclusion this supports is speed of learning, not size of commitment. Being late to a capability is recoverable in quarters; being three years into an unmeasured, ungoverned portfolio is far more expensive to unwind. Move fast on pilots and instrumentation, and let scale spending follow evidence.

7. The read

The business case, the board deck, and the portfolio review are the same argument at three altitudes, and the argument is about calibration. The trial evidence is real: measured, replicated task-level gains from 15% to over 50% on affected work.[2][3] The aggregate evidence is also real: most firms are still not using AI, and national productivity is growing at ordinary rates.[4] The enterprises that navigate the gap are the ones that forecast between the floor and the ceiling, price from published rates, stage their asks behind measured gates, and keep an honest register of what they chose not to build. That posture wins approvals slower on the first ask and faster on every ask after.

How to apply this

  • State the business problem in a metric the company already tracks — if you have to invent the metric, stop and reread /guides/measuring-ai-roi-guide first
  • Build the cost model in three layers — metered usage from published vendor pricing, delivery, and ongoing operations — and price the production state, not the pilot
  • Forecast benefits as floor / base / ceiling, anchored to the trial evidence as a task-level ceiling, with the discount arithmetic shown
  • Keep intangible benefits out of the ROI number — name them, argue them, never multiply them
  • Stage the funding ask behind named metric thresholds and a named review date
  • Bring adoption context to the board, and lead the deck with the decision you want made
  • Present risk as a one-page register with named controls and named owners
  • Run AI investments as a four-lane portfolio — workhorses, scaling pilots, options, and the not-built register — rebalanced quarterly
  • Track opportunity cost explicitly: review what was deferred to fund AI in the same meeting that reviews AI
  • Kill on schedule: an options bet that misses its time box ends, and the portfolio reports what it stopped as proudly as what it scaled

Sources

Every quantitative or attributed claim above is linked to a primary source. Last verified at publication.

  1. [1]
    AI Use Grew Between December 2025 and May 2026 Across Firm Sizes and Sectors
    U.S. Census Bureau · · accessed
  2. [2]
    Generative AI at Work
    arXiv (Brynjolfsson, Li & Raymond) · · accessed
  3. [3]
    The Impact of AI on Developer Productivity: Evidence from GitHub Copilot
    arXiv (Peng, Kalliamvakou, Cihon & Demirer) · · accessed
  4. [4]
    Productivity — Latest Numbers
    U.S. Bureau of Labor Statistics · accessed
  5. [5]
    Experimental evidence on the productivity effects of generative artificial intelligence
    Science (via PubMed / NCBI) · · accessed
  6. [6]
    Anthropic API pricing
    Anthropic · accessed
  7. [7]
    OpenAI API pricing
    OpenAI · accessed
  8. [8]
    Business Trends and Outlook Survey (BTOS)
    U.S. Census Bureau · accessed
Steps7