Skip to content
Use CaseBusiness Functions
Xither Staff14 min read

Business Functions · Use-case guide

AI for Sales Teams: SDR Agents, Conversation and Revenue Intelligence, Forecasting, and Proposals

Sales and marketing is the business function where firms deploy AI most often — and the one where vendor claims outrun public evidence by the widest margin. This guide maps the five sales-AI workloads, states what verifiable evidence supports each, and gives a sequencing plan: adopt assistive coaching and drafting first, pilot scoring and forecasting against your own baseline, and treat autonomous outreach as an experiment with kill criteria.

52%

Share of AI-using firms that deploy AI in sales and marketing — the most common business function, ahead of strategy and business development (45%) and IT (41%), per the Census Bureau's 2026 BTOS AI supplement[^census-ces-wp-26-25].

U.S. Census Bureau

15%

Average productivity lift (issues resolved per hour) when 5,172 customer-support agents got a generative-AI conversation assistant — the most rigorous published evidence for AI-assisted selling-adjacent conversations, with the largest gains going to the least experienced agents[^arxiv-2304-11771].

Brynjolfsson, Li & Raymond (RCT)

$25M+

Consumer losses alleged in a single deceptive "AI-powered" storefront scheme charged in the FTC's Operation AI Comply sweep — a marker of how directly AI sales claims now draw enforcement attention[^ftc-ai-comply-2024].

Federal Trade Commission

Sales is where enterprise AI budgets are actually landing: among firms using AI, sales and marketing is the most common business function for deployment, at 52%[1]. It is also the category with the loudest unverifiable claims — invented reply rates, "forecast accuracy" percentages with no methodology, and autonomous-rep demos that skip the compliance question. This guide separates the five sales-AI workloads a revenue leader can buy today, weighs the evidence for each, and turns the vendor noise into an evaluation you can actually run.

By the numbers

Read those three numbers together and the shape of the market appears. Adoption is broad but shallow: 57% of AI-using firms deploy it in three or fewer business functions[1], and overall business AI use hovered between 17% and 20% from December 2025 to May 2026[4]. The strongest controlled evidence comes from AI assisting human conversations, not replacing them. And the regulator has already shown up — for deceptive AI capability claims and for AI-fabricated social proof alike. Your adoption sequence should follow the evidence, not the demo.

The five workloads, and where the evidence actually stands

"AI for sales" is not one purchase. It is five distinct workloads with different data dependencies, different failure modes, and very different evidence bars. Conflating them is how teams end up buying an autonomous outreach agent when what their pipeline problem needed was forecast hygiene.

WorkloadWhat it automatesEvidence bar todayPrimary risk
SDR agents / automated outreachProspect research, enrichment, first-touch sequencing, initial qualificationVendor claims only; no public controlled studiesBrand and deliverability damage at scale; deceptive-claim exposure[^ftc-ai-comply-2024]
Conversation intelligenceCall recording, transcription, talk-pattern analysis, coachingStrong analog: RCT showing a 15% average productivity lift from AI conversation assistance in support[^arxiv-2304-11771]Recording-consent and data-governance obligations[^govinfo-18usc2511-2023]
Revenue intelligence & deal scoringActivity capture, win-probability scoring, competitor-mention and win/loss analysisSound statistical footing, but validity depends entirely on your own label and activity data qualityCorrelation mistaken for causation; reps gaming the inputs
AI forecastingPipeline roll-ups, risk-adjusted forecasts, scenario viewsNo verifiable public accuracy comparisons; must be backtested in-houseAdopting a model you never benchmarked against your human baseline
Proposal / RFP / SOW generationFirst drafts from approved content, question parsing, clause suggestionSame mechanism as general LLM drafting; well understoodConfabulation — confidently stated false content — reaching a customer document[^nist-ai-600-1]
The five sales-AI workloads. "Evidence bar" reflects publicly verifiable evidence, not vendor-reported outcomes.

SDR agents: automating prospect research and first-touch outreach

The SDR-agent pitch is straightforward: the top of the funnel is high-volume, templated work — account research, contact enrichment, personalized first-touch email, meeting booking — so let an agent do it. A crowded market has formed around that pitch, from dedicated AI-SDR startups such as 11x, Artisan, and Regie.ai to the sales-engagement incumbents adding agentic features. Naming them is where the reliable public information ends: none of these vendors publishes controlled, third-party-verifiable performance data, so any comparison built on their reported reply rates or "pipeline generated" figures is a comparison of marketing copy.

That does not mean the category is unbuyable. It means the evaluation has to be yours. Treat an SDR-agent selection like a model evaluation, not a feature checklist: run two or three candidates head-to-head on your own ideal customer profile for a fixed window, against a holdout worked by your human SDRs. Measure qualified meetings held and downstream opportunity conversion — never raw replies, which agents can inflate with volume. Instrument deliverability from day one: sender-domain reputation is a shared asset, and an agent that burns it has destroyed something that takes quarters to rebuild. And read the agent's actual output weekly, because the failure mode that kills these programs is not low volume — it is a plausible-sounding email that misstates what your product does, sent ten thousand times.

The compliance dimension is no longer hypothetical. In September 2024 the FTC announced Operation AI Comply, a five-case enforcement sweep against companies "that rely on artificial intelligence as a way to supercharge deceptive or unfair conduct"[3]. Two of the cases — Ascend Ecom, with alleged consumer losses over $25 million, and Ecommerce Empire Builders, which charged up to $35,000 for "AI-powered" storefronts — were built on deceptive AI business-opportunity claims[3]. A third, against the AI writing service Rytr, targeted a tool whose review-generation feature let subscribers mass-produce fake consumer reviews with fabricated details; the proposed order bars the company from selling any service dedicated to generating reviews[3].

The lesson cuts both ways for a sales organization. As a buyer, it tells you the regulator has found AI capability claims deceptive enough to prosecute — which is exactly the skepticism to bring to an AI-SDR vendor's own numbers. As an operator, it tells you that AI-generated social proof is a legal exposure, not a growth hack: the FTC's final rule banning fake reviews and testimonials, announced in August 2024, explicitly covers reviews and testimonials attributed to someone who does not exist, including AI-generated fake reviews, and strengthens the agency's ability to seek civil penalties[7]. An outreach agent that invents a customer quote or a fictitious happy user to warm up a cold email is generating enforcement risk under that rule.

Guardrail before the first send

Put a hard content policy between any outreach agent and your sending domain: no claims about product capability that are not in an approved claims library, no invented testimonials or named references, no fabricated personalization details. The FTC has already acted against AI-generated fake reviews and deceptive AI claims[3] — and separately, everything the agent sends is your brand's writing sample at scale.

Conversation intelligence: the workload with real evidence behind it

Conversation-intelligence platforms record and transcribe sales calls, analyze talk patterns and topics, and turn the corpus into coaching: what your best reps do differently, surfaced to everyone else. Of the five workloads, this is the one whose core mechanism has rigorous published support — just from an adjacent function. In a randomized rollout across 5,172 customer-support agents, Brynjolfsson, Li, and Raymond found that access to a generative-AI conversation assistant raised productivity, measured as issues resolved per hour, by 15% on average[2].

The distribution of that gain matters more than the average. The study found the improvements concentrated among less-experienced and lower-skilled agents, who got faster and better, while the most experienced, highest-skilled workers saw small speed gains and small quality declines[2]. The authors' interpretation is the exact mechanism conversation-intelligence vendors sell: the AI system captured and disseminated the patterns of high performers, and it facilitated learning — agents improved even in how they communicated, particularly newer ones[2]. If your sales team has a wide performance spread between top and bottom quartile, this evidence says AI-assisted coaching compresses it from below. If you are buying it to make your best closer better, the same evidence says to expect little.

Be honest about the transfer distance: this is customer support, not quota-carrying enterprise sales, and issue resolution is a cleaner outcome metric than a nine-month deal cycle. No comparably rigorous public study exists for B2B selling itself. But among all the mechanisms sales-AI vendors claim, "capture what top performers do and diffuse it to the rest" is the one with a real randomized trial behind it — which is why coaching-oriented conversation intelligence sits first in the adoption sequence at the end of this guide, and why the change-management side of that rollout (reps hearing "coaching," fearing "surveillance") deserves the treatment in /guides/ai-change-management-adoption.

Recording customer calls also creates a legal architecture problem before it creates an analytics opportunity. The federal baseline is one-party consent: under 18 U.S.C. § 2511(2)(d), interception is lawful where the person "is a party to the communication or where one of the parties to the communication has given prior consent"[5]. But state wiretap statutes layer on top of that baseline, and some states require the consent of all parties to a conversation. A national sales team cannot know at dial time which jurisdiction's rule a given call will touch, so the only sane enterprise posture is to design for the strictest case: affirmative disclosure and consent capture on every recorded call, logged alongside the recording itself.

Consent is an architecture decision, not a checkbox

Build the consent flow into the calling stack — announcement, opt-out path, and a consent record stored with each recording — rather than relying on rep behavior. Then govern the corpus like the sensitive dataset it is: retention limits, role-based access, and a clear policy on whether recordings feed employment decisions, which carries obligations well beyond wiretap law.

Revenue intelligence and deal scoring: linking activity to outcomes

Revenue-intelligence platforms ingest the activity exhaust of a sales organization — email metadata, call logs, meeting cadence, CRM stage changes — and try to connect it to outcomes: which deals are healthy, which are stalling, which behaviors correlate with winning. Deal scoring is the sharpest expression of the idea: a model trained on historical won/lost deals assigns each open opportunity a win probability from its activity signature.

The statistical machinery here is unglamorous and well understood. Gradient-boosted tree models handle the tabular, structured nature of activity data well and stay explainable; sequence models can capture temporal patterns — momentum, engagement bursts, multi-threading across stakeholders — at the cost of more data and less interpretability. The technology is rarely the constraint. Three data problems are.

  1. Label quality. The training signal is your own CRM's won/lost history, complete with deals marked closed-lost months late, wins recorded against the wrong opportunity, and stage definitions that drifted across three sales-process revisions. A scoring model faithfully learns all of it.
  2. Capture completeness. Activity that happens outside instrumented channels — a champion's text message, a hallway conversation at a conference — is invisible to the model, and its absence is systematically unequal across deal types and rep working styles.
  3. Correlation posing as cause. "Deals with a meeting cluster in a two-week window close more often" is a real pattern and a terrible instruction. Tell reps the model rewards meeting clusters and you will get meeting clusters — the input gamed, the signal destroyed, the score untethered from the outcome it once predicted.

The same caution governs deal intelligence — the competitor-mention detection and win/loss pattern analysis these platforms layer on top of the conversation corpus. Knowing that a specific competitor's appearance in second-call transcripts correlates with losses is genuinely useful input to competitive strategy. It becomes dangerous the moment it is treated as a verdict rather than a hypothesis, because mention patterns confound with segment, deal size, and rep skill. Route these outputs to the people equipped to interrogate them — competitive and enablement teams — not straight into rep-facing alerts.

The buying implication: evaluate revenue-intelligence platforms on data-plane fit, not model mystique. Depth of automated activity capture across your actual channel mix, transparency of the scoring features, and the ability to backtest scores against your own historical outcomes matter more than any claimed model sophistication. Score explanations that name the contributing signals are what make a score coachable rather than merely ominous — and they are your main defense against the gaming problem, because you can see what behavior the model is about to incentivize.

AI forecasting: the accuracy comparison you cannot buy

Forecasting is where the market's claim inflation is most visible. Buyers routinely ask which platform — a Salesforce Einstein, a Clari, a Gong — forecasts "most accurately," and content across the industry happily supplies percentage answers. Those answers are not verifiable. Forecast accuracy is a property of a model applied to a specific business — its deal mix, cycle length, seasonality, data hygiene, and override culture — measured at a specific snapshot cadence against a specific definition of error. No public, methodologically transparent benchmark exists that compares these platforms on equal footing, and vendor-reported accuracy figures never disclose the baseline they beat. Any cross-vendor accuracy percentage you encounter, including in earlier drafts of content this guide replaces, should be treated as fabricated until it names a methodology.

The good news is that the only benchmark that matters is one you already own: your historical CRM snapshots and actuals. Before any forecasting-AI purchase, backtest — feed candidates your history and compare their week-by-week predictions, at the same snapshot points your forecast calls actually happen, against what really closed. Compare them not to each other first, but to your incumbent process: the human roll-up with all its sandbagging and happy ears. An AI forecast that cannot beat your sales directors' committed number on your own data is expensive theater.

Evaluation dimensionWhat to pin downWhy it changes the answer
BaselineYour current human forecast's error, measured the same way"Accurate" is meaningless except relative to what you do today
Error definitionAbsolute percentage error vs. signed bias, at which aggregation levelA model can look great on total revenue while being wrong on every segment that matters
Snapshot timingAccuracy at week 1 vs. mid-quarter vs. final weekLate-quarter accuracy is easy; early-quarter accuracy is what changes decisions
GranularityCompany, segment, team, or rep-level forecastsAggregate accuracy often masks offsetting segment errors
Override behaviorWhat happens when managers adjust the model's numberIf overrides silently dominate, you are buying a very costly opinion-recording system
The forecast backtest you run before buying — on your own historical snapshots, against your own human baseline.

Playbooks and next-step suggestions: when the machine should speak

Between passive scoring and autonomous action sits the guided-selling layer: AI that suggests the next step — schedule the security review, multi-thread to the economic buyer, send the case study — at the moment it judges the deal needs it. The design question is not whether the model can generate suggestions; it is when a suggestion earns the interruption. Three conditions should gate every prompt: enough signal to make the recommendation confident, a trigger event that makes it timely rather than nagging, and a rep-facing explanation of why — the same explainability requirement as deal scoring, because an unexplained instruction from a machine is the fastest route to the dismiss-all reflex.

Calibrate the autonomy boundary deliberately. Suggestions that reps can accept, modify, or override keep the human accountable for the deal and generate the feedback data — which suggestions get accepted, and whether accepted suggestions correlate with better outcomes — that tells you whether the playbook logic is earning its place. A useful discipline is to track suggestion acceptance rate by rep tenure: the support-agent evidence predicts novices should gain most from encoded best practice[2], so if your most junior reps are the ones dismissing the prompts, the problem is trust or explanation quality, not the model.

Proposals, RFP responses, and SOW drafting: fast drafts, gated truth

Proposal and RFP work is the most naturally LLM-shaped workload in sales: parse a structured question set, retrieve relevant approved answers, assemble a coherent first draft in the house voice. The productivity case does not need vendor numbers — anyone who has run an RFP response knows most of the elapsed time is assembly, not judgment. The risk case is equally clear and has a name. NIST's Generative AI Profile defines confabulation as "the production of confidently stated but erroneous or false content (known colloquially as 'hallucinations' or 'fabrications') by which users may be misled or deceived," and lists it among the risks unique to or exacerbated by generative AI[6].

In most internal AI uses, confabulation costs you rework. In a proposal, it costs you a contract term: an invented compliance certification, a hallucinated SLA figure, or an overstated integration capability in a signed RFP response is a commitment your delivery and legal teams now own. The mitigation pattern is grounding plus gates. Ground generation strictly in a maintained, versioned library of approved answers, claims, and clauses — the model assembles and adapts, it does not invent. Then gate by content class: boilerplate flows through with spot checks, while pricing, legal terms, security and compliance claims, and anything with a number require named-owner review before the document leaves the building. Keep the provenance trail — which library entries fed which sections — so review is a diff, not a re-read.

The tell that your proposal AI is ungrounded

Ask the system a question your content library cannot answer — a certification you do not hold, an integration you have never built. A well-architected system returns "no approved content" and routes to a human. A system that produces a fluent, confident answer anyway[6] has just shown you exactly what it will one day put in front of a customer.

Honest objections

The steelman against aggressive sales-AI adoption deserves a fair hearing, because most of it is true. First, the evidence gap is real: the flagship RCT is about customer support, and this guide leans on it precisely because nothing of comparable rigor exists for quota-carrying B2B sales. The honest statement is that the mechanism — encoding top-performer behavior and diffusing it — is proven in an adjacent function, not that AI-assisted selling has been proven to lift revenue. Second, the same study cuts against the fantasy of across-the-board gains: the best workers saw small speed gains and small quality declines[2], so a tool sold as making everyone better is, on the best available evidence, mostly a floor-raiser.

Third, adoption statistics flatter the category. Sales and marketing leads all business functions at 52% among AI-using firms[1] — but only 18% of firms used AI at all in the survey's reference window (32% employment-weighted)[1], and adoption skews heavily to scale: 37% of firms with at least 250 employees reported using AI, against under 20% of the smallest firms[4]. Broad function-level adoption is compatible with most deployments being a copilot license and an email drafter. Fourth, automated outreach has a genuine tragedy-of-the-commons problem: every additional AI-personalized cold email trains buyers to ignore the channel, and your agent's marginal send degrades the same domain reputation your revenue depends on. None of these objections argues for doing nothing. They argue for sequencing — which is the decision this guide exists to support.

The read: sequence by evidence, not by demo quality

Put the five workloads on two axes — strength of evidence and blast radius when it fails — and the adoption order writes itself. Start where evidence is strongest and failure is cheapest: conversation-intelligence coaching with a consent architecture in place, and proposal drafting grounded in an approved library behind human gates. Both keep humans accountable for every customer-facing output, and both have a clear mechanism for why they should work[2][6].

Second tier: deal scoring and AI forecasting — pilot them, but only against your own backtested baseline, with explainable scores and a plan for the gaming problem. Third tier: autonomous SDR outreach, run as a bounded experiment with a claims-library guardrail, deliverability instrumentation, weekly output review, and pre-agreed kill criteria — because its failure mode is public, compounding, and now carries regulatory precedent[3]. Across all three tiers, resist buying each workload as an island: scoring, forecasting, conversation data, and outreach compound only when they share a data plane, which is the architecture argument made in /guides/unified-gtm-ai-stack. And define the measurement design before the first pilot — /guides/measuring-ai-roi-guide covers how to set baselines that survive contact with a renewal negotiation.

How to apply this: the sales-AI adoption checklist

  • Inventory which of the five workloads (outreach, conversation intelligence, revenue intelligence/scoring, forecasting, proposals) you are actually evaluating — refuse bundles that blur them.
  • For each candidate, write down the vendor's headline performance claim and whether any public, methodology-disclosed evidence supports it. Treat unsupported percentages as marketing.
  • Stand up the consent architecture before recording: disclosure on every call, consent logged with the recording, retention and access policies set for the strictest jurisdiction you sell into[^govinfo-18usc2511-2023].
  • Sequence coaching first: deploy conversation intelligence for rep development, and measure whether it compresses the top-to-bottom performance spread, where the RCT evidence predicts the gain[^arxiv-2304-11771].
  • Backtest any forecasting or deal-scoring model on your own historical snapshots against your human baseline before contract signature — accuracy claims you did not measure do not exist.
  • Require score and suggestion explanations naming the contributing signals; monitor for reps gaming the flagged behaviors.
  • Ground proposal generation in a versioned approved-content library; gate pricing, legal, security, and compliance claims behind named-owner review[^nist-ai-600-1].
  • For any outreach agent: enforce an approved-claims policy (no invented testimonials or capabilities — the FTC's fake-review rule reaches AI-generated personas[^ftc-fake-reviews-rule-2024]), instrument sender-domain health, review live output weekly, and set kill criteria in advance.
  • Measure qualified meetings held and downstream conversion, never raw replies or activity volume.
  • Fund the change-management work — coaching framing, not surveillance framing — as part of the purchase, not an afterthought (see /guides/ai-change-management-adoption).

Sources

Every quantitative or attributed claim above is linked to a primary source. Last verified at publication.

  1. [1]
    The Microstructure of AI Diffusion: Evidence from Firms, Business Functions, and Worker Tasks (CES-WP-26-25)
    U.S. Census Bureau, Center for Economic Studies · · accessed
  2. [2]
    Generative AI at Work
    arXiv (Brynjolfsson, Li & Raymond) · · accessed
  3. [3]
    FTC Announces Crackdown on Deceptive AI Claims and Schemes (Operation AI Comply)
    Federal Trade Commission · · accessed
  4. [4]
    Large Firms With at Least 20 Employees Biggest AI Users
    U.S. Census Bureau · · accessed
  5. [5]
  6. [6]
  7. [7]
    Federal Trade Commission Announces Final Rule Banning Fake Reviews and Testimonials
    Federal Trade Commission · · accessed