Financial Services · Use case
AI in Banking and Capital Markets: Underwriting, Wealth, Trading, and the Vendor Landscape
Banking runs four very different AI workloads — credit underwriting, wealth management, trading, and stress testing — and each answers to a different regulatory gatekeeper. The stack decision is not one platform choice but four: where explainability is a legal requirement, where fiduciary duty constrains personalization, where alpha decays on contact, and where the supervisor runs the model that matters.
US finance and insurance firms reporting AI use as of May 3, 2026 — well above the 19.8% national rate, with the Information sector at 39.7%. Roughly 39% of finance and insurance firms expected to use AI in their business functions over the following six months.[^census-btos-ai-2026]
US Census Bureau BTOS, May 2026
Aggregate losses projected for the 22 banks in the Federal Reserve's 2025 severely adverse stress scenario — $472 billion of it loan losses — the modeling exercise every large-bank risk stack must ultimately reconcile against.[^frb-dfast-2025]
Federal Reserve, 2025 DFAST results
Deadline under Regulation B for a creditor to notify an applicant of action taken on a completed application — and an adverse-action notice must state specific principal reasons, a constraint that binds every underwriting model you deploy.[^ecfr-12-1002-9]
12 CFR 1002.9
Banks and capital-markets firms are among the heaviest AI adopters in the US economy, but 'AI in banking' is not one decision. It is at least four — underwriting, wealth, trading, and risk — each with its own regulator, latency profile, and build-versus-buy default. Treat them as one platform purchase and you will satisfy none of the four gatekeepers.
Portfolio-day hit rate at which GPT-4 captured the initial (non-tradable) market reaction to news headlines in the Lopez-Lira and Tang study — with strategy returns declining as LLM adoption rises.[^arxiv-2304-07619]
Lopez-Lira & Tang, arXiv
Four workloads, four gatekeepers
The single most useful framing for a CIO or chief risk officer is that each banking AI workload answers to a different authority with a different tolerance for opacity. Underwriting answers to fair-lending law, where an unexplainable decision is an illegal one. Wealth answers to fiduciary duty under the Investment Advisers Act. Trading answers mostly to the market itself, which arbitrages away any signal that works. Stress testing answers to the Federal Reserve, which runs its own models and does not care how good yours are. Your stack architecture, vendor shortlist, and governance budget should differ accordingly.
| Dimension | Credit underwriting | Wealth & robo-advice | Trading & markets | Stress testing |
|---|---|---|---|---|
| Gatekeeper | ECOA / Regulation B (12 CFR part 1002)[^ecfr-12-1002] | SEC, Investment Advisers Act fiduciary duty[^sec-im-robo-2017] | Market efficiency itself; exchange and conduct rules | Federal Reserve supervisory stress test[^frb-stress-tests-overview] |
| What kills a deployment | Adverse-action notices that cannot state specific reasons[^ecfr-12-1002-9] | A conflict of interest embedded in the algorithm[^sec-pda-2023-140] | Alpha decay and regime shifts | A satellite model that cannot survive effective challenge[^frb-sr117-guidance-2011] |
| Latency profile | Seconds to days | Minutes to quarterly rebalance | Milliseconds to minutes | Weeks per cycle |
| Buy vs. build default | Buy decisioning platform, own the model governance | Buy the platform, own suitability logic | Build; buy data and models as features | Build satellites around vendor risk engines |
| Explainability bar | Legal requirement per decision | Fiduciary disclosure obligation | Internal risk control | Examiner-grade documentation |
Credit underwriting: explainability is law, not a feature
Machine-learning underwriting — often enriched with alternative data such as cash-flow, rent, and utility payment history — is the most mature of the four workloads and the most legally constrained. Regulation B, which implements the Equal Credit Opportunity Act[5], requires a creditor to notify an applicant of action taken within 30 days of receiving a completed application, and an adverse-action notification must carry either a statement of specific reasons or a disclosure of the applicant's right to one[3].
The regulation is blunt about the quality of those reasons: the statement 'must be specific and indicate the principal reason(s) for the adverse action,' and statements that the action was based on the creditor's internal standards or policies, or that the applicant 'failed to achieve a qualifying score on the creditor's credit scoring system,' are explicitly insufficient[3]. That sentence is the design spec for your underwriting stack. A gradient-boosted or deep model whose feature attributions cannot be mapped to specific, accurate principal reasons is not a compliance gap you can paper over with a disclosure template — it is a model you cannot lawfully use for consumer credit decisions.
Layered on top of fair lending is model risk management. The Federal Reserve's SR 11-7 guidance, issued April 4, 2011[10], defines model risk as the potential for adverse consequences from decisions based on incorrect or misused model outputs, and makes 'effective challenge' — critical analysis by objective, informed parties who can identify model limitations and produce appropriate changes — the guiding principle of managing it[9]. In practice that means every underwriting model, vendor-supplied or in-house, needs independent validation, documented assumptions, and a challenger process. Specialist vendors — nCino on the loan-origination side, Zest AI and Upstart in ML-driven decisioning — compete in this segment, but the validation obligation stays with the bank regardless of whose model runs. The full supervisory frame is covered in /guides/model-risk-management-guide, and the sector rulebook in /guides/sectoral-ai-regulation-regtech.
Alternative data cuts both ways
Alternative data can extend credit to thin-file applicants, but any variable correlated with a protected characteristic can become a proxy for it. Before a feature enters the model, decide how you would defend it as a specific, accurate principal reason on an adverse-action notice[3] — and how it would fare under disparate-impact testing. If you cannot answer both, the predictive lift is not worth the exposure.
Wealth management: personalization under fiduciary duty
Robo-advisors are not a regulatory gray zone and have not been for years. The SEC staff's February 2017 guidance (IM Guidance Update 2017-02) treats robo-advisors as what they typically are — registered investment advisers under the Investment Advisers Act of 1940 — and flags three areas where automation strains the traditional obligations: the substance and presentation of disclosures, the obligation to gather enough client information to support suitable advice, and compliance programs designed for the particular risks of automated advice[6]. A questionnaire-driven onboarding flow that under-collects client information does not shrink the adviser's duty; it just makes the duty harder to meet.
The frontier issue is conflicts of interest embedded in the optimization itself. On July 26, 2023, the SEC proposed rules that would require broker-dealers and investment advisers to evaluate whether their use of predictive data analytics and similar technologies in investor interactions creates a conflict that places the firm's interest ahead of investors' — and to eliminate, or neutralize the effect of, any such conflict, backed by written policies and recordkeeping[8]. Whatever the proposal's final fate, it names the design question every wealth platform now has to answer: what is your recommendation engine actually optimizing for, and can you prove it?
Today's predictive data analytics models provide an increasing ability to make predictions about each of us as individuals. This raises possibilities that conflicts may arise to the extent that advisers or brokers are optimizing to place their interests ahead of their investors' interests.[^sec-pda-2023-140]
For the buyer, the wealth segment splits into retail-direct platforms (Betterment and Wealthfront are the familiar names), advisor-facing personalization layers, and institutional portfolio-and-risk platforms such as BlackRock's Aladdin. The differentiation that matters in an RFP is rarely the optimizer — mean-variance math is a commodity — but the auditability of the advice chain: can the platform reconstruct, for any client on any date, what was recommended, from what inputs, under which fiduciary logic? That is the artifact an examiner will ask for, and the one a conflicts rule would make existential.
Trading: LLM signals are real, and really perishable
The honest evidence on LLMs in markets is more interesting than the hype. Lopez-Lira and Tang, in a study first posted in April 2023 and since revised, found that GPT-4 could score news headlines for stock implications without financial fine-tuning, capturing the initial market response at roughly 90% portfolio-day hit rates — but that initial reaction is non-tradable, and the exploitable part is the subsequent drift, strongest in small stocks and after negative news[4]. Two of their other findings should anchor your expectations: forecasting ability generally increases with model size, and strategy returns decline as LLM adoption rises, consistent with the market simply becoming more efficient as everyone deploys the same tools[4].
That decay dynamic is the strategic point. An LLM sentiment signal is not a durable asset; it is a temporary information-processing advantage that erodes as the capability commoditizes. Domain-specific pretraining follows the same curve at higher cost: BloombergGPT, a 50-billion-parameter model trained on a 363-billion-token financial dataset augmented with 345 billion general-purpose tokens, reported in 2023 that it outperformed existing models on financial tasks by significant margins without sacrificing performance on general benchmarks[11] — and general frontier models have been competing for exactly that ground ever since. So the build-versus-buy call on trading NLP is really a refresh-rate call: whatever you build, budget to re-evaluate it against the current frontier model every quarter.
Architecturally, desks that use LLM-derived signals treat them as features feeding conventional quantitative models, not as autonomous decision-makers — the latency of a large model is incompatible with the execution path, and its opacity is incompatible with trade-surveillance and best-execution obligations. Keep the LLM in the research and signal-generation loop, keep deterministic risk checks in the order path, and log the prompt-to-signal lineage the same way you log any other model input. Do not expect vendor backtests to survive contact with your own data; and treat any pitch quoting a specific alpha figure without a published methodology as marketing, not evidence.
Stress testing: the supervisor runs the model that matters
The Federal Reserve's stress test assesses whether large banks are sufficiently capitalized to absorb losses during stressful conditions while continuing to lend; it runs annually with a minimum of two scenarios, its results set each bank's stress capital buffer requirement, and bank-level results are publicly disclosed[7]. The 2025 exercise covered 22 banks and projected, under the severely adverse scenario, $549 billion in aggregate losses and cumulative pre-tax net income of negative $80 billion over nine quarters[2].
Aggregate CET1 capital ratio, 2025 severely adverse scenario (22 banks)
Under that scenario the aggregate common equity tier 1 ratio falls from an actual 13.4 percent to a minimum of 11.6 percent before recovering to 12.8 percent[2] — and the dispersion across banks is wide, driven by differences in portfolio composition and business mix. The strategic implication for an AI roadmap: you cannot 'AI your way' out of the supervisory number, because the Fed's projections come from the Fed's models. Where machine learning earns its place is in the surrounding work — internal scenario expansion beyond the supervisory pair, satellite models that translate macro scenarios into portfolio-level losses, challenger models that pressure-test the incumbent regression stack, and NLP over macro and geopolitical text to inform scenario design. Every one of those models lands inside the SR 11-7 perimeter[10], so each needs examiner-grade documentation and effective challenge, which is why neural-network satellites often lose to simpler models that validators can actually interrogate.
Customer service: the highest-volume, lowest-drama workload
Banking chatbots and voice agents are the workload where hyperscaler platforms are genuinely the default: Amazon Lex V2 provides natural language understanding and automatic speech recognition for building voice and text interfaces[12], and Google's Dialogflow CX models a virtual agent that translates end-user text or audio into structured data, with conversations designed as explicit state machines of flows and pages[13]. Microsoft and specialist banking-conversation vendors such as Kasisto and Personetics round out the shortlists. The state-machine detail matters more than it looks: in a regulated contact center you want the deterministic flow to own account actions and disclosures, with generative language confined to intent understanding and phrasing.
Evaluate on your own transcripts, not vendor demos: containment rate on your top intents, escalation quality, and auditability of every automated utterance. And treat the voice channel as a fraud surface — the same conversational interface that authenticates customers is the one that deepfake audio attacks; the adjacent defenses are mapped in /use-cases/ai-fraud-financial-crime.
Reading the vendor landscape by segment
The financial-services AI market does not reward a single-vendor strategy, because the four workloads buy differently. A more durable map is by segment and by what actually differentiates within it.
Lending & decisioning
Origination platforms and ML decisioning specialists. Differentiator: adverse-action reason quality and fair-lending tooling, not raw model lift.
Wealth & advice
Retail robo platforms, advisor personalization layers, institutional portfolio/risk platforms. Differentiator: auditability of the advice chain.
Markets & research NLP
Data vendors with embedded models, finance-tuned LLMs, in-house signal stacks. Differentiator: data exclusivity and refresh cadence, since model capability commoditizes.
Risk & finance
Stress-testing and balance-sheet engines plus in-house satellite models. Differentiator: validation artifacts an examiner will accept.
Engagement & servicing
Hyperscaler conversational stacks and banking-specific assistants. Differentiator: deterministic control of regulated actions.
Two cross-cutting lessons from firms that run this at scale. First, the operational pattern that works is boring: one fintech pattern worth copying is a lender running more than 50 models in production on a shared container-orchestrated platform, with DAG-based pipelines for training and promotion, event-driven retraining triggered by drift signals, and monitoring and governance artifacts generated by the same pipeline that deploys the model. The specifics vary; the principle — one paved road for all models, with governance as a pipeline output rather than an afterthought — is portable to any institution juggling underwriting, marketing, and risk models simultaneously. Second, vendor concentration is a real risk in this sector: when your underwriting, servicing, and risk vendors all resell the same upstream foundation model, a single deprecation or repricing event propagates across workloads you thought were diversified.
Make the governance question the first RFP question
For every segment, open the vendor conversation with the gatekeeper's artifact, not the feature list: adverse-action reason codes for lending[3], the advice-chain audit trail for wealth[6], model documentation that survives effective challenge for risk[9]. Vendors that lead with accuracy claims and defer governance to 'roadmap' are optimized for a different buyer than you.
Honest objections
- 'The adoption number overstates reality.' Fair: the BTOS asks whether a business used AI in any business function in the past two weeks — not whether it runs AI at production scale — and even in finance roughly two-thirds of firms reported no AI use as of May 2026[1]. The competitive question is whether the third that does includes your direct competitors.
- 'LLM trading signals are already dead.' Partly right — the same study that documented the signal documented its decay with adoption[4]. That argues for treating LLM capability as infrastructure with a depreciation schedule, not for ignoring it while competitors compress their research cycles.
- 'Regulatory direction is uncertain, so wait.' The proposals may move — the SEC's predictive-analytics rule remains a proposal[8] — but the binding constraints in this article are not proposals: Regulation B's notice requirements[3], the Advisers Act duties the 2017 guidance restates[6], and SR 11-7[10] are current law and guidance. Waiting buys no relief from any of them.
- 'Buying governed vendor models transfers the risk.' It does not. Under SR 11-7's frame, vendor models sit inside your model inventory and your validation obligation[9]; a vendor's fair-lending white paper is an input to your testing, not a substitute for it.
The read
Fund the four workloads as four programs with one shared platform layer. Underwriting and risk deserve the heaviest governance investment because their constraints are statutory and supervisory; wealth deserves conflict-of-interest engineering before the rules force it; trading deserves fast, disposable experimentation with a quarterly re-benchmark against frontier models; servicing deserves a deterministic core with generative language at the edges. Across all four, the durable in-house asset is not any model — models commoditize and decay — but the governance pipeline that lets you swap models quickly and prove, per decision, what ran and why. That capability is what regulators examine, what M&A diligence prices, and what the next model generation cannot obsolete.
How to apply this
- Inventory every AI use across the four workloads and assign each its gatekeeper: Regulation B, Advisers Act, market/conduct rules, or supervisory stress testing.
- Test every underwriting model — vendor or in-house — for its ability to produce specific principal reasons per 12 CFR 1002.9 before it touches a live decision[^ecfr-12-1002-9].
- Run disparate-impact analysis on every alternative-data feature at intake, and document the business justification for each retained variable.
- Map your wealth platform's optimization objective and prove the advice chain is reconstructable per client, per date — the artifact both fiduciary duty and the SEC's conflicts proposal point at[^sec-pda-2023-140].
- Confine LLMs on the trading side to research and signal generation; keep deterministic risk controls in the order path and log prompt-to-signal lineage.
- Re-benchmark any finance-tuned or in-house NLP model against the current general frontier model on a fixed cadence.
- Bring every satellite, challenger, and scenario model into the SR 11-7 inventory with examiner-grade documentation and independent validation[^frb-sr117-2011].
- Consolidate model deployment onto one pipeline that emits monitoring and governance artifacts automatically, so adding the next model costs process, not heroics.
- Map upstream foundation-model dependencies across all vendors to surface hidden concentration risk.
- Pressure-test the adjacent financial-crime stack — see /use-cases/ai-fraud-financial-crime — and the governance frame in /guides/model-risk-management-guide and /guides/sectoral-ai-regulation-regtech.
Sources
Every quantitative or attributed claim above is linked to a primary source. Last verified at publication.
- [1]Large Firms With at Least 20 Employees Biggest AI UsersU.S. Census Bureau · · accessed
- [2]2025 Federal Reserve Stress Test Results: Results for Banks under the Severely Adverse ScenarioFederal Reserve Board · · accessed
- [3]12 CFR 1002.9 — Notifications (Regulation B adverse action requirements)eCFR / Consumer Financial Protection Bureau · accessed
- [4]Can ChatGPT Forecast Stock Price Movements? Return Predictability and Large Language ModelsarXiv (Lopez-Lira & Tang) · · accessed
- [5]12 CFR Part 1002 — Equal Credit Opportunity Act (Regulation B)eCFR / Consumer Financial Protection Bureau · accessed
- [6]IM Guidance Update 2017-02: Robo-AdvisersSEC Division of Investment Management · · accessed
- [7]Stress Tests and Capital PlanningFederal Reserve Board · accessed
- [8]SEC Proposes New Requirements to Address Risks to Investors From Conflicts of Interest Associated With the Use of Predictive Data Analytics by Broker-Dealers and Investment AdvisersU.S. Securities and Exchange Commission · · accessed
- [9]Supervisory Guidance on Model Risk Management (SR 11-7 attachment)Federal Reserve Board / OCC · · accessed
- [10]SR 11-7: Guidance on Model Risk ManagementFederal Reserve Board · · accessed
- [11]BloombergGPT: A Large Language Model for FinancearXiv (Wu et al., Bloomberg) · · accessed
- [12]What is Amazon Lex V2?Amazon Web Services · accessed
- [13]Dialogflow CX basicsGoogle Cloud · accessed