Advisors: Claude Can't Pass an SEC Audit

Written by: Toby Wade, | CEO DeepVest

Picture this: an advisor runs a Sharpe ratio calculation on Claude twice in the same week. First answer: 0.82. Second answer: 1.14. She uses the second one in a client presentation. Three months later, during an SEC audit, the examiner asks her to reproduce the calculation. She can't explain why the number changed.

That story is more troubling than it might seem.

When Anthropic announced Claude's enterprise rollout for financial services, the response was predictable: enthusiasm mixed with "why not just build our own?" The temptation to cobble together a ChatGPT wrapper or Claude API integration is strong.

But finance is different. The characteristics that make AI transformative elsewhere make it dangerous in wealth management. The gap between "works most of the time" and "meets fiduciary standards" is wider than most firms realize.

Why Finance Is Different

Financial advisory has three requirements that general-purpose AI fundamentally cannot meet: regulatory auditability, deterministic calculations, and persistent client context.

The SEC's 2026 examination priorities require firms using AI to explain and document how investment decisions were made. An AI that produces different portfolio recommendations on different days cannot generate a defensible audit trail. The question isn't whether the answer is right today, but whether you can explain and reproduce it under examination.

General-purpose AI models are probabilistic by design. They generate answers based on patterns in training data, not by retrieving verified facts and applying documented formulas. That architecture works fine for drafting emails or summarizing research, but it breaks down the moment regulatory accountability enters the picture.

DeepVest’s recent study tested six AI systems on ten financial calculations, running each question five times in fresh sessions. Claude Opus 4.7 refused to answer 32% of all attempts, with complete failure on two moderately complex questions requiring portfolio tracking and monthly rebalancing. On questions it did answer, results varied: Sharpe ratio calculations showed 10.1% variability across identical runs, and several answers deviated materially from ground truth. An advisor preparing a client presentation cannot work with a system that intermittently provides unreliable results.

The Calculation Problem

Deterministic calculations require institutional-grade data sources and transparent methodology. General-purpose models lack both. When you ask Claude to calculate a rolling drawdown, it estimates based on training data. When you ask a purpose-built system, it retrieves adjusted close prices from a verified institutional database, applies the documented formula, and returns a number you can audit.

The difference matters because financial calculations compound. A 2% error in volatility estimation becomes a 15% error in portfolio optimization. A miscalculated correlation changes the entire efficient frontier. You don't discover these errors by reading the AI's output. You discover them when a client's accountant catches the mistake or when a regulator asks you to show your work.

The Context Gap

The third problem is persistent client context. Financial advice is not transactional. Every recommendation builds on a relationship: prior conversations, documented risk tolerance, family circumstances, tax situation, estate plans, and investment history. General-purpose AI treats each prompt as isolated. You get an answer, but not one that accounts for the full client relationship.

Some firms think they can solve this by fine-tuning or building retrieval systems on top of Claude. That misunderstands the problem. Data access is only part of the equation. The bigger question is whether the system was designed from the ground up to maintain compliance, manage permissions, and produce auditable outputs. Bolting those requirements onto a general-purpose model creates exactly the kind of fragile system that fails during an audit.

What Actually Works

The industry is learning a hard lesson: AI that works for consumer applications doesn't work for regulated financial advice. The companies succeeding aren't trying to make ChatGPT or Claude into something they were never designed to be. They're building systems where deterministic data retrieval, transparent calculations, and full audit trails are architectural requirements, not afterthoughts.

Anthropic has built an extraordinary model. But finance is not general-purpose. Neither is fiduciary responsibility or regulatory compliance.

Before investing six months building an AI stack on top of Claude, advisors should ask themselves a question: when the SEC examiner asks me to reproduce that calculation, can I explain why the number is different this time?

If the answer is no, Claude is not your friend. He's a liability.

Related: AI Is Giving Dangerous Life Insurance Advice. Who Pays When It's Wrong?