A 90.9% accuracy score in legal AI looks good until it's your client's motion rejected because the AI invented case law. The story of Big Law's Claude deployment is not about AI capability—it's about the governance vacuum that precedes every regulatory crackdown.
AI Governance  Trovix AuditLegal · Financial Services · Insurance

Anthropic's announcement of 20+ legal integrations and 12 new Claude plugins is not news about AI working. It is news about the market's collective decision to ignore the news. Courts have already sanctioned law firms for AI hallucinations in live filings. Harvey's benchmark—the tool being cited as proof Claude is ready—measures raw model performance in isolation. It does not measure what happens when a fee-earner uses Claude on a real file, with real client pressure, real time constraints, and real stakes. Freshfields, Quinn Emanuel and Holland & Knight are running Claude on live matters. Their adoption signals to the market that AI risk is acceptable. For mid-market UK regulated firms, this creates immediate pressure: match Big Law's speed or lose clients who expect it. The SRA Code and FCA Consumer Duty PS22/9 do not exempt you from that pressure. They do exempt you from responsibility if things go wrong.

This is the third wave of this pattern. First, chatGPT was a toy. Then GPT-4 became real and everyone had to have a policy. Now Claude and Harvey and Luminance are in the market, benchmarks are published, and the conversation has shifted from 'should we?' to 'how fast?'. The missing conversation is 'who is accountable?'. When Harvey publishes 90.9% accuracy, it is measuring how often the model gets it right on a standardised legal Q&A test. It is not measuring hallucination rate in document review under time pressure, or the probability that a fee-earner will miss the 9.1% error because they trusted the benchmark. The EU AI Act is coming. The ICO's AI guidance is already here. Neither will credit firms with good intentions. Both will look at governance frameworks, audit trails, and whether controls were in place before incidents, not after.

Trovix's view is direct: deployment speed and model accuracy are not the same as safe deployment. Claude may be excellent. The issue is not the model—it is the implementation gap. Most firms deploying Claude today are running it like a research assistant: point it at documents, get an output, use it. That works if the model is right. It fails catastrophically if it is wrong and nobody catches it. The alternative is to build governance around deployment. That means: (a) defining what counts as hallucination in your domain and how you measure it; (b) building human review checkpoints into workflows; (c) maintaining an audit trail of what the AI was told and what it concluded; (d) testing the model against real client files before going live, not after. Harvey's benchmark does not do this. Luminance's approach includes more human-in-loop design. Trovix's Audit product exists precisely because benchmarks tell you what a model can do, not what it does in your firm with your data and your people. The difference is real.

For a mid-market law firm, insurer, financial services firm or accountancy practice, the practical move is this: Do not wait for your regulator to tell you Claude is unsafe. Do not assume Big Law's adoption means it is proven. Instead, treat any new AI deployment—Claude, Luminance, Microsoft Copilot, Harvey, or anything else—as a controlled pilot. Run it on a closed set of files. Document what it does. Measure its errors specifically in your domain. Keep the human fee-earner or underwriter in the loop explicitly. Use tools like Trovix Audit to log decisions and model inputs so that when something does go wrong, you have evidence of governance, not just good faith. This is not slower than Big Law. It is faster to recover from regulatory scrutiny, and it keeps you compliant with the SRA Code and FCA Consumer Duty PS22/9. The firms that will win over the next 18 months are not the ones that adopted Claude first. They are the ones that adopted it safely first.

Source: Fortune

Related Trovix product:

Trovix Audit →Book a demo →