Anthropic's new legal AI tools are sophisticated and well-benchmarked. But UK regulated firms need to stop chasing Big Law's speed-to-market and start building AI governance frameworks first. The story Fortune buried—hallucinations are already appearing in real legal filings—is the one that should c
AI Governance  Trovix WatchLegal · Financial Services · Insurance

Anthropic's release of 22 legal integrations and Claude Opus 4.7's 90.9% score on Harvey's BigLaw Bench look impressive in isolation. But the story buried in this announcement is the one Fortune deliberately flagged: hallucinations are already appearing in legal filings. That gap—between benchmark performance and real-world failure—should be the actual headline for any UK mid-market firm considering rapid AI deployment. A 90.9% success rate sounds strong until you realize that even a single hallucinated case citation in a submission to the SRA, FCA or ICO could trigger a conduct investigation, not to mention the liability exposure to clients. Big Law can absorb reputational damage and legal costs that mid-market practices cannot.

This story is part of a larger pattern: AI vendors are now embedding legal-specific tools directly into law firm workflows at pace, betting that volume and specialization will solve the accuracy problem. Harvey's benchmarking framework has become the de facto industry standard for measuring billable work substitution. But benchmarks measure performance on clean datasets, not on the edge cases, archaic precedent and jurisdiction-specific nuances that real legal work demands. Meanwhile, firms like Luminance and Legora have taken different routes—focusing on transparency and human-in-the-loop workflows rather than end-to-end automation. Anthropic's approach assumes that more legal-specific training data and role-specific plugins will close the hallucination gap. The Fortune story suggests they haven't.

Trovix's view is this: UK regulated firms face a different risk calculus than American Big Law. The SRA's recent conduct guidance, the FCA's Consumer Duty PS22/9 expectations, and the ICO's ongoing position on AI systems and UK GDPR mean that deploying an AI system without documented governance, explainability and human sign-off is not just risky—it's potentially a breach. Anthropic's plugins are well-engineered, but they are tools, not governance frameworks. Installing Claude into your M&A workflow solves the speed problem and creates a compliance problem. We've seen this pattern before: firms adopt the technology, then scramble to build the guardrails. That approach is backwards. Trovix Audit exists precisely because mid-market regulated practices need to demonstrate that their AI use is compliant, auditable and defensible—before they scale it, not after. Microsoft Copilot and other embedded LLMs face the same governance gap; they don't solve it.

If your firm is considering Anthropic's legal plugins or any comparable AI system, your first step should not be implementation. It should be governance mapping: document which decisions AI will and will not make, establish human sign-off requirements, define your audit trail, and clarify your liability chain. Only then should you pilot the technology. Use Trovix Watch to stay ahead of regulatory changes that will inevitably tighten AI accountability rules over the next 18 months. The firms that move fastest are also the firms that expose themselves first. Being second to market with a compliant, documented, defensible AI workflow is better than being first with a hallucination in a court filing.

Source: Fortune

Related Trovix product:

Trovix Watch →Book a demo →