Anthropic's release of 23 new legal integrations and Claude Opus 4.7's 90.9% score on Harvey's BigLaw Bench marks a genuine technical milestone. For UK legal practices regulated under the SRA Code of Conduct, this matters: AI that performs at near-human level on M&A due diligence, employment contracts and document review is genuinely useful. But the story buried in Fortune's coverage is the one that should concern mid-market firms most. The same month Anthropic released these tools, courts across three jurisdictions were still finding AI-generated hallucinations in actual case filings. A 90% benchmark does not mean 10% of critical errors are acceptable in a contract review or a compliance opinion.
This pattern — vendors shipping powerful models while hallucinations persist in production — has become routine. Harvey, Legora and Luminance have all released similar high-benchmark products to similar acclaim. What's changed is not the underlying risk; it's the pace at which law firms are being asked to bet their fee-earner allocation and pricing models on tools that still fail in predictable, catastrophic ways. The FCA's Consumer Duty (PS22/9) and emerging ISA UK audit standards increasingly require firms to understand and document the failure modes of tools they use in regulated work. Big Law can absorb a few hallucination incidents; mid-market practices serving SME clients cannot.
Here's Trovix's honest view: the benchmark chase is real but it's not the same as safe implementation. A model that scores 90% on a controlled test set behaves differently when it encounters novel contract language, emerging regulatory requirements, or the edge cases that actually generate fee disputes. The firms winning with AI — and there are genuinely productive ones — are not those using single-vendor solutions deployed at scale. They're the ones treating AI as part of a compliance workflow, not a replacement for it. That means: governance dashboards that flag when AI output needs human review (Trovix Audit), document intelligence that extracts data reliably but doesn't make interpretive leaps (Trovix Sift), and fee-earner tools that assist rather than substitute (Trovix Aria). It means asking what happens when the tool fails, not assuming it won't.
If you run a mid-market legal practice, insurance broker, financial advisory firm or accountancy practice, the practical move right now is not to buy the new Anthropic release because Harvey or Luminance have 85% benchmarks. The move is to audit what's already deployed. How are hallucinations being caught? Who documents the decision to use AI on a given task? Are your governance controls fit for SRA Code compliance, FCA Consumer Duty expectations, or ICO UK GDPR obligations around automated decisions? The firms that will own the next three years are those that implement AI slowly, with documented reasoning, and with a hard stop when the stakes are high. That's not a constraint. That's a competitive moat.
Source: Fortune