Anthropic's new legal plugins and 90.9% benchmark score are genuinely impressive. But they tell only half the story — and UK regulated firms are about to learn the costly difference between what AI can do and what it should do in practice.
Legal Tech  Trovix SiftLegal · Financial Services · Insurance · Accountancy

On 12 May, Anthropic released over 20 legal workflow integrations, including 12 role-specific plugins for M&A due diligence and employment drafting, with Claude Opus 4.7 scoring 90.9% on the BigLaw Bench benchmark. For big magic circle practices, this is a credible productivity signal. For UK mid-market law firms, insurers, financial services firms and accountancy practices — it's a trap. A 90.9% accuracy rate sounds reassuring until you remember that legal AI hallucinations have already appeared in actual court filings, regulatory submissions and transaction documents. In a sector bound by the SRA Code, FCA Consumer Duty (PS22/9), and increasingly strict ICO UK GDPR interpretations, the question is never whether AI can do the work. The question is whether your firm can explain to regulators, clients and courts why you delegated a £2m transaction to a system that gets 9% of things wrong — and can't tell you which 9%.

This story is part of a wider industry pattern: major AI vendors (Harvey, Legora, Luminance, Microsoft Copilot, now Anthropic) are racing toward feature completeness and benchmark performance, while the boring, unglamorous work of safe implementation lags behind. Big Law can absorb the liability risk and train senior lawyers to catch errors. Mid-market firms cannot. They have tighter margins, fewer quality gates, and clients who expect Big Law service at mid-market pricing. The result: firms adopt these tools too fast, cut verification processes to save time, and find themselves facing SRA investigations, FCA breaches of Consumer Duty, or PRA SS1/23 governance failures because their AI governance framework (if they have one) is template-based theatre rather than real controls. The EU AI Act classification of legal AI as high-risk and the emerging ICO guidance on AI and data protection have made compliance a liability issue, not just a nice-to-have.

Trovix's view is simple: benchmark performance is marketing. Real safety is engineering. We have always believed that AI tools must be implemented within a governance layer that treats the system as a draft-stage assistant, not a decision-maker. We don't sell you a plugin and hope for the best. We sell you governance through Trovix Sift, which gives you explainable data extraction and document intelligence with full audit trails and human validation checkpoints built in. We pair that with Trovix Aria, which provides RAG-backed knowledge retrieval where you can trace every source document and verify every claim before relying on it. This is fundamentally different from Anthropic's plugin approach, which assumes the underlying model can be trusted to apply workflow rules without external verification. It can't. Not yet. Maybe not ever in high-stakes legal work. Harvey and Luminance have learned this the hard way — their early wins have been in document review and predictive due diligence, not in autonomous drafting. Legora has positioned itself as a compliance layer on top of LLMs, which is closer to the right architecture. But none of them have built the kind of human-in-the-loop governance that UK regulated firms actually need to stay compliant.

If you're a mid-market firm deciding whether to adopt Anthropic's legal suite right now, don't. Or if you do, treat it as a fast-first-draft tool and build a verification process around it. Specifically: (1) demand an audit trail for every AI decision (who ran it, when, what were the inputs, what was the output); (2) require mandatory human review before any output touches a client or a filing; (3) use Trovix Watch to monitor regulatory guidance on AI liability as it evolves through 2026 and beyond; (4) document your AI governance controls in writing and make sure your partners and the SRA (if you're a law firm) can see that you have a real framework, not just a policy. The firms that will win are not the ones that move fastest to AI. They are the ones that move safest. Anthropic's benchmark is real. But your liability is realer.

Source: Fortune

Related Trovix product:

Trovix Sift →Book a demo →