Anthropic's new legal plugins are impressive on paper. But the industry's rush to deploy specialist AI before solving hallucination and accountability problems puts mid-market UK firms at real SRA and FCA risk.
AI Governance  Trovix SiftLegal · Financial Services

Anthropic released 20+ legal workflow integrations in May 2026, with Claude Opus 4.7 scoring 90.9% on Harvey's BigLaw Bench. The headline is bullish: enterprise-grade AI for M&A due diligence, employment handbooks, contract review, document drafting. But Fortune's story buried the real news—even as these tools roll out across Big Law, hallucinations are still appearing in actual legal filings. A 90.9% benchmark score means 9.1% of the time, the model gets it wrong. In legal services, 9% is not acceptable. The SRA Code of Conduct requires that regulated firms act with integrity and comply with the law. An AI tool that generates plausible-sounding but false case citations, contractual provisions, or due diligence findings does not meet that standard. The question for mid-market UK law firms, insurers, and accountancy practices is not whether Anthropic's model is clever—it clearly is. The question is whether your firm can defend its use of it to the regulator if something goes wrong.

This story is the third act in a pattern we are watching closely. First came the hype cycle: Claude and GPT-4 will revolutionise legal work. Second came the benchmark wars: Harvey, Legora, Luminance all published impressive accuracy scores. Third comes the reality check: tools that score well in controlled lab conditions still fail in production, still hallucinate, and still expose firms to liability. Big Law is moving fast because it has the resources—insurance, compliance teams, brand protection budgets—to absorb early mistakes. Mid-market firms do not. Yet mid-market firms are under the same regulatory pressure to compete and innovate. The Lloyd's Blueprint Two framework, the FCA Consumer Duty PS22/9, and the incoming EU AI Act all demand that firms understand, govern, and justify the AI they use. A 90% accurate AI tool is not justified if you cannot explain to your regulator how you caught the 10% that was wrong.

Here is Trovix's honest view: specialist AI plugins are only one piece of the puzzle. The other piece—and the one most vendors skip over—is governance, validation, and human oversight at the point of output. Anthropic's approach is model-centric: build a better model, integrate it deeper into workflows, let benchmarks prove quality. That works for customer service. It does not work for legal, insurance, or financial services. At Trovix, we start from the opposite end: assume the AI will make mistakes, and build systems to catch them before they reach a filing, a policy, or a regulator. Our approach to document intelligence through Trovix Sift is not to replace human review with an AI score—it is to identify what matters, flag what is uncertain, and make the fee-earner's job faster and safer. Our Trovix Aria RAG knowledge assistant does not generate new legal theory; it retrieves and ranks existing precedent so the lawyer stays in control. The difference between Anthropic's plugins and our approach is the difference between saying 'the AI will handle this' and saying 'the AI will support this while you remain accountable.' Legora and Harvey have built similar governance layers, and they are the vendors that will survive regulatory scrutiny. Luminance focused early on transparency; that positioning is now paying dividends. Microsoft Copilot for legal is too generic; it will not work for this sector without firm-specific guardrails. The market will sort this out, but it will sort it out the hard way—through complaints to the SRA, investigations by the ICO under UK GDPR, and eventual enforcement.

If you are a mid-market law firm, insurer, accountancy practice, or financial services firm evaluating AI tools right now, do not ask 'What is the benchmark score?' Ask instead: 'How do I audit every output before it leaves my firm? How do I explain to my regulator why this tool was the right choice? What happens when it fails?' Anthropic's plugins are technically impressive. But they are not a substitute for governance. If you are deploying them, you need a validation layer—either built into the tool or bolted on top of it. If you are evaluating alternatives, look for vendors who start from compliance and accuracy, not from model scale. The firms that win this transition are not the ones that moved fastest to AI. They are the ones that moved fastest to auditable, explainable, regulated AI.

Source: Fortune

Related Trovix product:

Trovix Sift →Book a demo →