Anthropic has released 20+ integrations and new role-specific plugins for legal work, built on Claude Opus 4.7, which scored 90.9% on Harvey's BigLaw Bench. This matters to UK law firms, insurers and financial services firms because it represents the industry's clearest signal yet: generalist large language models are being packaged as specialist legal tools. The timing is deliberate. Big Law is moving fast, and Big Law's infrastructure choices shape what mid-market firms can access and afford. But here is the uncomfortable truth: a 90.9% benchmark score on a curated legal task set tells you almost nothing about hallucination risk in real client work, where the cost of a single confabulated case reference or invented statute is not a benchmark failure—it is a regulatory breach.
This release is part of a pattern. Harvey, Luminance and now Anthropic are all competing on model performance and integration breadth. Each claims to solve the hallucination problem through fine-tuning, retrieval-augmented generation or prompt engineering. What they actually demonstrate is that the industry has accepted hallucinations as inevitable and is now optimising for detection and mitigation rather than prevention. That is not a bad strategy—it is honest—but it is a strategy, not a solution. The FCA's Consumer Duty (PS22/9), the SRA Code and the ICO's UK GDPR principles all require firms to understand the tools they use and to take responsibility for the output they rely on. A 90.9% score does not discharge that responsibility. The regulatory framework has not caught up with the AI adoption rate, and firms that race ahead without governance are taking on unquantified tail risk.
Trovix's view is this: capability without governance is liability dressed as innovation. Anthropic's plugins are competent. But competence at M&A due diligence or employment handbook drafting is not the same as competence in a regulated environment where document intelligence must be auditable, traceable and defensible. The difference between Harvey's approach (which prioritises legal domain datasets) and Microsoft Copilot's approach (which leans on general knowledge) is real, but both require the same thing: a firm that can prove it validated the tool before deployment, monitored its performance in live work, and documented why it trusts the output. Trovix Watch and Trovix Audit exist because that governance layer is missing from the product marketing. The 90.9% score is the easy part. The hard part—the part Anthropic does not sell—is knowing when and why your firm can rely on it.
If you are a mid-market law, insurance, financial services or accountancy firm considering these new plugins, do not ask whether Claude scores well on a benchmark. Ask: can we audit every output? Do we understand the failure modes? Can we explain to our regulator why we chose this over alternatives? Can we quantify the cost of a hallucination in our business? If the answer to any of these is no, you are not ready to deploy—regardless of the benchmark score. Start with a governance framework: Trovix Audit gives you the dashboard to monitor AI outputs in live matter work, so you can see failure modes before they become complaints. Then deploy capabilities. Then publish your confidence intervals. The firms that will win the next three years are not the ones that adopt AI fastest. They are the ones that adopt it most defensibly.
Source: Fortune