Anthropic's new legal AI plugins are technically impressive but deliberately ignore the governance problem. Benchmark scores don't equal regulatory compliance—and hallucinations are still hallucinating.
AI Governance  Trovix SiftLegal · Financial Services · Insurance

Anthropic's announcement in May—20 new integrations and 12 role-specific plugins for M&A, employment law and due diligence, wrapped around Claude Opus 4.7's 90.9% score on Harvey's BigLaw Bench—is a genuine technical achievement. For mid-market UK regulated firms subject to the FCA's Consumer Duty PS22/9 and the SRA's Code of Conduct for Solicitors, this matters directly: if your competitors deploy AI-assisted due diligence or document review faster, and you don't, your unit economics and matter turnaround will suffer. But the timing of this release is troubling. The same month Anthropic shipped these plugins, hallucinations were appearing in live legal filings. The gap between what AI can do and what it should be allowed to do in a regulated context remains unfilled.

What this story reveals is a pattern we see everywhere in professional services AI right now: vendors chase benchmark scores and workflow integration depth as if those two things alone guarantee safe deployment. They don't. Harvey's BigLaw Bench is useful but measures accuracy on curated test sets—not resilience to edge cases, not transparency when the model is uncertain, not compliance with data residency requirements or the UK GDPR's Article 32 standards. The legal sector is now in a technology arms race where the fastest movers win clients, and the slowest get left behind. But speed without governance kills reputation. The same firms buying Anthropic plugins are simultaneously exposed to SRA disciplinary risk and FCA complaints if those plugins fail quietly.

Trovix's view is clear: integration breadth without data discipline is professional malpractice in the making. The real problem isn't whether Claude can score 90% on a benchmark. It's whether your firm knows what data it's feeding the model, where that data lives, who can see the model's reasoning, and what happens when it fails. Most Big Law implementations treat AI as a black box intelligence layer bolted on top of existing workflows. That's backwards. Before you deploy Anthropic plugins—or Harvey, Legora, Luminance, Microsoft Copilot for legal or any other product—you need document intelligence that understands what you're actually working with: Trovix Sift does this. You need a disciplined knowledge layer that surfaces confidence scores and source citations so fee-earners can verify the model's work: Trovix Aria does this. Without those foundations, you're just moving the risk around.

If you're a mid-market law firm, insurer, financial services firm or accountancy practice right now, don't wait for your Big Law competitors to hit a regulatory wall before you act. Don't assume that a 90% benchmark score translates to safe, auditable, compliant deployment in your firm. Instead: (1) map what sensitive data you're currently processing in AI workflows, (2) identify which plugins and integrations require explicit client consent under your engagement terms, (3) build a human review gate for any output that will be shared with third parties or filed in court, (4) document your AI governance approach in writing and share it with your audit and compliance teams. The SRA Code requires competence and honesty. That means knowing what your AI is doing and why. The firms that win the next three years will be the ones that moved fast *and* stayed compliant, not the ones that moved fastest.

Source: Fortune

Related Trovix product:

Trovix Sift →Book a demo →