Anthropic's legal integrations are fast and impressive—but the judicial system is already punishing hallucinations. The conversation about AI in law has moved past whether it works to whether your firm can afford the professional indemnity claim when it fails.
AI Governance  Trovix AuditLegal · Insurance · Financial Services · Accountancy

Anthropic released 20+ legal plugins in May 2026, with Claude Opus 4.7 scoring 90.9% on Harvey's benchmark. This is objectively good. But it is not good enough for a mid-market law firm facing sanctions from the High Court, FCA enforcement action, or an SRA investigation into competence under the SRA Code of Conduct for Solicitors. The story is not the benchmark score. The story is that courts are already issuing sanctions for AI-generated errors in legal filings, and law firms are still moving into production workflows that depend on 90.9% accuracy. That gap between capability and regulatory accountability is where UK regulated firms face real risk.

This is not a story about Anthropic's engineering or Legora's or Harvey's. This is a story about the industry's collective choice to treat AI as a productivity tool first and a governance problem second. When benchmarks matter more than audit trails, when speed matters more than verifiability, and when a vendor's confidence in their own product becomes the primary risk management strategy, you get exactly this situation: impressive capability, mounting sanctions, and no clear owner of the liability. The pattern is clear across legal tech, insurance operations, and accountancy—vendors release products faster than firms can implement them safely, and insurance covers gaps rather than controls creating them.

Trovix does not believe 90.9% accuracy is a compliance statement. It is a starting point for governance. The firms managing AI safely right now are the ones with three things: first, they know where AI is being used and what it touches (matter data, client confidentiality, regulatory filings). Second, they have audit capability that goes beyond vendor benchmarks—they test their own models against their own workflows and document failures. Third, they can explain their AI decisions to the SRA, FCA or ICO without referencing a vendor's marketing material. This requires more than plugins. It requires Trovix Audit, which gives you real-time governance of AI systems in use, not theoretical performance on someone else's test set. Harvey's benchmark tells you what Claude can do. Trovix Audit tells you what Claude is actually doing in your firm.

Here is what a mid-market law firm should do this week: stop treating AI plugins as feature releases and start treating them as policy decisions. Before you deploy Anthropic's M&A due diligence integration or any other vendor's legal automation, run a three-part test. First, can you audit what the system did on a specific file? Second, can you explain why your firm chose this tool rather than a qualified solicitor for this task—and is the answer 'cost' or 'quality'? Third, what happens when it fails, and who tells the client? If you cannot answer these questions cleanly, do not install the plugin. The speed gain is not worth the professional indemnity exposure. This is not a Trovix sales point. This is basic SRA competence.

Source: Fortune

Related Trovix product:

Trovix Audit →Book a demo →