Anthropic's May release of 20+ legal integrations, including Harvey-branded plugins for M&A due diligence and employment drafting, arrives with a problem Big Law is choosing to ignore. Courts are already issuing sanctions for AI-generated false citations. The underlying Claude Opus 4.7 scores 90.9% on Harvey's BigLaw Bench — which sounds impressive until you realise a 9.1% error rate in legal citations is not a benchmark; it is a liability admission. For UK-regulated law firms operating under SRA Code standards and facing FCA Consumer Duty PS22/9 scrutiny, this is not a marginal issue. False citations in due diligence packs, contracts or regulatory submissions expose firms to professional negligence claims and disciplinary action. The story matters because it shows the market's preferred vendors are prioritising speed and integration breadth over the governance architecture that regulated firms actually need.
This pattern reflects a wider industry misconception: that AI accuracy improves linearly with model size and benchmark scores, so firms should adopt the best-performing tool available. Harvey, Luminance, and now Anthropic are racing to embed AI deeper into workflow-critical tasks. But the courts' hallucination rulings expose a harder truth. Legal work is not a task where 90% accuracy is defensible. A 9% hallucination rate might distribute randomly across thousands of documents, but when a single fabricated case citation ends up in a high-value transaction or regulatory filing, the firm — not the software vendor — faces the SRA, the FCA, and the judge. UK regulated firms are watching competitors adopt Claude, GPT-4, and purpose-built legal models, creating competitive pressure to move fast. That pressure is exactly when governance lapses occur.
Trovix's approach inverts the industry's implicit hierarchy. Instead of 'find the best-performing model, then add compliance guardrails,' we start with governance architecture and then connect the right models to it. Our Trovix Audit platform enforces traceability, confidence scoring, and output validation before any AI-generated content reaches a fee-earner or client. For document-heavy work like M&A due diligence, Trovix Sift uses hybrid extraction — combining Claude, open-source models, and symbolic logic — with human-in-the-loop gates on high-risk findings. This is less glamorous than embedding a 4.7-generation LLM directly into a plugin and releasing it. It is also the difference between defensible AI and AI that courts are already penalising. Luminance and Harvey are optimising for speed. We optimise for the question a court will ask: 'How did you validate this output before it left your firm?'
If you are a mid-market law firm, accountancy practice, or financial services firm evaluating the Anthropic release or similar products, do three things now. First, audit how you currently validate AI-generated citations, data extractions, and factual claims — if your answer is 'we trust the model,' you are already exposed. Second, demand governance documentation, not just accuracy benchmarks, from any AI vendor. ISA UK (UK 260) requires audit committees to understand material AI risks; vendors must explain their validation and rollback procedures. Third, run a pilot on non-critical work before deploying to client-facing outputs. The firms issuing sanctions against themselves are the ones that deployed Big Law's newest tools to their biggest deals without an intervening validation layer. Competitive pressure to adopt is real. Regulatory reality is harder.
Source: Fortune