Private Intelligence: Will On-Prem Legal AI Strategy Work?
AI for legal work faces cost, bias, and ROI challenges. On-premise AI is expensive, but cloud AI has unpredictable costs as well. I see 3 strategies evolving where vendors prep AI for legal work:
-
training open models
-
orchestrating multiple models in a harness and /or
-
designing benchmarks, evals and outputs to deliver work product.
Which of these strategies will work?
Harvey has Tenet and Thomson Reuters (TR) has Thomson One. Training Thomson One costs TR $40 million and 20 months. With new models appearing roughly every 43 days, this is an expensive endeavor. That's what I said when Tenet was released. However, the world needs AI to run locally and regionally if we want democracy. A free society needs decentralized AI. So I started working on Sabaio in 2023 and showed it to a few friends.
In 2024, I stopped Sabaio because running on-premise AI thus Private Intelligence is expensive. I would be stuck trying to convert a small , demanding market of top‑tier firms. I moved on to work on Private AI plus AI Governance with the Dutch government. Now in 2026, Kirkland commits $500 million, Latham buys GPU's and the Sabaio LinkedIn page grew followers. So what changed? Private AI remains expensive, but cloud AI is expensive and unpredictable.
In June 2026, I calculated that Harvey would go broke serving state-of-the-art (SOTA) models. In September Harvey admitted going to negative 50% gross margin due to the explosive growth in tokens use. Now one might say that taking a subscription at cloud or one of the major vendors would make costs more manageable. However, usage-based pricing is severely unpredictable for customers in the legal sector. When Westlaw switched legal research in 2021 from fixed-fee subscriptions to usage-based pricing, the backlash was immense. Legal business relies largely on transactional work so the key question: will AI on legal work deliver a return on investment?
100% AI on Legal Work
Each of the three strategies has a different return on investment (ROI). You do not need to grammar check your emails with Fable 5 from Anthropic. Why? It costs $15 per million tokens and was designed to complete a S-10 filing unattended. You may not be allowed to let Anthropic read your emails so opt for Palantir Technologies to check grammar. Still overkill, but if you are in a highly secure Defense environment, you may have little choice. Hence, the AI ROI is a tricky question which depends on context and sensitivity.
Back in 2007, I made a mind map to explore how any legal professional operates. I found that every legal case started with a single question: is this case unique? Since legal operates in silos of jurisdictions, practice areas, and knowledge fiefdoms, we are forced to perceive cases as unique. We do so because our tools fail to find similar cases in either other jurisdictions, areas, or your peers' case folders.
Let's say you have solved your case. Then delivery starts, e.g., drafting with custom advice and sharing for negotiations or litigation. If you're in-house, the administration is slightly different since you do not bill by the hour. Another distinction for in-house is that the facts in one company more or less remain stable over time, even while the legislative landscape may shift. For example, Google would not start manufacturing shoes as their main product.
20% Legal AI Use Cases
So where does AI not work yet? Cases where data exceeds a 1 million context window and we need to ask particular questions like: Is this company worth $11.5 Billion? This was the case with HP v Autonomy, which I work out here: is buying Autonomy for $11.1 billion by HP a good deal? Now a lawyer in Mozambique would not get the call from Elon to help negotiate the acquisition of Twitter. Even though with powerful AI, all lawyers are now equal in access to knowledge and, of course, coding skills.
The Twitter example is not by chance but chosen: if Elon could prove over 20% of Twitter users were bots, we would be living in an entirely different world right now. That was another big data legal problem no technology could solve at that moment in time. Primer for newbies: there are big non-legal data problems for legal and big legal data problems. Meaning access to codes, cases, and commentary (3C) on law. Commentary, broadly speaking, also includes the policies and rules not publicly available from governments or inside companies. One needs all 3 C's to be in your AI in order for it to be legally accurate in every language and jurisdiction.
Now even if we have applications, access, and accuracy, there is still this little thing called legal objectivity or bias. That thing that lawyers call 'quality', a scary word to me. Quality is hard to measure hence lawyers like using it as leverage. Is it the win rate of a litigation lawyer? Is it the creativity in extracting maximum contract value? Or is it just access to more legal data than your counterparty?
1% Legal ROI
Which brings me to little nuance which actually bring value and can use a specific legal AI strategy. On February 13, 2024, Nvidia released chat with RTX and journalists used it to use on local legal AI research on court cases by journalist. A more specific use case came from Casemark where a lawyer won $20 million settlement in Ku Klux Klan case. This was peculiar: no model would work on such a racially sensitive matter. Only after jailbreaking an older Llama model could the lawyer use AI for depositions.
More recently, Ted Theodoropoulos noted Chinese models like DeepSeek, Kimi, and Qwen show political refusal and bias patterns, around China-related topics. Making these models unreliable around contract drafting with Chinese counterparties. By the way, refusals are not only a Chinese‑model problem. American models also refuse due to safety concerns. Chief of Staff complained about this on X.com. Legal professionals working on biotech patents, chemical regulations, export-control compliance, pharmaceutical contracts, cyber-liability clauses, or privacy/surveillance law can hit refusals. Or they face silent fallbacks to weaker models.
I personally cannot testify to any of the accusations above because I have not tested the models for these behaviors. This brings me to my last and perhaps most important challenge in figuring out returns from AI use in legal. While I love all above, I cannot rely on any without a trusted benchmark. As I said to Anna Guo at The Beacon Collective this September that her legal benchmark work is crucial to our profession.
0% AI Strategy
Recap: legal work does not always require AI to be useful. A point further enforced by products such as Google's NotebookLM and, more recently, Meta's Muse. Confidentiality has driven some to spend big on GPU's or on-premise AI respectively Latham and Kirkland. Moreover, AI hallucinations makes traditional search still more reliable. Curious to know how pervasive hallucinations still are? Check Kyle Bahr's database.
New AI innovations like Jev offer alternative solutions for related problems like classifying large datasets. In some industries, like defense, I recommend a zero-cloud AI policy. Other heavily regulated industries should carefully assess when AI is safe and useful to apply. Now some of you eagle-eye readers noticed I did not address the 3rd strategy above. Well, I ran out of reading time (my max is 8 minutes) so I promise to address AI design at the Monocle launch.
In closing, I recommend legal professionals generate a map of activities before deciding on what AI works for them. Measure the value of the activities to gauge if the action requires AI. If so, what level of effort is required from AI?
Here is how: below are three cards you can click to generate an AI Activity Map™ for a quick check. There is also a more personalized AI Orchestration Map™ which looks at your practice area, industry, operation size, and existing tools and services. Feel free to try or leave feedback.







