Blog • Science

Benchmarking Litigation Outcome Prediction

A point prediction for a complex liability claim is mathematically meaningless. True litigation forecasting requires separating the extraction of text from the calculation of risk, delivering calibrated ranges rather than brittle guesses.

TL;DR — Stop evaluating litigation models on pinpoint accuracy. The correct metrics are calibration and coverage, which use conformal prediction to output bounded settlement ranges that accurately reflect venue volatility and social inflation.

A claims dashboard displays a projected settlement value: $843,211. The precision implies certainty. A claims executive reading that number will adjust a reserve or authorize a settlement offer based on it. Mathematically, that single number is a fiction. Litigation is an adversarial system governed by human volatility. Any model offering a point prediction for a complex liability claim is lying about its own confidence.

The industry obsession with prediction accuracy stems from a misunderstanding of probability. Vendors often benchmark their models using mean absolute error, measuring the average distance between the model's point prediction and the final settlement amount. This metric works for predicting factory output or crop yields. It fails catastrophically in litigation. Claim outcomes are not normally distributed. They are heavily skewed by social inflation, third-party litigation funding, and the rising frequency of nuclear verdicts. When a distribution is heavy-tailed, a point estimate collapses the exact uncertainty you need to measure.

The Boundary Between Reading and Reasoning

The recent fascination with large language models has blurred the fundamental distinction between generation and prediction. An LLM is a sequence generator. Its architecture is optimized to predict the next token in a string of text based on learned probabilistic associations. This makes it an extraordinary tool for reading. A modern language model can ingest thousands of pages of unstructured pleadings, medical records, and legal correspondence, extracting the date of a spinal fusion or identifying a history of substance abuse.

Reading a file is not the same as calculating its financial exposure. A sequence generator has no internal representation of mathematics, geometric distance, or causal inference. When forced to predict a settlement value, an LLM merely hallucinates a statistically plausible number based on text patterns. It cannot calibrate its confidence. It cannot provide a mathematical proof of its reasoning.

At Canotera, we enforce a strict neuro-symbolic boundary between these tasks. Generative AI does the reading. It parses the chaotic source documents and structures the facts into a standardized ontology. Then its job is done. The actual forecasting is handed off to separate mathematical and geometric machine-learning models. These models are trained strictly on large datasets of resolved cases with known outcomes. They operate exclusively on the structured facts, not the raw text.

Benchmarking Calibration and Coverage

If point accuracy is a structural failure, how do we benchmark a litigation model? The correct metrics are calibration and coverage. A calibrated model tells the truth about its own uncertainty. If it calculates a 30 percent probability that a claim will escalate to trial, we expect exactly three out of ten such claims to go to trial across a large validation set. Calibration ensures that a probability score is an empirical fact, not a subjective guess.

In financial forecasting, calibration is expressed through conformal prediction ranges. Instead of returning a single dollar figure, our mathematical models output a mathematically bounded settlement range. The model states, with a specified confidence level, that the claim will resolve between $600,000 and $1.4 million.

The width of this conformal range is a highly informative feature. It quantifies the precise volatility of the venue, the judge, the specific injury profile, and the plaintiff's counsel. If the facts are standard and the jurisdiction is predictable, the range tightens. If the case involves a novel mechanism of injury in a venue known for nuclear verdicts, the range widens automatically to maintain the guaranteed coverage level.

Claims executives can build actual strategy around a conformal range. You set an initial reserve against the upper bound to protect the balance sheet on day one. You allocate aggressive defense spending when the uncertainty range is exceptionally wide, knowing that investigation will collapse the variance. You detect escalation early when a newly ingested medical report pushes the lower bound of the prediction higher than your current reserve. You negotiate from hard data instead of gut instinct. A point estimate offers none of these operational levers.

Traceability and Honest Error Reporting

The final requirement for a valid benchmarking framework is traceability. A forecast is only useful if its logic can be audited by a domain expert. When our mathematical models output a reserve delta or flag an escalation probability, they also output the specific drivers behind the calculation. The model isolates the variables driving the risk, whether it is a specific comorbid condition, a shift in the plaintiff firm's recent trial history, or a pattern in comparable resolved cases.

Because we separate the reading layer from the prediction layer, these variables are linked directly back to the source documents. The user clicks a driver and views the exact paragraph in the medical file that triggered the risk calculation. This architecture enforces honest error reporting. If a forecast misses the mark, we isolate the point of failure. We determine whether the generative model missed a critical doctor's note or whether the geometric model misweighed the impact of a specific venue dynamic. You cannot debug a monolithic language model this way.

We do not benchmark models to prove they are omniscient. We benchmark them to ensure they map the boundaries of the unknown accurately. The goal of litigation forecasting is not to guess a magic number.

The goal is to price the uncertainty before the plaintiff does.

Want to talk to an executive?

Press, partners, investors, candidates — the inbox is monitored. Tell us who you are and we'll route it to the right person within two business days.