The New Arbitrator-Selection Problem in the Age of AI: Choosing Which Model Decides Your Dispute

AI and arbitrator

Artificial intelligence (“AI”) is usually described as the next frontier for dispute resolution: a technology that will make adjudication fast, cheap, and consistent. Feed the facts and the applicable rules to a large language model, and out comes a verdict, free from the cost, delay, and idiosyncrasies of human arbitration.

The reality is more complicated, and the complication starts with a simple question that the “AI as arbitrator” narrative tends to skip: which AI to use? To answer that question, we put it to the test on a real corpus of disputes.

 

The Experiment Design

The source of our corpus is Lemon, a cryptocurrency exchange from Latin America. As in any fintech company, it has customer disputes about unauthorised withdrawals, failed transfers, and disputed charges. For a portion of these claims, they use Kleros, a decentralised dispute resolution system. Claims are submitted to the platform and resolved by a randomly selected panel of crowdsourced community jurors.

With members of the Kleros research team, we took a sample of 99 cases and put them through several tests to see how two different large language models (“LLMs”) would solve them, favouring either the consumer or the company.

Both went through an identical set-up: the same instruction files, the same process, and the same decision rules the human jurors had been given.

 

First Finding: Same Facts, Different Temperaments

The first comparison was between ChatGPT 5.5 and Claude Opus 4.7, each acting as a juror. We found a striking difference in disposition.

Out of the 99 cases, ChatGPT ruled for the consumer in only three; Claude ruled for the consumer in 14. The latter is closer to what the human jurors themselves produced. In other words, different AI models can have opposite leanings on who wins.

 

Second Finding: The Same Model Does Not Stay the Same

The second part of the experiment held the model family constant and varied only the version. We ran the identical 99 disputes through two successive versions of Claude: Opus 4.7 (which we used in the first comparison), and the newer version Opus 4.8.

The two versions produced the same decision on 88 out of 99 cases, an 88.9% agreement rate. That sounds reasonably high, until one looks at the 11 disagreements: ten moved from a consumer win to a platform win, and only one moved the other way. The platform’s overall win rate rose from 86% to 95%.

Reading the models’ reasoning on the flipped cases shows this was not random drift. The newer version applied a stricter and more consistent theory of where the burden of proof sits: it required claimants to prove their case on the documentary record actually submitted and read ambiguous contract terms in favour of the platform rather than the consumer. The same logic applied in both directions, which is why the one case that moved toward the consumer did so.

For instance, in one disputed card charge, a user challenged a roughly USD 60 payment to an online processor made on a Lemon card that had been dormant since 2022, noting that Lemon had produced no authentication evidence linking the charge to them.

The older version of Claude read this as a strong prima facie case of fraud and ruled for the user. The newer version accepted that Lemon’s defence was weak but held that a weak defence does not discharge the user’s own burden of proof. The charge had gone to a legitimate processor which, the model reasoned, was more consistent with a third-party data breach than with a proven failure on Lemon’s part.

 

Why This Matters: Picking the Decision-Maker, Again

Many legal technologists envision an expanding role for AI in arbitration, driven by the promise of speed and cost reduction. It’s worth thinking carefully — if only as a thought experiment for now — about what that trajectory actually implies for how disputes get decided.

If we put together the findings from the previous sections, we see a familiar problem from traditional arbitration reemerging in a new form. In commercial arbitration, a significant share of the effort (and the cost) goes into selecting the tribunal, because different arbitrators, applying the same law to the same facts, can and do reach different results. Reputations, prior awards, and known tendencies all factor into which names end up on a shortlist.

If AI comes to play a meaningful role in dispute resolution, the same dynamics that drive arbitrator selection will apply, just transformed into a model-selection problem.

A company that knows a particular model version tends to rule in its favour has every incentive to specify that model in its dispute-resolution clause. A consumer advocacy group, aware of the same pattern, has every incentive to push back. And because model versions update on a release cycle the parties do not control, even a clause that names a specific model today may behave differently in a few months.

The technical finding, that different models and different versions of the same model produce systematically different outcomes, is observable today, not in some hypothetical future. What remains to be seen is the institutional response: whether and how parties, institutions, and regulators choose to treat model selection as a procedurally significant choice.

 

Panels, Not Single Models

Historically, the arbitration community sought to solve the arbitrator-selection problem by building practices and institutions around it: lists of approved arbitrators, disclosure obligations, challenge procedures, and the use of tribunals rather than sole arbitrators.

The same toolkit, suitably adapted, applies to AI. Approved model lists would give parties and institutions a vetted set of options with known characteristics. Disclosure obligations would require that the model used and its version be identified in the procedural record, just as an arbitrator’s identity and any relevant conflicts must be. Challenge procedures would give parties a mechanism to object when a model’s known disposition creates a reasonable suspicion of imbalance.

And, perhaps most importantly, a “tribunal” of several models rather than a single model would distribute the decision across architectures with different training histories and default tendencies, doing what tribunals of human arbitrators have done for decades: average out idiosyncratic leanings and make the overall result more legitimate.

The analogy to a jury is perhaps more apt than the analogy to a three-member arbitral tribunal: because AI models are far cheaper than human arbitrators, a panel of several (maybe even dozens!) models becomes economically viable.

In that sense, AI dispute resolution may end up resembling not the tribunals of modern commercial arbitration, but the popular courts of ancient Athens, where hundreds of citizens selected by lot would hear a case. The Athenians understood that a large panel chosen through a robust random process was harder to game than any individual judge.

The arbitration community may be rediscovering the same insight for the long-term architecture of an algorithmic adjudication system.

 

Calculemus… But Whose Calculator?

The promise that opened this piece is not new. Philosopher and mathematician Gottfried Leibniz spent much of his career chasing a version of it: a universal symbolic language precise enough that any dispute could be settled by computation.

But this experiment shows that the dream has not come true. Leibniz’s calculus assumed that any two competent calculators would reach the same answer. Here, two AI models computing over the same case file systematically reach different answers, and the same model family reaches different answers across versions.

None of this is an argument against AI in dispute resolution. But those who envision a future of fast, automated arbitration might benefit from looking at why due process developed the way it did. The rules around arbitrator selection, disclosure, and challenge procedures were responses to a recurring problem: whoever controls the decision-maker controls the outcome. AI does not dissolve that problem. It compresses it into a model-selection decision made once, invisibly, at the system-design stage, long before any individual dispute arises.

This will confront dispute system designers with questions that current frameworks are not yet equipped to answer. How many models should resolve a case, and how should the trade-off between panel size and cost be governed? From what pool should those models be drawn, and who maintains that pool (especially considering that the new version of a model family might have a different temperament)? Should model identity and version be disclosed as a matter of procedural right, the way arbitrator identity is? Who has standing to challenge a model, and on what grounds? And if large panels can statistically cancel out individual model bias, do we even need challenge procedures, or can we simply let the law of large numbers do the work that institutional safeguards currently perform?

These are not hypothetical questions. Leibniz dreamed of a calculus that would make them unnecessary. The experiment described here suggests we are not nearly close enough to stop asking them.

 

Federico Ast is the founder and CEO of Kleros, a decentralised dispute resolution platform. William George is director of research at Kleros. Robert Dean is a member of Kleros’ supervisory board. The data used in the experiment can be accessed here.

Comments (0)
Your email address will not be published.
Leave a Comment
Your email address will not be published.
Clear all
Become a contributor!
Become a contributor Contact Editorial Guidelines