Document Type

Article

Publication Date

2026

Publication Title

UNT Dallas Law Review On the Cusp

Abstract

Can AI replace human jurors? More specifically, can large language models predict how jurors interpret evidence and reach decisions based on legally salient facts and demographic characteristics? As legal scholars and practitioners increasingly explore AI-generated jury simulations, this Article offers the first empirical test of whether models like GPT-4, Claude, and Gemini can faithfully replicate juror reasoning. The answer, for now, is no. Across a series of mock trial scenarios involving redacted confessions, GPT- 4, Claude, and Gemini repeatedly failed to replicate how real jurors interpret evidence or exercise judgment. Their errors were not random, but systematic. Hidden prompts, built-in content filters, and demographic flattening produced distortions that cut across sex, ethnicity, political affiliation, economic status, and education level.

Yet the promise of simulation remains within reach. In the second phase of the study, we fine-tuned an open-source model on actual mock juror data, achieving significant gains in accuracy and alignment. Although today’s LLMs fall short of simulating juror reasoning, models refined through transparent methods and real human data could assist judges in applying evidentiary standards, help researchers test doctrinal assumptions, and give trial lawyers new tools for strategic decision-making. This paper maps the risks and outlines a path toward responsible AI-based jury simulation.

Volume

8

Issue

2

First Page

73

Last Page

131

Share

COinS