Bye-bye, Bluebook? Automating Legal Procedure with Large Language Models
2025/05/05 by Matthew Dahl, Dahl, Matthew · 5 voices
Social Sciences · #Artificial Intelligence in Law #Comparative and International Law Studies
paper · pdf · doi:10.48550/arxiv.2505.02763
Abstract
Legal practice requires careful adherence to procedural rules. In the United States, few are more complex than those found in The Bluebook: A Uniform System of Citation. Compliance with this system's 500+ pages of byzantine formatting instructions is the raison d'etre of thousands of student law review editors and the bete noire of lawyers everywhere. To evaluate whether large language models (LLMs) are able to adhere to the procedures of such a complicated system, we construct an original dataset of 866 Bluebook tasks and test flagship LLMs from OpenAI, Anthropic, Google, Meta, and DeepSeek. We show (1) that these models produce fully compliant Bluebook citations only 69%-74% of the time and (2) that in-context learning on the Bluebook's underlying system of rules raises accuracy only to 77%. These results caution against using off-the-shelf LLMs to automate aspects of the law where fidelity to procedure is paramount.
Citations
Discussions
- New paper: Matthew Dahl, Bye-bye, Bluebook? Automating Legal Procedure with Large Language Models arxiv.org/pdf/2505.02763 #LegalWriting #Citation #LegalTech [bsky, 9 points, 1 comments]
- For what it’s worth, there’s a recent paper that benchmarks Bluebook citation formatting on frontier models and finds compliant/accurate citations only ~70% of the time. arxiv.org/abs/2505.02763 [bsky, 8 points, 1 comments]
- Ai "produces fully compliant bluebook citations only 69-74% of the time" bro that's far more than any human does and far far more than one ought to aspire to, this is a brief for immediately putting i [bsky, 6 points, 0 comments]
- Someone posted this link over at the other place— study on LLM ability to master the blue book, which is . . . Not great. Interesting. arxiv.org/pdf/2505.02763 [bsky, 3 points, 2 comments]
- ChatGPT for Bluebooking? Study shows that frontier models get citations right less than 75% of the time, raising questions on ability of LLMs to automate procedural tasks arxiv.org/abs/2505.02763 [bsky, 0 points, 0 comments]
Related