Skip to content

By Michael Blanding

We even have a word to describe law’s complexity: legalese. But where did such language come from and why does it persist? “Past studies asked lawyers their opinions about these laws, and lawyers don’t like them either—they also have difficulty reading legalese,” says cognitive scientist Sihan Chen PhD ’26. “So why do lawyers write this way?” Chen has looked to a surprising source for the answer: ancient Chinese law.

In a unique collaboration between MIT’s Brain and Cognitive Sciences department in the School of Science and the History section of the School of Humanities, Arts, and Social Sciences, Chen took the lead on scientifically analyzing the complexity of medieval and modern Chinese legal writing and comparing it to American legalese. As a result, he and colleagues determined that the complexity of legal writing may not be so fixed after all. “The Chinese case shows that legal complexity isn’t set in stone,” says Tristan Brown, associate professor of history and co–principal investigator on the project. “How lawmakers write laws is, at least in part, a choice.”

56

Since MITHIC’s inception, it has provided $4.7M in funding to 56 projects across 26 MIT units.

The project, “A Cross-Cultural Study of the Complexity of Legal Language,” was supported by a grant from the MIT Human Insight Collaborative (MITHIC), a Presidential Strategic Initiative that is elevating human-centered research and teaching, and bringing together scholars in the humanities, arts, and social sciences with their colleagues across the Institute. The study was co-led by Brown and Edward “Ted” Gibson, professor of brain and cognitive sciences and Chen’s advisor.

Unresolved questions in language development

Various hypotheses explain how legalese developed. One possibility is that complicated language is a historical accident: a relic of flowery royal language dating back to the Code of Hammurabi that has been repeatedly copied throughout Roman and English law. An alternative possibility is that something about the content of legal language, which requires detailed descriptions for offenses and punishments, makes it convoluted.

“A transparent way to express conditional relationships in law is to use conditional statements,” Chen explains. “What’s more, the preconditions and punishments in a law are usually complex events themselves.” To examine that complexity, says Chen, cognitive scientists can measure something called syntactic dependency, which takes related words in a sentence, such as the subject and verb, and determines how far they are separated by other clauses, increasing comprehension difficulty. Comparing Chinese law to American law, the researchers surmised, could help show whether linguistic complexity is due to complexity in legal concepts or is a historical accident. “It’s a system with very different origins from the legal tradition of the West,” Chen says. “We wanted to see if it has the same structure or not.”

Intrigued by the language of music

Chen has been fascinated by language since singing in a choir in China at age 9, when he saw how different sounds are used in different languages. He studied mechanical engineering at the University of Miami but was inspired by linguistics classes to switch fields, examining language as a code of communication that could be quantitatively analyzed to see how it varies across culture. “It’s possible that the commonalities we see in language say something about the human mind or physiology, while the differences show how different cultural and societal factors drive language in different ways,” he says.

In 2020, he joined TedLab, Gibson’s laboratory investigating human language and communication. Central to Chen’s work is the delicate balance between efficiency (expressing concepts in the simplest words possible) and accuracy (expressing them in the most specific form possible). Legal language tends toward the latter, sacrificing simplicity for specificity. Previous research by Ted Lab colleagues Eric Martinez PhD ’24 and Frank Mollica measured the long-distance dependencies in legalese to show how it was much more complicated than everyday language.

Chen applied the same methods to the Tang Code, the legal code of Imperial China’s Tang Dynasty, dating to the 7th century. By comparing the legal language to other writings of the time, he was able to show that the medieval code was substantially more complex than everyday prose of its era. When he applied the same methods to modern Chinese legal language, however, he found something quite different. In the decades after 1949, and especially after economic reforms began in 1978, the People’s Republic of China (PRC) rebuilt its legal system largely from scratch, using much plainer language.

When he repeated the analysis, Chen found current mainland Chinese legal language is only modestly more complicated than contemporary nonlegal prose, including the language used in sources such as Wikipedia. “Modern mainland Chinese law is still more complex than everyday prose, but the gap is far smaller than in American legal English. The PRC case suggests that drafting choices can reduce legal complexity when accessibility is a goal,” says Brown, who worked with Chen to understand those historical aspects. “The PRC case is especially interesting because language reform and legal reform overlapped. Linguists were consulted on constitutional drafting in the 20th century, and in 2007 the [National People’s Congress] created an expert committee to review legislative language. That suggests legal complexity is ancient, but not inevitable: Institutions can act to reduce it.”

Chen notes that efforts to simplify legal language in the United States have largely failed, perhaps in part because of a misunderstanding of the problem. “We hypothesize that people didn’t realize this long dependency as the core problem in comprehension difficulty,” he says. “If you don’t know what makes something difficult to read, you cannot fix it effectively.”

Chen plans to continue to examine this phenomenon, designing experiments to further test what makes legal language difficult to understand, and how it might be simplified. “My goal is to figure out what’s an effective way to communicate these different concepts,” he says. The research might have implications that help governments write better laws, perhaps one day giving legalese the boot for good.


SUPPORT MITHIC

Sihan Chen’s research was funded by the MIT Human Insight Collaborative (MITHIC), a Presidential Strategic Initiative designed to identify and elevate scholars at the frontiers of human-centered research and education, providing them with resources to pursue their most innovative ideas. The initiative is led by Seth Mnookin, professor of science writing and head of Comparative Media Studies/Writing, and Heather Paxson, the William R. Kenan, Jr. Professor of anthropology and associate dean for faculty for the School of Humanities, Arts, and Social Sciences “MITHIC affirms that humanistic inquiry, grounded in interpretation, creativity, and ethical reasoning, is central to addressing the most consequential challenges facing society today,” says Mnookin. Make your gift to MITHIC.

Tell Us What
You Care About