Using Topic-Modeling in Legal History, with an Application to Pre-Industrial English Case Law on Finance

Authors

  • Peter Grajz Department of Economics, The Williams School of Commerce, Economics, and Politics, Washington and Lee University , CESifo, Munich, Germany Author
  • Peter Murrel Author

Keywords:

Topic modeling Legal history Machine learning English case law Law and finance Common law and equity

Abstract

Abstract

We argue that topic-modeling, an unsupervised machine-learning technique for analysis of large corpora, can be a powerful tool for legal-historical research. We provide a non-technical introduction to topic-modeling driven by the presentation of an example of how researchers can use the data that topic-modeling produces. The context of the example is pre-industrial English caselaw on finance. We generate new insights on the timing of pertinent legal developments, the linkages of law on finance to other areas of law, and the relative importance of common-law and equity in the emergence of law and legal ideas relevant to finance. We argue that topic-modeling has the potential to bridge traditional legal history and economics, increasing the influence of the former on the latter, which is overdue. The output of topic-modeling includes the data required to generate a quantitative macroscopic overview of the flow of legal history. These data can be used in many ways in subsequent legal-historical research. Epistemologically, topic-modeling offers an escape from the temptations of Whig history and opens up new avenues for inductive analysis characteristic of traditional historical research.

References

Robertson, S., “Searching for Anglo-American Digital Legal History,” Law and History Review 34 (2016): 1047–69CrossRefGoogle Scholar, noting that “as the fields of digital humanities and digital history have grown in scale and visibility since the 1990s, legal history has largely remained on the margins of those fields.” There are some important very recent examples for the nineteenth century, such as Funk, K. and Mullen, L.A., “The Spine of American Law: Digital Text Analysis and U.S. Legal Practice,” American Historical Review 123 (2018): 132–64CrossRefGoogle Scholar. In recent years, a number of empirical articles are appearing that use data from the eighteenth century made available by the Old Bailey Proceedings project. See T. Hitchcock, R. Shoemaker, C. Emsley, S. Howard, and J. McLaughlin, “The Proceedings of the Old Bailey, 1674-1913” www.oldbaileyonline.org (accessed April 2021). Both of these works rely on the types of computational advances that we highlight in this article and that we feel will lead to a quiet revolution in legal historical studies. Existing, more traditional studies on the period before the nineteenth century usually contain very small samples or few variables, implying that there is a limited ability to apply the types of empirical methods that are now commonplace in economic history. Cavell, E., “The Measure of Her Actions: A Quantitative Assessment of Anglo-Jewish Women's Litigation at the Exchequer of the Jews, 1219-81,” Law and History Review 39 (2021): 135–72CrossRefGoogle Scholar provides a recent example of a very interesting exercise in early legal history that is, understandably, limited by a small sample with few variables. Klerman, D., “Settlement and the Decline of Private Prosecution in Thirteenth-Century England,” Law and History Review 19 (2001): 1–65CrossRefGoogle Scholar is notable in providing a very early example of pre-industrial legal history that is exceptional for the centrality of empirical methods in its contribution.

2

The general problem is usefully captured as “How do you write a national history that was the product of lawmaking in 50 separate jurisdictions?” as posited by Nystrom, E. and Tanenhaus, D., “The Future of Digital Legal History: No Magic, No Silver Bullets,” American Journal of Legal History 56 (2016): 150–67CrossRefGoogle Scholar. This problem is multiplied in case law where one is studying hundreds of years and thousands of cases. The methods we describe in this article almost completely remove the sample-size and limited-observations constraint referred to in the previous footnote. Notably, the general project that includes the current article did not rely on any extramural funding, emphasizing that the techniques we describe are within the reach of all scholars.

3

See, for example, Grimmer, J. and Stewart, B. M., “Text as Data: The Promise and Pitfalls of Automatic Content Analysis Methods for Political Texts,” Political Analysis 21 (2013): 267–97CrossRefGoogle Scholar; Gentzkow, M., Kelly, B., and Taddy, M., “Text as Data,” Journal of Economic Literature 57 (2019): 535–74CrossRefGoogle Scholar; and Livermore, M.A. and Rockmore, D. N., eds., Law as Data: Computation, Text, and the Future of Legal Analysis (Santa Fe: SFI Press, 2019)CrossRefGoogle Scholar.

4

In legal history, two early examples of the use of the new sets of computational tools are provided by Tanenhaus, D. and Nystrom, E., “Let's Change the Law: Arkansas and the Puzzle of Juvenile Justice Reform in the 1990s,” Law and History Review 34 (2016): 957–97CrossRefGoogle Scholar; and Romney, C., “Using Vector Space Models to Understand the Circulation of Habeas Corpus in Hawai'i, 1852–92,” Law and History Review 34 (2016): 999–1026CrossRefGoogle Scholar. In contrast to the exercise reported in this article, these two examples do not use the computational methods to drive an empirical exercise but rather use these methods as search procedures to find those legal materials on which a more traditional analysis should be focused.

5

On the popularity of topic-modeling, see Gentzkow et al., “Text as Data” and J. Guldi and B. Williams “Synthesis and Large-Scale Textual Corpora: A Nested Topic Model of Britain's Debates over Landed Property in the Nineteenth Century,” Current Research in Digital History 1 (2018), https://doi.org/10.31835/crdh.2018.01 (accessed April 2021) Several other machine-learning and related computational approaches have been utilized to investigate law-as-data. Machine-learning methods have been used, for example, to predict court outcomes; see for example, D.M. Katz, M.J. Bommarito, and J. Blackman, “A General Approach for Predicting the Behavior of the Supreme Court of the United States,” PLoS One 12 (2017): e0174698. Word and document embedding models represent words and documents as numerical scores for a long list of variables, thereby helping to quantify the meaning of words and documents on the basis of their proximity to other words and documents in the corpus; see, for example, E. Ash and D.L. Chen, “Case Vectors: Spatial Representations of the Law Using Document Embeddings,” in Law as Data, ed. M.A. Livermore and D.N. Rockmore (Santa Fe: SFI Press, 2019), 313–37. Embedding approaches have been employed, for example, to investigate the presence of racial bias in judicial opinions; see, for example, D. Rice, J.H. Rhodes, and T. Nteta, “Racial Bias in Legal Language,” Research & Politics April-June (2019), 1–7. For an overview of the use of machine-learning and computational methods in the emerging research field of computational analysis of law-as-data, see J. Frankenreiter and M.A. Livermore, “Computational Methods in Legal Analysis,” Annual Review of Law and Social Science 16 (2020): 39–57. For innovative applications of computational methods to legal-historical themes, but not focusing on English case law, see, for example, S. Klingenstein, T. Hitchcock, and S. DeDeo, “The Civilizing Process in London's Old Bailey,” Proceedings of the National Academy of Sciences 111 (2014): 9419–24 and Funk and Mullen, “The Spine of American Law”.

6

J. Grimmer, M. E. Roberts, and B. Stewart, “Machine Learning for Social Science: An Agnostic Approach”, Annual Review of Political Science 24 (2021): 395–419.

7

R. Harris, “The Encounters of Economic History and Legal History,” Law and History Review 21 (2003): 297–346 identified this separation of these fields, and his conclusions still seem to apply today.

8

On these points more generally, see S. Robertson and L. Mullen, “Arguing with Digital History: Patterns of Historical Interpretation,” Journal of Social History 54 (2021): 1005–22, who argue that “Digital history has only rarely contributed interpretative or argumentative scholarship that contributes to disciplinary understandings of the past,” largely because of its focus on the methodological. Beyond the example appearing here, the use of the methods outlined in this article and of the data set discussed here are provided in several additional articles that contribute to disciplinary understandings of the past: See P. Grajzl and P. Murrell, “A Machine-Learning History of English Caselaw and Legal Ideas Prior to the Industrial Revolution II: Applications,” Journal of Institutional Economics 17 (2021): 201–16; P. Grajzl and P. Murrell, “A Macrohistory of Legal Evolution and Coevolution: Property, Procedure, and Contract in Pre-Industrial English Caselaw” https://dx.doi.org/10.2139/ssrn.4005612; and P. Grajzl and P. Murrell “Of Families and Inheritance: Law and Development in Pre-Industrial England” https://dx.doi.org/10.2139/ssrn.3975015

9

The useful distinction between close and distant reading arose among scholars of literature in what has become known as the “digital humanities,” in which debates about the usefulness of computational methods, particularly topic-modeling, were both early and very spirited. For the digital humanities, see, for example, A. Goldstone and T. Underwood, “The Quiet Transformations of Literary Studies: What Thirteen Thousand Scholars Could Tell Us,” New Literary History 5 (2014): 359–84. For history, see the very early study by S. Block “Doing More with Digitization: An Introduction to Topic Modeling of Early American Sources,” Commonplace 6 (2006), http://commonplace.online/article/doing-more-with-digitization/, accessed April 2006 and more recently J. Guldi, “Critical Search: A Procedure for Guided Reading in Large-Scale Textual Corpora,” Journal of Cultural Analytics 3 (2018), https://doi.org/10.22148/16.030, (accessed April 2021). For the same emphasis in political science, see Grimmer and Stewart, “Text as Data” and, in a joint product of a sociologist and two computer scientists, P. DiMaggio, M. Nag, D. Blei, “Exploiting Affinities Between Topic Modeling and the Sociological Perspective on Culture: Application to Newspaper Coverage of U.S. Government Arts Funding,” Poetics 41 (2013): 570–606. For legal history, see Robertson, “Searching”.

10

See P. Grajzl and P. Murrell, “A Machine-Learning History of English Caselaw and Legal Ideas Prior to the Industrial Revolution I: Generating and Interpreting the Estimates,” Journal of Institutional Economics 17 (2021): 1–19; and P. Grajzl and P. Murrell, “A Machine-Learning History of English Caselaw and Legal Ideas Prior to the Industrial Revolution II: Applications,” Journal of Institutional Economics 17 (2021): 201–16.

11

The work on English case reports is part of a much larger project on using computational and statistical techniques to understand English history. The earliest products of this project combined legal history and intellectual history, with two articles addressing understanding more general sets of ideas, focused on Francis Bacon and Edward Coke. See P. Grajzl and P. Murrell, “Toward Understanding 17th Century English Culture: A Structural Topic Model of Francis Bacon's Ideas,” Journal of Comparative Economics 47 (2019): 111–35; and P. Grajzl and P. Murrell, “Characterizing a Legal-Intellectual Culture: Bacon, Coke, and Seventeenth-Century England,” Cliometrica 15 (2021): 43–88.

12

The digitized copies of the English Reports were purchased from a publishing company domiciled in South Africa. It is beyond the scope of this article to provide the many details of the initial processing of these digital copies, and the cleaning of them. Suffice it to say that a very large proportion of GM's labor time devoted to the pertinent research projects was consumed in all of these tasks. For more details, see GM and M. Schmidt, “Institutional Persistence and Change in England's Common Law: 1700-1865” (PhD diss., University of Maryland, 2015). Because a central objective of GM was to include as many reports as possible, which necessarily implied computational processing of all reports, there was a need to exclude a small percentage of reports with too many words that did not have a counterpart in either modern English or standard Latin. Chiefly, this had the effect of excluding reports in Law French. There is no doubt that this is a blemish on the application of the computational methods. Initially, there were 60,249 pre-1765 reports in the data set, but 6,917 were dropped because they were in Law French and a further 383 were removed because they contained too many unrecognizable words. This left 52,949 reports.

13

The summary does not rely on any existing classifications: we return to this point in the conclusion.

14

See J. Guldi and D. Armitage, The History Manifesto (Cambridge: Cambridge University Press, 2014), emphasizing that “Over the last decade, the emergence of the digital humanities as a field has meant that a range of tools are within the grasp of anyone, scholar or citizen, who wants to try their hand at making sense of long stretches of time. Topic modeling software can machine read through millions of government or scientific reports and give back some basic facts about how our interest in ideas have changed over decades and centuries.”

15

D.M. Blei, A.Y. Ng, and M.I. Jordan, “Latent Dirichlet Allocation,” Journal of Machine Learning Research 3 (2003): 993–1022. One measure of the prominence of this contribution is that this is the seventh most cited article in computer science that was produced this millennium (https://citeseerx.ist.psu.edu/stats/articles). To be sure, there were a number of similar algorithms developed before the Blei et al. contribution, but early excitement about these methods seems to have focused on Blei et al., perhaps because of the accessible software developed for implementation. See A.K. McCallum, “MALLET: A Machine Learning for Language Toolkit,” (http://mallet.cs.umass.edu, accessed April 2021). In their note explaining this software, S. Graham, S. Weingart, and I. Milligan, “Getting Started with Topic Modeling and MALLET,” (https://programminghistorian.org/en/lessons/topic-modeling-and-mallet, accessed April 2021). state: “You will sometimes come across the term ‘LDA' when looking into the bibliography of topic modeling. LDA and Topic Model are often used synonymously, but the LDA technique is actually a special case of topic modeling created by David Blei and friends… . It was not the first technique now considered topic modeling, but it is by far the most popular…They all work in much the same way.” One such earlier algorithm was used in study by Newman and Block in the first history publication to use topic-modeling. See D. J. Newman and S. Block, “Probabilistic Topic Decomposition of an Eighteenth-Century American Newspaper,” Journal of the American Society for Information Science and Technology 57 (2006): 753–67.

16

See M.E. Roberts, B.M. Stewart, and E.M. Airoldi, “A Model of Text for Experimentation in the Social Sciences,” Journal of the American Statistical Association 111 (2016): 988–1003, whose general approach is very similar to that of Blei et al., but has an emphasis on incorporating document meta-information (such as date of publication) directly into the analysis. Small details would have changed had we used LDA, but we are sure the overall picture would have remained the same. For copious detail on the structural-topic-model approach to topic-modeling, including how to get started on implementation, see https://www.structuraltopicmodel.com (accessed June 2019).

17

On this point, see H. Bonin, “From Antagonist to Protagonist: ‘Democracy' and ‘people' in British Parliamentary Debates, 1775–1885,” Digital Scholarship in the Humanities 35 (2020): 759–75. One example using a topic-model-type method is L. Blaydes, J. Grimmer, and A. McQueen, “Mirrors for Princes and Sultans: Advice on the Art of Governance in the Medieval Christian and Islamic Worlds,” Journal of Politics 80 (2018): 1150–67.

18

Applications in sociology have also lagged, perhaps because the hypothetico-deductive method has had increasing sway in that field as well. For the lag in sociology, see N.C. Lindstedt, “Structural Topic Modeling for Social Scientists: A Brief Case Study with Social Movement Studies Literature, 2005–2017,” Social Currents 6 (2019): 307–18.

19

For example, S. Hansen and M. McMahon, “Shocking Language: Understanding the Macroeconomic Effects of Central Bank Communication,” Journal of International Economics 99 (2016): S114–S133.

20

Stephen Robertson emphasizes that text-analysis in history has been held back simply by the availability of a large stock of digital texts. S. Robertson, “The Differences between Digital Humanities and Digital History,” Debates in Digital Humanities (2016) (https://dhdebates.gc.cuny.edu/read/untitled/section/ed4a1145-7044-42e9-a898-5ff8691b6628#ch25m, accessed March 2021). This constraint is rapidly being relaxed. Indeed, one of the contributions of GM is to make machine readable, cleaned versions of the English Reports available for scholars in general. See GM and the concluding section of this article for more details.

21

For interesting articles of this kind, see A. Barron, J. Huanga, R. Spang, and S. DeDeo, “Individuals, Institutions, and Innovation in the Debates of the French Revolution,” Proceedings of the National Academy of Sciences 115 (2018): 4607–12; and A. Rule, J. Cointet, and P. Bearman, “Lexical Shifts, Substantive Changes, and Continuity in State of the Union Discourse, 1790–2014,” Proceedings of the National Academy of Sciences 112 (2015): 10,837–44.

22

As counterpoint to this apology for simplification, see S. Robertson, “Digital Humanities” in The Oxford Handbook of Law and Humanities , ed. S. Stern, M. Del Mar, and B. Meyler, (Oxford: Oxford University Press, 2019), emphasizing that “If humanities scholars chafe at such simplification, it is worth noting that narrative, the favored representational model of humanities scholars, is a deliberately simplified account that is illuminating because of, not despite, its simplification.”

23

The productive use of machine-learning to detect style was emphasized by Matthew L. Jockers, one of the most forceful advocates of machine-learning in the digital humanities; see M.L. Jockers, Macroanalysis: Digital Methods and Literary History (Urbana: University of Illinois Press, 2013).

24

J. W. Mohr and P. Bogdanov, “Introduction−Topic Models: What They Are and Why They Matter,” Poetics 41 (2013): 545–69.

25

DiMaggio et al., “Exploiting Affinities”.

26

One could instead view documents as collections of two- or three-word chunks, or even larger phrases. But the required processing power increases proportionately with the number of distinct phrases, which increases exponentially with the number of words allowed to be in a phrase.

27

See Robertson, “Searching”. On this point, see also P. Grajzl and P. Murrell, “Lasting Legal Legacies: Early English Legal Ideas and Later Caselaw Development During the Industrial Revolution,” Review of Law & Economics (2022), pre-publication online version, https://doi.org/10.1515/rle-2021-0070 (accessed April 16, 2022).

28

As evidenced by the large, related literature, computational scientists and statisticians usually emphasize rule-based criteria for model choice, relying solely on numerical information derived from the estimating process or the output data. In contrast, practitioners emphasize the element of subjective judgment, which would take into account the perceived quality of the topics reflecting both the uses to which they are to be put and the nature of the text data that is used in estimation. See, for example, Gentzkow et al., “Text as Data”; DiMaggio et al., “Exploiting Affinities”; and Mohr and Bogdanov, “Introduction−Topic Models.”

29

See Mohr and Bogdanov, “Introduction−Topic Models,” 560.

30

See Grimmer et al., “Machine Learning for Social Science,” stating: “Rather than place our trust fully in models and fit statistics, we argue that human feedback is essential for judging the quality of model results used for discovery.”

31

In order to distinguish our topic names clearly in the remainder of this article, we capitalize them.

32

An additional type of topic is identified by DiMaggio et al. who in “Exploiting Affinities,” argue that “Topic models often shunt noisy data into uninterpretable topics in ways that strengthen the coherence of topics that remain.” In fact, our experience is not that the topics are uninterpretable, per se, but rather that the interpretation means that the topic tells one nothing about the substantive inquiry in question. For example, GM find a topic that they call Non-Translated Latin. Sixteenth and seventeenth century lawyers not only had their own version of English, but their Latin was also highly idiosyncratic. The text preparation procedures were able to handle idiosyncratic English and standard Latin, but not idiosyncratic Latin.

33

This point is much emphasized in the literature, something to which we return more fully in the Conclusion. See, for example, DiMaggio et al., “Exploiting Affinities”; Mohr and Bogdanov, “Introduction−Topic Models”; A. Goldberg, “In Defense of Forensic Social Science,” Big Data & Society (2015), July-Dec: 1–3 and L.K. Nelson, “Leveraging the Alignment Between Machine Learning and Intersectionality: Using Word Embeddings to Measure Intersectional Experiences of the Nineteenth Century U.S. South,” Poetics 88 (2021): 101539, 1–18.

34

Since the topic-modeling algorithm produces unlabeled topics and since the data output from that algorithm can easily be transmitted, other researchers could easily produce their own set of names for the 100 topics. Indeed, doing so could inspire much more research that benefits from topic models. One research team produces the output data from the topic model, which can then easily be the input data for the work of other researchers.

35

Some economic historians have been very aware of detailed developments in the legal sphere, but it seems to be the case that such economic historians have had little effect on the perspectives on English legal history that are dominant in the mainstream of economic analysis, as exemplified in the works to be discussed in the ensuing paragraphs. R. Harris, in “The Encounters,” was early in making a case for productive exchange between legal history and economics, stressing that legal historians did not pay sufficient attention to the economic history literature. We are more concerned here with the lack of interchange in the reverse direction.

36

The word “finance” appears only twice in J.H. Baker, An Introduction to English Legal History, fifth edition (Oxford: Oxford University Press, 2019) and the pertinent issues are in separate discussions, included under property and contract.

37

Some salient critiques are N. Sussman and Y. Yafeh, “Institutional Reforms, Financial Development and Sovereign Debt: Britain 1690–1790,” Journal of Economic History 66 (2006): 906–35; P. Murrell, “Design and Evolution in Institutional Development: The Insignificance of the English Bill of Rights,” Journal of Comparative Economics 45 (2017): 36–55; L. Neal, “How It All Began: The Monetary and Financial Architecture of Europe During the First Global Capital Markets, 1648–1815,” Financial History Review 7 (2000): 117–40; P. O'Brien, “The Nature and Historical Evolution of an Exceptional Fiscal State and Its Possible Significance for the Precocious Commercialization and Industrialization of the British Economy from Cromwell to Nelson,” Economic History Review 64 (2011): 408–46; S. Ogilvie and A.W. Carus, “Institutions and Economic Growth in Historical Perspective,” in Handbook of Economic Growth, ed. P. Aghion and S.N. Durlauf (Amsterdam: Elsevier, 2014), 403–513; D. Coffman, A. Leonard, and L. Neal (ed.), Questioning Credible Commitment: Perspectives on the Rise of Financial Capitalism (Cambridge: Cambridge University Press, 2013); and G.M. Hodgson, “1688 and All That: Property Rights, the Glorious Revolution and the Rise of British Capitalism,” Journal of Institutional Economics 13 (2017): 79–107.

38

NW, “Constitutions and Commitment: The Evolution of Institutions Governing Public Choice in Seventeenth-Century England,” Journal of Economic History 49 (1989): 803–32.

39

LLSV, “Legal Determinants of External Finance,” Journal of Finance 52 (1997): 1131–50; and LLSV, “Law and Finance,” Journal of Political Economy 106 (1998): 1113–55.

40

A computational search of JSTOR reveals how unusual these two works are in their spread across the whole of economics. NW, “Constitutions and Commitment,” appears in the Journal of Economic History and is referred to in JSTOR thirteen times as often as the typical article published at the same time in that journal. The references to NW, “Constitutions and Commitment,” are twice as common in the journals outside economic history as in economic history journals, while for the typical article published in the same journal at the same time, the ratio is 0.6. Similarly, LLSV, “Legal Determinants,” appears in the Journal of Finance and is referred to in JSTOR twenty times as often as the typical article published at the same time in the same journal. The references to LLSV, “Legal Determinants,” are twice as common in the journals outside of finance as in finance journals, while for the typical article at the same time in the same journal, the ratio is 0.24.

41

NW, “Constitutions and Commitment,” 830.

42

D. Acemoglu and J.A. Robinson, Why Nations Fail: The Origins of Power, Prosperity and Poverty (New York: Crown Business, 2012). See, for example, reiteration that “The Glorious Revolution limited the power of the king and the executive, and relocated to Parliament the power to determine economic institutions …The Glorious Revolution was the foundation for creating a pluralistic society…The government…steadfastly enforced property rights… Historically unprecedented was the application of English law to all citizens. Arbitrary taxation ceased, and monopolies were abolished almost completely…” at 102.

43

D.C. North, J.J. Wallis, and B.R. Weingast, Violence and Social Orders: A Conceptual Framework for Interpreting Recorded Human History (Cambridge: Cambridge University Press, 2009).

44

See LLSV, “Legal Determinants”; and LLSV, “Law and Finance”.

45

R. La Porta, F. Lopez-de-Silanes, and A. Shleifer, “The Economic Consequences of Legal Origins,” Journal of Economic Literature 46 (2008): 285–332.

46

See Baker, An Introduction.

47

Many more details on these topics can be found in GM and the corresponding appendices. See note 10.

48

Of course, despite the large number of reports used to produce the data, any given year might have only a few cases. Therefore, the figures are moving averages, producing smoothness, especially removing prominent idiosyncrasies arising in years when the data are sparse. Additionally, such figures are usually accompanied by confidence intervals that indicate how imprecise the estimate of the timeline is in any given year. In our applications, those intervals are very narrow for all the timelines. Thus, it is sufficient to focus only on the averages that appear in the diagram.

49

Goldstone and Underwood, “The Quiet Transformations,” 379. This resonates with comments in Guldi and Armitage, The History Manifesto, and in Flanders on what computers can do: J. Flanders, “Detailism, Digital Texts, and the Problem of Pedantry,” TEXT Technology 2 (2005): 41–70.

50

Note that it is entirely possible that the topic Assumpsit vanished from cases in the early eighteenth century while the word “assumpsit” was used in a considerable number of case reports from that era. This is possible because topics reflect the co-occurrence of related words rather than only the frequency of single words. When the word “assumpsit” is used in later cases it might be invoked very briefly to reference a huge area of law without being accompanied by many words that were necessary to use in earlier cases, before the notion of assumpsit became readily accepted.

51

DiMaggio et al., “Exploiting Affinities,” 582, point out, in a rather different context, that topic-modeling's assumption of many ideas mixed in a single text provides a significant advantage: “[A] virtue of topic modeling is its deep affinity to the central insight in the sociology of culture that texts do not necessarily reflect a single perspective but are often characterized by heteroglossia, the co-presence of competing ‘voices'—perspectives or styles of expression—within a single text.”

52

Alcock v. Blowfield (1627) 95 E.R. 74, 1061.

53

Flanders, “Detailism,” 57.

54

For example, if one were interested in the workings of the Poor Laws one might want to examine topics related to Geographic Settlement of Children. Then one would be led to examine a narrow but interesting set of topics: Reviewing Local Orders, Employment of Apprentices and Servants, Decisions after Criminal Conviction, and Clarifying Legislative Acts.

55

For an understanding of what the finance topic names signify, the reader is directed to Table 1. For reasons of brevity, similar discussions of the topic names for non-finance topics are omitted, with the reader referred to the relevant elements of GM. After the publication of GM, one topic name that appears in Figure 2 was reconsidered and changed. Interacting in Court has been changed to Decisional Logic, with the renaming prompted by a further reading of the case reports that most use this topic.

56

Our interest in examining law versus equity was stimulated by J. Morley, “The Common Law Corporation: The Power of the Trust in Anglo-American Business History,” Columbia Law Review 116 (2016): 2145–98. Morley describes the crucial role of equity in the development of trusts and argues that trusts afforded many of the legal properties now associated almost uniquely with the modern, legislated corporate form.

57

In fact, Exchequer considered both common-law and equity cases. See W. H. Bryson, The Equity Side of The Exchequer: Its Jurisdiction, Administration, Procedures and Records (Cambridge: Cambridge University Press, 1975). However, the equity side accounted for many fewer cases than the common-law side. For example, in the mid-seventeenth-century Exchequer reports of Hardres, fewer than one quarter of the cases are equity cases. Moreover, Exchequer reports as a whole are small in number compared to the number of reports from the other three major courts. Lastly, the decision to not take into account the mixed set of cases in Exchequer actually biases our results against finding the conclusions that we reach in this section; that is, a more refined treatment of the division between common-law and equity Exchequer reports would actually strengthen our conclusions.

58

Usually such figures are accompanied by confidence intervals that indicate the imprecision of the estimates. Yet in the present context (of abundance of data), those intervals are very small, so focusing on the sizes of the bars alone is sufficient.

59

The reader might be tempted to conclude that the results in Figure 3 are simply due to the fact that Chancery reports became relatively more numerous as the seventeenth century proceeded. But the pertinent statistical procedures control for year. Therefore the results summarized in Figure 3 do not reflect the relationship between the time period in which a particular issue is prominent and the relative number of all the case reports emanating from the different courts during that time period.

60

For example, one reviewer of an earlier version of this article suggested that the patterns in Figure 3 are also consistent with the rise of early corporate law and the increasing use of equity for purposes of resolution of especially complex creditor–debtor arrangements.

61

Sentiment analysis is another vibrant area of machine-learning research; see Frankenreiter and Livermore, “Computational Methods”. Sentiment analysis, which focuses on specific types of sentiments, was not suitable for our analysis given our objective in the present exercise: to obtain a broad historical overview from a large corpus.

62

For spirited criticisms of topic-modeling and related computational techniques, see N. Z. Da, “The Computational Case against Computational Literary Studies,” Critical Inquiry 45 (2019): 601–39; and G. Brookes and T. McEnery, “The Utility of Topic Modelling for Discourse Studies: A Critical Evaluation,” Discourse Studies 21 (2019): 3–21. Da primarily focuses on the problems of machine learning techniques when applied to literature. Brookes and McEnery criticize a particular implementation of topic-modeling from the perspective of corpus linguistics. We view these contributions as interesting in detecting pitfalls in applications of topic-modeling but not dispositive on its value in general.

63

P. Murrell, “Did the Independence of Judges Reduce Legal Development in England, 1600–1800?” Journal of Law and Economics 64 (2021) 539–65 extends this conclusion more generally to areas outside finance using citation analysis to show that granting independence to judges in England might have had deleterious effects on the development of caselaw in the 1600–1800 time period.

64

In the articles that are seminal for economics (see note 39), LLSV do not mention or allude to the distinction between common-law and equity in the context of the English legal tradition; the authors only emphasize the importance of overall “legal style,” a construct perhaps intended to encompass more than just the common-law system itself. The empirical work in economics that follows LLSV has focused on easily discernible, legal-system-wide attributes, likely because it is easier to put those in numerical form than to make data out of texts. This often results in the use of data that reflect the operations of common-law courts and not those of equity courts. For example, the economics-oriented empirical research following LLSV (see note 45) emphasizes judge-made law, adaptability, and judicial independence as institutional sources of superiority of common-law legal systems over civil-law legal systems. Yet it was Lord Chancellors who made law in instances when the common-law judges, bound by strict procedural rules, could not: equity was far more flexible than the common law. And the Lord Chancellor was a government official.

65

For one outstanding example in British economic history, see S. Broadberry, B.M.S. Campbell, A. Klein, M. Overton, and B. van Leeuwen, British Economic Growth, 1270–1870 (Cambridge: Cambridge University Press, 2015).

66

GM's data set is freely available at http://www.econweb.umd.edu/~murrell/Data/ER/ER.html (accessed February 2022). For any help needed to process this data set, please contact the authors of the current article directly.

67

This data set is the easiest to one to convey (in a spreadsheet) to those who have no intention of using the topic-modeling software itself. Use of that software implies that more data are immediately available.

68

Goldberg, “In Defense,” 1.

69

See J.W.F. Allison, “History to Understand, and History to Reform, English Public Law,” Cambridge Law Journal 72 (2013) 526–57. Allison emphasizes the selective invocations to which some uses of legal history have fallen prey, and suggests as one remedy the widening of sources.

70

Topic-modeling can help scholars circumvent the limitations of existing theories and look at the data anew; see R.S. Buurma, “The Fictionality of Topic Modeling: Machine Reading Anthony Trollope's Barsetshire Series,” Big Data & Society (2015), July-Dec., 1–6.

71

Goldstone and Underwood, “The Quiet Transformations,” 370.

72

Allison, “History to Understand,” makes clear that traditional legal history is not immune to such problems.

73

Robertson, “Digital Humanities”, also emphasizes the positive: finding legal ideas where they are not expected by working outside traditional legal classifications.

74

Goldberg, “In Defense,” 1.

75

Grimmer et al., “Machine Learning for Social Science,” 2.

76

P. DiMaggio, “Adapting Computational Text Analysis to Social Science (and Vice Versa),” Big Data & Society (2015), July-Dec., 1–5.

77

Nelson, “Leveraging the Alignment,” 2.

78

Gentzkow et al., “Text as Data,” 549, 555, 556.

Published

2026-07-03