2019/10/21 by Tereza Novotná, Novotná, Tereza, Jakub Harašta +1 · 4 citations
Computer Science · Social Sciences · #Artificial Intelligence in Law #Comparative and International Law Studies #Computation and Language (cs.CL) #Computer science #Constitutional court #Court decision #Court of record #Czech #Documentation #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Law #Law of the case #Legal Language and Interpretation #Linguistics #Majority opinion #Original jurisdiction #Plain language #Political science #Supreme court #cs.CL #cs.IR
paper · pdf · doi:10.48550/arxiv.1910.09513
published in arXiv (Cornell University) (Cornell University)
arxiv created 2019/10/21 · openalex publication_date 2019/10/21 · arxiv updated 2019/10/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In this paper, we describe the Czech Court Decision Corpus (CzCDC). CzCDC is a dataset of 237,723 decisions published by the Czech apex (or top-tier) courts, namely the Supreme Court, the Supreme Administrative Court and the Constitutional Court. All the decisions were published between 1st January 1993 and 30th September 2018. Court decisions are available on the webpages of the respective courts or via commercial databases of legal information. This often leads researchers interested in these decisions to reach either to respective court or to commercial provider. This leads to delays and additional costs. These are further exacerbated by a lack of inter-court standard in the terms of the data format in which courts provide their decisions. Additionally, courts' databases often lack proper documentation. Our goal is to make the dataset of court decisions freely available online in consistent (plain) format to lower the cost associated with obtaining data for future research. We believe that simplified access to court decisions through the CzCDC could benefit other researchers. In this paper, we describe the processing of decisions before their inclusion into CzCDC and basic statistics of the dataset. This dataset contains plain texts of court decisions and these texts are not annotated for any grammatical or syntactical features.