Exploring the Path of Judicial Big Data to Enhance Data Governance Capability

In ample data justice, predicting legal outcomes and identifying similar cases hold significant value. This paper presents an advanced legal prediction algorithm that integrates the specific features of legal texts. Utilizing the Text Rank model, it extracts essential text features from legal provisions and facts, enabling the precise deployment of legal requirements based on detailed case analyses and legal knowledge. To overcome the hurdles of scant training data and the challenge of distinguishing similar legal documents, we developed a similar case matching model employing twin Bert encoders. Our empirical study reveals theft, intentional injury, and fraud as the predominant crimes, with sample counts of 335,745, 174,526, and 47,677, respectively. These top offenses, correlating with the most frequently cited laws, account for 85.79% of our dataset. The analysis further indicates “RMB” as the most recurring word in theft and fraud cases, and “minor injury” in intentional injury instances. Notably, our findings show that categories such as “misappropriation” are prone to misclassification as “embezzlement,” and “robbery” often gets confused with “theft,” highlighting the complexities of legal classification.

Lingua:: Inglese

Frequenza di pubblicazione:: 1 volte all'anno
Argomenti della rivista:: Scienze biologiche, Scienze della vita, altro, Matematica, Matematica applicata, Matematica generale, Fisica, Fisica, altro

Feed RSS della rivista

Exploring the Path of Judicial Big Data to Enhance Data Governance Capability

Yingshuai Liang

Pubblicato online: 03 mag 2024

Ricevuto: 08 apr 2024

Accettato: 26 apr 2024

DOI: https://doi.org/10.2478/amns-2024-1015

Parole chiaveLegal Prediction, Big Data Justice, Similar Cases, Text Rank Model, Twin Bert

© 2024 Yingshuai Liang, published by Sciendo

This work is licensed under the Creative Commons Attribution 4.0 International License.

Parole chiave
Legal Prediction, Big Data Justice, Similar Cases, Text Rank Model, Twin Bert