Joint Span Segmentation and Rhetorical Role Labeling with Data Augmentation for Legal Documents

T. Y.S.S. Santosh, Philipp Bock, Matthias Grabmair

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

1 Scopus citations

Abstract

Segmentation and Rhetorical Role Labeling of legal judgements play a crucial role in retrieval and adjacent tasks, including case summarization, semantic search, argument mining etc. Previous approaches have formulated this task either as independent classification or sequence labeling of sentences. In this work, we reformulate the task at span level as identifying spans of multiple consecutive sentences that share the same rhetorical role label to be assigned via classification. We employ semi-Markov Conditional Random Fields (CRF) to jointly learn span segmentation and span label assignment. We further explore three data augmentation strategies to mitigate the data scarcity in the specialized domain of law where individual documents tend to be very long and annotation cost is high. Our experiments demonstrate improvement of span-level prediction metrics with a semi-Markov CRF model over a CRF baseline. This benefit is contingent on the presence of multi sentence spans in the document.

Original languageEnglish
Title of host publicationAdvances in Information Retrieval - 45th European Conference on Information Retrieval, ECIR 2023, Proceedings
EditorsJaap Kamps, Lorraine Goeuriot, Fabio Crestani, Maria Maistro, Hideo Joho, Brian Davis, Cathal Gurrin, Annalina Caputo, Udo Kruschwitz
PublisherSpringer Science and Business Media Deutschland GmbH
Pages627-636
Number of pages10
ISBN (Print)9783031282379
DOIs
StatePublished - 2023
Event45th European Conference on Information Retrieval, ECIR 2023 - Dublin, Ireland
Duration: 2 Apr 20236 Apr 2023

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume13981 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference45th European Conference on Information Retrieval, ECIR 2023
Country/TerritoryIreland
CityDublin
Period2/04/236/04/23

Keywords

  • Data augmentation
  • Rhetorical Role Labeling
  • Semi-Markov CRF

Fingerprint

Dive into the research topics of 'Joint Span Segmentation and Rhetorical Role Labeling with Data Augmentation for Legal Documents'. Together they form a unique fingerprint.

Cite this