MACHINE LEARNING · EMPIRICAL
Original research: LexIssue: Benchmarking Legal Issue Identification in Chinese Civil Litigation · 2609.02954v1
Paper authors: Huiyuan Xie, Yuqin Huang, Zhicheng Hao, Yida Cai, Shaochun Wang, Zhenghao Liu, Yuxiao Ye
Source license: CC BY 4.0. This article summarizes and interprets the source using AI. Attribution does not imply endorsement by the source authors.
This adapted analysis is shared under the same CC BY 4.0 license. Semantic status: supported by automated evidence review. Human scientific review and independent replication have not been established.
TL;DR
The work introduces a legally grounded hierarchy that combines free-form descriptions of disputed matters with structured legal categories.
It frames the task as complementary generation of issue descriptions and classification of their legally grounded attributes.
Retrieval augmentation generally improves performance relative to corresponding settings without retrieval, although reported exceptions remain.
Under direct instruction, a proprietary general-purpose system achieved the strongest overall reported performance across generation and hierarchical classification.
Source: E005
Legally informed prompting does not reliably improve performance and frequently harms it relative to the corresponding setting without that prompting.
Source: E012
Among predictions aligned to reference issues, classification performance declines markedly as hierarchical specificity increases, including for the strongest reported systems.
Source: E002
Significance
The work introduces a legally grounded hierarchy that combines free-form descriptions of disputed matters with structured legal categories.
Research Question
The study examines computational identification of disputed matters in litigation together with their legal attributes.
Source: E014
Contribution
The work introduces a legally grounded hierarchy that combines free-form descriptions of disputed matters with structured legal categories.
Datasets
The evaluation resource comprises real-world civil litigation cases with expert annotation of disputed matters.
Source: E013
Each annotated matter includes both a descriptive summary and structured category labels, and the cases cover diverse civil dispute types.
Source: E013
The benchmark uses synthetic claim and defence materials reconstructed from judicial accounts of party pleadings.
This reconstruction approach was used because original pleadings are difficult to obtain at scale from public sources.
Source: E021
Limitations
Synthetic inputs cannot fully reproduce the form, style, strategic framing, or evidential detail of materials used in actual litigation.
Source: E021
Assessment is challenging because equivalent disputed matters can differ substantially in wording and abstraction.
Source: E001
Reliable prediction-reference alignment is therefore needed before computing evaluation metrics.
Source: E001
Method
An issue-centred legal knowledge resource was curated from doctrinal and practical materials alongside litigated cases.
Expert integration consolidated overlapping entries and normalized their granularity and style, producing reusable abstractions rather than case-specific annotations.
Source: E015
Models are tested in direct instruction, legally guided prompting, and retrieval-augmented settings.
For retrieval augmentation, candidate issue entries are reranked for relevance and the selected entries are placed into the generation prompt.
Source: E006
Baselines
Metrics
Issue generation is scored through average case-level weighted F-score over the test set.
Source: E024
Metrics Conditions
Prediction-reference matching proceeds through coarse semantic candidate mapping followed by stricter rubric-based verification of issue equivalence.
Only candidate pairs passing the predefined acceptance criterion are retained as aligned.
Source: E016
Findings
Retrieval augmentation generally improves performance relative to corresponding settings without retrieval, although reported exceptions remain.
Under direct instruction, a proprietary general-purpose system achieved the strongest overall reported performance across generation and hierarchical classification.
Source: E005
Failure Modes
Legally informed prompting does not reliably improve performance and frequently harms it relative to the corresponding setting without that prompting.
Source: E012
Among predictions aligned to reference issues, classification performance declines markedly as hierarchical specificity increases, including for the strongest reported systems.
Source: E002
Evidence and source
Show evidence locators
Evidence labels locate support in the original paper; they do not establish independent replication.
- E001 · page 5 — Introduction: Evidence E001
- E002 · page 8 — Introduction: Evidence E002
- E003 · page 4 — Introduction: Evidence E003
- E004 · page 4 — Introduction: Evidence E004
- E005 · page 7 — Introduction: Evidence E005
- E006 · page 7 — Introduction: Evidence E006
- E007 · page 6 — Introduction: Evidence E007
- E008 · page 4 — Introduction: Evidence E008
- E009 · page 7 — Introduction: Evidence E009
- E010 · page 2 — Introduction: Evidence E010
- E011 · page 6 — Introduction: Evidence E011
- E012 · page 7 — Introduction: Evidence E012
- E013 · page 5 — Introduction: Evidence E013
- E014 · page 1 — Abstract: Evidence E014
- E015 · page 6 — Introduction: Evidence E015
- E016 · page 5 — Introduction: Evidence E016
- E017 · page 7 — Introduction: Evidence E017
- E018 · page 5 — Introduction: Evidence E018
- E019 · page 6 — Introduction: Evidence E019
- E020 · page 6 — Introduction: Evidence E020
- E021 · page 4 — Introduction: Evidence E021
- E022 · page 6 — Introduction: Evidence E022
- E023 · page 7 — Introduction: Evidence E023
- E024 · page 5 — Introduction: Evidence E024