0% found this document useful (0 votes)
2 views30 pages

Legal Graph Rag

LegalGraphRAG is a proposed framework that enhances legal reasoning by integrating graph-based retrieval with a multi-agent system for evidence verification. It addresses challenges in traditional retrieval-augmented generation methods, such as handling heterogeneous legal documents and ensuring reliable, evidence-based reasoning. The framework's hierarchical legal graph and structured decision-making process significantly improve the accuracy and trustworthiness of legal analysis, outperforming existing models in experimental evaluations.

Uploaded by

omraut41105
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views30 pages

Legal Graph Rag

LegalGraphRAG is a proposed framework that enhances legal reasoning by integrating graph-based retrieval with a multi-agent system for evidence verification. It addresses challenges in traditional retrieval-augmented generation methods, such as handling heterogeneous legal documents and ensuring reliable, evidence-based reasoning. The framework's hierarchical legal graph and structured decision-making process significantly improve the accuracy and trustworthiness of legal analysis, outperforming existing models in experimental evaluations.

Uploaded by

omraut41105
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

LegalGraphRAG: Multi-Agent Graph Retrieval-Augmented Generation

for Reliable Legal Reasoning


Zerui Chen1,3 , Qinggang Zhang4† , Zhishang Xiang2,3 , Zhimin Wei1,3 , Linfeng Gao1,3 ,
Xiao Huang5 , Zhihong Zhang1† , Jinsong Su1,3†
1
School of Informatics, Xiamen University
2
Institute of Artificial Intelligence, Xiamen University
3
Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural
Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism, China
4
School of Artificial Intelligence, Jilin University 5 The Hong Kong Polytechnic University
chenzerui1@[Link] qinggangzhang@[Link] {zhihong,jssu}@[Link]

Domain-Specific Query: Please adjudicate the following cases: Defendant


Abstract Zhang signed a labor contract ... serving as the head chef of the project
department. His main responsibility was handling the financial... Legal
Medical
Financail
Graph-based Retrieval-Augmented Generation
(GraphRAG) advances flat document retrieval (a) Heterogeneous Knowledge Base Mixed Granularity

by structuring knowledge as relational graphs, indexing


enabling more coherent and effective reason- Tree Vector DB Graph
ing. However, applying it to specific domains Facts/Rules/Principles Flat Knowledge Construction

like legal reasoning faces critical challenges.


(i) Legal corpora are heterogeneous, containing (b) The Semantic Similarity Trap Similarity ≠ Utility

multi-granular knowledge from cases, articles,


and interpretations. A flat knowledge graph
cannot adequately differentiate between factual
details, applied rules, and abstract principles, I missed the governing Rules
needed to solve this.
limiting accurate retrieval. (ii) Reliable legal
judgment demands transparent, evidence-based
reasoning. Traditional RAG passes retrieved (c) Trust & Verification Issues Blind Trust in Context
Answer: Referencing the similar case (Fragment A), such acts of utilizing...with the
context directly to an LLM without verification, provisions of Article 382...liable for the crime of Corruption.

resulting in opaque, error-prone reasoning. To Case A: defendant Bai has... (irrelavant)


Article 271: Employees of... (relavant)
How can I trust this unexplainable
verdict? This answer is risky.
this end, we propose LegalGraphRAG, a frame- Article 382: If a state functionary...(irrelavant)

work designed for reliable legal reasoning. Our


approach introduces two core components: a
hierarchical legal graph that hierarchically or- Figure 1: Challenges of Traditional RAG in Domain-
ganizes legal sources to enable retrieval at ap- Specific Tasks. (i) Flat Graph Structure: Struggles
propriate abstraction levels, and a multi-agent to handle heterogeneous documents. (ii) Unverified
system for reliable legal reasoning, where a Re- Retrieval: Contains excessive irrelevant information.
searcher retrieves candidate evidence, an Audi-
tor rigorously verifies its validity against source intelligent decision-making across various real-
documents, and an Adjudicator synthesizes the world tasks (Zhao et al., 2023; Naveed et al., 2025).
set of verified evidence to render a final judg- However, deploying these models in specialized,
ment. Extensive experiments show that Legal- knowledge-intensive fields like legal reasoning re-
GraphRAG achieves the state-of-the-art per- mains challenging due to the domain’s demand-
formance, outperforming existing GraphRAG ing standards of rigor and reliability (Lai et al.,
baselines in accurate and trustworthy legal anal-
2024; Hou et al., 2025; Siino et al., 2025). Domain-
ysis. Our code, datasets and implementation
details are available at [Link] specific tasks necessitate a comprehensive under-
XMUDeepLIT/LegalGraphRAG. standing and multi-step reasoning across a vast
knowledge base of specialized concepts, rigorous
1 Introduction rules, and complex dependencies (Wang et al.,
2023; Kim et al., 2025), which requires strict logi-
The rapid advancement of Large Language Models cal reasoning and domain expertise that exceed the
(LLMs), like GPT (Achiam et al., 2023), Gem- capabilities of general-purpose LLMs. While Su-
ini (Comanici et al., 2025) and Qwen (Yang et al., pervised Fine-Tuning (SFT) (Ouyang et al., 2022;
2025a) series, has driven significant progress in Hu et al., 2022) on domain corpora enables mod-
† Corresponding author. els to internalize that expertise, this approach in-
37455
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 37455–37484
July 2-7, 2026 ©2026 Association for Computational Linguistics
rate retrieval. ❷ Lack of verifiable, evidence-based
reasoning. Traditional RAG passes retrieved con-
text directly to an LLM without any verification.
This “retrieve-then-generate” pipeline often results
in opaque, error-prone reasoning, particularly when
the retrieved context is long and contains substan-
tial irrelevant information (Cao et al., 2024).
In this paper, we propose LegalGraphRAG,
a novel framework that synergizes graph-based
retrieval with the multi-agent reasoning system
Figure 2: Retrieval performance comparison revealing for reliable legal reasoning. Specifically, Legal-
that conventional RAG methods struggle with hetero- GraphRAG consists of two key components: (i)
geneous domain documents, suffering from high error Hierarchical legal graph (HierarGraph), which or-
rates and limited effectiveness. detailed experimental ganizes legal knowledge into a hierarchical graph
setup is introduced in Section 3.1 and Appendix A.3. to effectively decouple historical cases, relevant
statutes, and judicial interpretations, and (ii) a
curs substantial computational costs and often risks
multi-agent system for evidence-based reason-
critical catastrophic forgetting in many real-world
ing (Xi et al., 2025; Xiang et al., 2026), where the
scenarios (Yue et al., 2024; Luo et al., 2025).
legal judgment process is structured as a transpar-
Recently, Retrieval-Augmented Generation ent pipeline that retrieves, verifies, and reasons over
(RAG) (Lewis et al., 2020; Borgeaud et al., 2022; graph-grounded evidence to produce interpretable
Li et al., 2025; Zhang et al., 2025b) offers a prac- decisions. Generally, our contributions are summa-
tical solution to adapt LLMs for specific domains. rized as follows:
RAG systems enable LLMs to generate responses
by leveraging not only their parametric knowledge • We propose LegalGraphRAG, an evidence-
but also real-time retrieved domain knowledge, based legal reasoning framework driven by
thereby providing more accurate and reliable an- a multi-agent system operating on a hierar-
swers (Mallen et al., 2023; Zhang et al., 2025b; chical knowledge graph, which address legal
Zhou et al., 2025; Liu et al., 2026). However, heterogeneity and ensure reliable reasoning.
standard RAG systems typically retrieve informa- • We design a hierarchical legal knowledge
tion based on semantic similarity (Karpukhin et al., graph with Ontology, Fact, and Rule layers
2020; Chen et al., 2024), treating documents as to model multi-granular legal knowledge and
independent text segments. This hinders complex support accurate retrieval.
multi-hop reasoning over hierarchical legal con- • We establish a multi-agent system for
cepts and multiple documents, limiting effective- evidence-based reasoning that performs ad-
ness in legal analysis. judication through a transparent pipeline of
Graph-based Retrieval-Augmented Generation retrieval, validation, and synthesis, grounding
(GraphRAG) (Edge et al., 2024; Zhang et al., judgments in verifiable evidence chains.
2025a; Xiang et al., 2025; Liu et al., 2025; Yang • Extensive experiments show that Legal-
et al., 2026) advances this paradigm by organizing GraphRAG consistently outperforms existing
domain corpora into structured relational graphs. GraphRAG baselines and legal language mod-
This structural awareness captures hierarchical re- els in accurate and trustworthy legal analysis.
lationships between different concepts, thereby en-
2 Problem Statement
abling more precise retrieval and supporting the
multi-hop reasoning required for complex queries. Complex legal reasoning is formulated as an open-
However, directly applying standard GraphRAG to ended generation task evaluating the decision-
the legal domain faces critical challenges (as illus- making capabilities of LLMs within the legal do-
trated in Figure 1): ❶ A flat graph structure cannot main. Formally, given a criminal fact description f
capture the multi-granular hierarchies present in and a defendant d, a LLM is tasked with predicting
legal corpora, which span factual details, applied the applicable charges y. In this paper, we focus on
rules, and abstract principles across legal cases, integrating this reasoning framework with RAG to
articles, and interpretations, thereby limiting accu- assess the model’s ability to leverage external legal
37456
knowledge for judicial reasoning. This task can be improving retrieval performance by 25.3%. This
organized into the following stages: observation suggests that structural flatness consti-
Knowledge Organization. Given an offline corpus tutes a fundamental bottleneck for standard RAG
of legal documents D, including historical cases, when handling multi-granular knowledge.
articles and interpretations we construct a domain-
specific legal knowledge graph: 3.2 Investigation on Generation Quality

KG = Φ(D), (1) Reliable domain reasoning demands not only in-


formation retrieval but also evidence verification.
where Φ(·) denotes the organization function. Real-world legal environments often contain docu-
Knowledge Retrieval. For a legal query charac- ments that share similar keywords but differ funda-
terized by criminal facts f and a defendant d, we mentally in their domain applicability. To simulate
retrieve relevant evidence from KG to form a con- this realistic challenge, we conduct a test (detailed
textual reference: in Appendix A.4). Specifically, we inject legally
plausible but factually irrelevant documents into
C = R(f, d, KG), (2)
the retrieval context to evaluate the model’s ability
where R(·) represents the retrieval operator. to focus on relevant evidence.
Judgment Generation. Finally, the legal judgment
(e.g., charge) y is inferred by reasoning over the Charge Articles Term of Penalty
ACC↑ ACC↑ MAE↓
query and retrieved evidence: Method
(%)

(%)

(months)

RAG (Correct Context) 42.8 – 74.7 – 24.3 –


P (y | f, d, C) = G(f, d, C), (3) RAG + 2 Irrelevant Docs 34.9 ↓ 7.9 57.2 ↓ 17.5 27.7 ↑ 3.4
RAG + 4 Irrelevant Docs 32.9 ↓ 9.9 51.1 ↓ 23.6 28.4 ↑ 4.1
RAG + 6 Irrelevant Docs 29.8 ↓ 13.0 46.8 ↓ 27.9 31.7 ↑ 7.4
where G(·) denotes the generator LLM.
Table 1: Performance degradation under varying levels
3 Preliminary Study of simulated retrieval noise. ACC (↑) denotes Accuracy
for Charge and Articles prediction. MAE (↓) represents
Applying standard retrieval paradigms to the spe-
Mean Absolute Error for Term of Penalty.
cialized, knowledge-intensive legal domain faces
critical challenges due to the inherent structural
As summarized in Table 1, standard RAG mod-
complexity and rigorous standards of such fields.
els exhibit significant sensitivity to context purity.
To illustrate these challenges, we conduct two pre-
The inclusion of irrelevant information precipitates
liminary experiments to empirically investigate the
a sharp performance drop. This observation shows
specific limitations of existing methods regarding
that without a dedicated verification mechanism
knowledge granularity and generation quality.
to filter irrelevant content, the model struggles to
3.1 Investigation on Knowledge Granularity distinguish valid evidence from misleading infor-
mation, which undermines reasoning reliability.
Complex domain knowledge possesses an inherent
hierarchy. In the legal context, this necessitates dis-
3.3 Discussion and Motivation
tinguishing between abstract statutory principles
and concrete case facts. We hypothesize that stan- The findings from these two studies highlight fun-
dard retrieval strategies fail to distinguish between damental limitations in applying standard RAG to
these semantic granularities because they treat all complex domains: ❶ Flat retrieval mechanisms
text segments in the same way. To verify this, we fail to navigate the hierarchical nature of domain
compare a Flat Strategy against a naive Hierarchi- knowledge (e.g., distinguishing rules from facts),
cal Strategy that explicitly segregates articles from resulting in biased context. ❷ The lack of an ex-
case narratives (detailed in Appendix A.3). plicit verification step makes the system fragile to
As illustrated in Figure 2, the empirical results misleading information, which is unacceptable in
confirm our hypothesis. Flat Strategy exhibit a rigorous fields like law. These insights motivate
distinct “granularity bias”, frequently prioritizing the design of LegalGraphRAG, which incorporates
high-frequency factual details due to surface-level a Hierarchical Legal Graph to resolve granularity
semantic overlaps, often at the expense of essential conflicts and a Evidence-based Legal Reasoning
abstract principles. Conversely, Hierarchical Strat- (Researcher-Auditor-Adjudicator) framework to en-
egy aligns better with the domain’s logical structure, force rigorous verification.
37457
Stage1: Exploring Heterogeneous Legal Knowledge Stage2: Evidence-Based Reasoning

Researcher HierarGraph Auditor


Searched Results A1 C2 C3 A2 C4
Query Defendant Zhang signed a labor Ontology Graph
contract with the Xiangjiaba Hydropower F5 Query The defendant Zhang Mou entered into a labor contract
Station Yaojiaba Tunnel Project Department of F6
the original Sichuan Road and ... Com2 with the ... serving as the head chef of the project department. His
F4
main responsibility was handling the financial reimbursement for the
Mapping to Predict monthly expenses of all canteens ...
Com1
legal essence directly F3
F1
F2 Fine-grained
Case Feature
Evidence Check Chef is not State
D: Employee... C: Misappropriation of funds... functionaries
V: Company funds... M: Subjective intent...
Fact Graph C4 Article 382

O2 State functionaries who, taking advantage of their office,


Candidate charges A2 C3 misappropriate, steal, swindle, or use other illegal means to take
Embezzlement, Misappropriation of funds... Judicial Interpretation
possession of public property are to be convicted of the crime of
..."state functionaries" refers to
A1 embezzlement... persons performing public service...
O1 C2
Article 271
C1

Symbols Prune Irrelevant Documents


Rule Graph
J4
: Forward Process
: Edges between nodes A2 Adjudicator
Final Context A1 C2 C3 A2 C4
: Community Node : Case Node
: Case Feature Node : Offence Node A1
: Article Node(Gi) : Article Node(Ge) J1 J3 Final Results Charge: job-related embezzlement Articles: 271 Imprisonment: 18
: Judicial Interpretation Node Resoning: Zhang served as the head chef of the project department and was ...
J2
[1] corresponds to Article 271 Paragraph 1...
[2] corresponds to Cases1([Charges: job-related embezzlement] [Name: Xu 3Mou])...

Figure 3: The architecture of LegalGraphRAG. The framework consists of two main phases: (1) Hierarchical
Knowledge Construction, which builds a Hierarchical Legal Graph (HierarGraph) comprising an Fact Graph,
Ontology Graph and Rule Graph to organize heterogeneous legal knowledge; and (2) Evidence-based Legal
Reasoning, where a multi-agent system (Researcher, Auditor, and Adjudicator) performs structured retrieval,
validation, and synthesis over the HierarGraph to generate interpretable legal decisions.

4 The Framework of LegalGraphRAG a Hierarchical Legal Graph (HierarGraph) H that


organizes legal knowledge into distinct semantic
4.1 Overview layers, enabling explicit differentiation among le-
Traditional GraphRAG approaches face limitations gal concepts and providing a structured basis for
in legal judgment due to the heterogeneous and reliable reasoning. The HierarGraph is composed
multi-granular nature of legal corpora. To address of three specialized subgraphs:
this challenge, We propose LegalGraphRAG, an Fact Graph (Gf ac ), which serves as a structured
evidence-based legal reasoning framework driven collection of verified legal precedents, providing
by a multi-agent system operating on a hierarchical the essential factual basis for ensuring legally
knowledge graph. The framework operates in two grounded judgments. Accordingly, Gfac models the
distinct phases: (i) Hierarchical Knowledge Con- natural structure of legal documents by explicitly
struction, which organizes legal knowledge into connecting Cases (C), Articles (A), and Offense
a layered graph structure to effectively decouple (O) nodes. Relationships are established via eca ,
historical cases, relevant statutes, and judicial in- linking a case c to its cited article a, and eco , linking
terpretations, and (ii) Evidence-based Legal Rea- a case c to its convicted offense o. This structure
soning, structures the legal judgment process as provides the factual granularity required for evi-
a transparent pipeline that retrieves, verifies, and dence gathering. Formally, it is defined as:
reasons over graph-grounded evidence to produce 
|Gfac |

Gf ac = (Vf ac , Ef ac ) = ci , a i , o i . (4)
interpretable decisions. The whole framework is i=1

illustrated in Figure 3. Ontology Graph (Gont ), which bridges the seman-


tic gap and mitigates noise by abstracting case
4.2 Hierarchical Knowledge Construction features. Gont distills raw narratives containing
Legal reasoning involves heterogeneous informa- instance-specific details (e.g., dates and locations)
tion sources, including historical cases, abstract into a purified semantic space that reflects the “legal
legal articles, and interpretations. Employing a flat essence”. Specifically, we design a domain-specific
storage structure is not enough to handle the in- legal ontology based on legal theory (Rüthers et al.,
herent structural differences of these data sources, 2013), encompassing four key dimensions: Defen-
leading to disorganized information and inefficient dant Attributes, Criminal Behaviors, Victim Char-
retrieval. To address this challenge, we construct acteristics and Subjective Mental States. Keywords
37458
and entities are extracted and aligned with these analysis, the framework resolves the raw case query
properties to form structured embeddings, serving by constructing a final, verifiable judgment.
as indices for Case Feature Nodes (F).
4.3.1 Evidence Retrieval
To reveal hidden connections between different
cases, we employ the k-Nearest Neighbors (k-NN) A reliable evidence-based reasoning process begins
algorithm to connect nodes with high semantic sim- with grounding a raw case description in relevant
ilarity. We then apply the Leiden algorithm (Traag legal evidence. To this end, Researcher perform
et al., 2019) to group related cases into communi- structured evidence retrieval over the Gont and the
ties, each treated as a Community Node (K). Each Gfac , transforming unstructured case narratives into
k contains the summarized information of the cases a coherent set of related Cases (C) and Articles (A).
inside it, facilitating hierarchical retrieval that navi- Specifically, the Researcher aligns the case de-
gates from broad contexts to specific details. For- scription with the ontological dimensions defined
mally, this subgraph is defined as: in Section 4.2. Based on these features, We for-
 
mulate the evidence retrieval process R(q) as the
Gont = (Vont , Eont ) = ci , k j
|Gont |
i=1,j=1
. (5) union of three operators, where q is the legal query:
R(q) = Rsem (q) ∪ Rcom (q) ∪ Rchg (q) (8)
Rule Graph (Grul ), which resolves statutory am-
biguities by systematically linking Articles (A) First, we employ Semantic Match Retrieval to lo-
with its corresponding Judicial Interpretations (J ). cate direct evidence via semantic similarity, where
This explicit alignment establishes the contextual ϕ(·) denotes ontology-aligned embeddings:
grounding necessary for precise legal reasoning.

Moreover, applying the correct article often de- Rsem (q) = Top-k sim ϕ(q), ϕ(c) (9)
c∈Gont
pends on specific conditions. A small difference
can lead to a completely different judgment for the Next, to capture structural context, we conduct
same crime. (e.g. whether the defendant is an adult Community Expansion Retrieval. We first identify
or a minor). Simple semantic matching often fails the top-ranked communities by topic SK aligned
to distinguish these subtle differences. To address with the query, and then retrieve the most similar
this, we equip each a with a Diagnostic Checklist cases within these communities:
(D). This mechanism breaks down complex legal 
K∗ = argmax sim ϕ(q), ϕ(K)
rules into specific verification steps. Formally, this K∈Gont
 (10)
subgraph is defined as: Rcom (q) = Top-k sim ϕ(q), ϕ(c)
  c∈K∗
|Grul |
Grul = (Vrul , Erul ) = ai , ji i=1
. (6)
Finally, we implement Charge-Anchored Re-
where trieval to anchor the legal basis by collecting cases
D(ai ) = {d1 , . . . , d|C| } (7)
linked to inferred charges. Here, O(q) denotes
By integrating these three layers, HierarGraph the set of predicted charges and N represents the
H transforms heterogeneous legal corpora into a neighboring cases connected to charge o in Gf ac :
structured ecosystem. This architecture directly [
addresses the limitations of flat retrieval by offer- Rchg (q) = NGf ac (o) (11)
ing multi-granular support for following evidence- o∈O(q)

based legal reasoning. The detailed construction The specific retrieval algorithms and parameter
procedures are provided in Appendix B.1. settings are detailed in Appendix B.2.
4.3 Evidence-based Legal Reasoning 4.3.2 Evidence Validation
To leverage the multi-granular knowledge encoded Given the candidate evidence retrieved in the Ev-
in our HierarGraph, we propose a multi-agent sys- idence Retrieval, this stage focuses on validating
tem for evidence-based reasoning, in which spe- whether the case facts genuinely satisfy the con-
cialized agents sequentially traverse the graph to ditions required by the law, rather than relying on
perform evidence retrieval, validation, and synthe- surface-level semantic relevance.
sis. Specifically, the workflow consists of three Specifically, for each candidate article, we verify
agents:1) Researcher, 2) Auditor, and 3) Adjudica- its applicability by evaluating the case facts using
tor. Through structured graph traversal and logical the associated Diagnostic Checklist and Judicial
37459
CAIL CMDL Average
Public Safety Economic Social Order Person Rights Public Safety Economic Social Order Person Rights
Model Size ACC / F1 ACC / F1 ACC / F1 ACC / F1 ACC / F1 ACC / F1 ACC / F1 ACC / F1 All ∆

Open-Source Models
Qwen-2.5-7B-Instruct 7B-Inst 24.0 45.8 23.1 42.5 22.9 36.7 27.4 46.0 25.8 32.4 28.7 35.8 27.2 42.1 32.8 49.6 26.7 ↑ 22.8
Qwen-3-8B 8B-Inst 31.7 49.2 25.8 42.7 26.3 39.8 27.6 47.8 44.0 52.3 44.7 53.1 42.7 51.9 53.0 57.7 35.2 ↑ 19.9
Internlm3-8b-instruct 8B-Inst 29.8 49.1 26.7 42.0 25.2 34.3 28.1 47.3 25.4 32.1 35.7 37.0 27.5 36.2 34.1 53.6 26.6 ↑ 22.9
Glm-4-9b-chat 9B-Inst 18.4 33.7 19.7 36.1 15.8 32.1 26.0 44.5 23.5 34.2 23.6 40.8 19.1 37.0 41.5 47.0 21.2 ↑ 28.2
Advanced Models
GPT-4o-mini ∼8B 19.7 35.5 19.6 33.3 15.5 35.2 29.0 46.3 18.0 28.0 22.7 31.9 21.9 32.0 35.9 50.3 28.4 ↑ 21.1
DeepSeek-V3.1 ∼200B 31.0 51.3 29.0 48.4 29.8 50.2 35.2 54.8 35.0 64.0 54.7 62.7 58.2 61.9 62.5 71.6 42.8 ↑ 6.7
Legal Specific Methods
DISC-LawLLM-7B 7B-Inst 40.1 50.9 31.0 51.5 34.8 47.7 34.5 56.0 49.7 53.6 39.6 52.1 30.3 49.5 48.4 63.3 30.3 ↑ 19.1
ADAPT 7B-Inst 38.7 43.7 32.7 43.4 27.6 41.7 35.2 50.7 54.5 58.8 57.1 59.4 40.9 43.4 61.5 62.1 42.8 ↑ 6.7
Legal ∆ 7B-Inst 40.8 50.6 25.1 37.4 32.1 43.7 34.1 53.6 58.3 61.5 51.8 55.8 50.2 54.8 65.8 64.4 42.4 ↑ 7.1
RAG Based Methods
Naive RAG 8B-Inst 31.0 45.7 24.4 38.7 28.1 38.4 34.5 46.8 45.8 57.3 44.8 55.2 46.8 58.5 49.6 57.8 33.3 ↑ 16.1
G-retriever 8B-Inst 33.8 48.0 26.0 39.8 23.8 39.3 32.6 50.1 36.8 40.0 42.5 48.8 45.3 50.7 46.2 52.4 34.4 ↑ 13.2
LightRAG 8B-Inst 20.4 43.6 21.7 42.5 19.0 42.5 26.9 50.6 37.9 50.1 43.2 45.1 44.2 51.3 43.7 46.9 30.5 ↑ 19.0
RAPTOR 8B-Inst 34.6 50.4 31.6 43.9 32.1 45.6 32.4 45.7 53.8 62.6 53.6 60.1 52.5 62.8 52.1 66.9 43.1 ↑ 6.3
HippoRAG2 8B-Inst 34.5 38.2 24.0 33.5 28.8 35.0 31.0 36.3 53.5 56.5 50.6 52.7 53.5 55.0 62.4 62.8 43.1 ↑ 6.3
LegalGraphRAG (Ours) 8B-Inst 42.9 54.3 38.5 53.6 37.6 51.1 37.2 58.3 65.5 66.5 59.8 65.1 58.5 63.7 70.1 72.7 49.5 –

Table 2: Performance comparison on CAIL and CMDL. We employ Qwen3-8B as the default backbone
model. The best results are highlighted in bold, and the second-best are underlined. We visualize the gains of
LegalGraphRAG over each baseline in the ∆ columns .

Interpretations encoded in the Grul . The verifica- and synthesis, the system enforces stepwise verifi-
tion outcomes are then aggregated to produce a cation and ensures that every conclusion is explic-
definitive applicability judgment for each article. itly derived from and supported by verified legal
Based on these judgments, Auditor filters the evidence, resulting in reliable judicial decisions.
retrieval subgraph by pruning inapplicable articles
and their associated case and charge nodes. Finally, 5 Experiment
it organizes the remaining nodes into a legally con- This section presents a comprehensive evaluation
sistent and evidence-supported subgraph, which of LegalGraphRAG on two legal judgment bench-
serves as a validated knowledge basis for subse- marks. Our experiments are designed to answer the
quent decision-making. Further implementation following three questions. Q1 (Generation Accu-
details can be found in Appendix B.2. racy): Does LegalGraphRAG outperform SOTA
4.3.3 Evidence Synthesis GraphRAG methods and leading legal-domain
LLMs in generation quality? Q2 (Case Study):
In the final stage, the validated evidence produced
How does LegalGraphRAG handle specific legal
in the previous steps is synthesized to derive a
cases, and does it provide more interpretable out-
legally grounded judgment. Based on the verified
puts compared to baselines? Q3 (Ablation Study):
subgraph, Adjudicator integrates the confirmed ar-
What is the contribution of each core component to
ticles (Af ), cases (C f ), and offense information
the final performance of LegalGraphRAG? More
(Of ) to determine the applicable charges and their
additional experiments are provided in Appendix C
statutory basis. This process is formulated as:
5.1 Experiment Setup
J = Adjudicator(q ⊕ Af ⊕ C f ⊕ Of ) (12)
Datasets We evaluate on two widely used legal
Crucially, the judgment is not produced as a di- benchmarks: CAIL2018 (Xiao et al., 2018) and
rect verdict. Instead, it is accompanied by explicit CMDL (Huang et al., 2024), covering diverse crim-
citations to the statutory articles and judicial inter- inal sub-fields such as Public Safety, Social Order,
pretations used in the reasoning process, ensuring Economic Offenses, and Person Rights. The re-
that every conclusion is directly traceable to veri- trieval knowledge base is built from a collection of
fied evidence in the HierarGraph. authoritative legal sources, including case datasets
Overall, LegalGraphRAG formulates legal judg- and statutory texts. Further dataset details are pro-
ment as a transparent, evidence-based reasoning vided in Appendix D.1 & D.2.
pipeline rather than a black-box generation process. Baselines To ensure a comprehensive evalua-
Through sequential evidence grounding, validation, tion, we categorize our comparative experiments
37460
Query Please render a judgment against the defendant Zhang based on the following facts: “Defendant Zhang signed a Answer Charge: job-related embezzlement
labor contract ... On December 19, 2016, the Yaojiaba Tunnel Project Department transferred 98,160 yuan into Zhang's account Article: 271 Imprisonment: 18 months
after the canteen expense reimbursement. Zhang withdrew 95,000 yuan ... and took an additional 3,000 yuan the following day.”

RAG Based Methods GraphRAG Evidence-Based Reasoning


Indexing Retrieval and Inference HierarGraph Multi-Agent Reasoning
Similarity: 0.94 Content: defendant Bai has...
Similarity: 0.92 Content: From 2009 to 2011...
Similarity: ... Content: ...
Tree Graph Similarity: 0.79 Content: Article 271 ...
Vector DB Missing key Articles for reasoning
Researcher Auditor Adjudicator
Due to a lack of key knowledge, I am unable
Flat Knowledge Construction to perform credible reasoning.
Execution Process
Answer Charge:Embezzlement Article: 382 Imprisonment: 36 months Searching Cases and Articles in HierarGraph
Article 382: If a state functionary... Case 1: defendant Bai has...
Article 271: Employees of... Case 2: From 2009 to 2011...
Legal Syllogism Based Methods Auditor each Articles with Fact
Article 382: If a state functionary... Case 1: defendant Bai has...
Major premise Minor premise Article 271: Employees of... Case 2: From 2009 to 2011...
Considering relevant articles Reasoning with fact
Article 382: If a state functionary... According to Article 382, ​I need to first check Final Judge Based on the Auditor’s result
Article 383: Those who commit... whether the defendant is a state employee, but I Defendant Zhang is found guilty of ... He was an employee of the project
Article 385: A state functionary... do not know the definition of a state employee? department [2] and used his position to illegally take possession of ...
Article 271: Employees of... [1] Criminal Law, Article 271 [2] Criminal Law, Article 92 [3] Judici...
Missing Articles because of Obstacles encountered during reasoning
Model bias due to a lack of external knowledge.

Answer Charge:Embezzlement Article: 382 Imprisonment: 36 months Answer Charge: Job-Related Embezzlement Article: 271 Imprisonment: 18 months

Figure 4: A comparative case study illustrating the reasoning trajectories of different methods. While Naive
RAG fails due to missing legal articles and syllogism-based methods struggle with ambiguities, LegalGraphRAG
derives the correct judgment. By leveraging the HierarGraph and Evidence-based Legal Reasoning, our framework
demonstrates transparency and reliability, providing a verifiable reasoning chain grounded in legal evidence.

Evaluation Metrics We employ Accuracy and


Micro-F1 score to evaluate prediction performance.
Detailed definitions are provided in Appendix D.3.
Implementation Details We utilize GPT-4o-
mini for graph construction and BGE-m3 (Chen
et al., 2024) for embedding generation. Various
LLMs serve as backbone models for the reasoning
phase. We employ Qwen3-8B (Yang et al., 2025a)
as the default backbone model for our main experi-
Figure 5: Retrieval Performance Comparison. Legal- ments. Full hyperparameter settings and hardware
GraphRAG demonstrates superior retrieval effectiveness specifications are detailed in Appendix D.5.
and significantly lower error ratios compared to conven-
tional flat graph baselines. 5.2 Generation Accuracy (Q1)
To address Q1, we evaluate LegalGraphRAG
into four distinct groups: (i) Open-Source Mod- against SOTA RAG methods and specialized le-
els, utilizing Qwen-series (Yang et al., 2025a), In- gal LLMs on two legal judgment datasets. The
ternLM (Fei et al., 2025) and GLM (GLM et al., primary comparison results for charge prediction
2024) as foundational backbones; (ii) Advanced are reported in Table 2, with extended analyses in
Models, represented by GPT-4o-mini (Achiam Tables 4, 5, and 6 in Appendix. We summarize the
et al., 2023) and DeepSeek-V3.1 (Liu et al., key observations below.
2024). (iii) Legal-Specific Methods, which in- Obs.1. LegalGraphRAG consistently outper-
clude domain-specialized approaches such as Disc- forms baselines in legal datasets. Our method
LLM (Yue et al., 2024), Legal∆ (Dai et al., achieves the best results on most evaluation metrics
2025), and ADAPT (Deng et al., 2024b); and (iv) across both datasets. Notably, LegalGraphRAG de-
RAG-Based Methods, encompassing Naive RAG livers significant improvements ranging from 6.3%
and advanced graph-augmented strategies like G- to 19.1% over the strongest baselines. Unlike stan-
retriever (He et al., 2024a), RAPTOR (Sarthi et al., dard GraphRAG methods that struggle in the legal
2024), LightRAG (Guo et al., 2024), and Hip- domain, our approach effectively structures het-
poRAG2 (Gutiérrez et al., 2025). Detailed con- erogeneous knowledge, thereby enhancing legal
figurations are provided in Appendix D.4. reasoning capabilities and improving charge pre-
37461
Traceable Correct Untraceable Correct
45
CAIL
40.5
40 2.4 Settings ACC ∆
35 32.6 31.9
LegalGraphRAG (Full) 40.9 –
30 28.1
Accuracy (%)

25 w/o HierarGraph 33.7 ↓ 7.2


19.2 19.6
20
15.6 20.1
38.1 w/o Researcher 36.9 ↓ 4.0
15
w/o Semantic Match 39.1 ↓ 1.8
10
14.3 w/o Community Exp. 38.5 ↓ 2.4
13.4 12.3
5
8.0 w/o Charge-Anchored 39.3 ↓ 1.6
0 1.3
G-retriever LightRAG Raptor HippoRAG2 LegalGraphRAG
(Ours)
w/o Auditor 37.5 ↓ 3.4

Figure 6: Reliability Analysis. LegalGraphRAG signif- Table 3: Ablation study of LegalGraphRAG compo-
icantly increases the proportion of Traceable Correct nents on the CAIL dataset. Results underscore the in-
samples, effectively minimizing Untraceable Correct dispensable role of the HierarGraph for knowledge or-
predictions where the answer is correct but lacks sup- ganization and the synergy between the Researcher and
porting evidence in the retrieved context. Auditor agents in ensuring reasoning accuracy.
diction accuracy overall. chain. LegalGraphRAG significantly increases the
Obs.2. LegalGraphRAG substantially surpasses ratio of “Traceable Correct” samples (defined in
existing specialized legal LLMs. Our approach Appendix A.5). By enforcing strict verification, our
outperforms Legal ∆ and ADAPT by an average of system ensures that every statute cited in the judg-
7.1% and 6.7%, respectively. Moreover, as shown ment is explicitly present in the retrieved context,
in Table 4 in Appendix, LegalGraphRAG integrates transforming opaque predictions into transparent,
flexibly with different backbone models, achieving traceable decisions.
a peak performance of 78.7% on CMDL when com-
bined with strong backbones. This demonstrates 5.4 Ablation Study (Q3)
strong adaptability and robust reasoning compared To quantify the impact of each component, we per-
to specialized legal-domain baselines. formed a systematic ablation study by removing
specific modules from the full LegalGraphRAG
5.3 Case Study (Q2)
framework. Results are detailed in Table 3.
To demonstrate the superior interpretability of our Obs.5. Hierarchical structure is the cornerstone
framework, we present a qualitative analysis of of performance. Removing the hierarchical graph
a representative criminal case in Figure 4. More (w/o HierarGraph) causes the sharpest accuracy
cases are provided in Appendix E. drop of 7.2%. This confirms that separating con-
Obs.3. LegalGraphRAG retrieves significantly crete facts from abstract rules into distinct granular
more relevant and comprehensive evidence. As levels is essential, providing structural precision
illustrated in Figure 5, conventional flat graph struc- that flat indexing lacks.
tures (e.g., HippoRAG2) struggle to handle hetero- Obs.6. The multi-agent workflow guarantees
geneous legal documents, often failing to capture reasoning reliability. Excluding the Researcher
essential statutes. This structural limitation leads and Auditor degrades accuracy by 4.0% and
to fragmented context. In contrast, our hierarchical 3.4%, respectively. This validates their synergis-
organization effectively structures legal knowledge, tic roles: the Researcher maximizes evidence cov-
ensuring that the retrieved context is sufficient to erage through diverse retrieval strategies, while
support robust reasoning. the Auditor enforces rigorous verification, ensuring
Obs.4. LegalGraphRAG guarantees decision only validated evidence supports the judgment.
traceability through rigorous evidence ground-
ing. While baseline models often achieve cor- 6 Conclusion
rect predictions, our reliability analysis (Figure In conclusion, we have presented LegalGraphRAG,
6) reveals a critical issue of “unsupported correct- an evidence-based legal reasoning framework that
ness”, where the model predicts the right charge addresses the critical challenges of legal hetero-
but fails to retrieve the necessary supporting evi- geneity and reasoning reliability. By integrating a
dence. This implies that the prediction is not sup- hierarchical knowledge graph with a collaborative
ported by relevant evidence or a valid reasoning multi-agent system, our approach transforms the le-
37462
gal reasoning process into a transparent pipeline of Bias and Fairness We acknowledge that mod-
retrieval, verification, and synthesis. Extensive ex- els trained on historical legal judgment data may
periments on legal judgment benchmarks validate inadvertently capture or amplify inherent biases
that LegalGraphRAG establishes a new state-of- present in the judicial system, such as those related
the-art, significantly advancing accurate and trust- to region or gender. While our work focuses on
worthy AI for reliable and complex legal analysis. improving the logical reasoning and retrieval capa-
bilities of legal LLMs through GraphRAG, where
Limitation the outputs are interpreted with clear evidence.
While LegalGraphRAG demonstrates significant Intended Use and Misuse The proposed Legal-
proficiency in processing textual legal documents GraphRAG is designed as an assistive tool to sup-
and statutes, its current scope is confined to uni- port legal professionals and researchers in retriev-
modal textual inputs. Real-world judicial proceed- ing precedents and analyzing case facts. It is not
ings, however, often rely on a heterogeneity of intended to replace human judges or lawyers, nor
evidence types, including crime scene photogra- should it be deployed as a fully automated decision-
phy, surveillance footage, scanned handwritten doc- making system in real-world judicial scenarios.
uments, and audio recordings of court hearings. The “prison term” and “judgment” predictions gen-
Currently, our framework requires all non-textual erated by the model should be viewed as reference
evidence to be transcribed or described textually probabilities rather than enforceable verdicts.
before processing, which may result in the loss
of critical visual or auditory nuances essential for Acknowledgements
fact verification. For instance, distinguishing be-
tween “inten” and “negligence” might sometimes The project was supported by National Key
rely on visual cues in surveillance video that tex- R&D Program of China (No. 2022ZD0160501),
tual descriptions fail to capture fully. Extending Natural Science Foundation of Fujian Province
the Hierarchical Legal Knowledge Graph to incor- of China (No. 2024J011001), and the Public
porate multimodal nodes (e.g., embedding visual Technology Service Platform Project of Xiamen
evidence into the Fact Graph) represents a promis- (No.3502Z20231043). We also thank the reviewers
ing avenue for future research. Such an extension for their insightful comments.
would enable the model to perform cross-modal
reasoning, verifying textual testimony against vi-
References
sual evidence, thereby moving closer to a holistic
and robust “Smart Court” system. Josh Achiam, Steven Adler, Sandhini Agarwal, Lama
Ahmad, Ilge Akkaya, Florencia Leoni Aleman,
Ethics Statement Diogo Almeida, Janko Altenschmidt, Sam Altman,
Shyamal Anadkat, and 1 others. 2023. Gpt-4 techni-
We confirm that this study fully complies with the cal report. arXiv preprint arXiv:2303.08774.
ACL Ethics Policy. Below, we address specific eth- Sebastian Borgeaud, Arthur Mensch, Jordan Hoff-
ical considerations regarding the data and the ap- mann, Trevor Cai, Eliza Rutherford, Katie Milli-
plication of our proposed model, LegalGraphRAG. can, George Bm Van Den Driessche, Jean-Baptiste
Lespiau, Bogdan Damoc, Aidan Clark, and 1 others.
Data Privacy and Compliance Our experiments 2022. Improving language models by retrieving from
involve four publicly available datasets (CAIL2018, trillions of tokens. In International conference on
CMDL, JuDGE, and LeCaRDv2) and statutory machine learning. PMLR.
texts. These resources are established benchmarks Zhiwei Cao, Qian Cao, Yu Lu, Ningxin Peng, Luyang
in the legal NLP community. We emphasize that Huang, Shanbo Cheng, and Jinsong Su. 2024. Retain-
all court judgments utilized in this work have been ing key information under high compression ratios:
Query-guided compressor for llms. In Proceedings
pre-processed and anonymized by the original data
of the 62nd Annual Meeting of the Association for
providers. Private details, including the real names Computational Linguistics (Volume 1: Long Papers),
of defendants and victims, have been removed or pages 12685–12695.
masked to ensure no personally identifiable infor-
Zhiwei Cao, Baosong Yang, Huan Lin, Suhang Wu, Xi-
mation (PII) is exposed. We strictly use this data angpeng Wei, Dayiheng Liu, Jun Xie, Min Zhang,
for academic research purposes and adhere to their and Jinsong Su. 2023. Bridging the domain gaps in
respective data usage licenses. context representations for k-nearest neighbor neural
37463
machine translation. In Proceedings of the 61st An- Zhiwei Fei, Songyang Zhang, Xiaoyu Shen, Dawei
nual Meeting of the Association for Computational Zhu, Xiao Wang, Jidong Ge, and Vincent Ng. 2025.
Linguistics (Volume 1: Long Papers), pages 5841– Internlm-law: An open-sourced chinese legal large
5853. language model. In Proceedings of the 31st Interna-
tional Conference on Computational Linguistics.
Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu
Lian, and Zheng Liu. 2024. Bge m3-embedding: Yan Gao, Zhiwei Cao, Zhongjian Miao, Baosong Yang,
Multi-lingual, multi-functionality, multi-granularity Shiyu Liu, Min Zhang, and Jinsong Su. 2024. Ef-
text embeddings through self-knowledge distillation. ficient k-nearest-neighbor machine translation with
arXiv preprint arXiv:2402.03216. dynamic retrieval. In Findings of the Association for
Computational Linguistics: ACL 2024, pages 7990–
Pierre Colombo, Telmo Pessoa Pires, Malik Boudiaf,
8001.
Dominic Culver, Rui Melo, Caio Corro, Andre FT
Martins, Fabrizio Esposito, Vera Lúcia Raposo, Sofia Sudipto Ghosh, Devanshu Verma, Balaji Ganesan, Purn-
Morgado, and 1 others. 2024. Saullm-7b: A pioneer- ima Bindal, Vikas Kumar, and Vasudha Bhatnagar.
ing large language model for law. arXiv preprint 2024. Inlegalllama: Indian legal knowledge en-
arXiv:2403.03883. hanced large language model. In International Joint
Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Conference on Artificial Intelligence.
Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Mar-
cel Blistein, Ori Ram, Dan Zhang, Evan Rosen, and Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, Chen-
1 others. 2025. Gemini 2.5: Pushing the frontier with hui Zhang, Da Yin, Dan Zhang, Diego Rojas, Guanyu
advanced reasoning, multimodality, long context, and Feng, Hanlin Zhao, and 1 others. 2024. Chatglm: A
next generation agentic capabilities. arXiv preprint family of large language models from glm-130b to
arXiv:2507.06261. glm-4 all tools. arXiv preprint arXiv:2406.12793.

Jiaxi Cui, Munan Ning, Zongjian Li, Bohua Chen, Yang Zirui Guo, Lianghao Xia, Yanhua Yu, Tu Ao, and
Yan, Hao Li, Bin Ling, Yonghong Tian, and Li Yuan. Chao Huang. 2024. Lightrag: Simple and fast
2023. Chatlaw: A multi-agent collaborative legal retrieval-augmented generation. arXiv preprint
assistant with knowledge graph enhanced mixture- arXiv:2410.05779.
of-experts large language model. arXiv preprint
arXiv:2306.16092. Bernal Jiménez Gutiérrez, Yiheng Shu, Weijian Qi,
Sizhe Zhou, and Yu Su. 2025. From rag to memory:
Xin Dai, Buqiang Xu, Zhenghao Liu, Yukun Yan, Non-parametric continual learning for large language
Huiyuan Xie, Xiaoyuan Yi, Shuo Wang, and Ge Yu. models. arXiv preprint arXiv:2502.14802.
2025. Legal δ: Enhancing legal reasoning in llms via
reinforcement learning with chain-of-thought guided Zhang Han and Dou Zhicheng. 2023. Case retrieval
information gain. arXiv preprint arXiv:2508.12281. for legal judgment prediction in legal artificial intelli-
gence. In Proceedings of the 22nd Chinese National
Hudson de Martim. 2025. Graph rag for legal norms: A Conference on Computational Linguistics.
hierarchical and temporal approach. arXiv preprint
arXiv:2505.00039. Xiaoxin He, Yijun Tian, Yifei Sun, Nitesh Chawla,
Thomas Laurent, Yann LeCun, Xavier Bresson,
Chenlong Deng, Kelong Mao, and Zhicheng Dou. and Bryan Hooi. 2024a. G-retriever: Retrieval-
2024a. Learning interpretable legal case retrieval augmented generation for textual graph understand-
via knowledge-guided case reformulation. arXiv ing and question answering. Advances in Neural
preprint arXiv:2406.19760. Information Processing Systems, 37:132876–132907.
Chenlong Deng, Kelong Mao, Yuyao Zhang, and
Zhicheng Dou. 2024b. Enabling discriminative rea- Zhitao He, Pengfei Cao, Chenhao Wang, Zhuoran Jin,
soning in llms for legal judgment prediction. arXiv Yubo Chen, Jiexin Xu, Huaijun Li, Kang Liu, and
preprint arXiv:2407.01964. Jun Zhao. 2024b. Agentscourt: Building judicial
decision-making agents with court debate simula-
Darren Edge, Ha Trinh, Newman Cheng, Joshua tion and legal knowledge augmentation. In Find-
Bradley, Alex Chao, Apurva Mody, Steven Truitt, ings of the Association for Computational Linguis-
Dasha Metropolitansky, Robert Osazuwa Ness, and tics: EMNLP 2024.
Jonathan Larson. 2024. From local to global: A
graph rag approach to query-focused summarization. Mengzhe Hei, Qingbao Liu, Sheng Zhang, Honglin
arXiv preprint arXiv:2404.16130. Shi, Jiashun Duan, and Xin Zhang. 2024. A het-
erogeneous graph based on legal documents and le-
Zhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou, gal statute hierarchy for chinese legal case retrieval.
Zhuo Han, Alan Huang, Songyang Zhang, Kai Chen, IEEE Access, 12:93502–93516.
Zhixin Yin, Zongwen Shen, and 1 others. 2024. Law-
bench: Benchmarking legal knowledge of large lan- Justin Ho, Alexandra Colby, and William Fisher. 2025.
guage models. In Proceedings of the 2024 conference Incorporating legal structure in retrieval-augmented
on empirical methods in natural language process- generation: A case study on copyright fair use. arXiv
ing. preprint arXiv:2505.02164.
37464
Zhitian Hou, Zihan Ye, Nanli Zeng, Tianyong Hao, and legal consultation conversation. In Proceedings of
Kun Zeng. 2025. Large language models meet le- the 48th International ACM SIGIR Conference on
gal artificial intelligence: A survey. arXiv preprint Research and Development in Information Retrieval.
arXiv:2509.09969.
Haitao Li, Yunqiu Shao, Yueyue Wu, Qingyao Ai, Yix-
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan iao Ma, and Yiqun Liu. 2024. Lecardv2: A large-
Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, scale chinese legal case retrieval dataset. In Proceed-
Weizhu Chen, and 1 others. 2022. Lora: Low-rank ings of the 47th International ACM SIGIR Confer-
adaptation of large language models. ICLR, 1(2):3. ence on Research and Development in Information
Retrieval.
Wanhong Huang, Yi Feng, Chuanyi Li, Honghan Wu, Ji-
dong Ge, and Vincent Ng. 2024. Cmdl: A large-scale Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang,
chinese multi-defendant legal judgment prediction Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi
dataset. In Findings of the Association for Computa- Deng, Chenyu Zhang, Chong Ruan, and 1 others.
tional Linguistics ACL 2024. 2024. Deepseek-v3 technical report. arXiv preprint
arXiv:2412.19437.
Cong Jiang and Xiaolei Yang. 2023. Legal syllogism
prompting: Teaching large language models for legal Qichuan Liu, Chentao Zhang, Yuxuan Hu, Chenfeng
judgment prediction. In Proceedings of the nine- Zheng, Qinggang Zhang, and Zhihong Zhang. 2026.
teenth international conference on artificial intelli- Facilitating generative retrieval with logical denois-
gence and law. ing for interpretable conversational search. In Pro-
ceedings of the ACM Web Conference 2026, page
Hui Jiang, Ziyao Lu, Fandong Meng, Chulun Zhou, 2296–2307.
Jie Zhou, Degen Huang, and Jinsong Su. 2022. To-
wards robust k-nearest-neighbor machine translation. Qichuan Liu, Chentao Zhang, Chenfeng Zheng, Gu-
In Proceedings of the 2022 Conference on Empiri- osheng Hu, Xiaodong Li, and Zhihong Zhang. 2025.
cal Methods in Natural Language Processing, pages Beyond the answer: Advancing multi-hop QA with
5468–5477. fine-grained graph reasoning and evaluation. In Pro-
ceedings of the 63rd Annual Meeting of the Associa-
Vladimir Karpukhin, Barlas Oguz, Sewon Min, tion for Computational Linguistics (Volume 1: Long
Patrick SH Lewis, Ledell Wu, Sergey Edunov, Danqi Papers), pages 23433–23456.
Chen, and Wen-tau Yih. 2020. Dense passage re-
trieval for open-domain question answering. In Antoine Louis, Gijs Van Dijck, and Gerasimos Spanakis.
EMNLP (1), pages 6769–6781. 2023. Finding the law: Enhancing statutory article
retrieval via graph neural networks. arXiv preprint
Zixuan Ke, Fangkai Jiao, Yifei Ming, Xuan-Phi Nguyen, arXiv:2301.12847.
Austin Xu, Do Xuan Long, Minzhi Li, Chengwei Qin,
Peifeng Wang, Silvio Savarese, and 1 others. 2025. Antoine Louis, Gijs van Dijck, and Gerasimos Spanakis.
A survey of frontiers in llm reasoning: Inference scal- 2024. Interpretable long-form legal question answer-
ing, learning to reason, and agentic systems. arXiv ing with retrieval-augmented large language models.
preprint arXiv:2504.09037. In Proceedings of the AAAI Conference on Artificial
Intelligence, volume 38.
Hyunjae Kim, Jiwoong Sohn, Aidan Gilson, Nicholas
Cochran-Caggiano, Serina Applebaum, Heeju Jin, Yun Luo, Zhen Yang, Fandong Meng, Yafu Li, Jie Zhou,
Seihee Park, Yujin Park, Jiyeong Park, Seoyoung and Yue Zhang. 2025. An empirical study of catas-
Choi, and 1 others. 2025. Rethinking retrieval- trophic forgetting in large language models during
augmented generation for medicine: A large-scale, continual fine-tuning. IEEE Transactions on Audio,
systematic expert evaluation and practical insights. Speech and Language Processing.
arXiv preprint arXiv:2511.06738.
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das,
Jinqi Lai, Wensheng Gan, Jiayang Wu, Zhenlian Qi, and Daniel Khashabi, and Hannaneh Hajishirzi. 2023.
Philip S Yu. 2024. Large language models in law: A When not to trust language models: Investigating
survey. AI Open, 5:181–196. effectiveness of parametric and non-parametric mem-
ories. In Proceedings of the 61st Annual Meeting of
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio the Association for Computational Linguistics (Vol-
Petroni, Vladimir Karpukhin, Naman Goyal, Hein- ume 1: Long Papers).
rich Küttler, Mike Lewis, Wen-tau Yih, Tim Rock-
täschel, and 1 others. 2020. Retrieval-augmented gen- Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad
eration for knowledge-intensive nlp tasks. Advances Saqib, Saeed Anwar, Muhammad Usman, Naveed
in neural information processing systems, 33:9459– Akhtar, Nick Barnes, and Ajmal Mian. 2025. A com-
9474. prehensive overview of large language models. ACM
Transactions on Intelligent Systems and Technology,
Haitao Li, Yifan Chen, Hu YiRan, Qingyao Ai, Jun- 16(5):1–72.
jie Chen, Xiaoyu Yang, Jianhui Yang, Yueyue Wu,
Zeyang Liu, and Yiqun Liu. 2025. Lexrag: Bench- Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida,
marking retrieval-augmented generation in multi-turn Carroll Wainwright, Pamela Mishkin, Chong Zhang,
37465
Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 Zhen Wan, Yating Zhang, Yexiang Wang, Fei Cheng,
others. 2022. Training language models to follow in- and Sadao Kurohashi. 2024. Reformulating domain
structions with human feedback. Advances in neural adaptation of large language models as adapt-retrieve-
information processing systems, 35:27730–27744. revise: A case study on chinese legal domain. In
Findings of the Association for Computational Lin-
Xiao Peng and Liang Chen. 2024. Athena: Retrieval- guistics: ACL 2024.
augmented legal judgment prediction with large lan-
guage models. arXiv preprint arXiv:2410.11195. Cunxiang Wang, Xiaoze Liu, Yuanhao Yue, Xiangru
Tang, Tianhang Zhang, Cheng Jiayang, Yunzhi Yao,
Nicholas Pipitone and Ghita Houir Alami. 2024. Wenyang Gao, Xuming Hu, Zehan Qi, and 1 others.
Legalbench-rag: A benchmark for retrieval- 2023. Survey on factuality in large language models:
augmented generation in the legal domain. arXiv Knowledge, retrieval and domain-specificity. arXiv
preprint arXiv:2408.10343. preprint arXiv:2310.07521.
Xuran Wang, Xinguang Zhang, Vanessa Hoo, Zhouhang
B. Rüthers, C. Fischer, and A. Birk. 2013. Rechtstheo- Shao, and Xuguang Zhang. 2024. Legalreasoner: A
rie mit juristischer Methodenlehre. Grundrisse des multi-stage framework for legal judgment prediction
Rechts. C.H. Beck. via large language models and knowledge integration.
IEEE Access.
Pranab Sahoo, Ayush Kumar Singh, Sriparna Saha,
Vinija Jain, Samrat Mondal, and Aman Chadha. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten
2024. A systematic survey of prompt engineering in Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou,
large language models: Techniques and applications. and 1 others. 2022. Chain-of-thought prompting elic-
arXiv preprint arXiv:2402.07927. its reasoning in large language models. Advances
in neural information processing systems, 35:24824–
Parth Sarthi, Salman Abdullah, Aditi Tuli, Shubh 24837.
Khanna, Anna Goldie, and Christopher D Manning.
2024. Raptor: Recursive abstractive processing for Hannes Westermann. 2024. Dallma: Semi-structured
tree-organized retrieval. In The Twelfth International legal reasoning and drafting with large language mod-
Conference on Learning Representations. els. In 2nd Workshop on Generative AI and Law.
Nirmalie Wiratunga, Ramitha Abeyratne, Lasal Jayawar-
Jeffrey A Segal. 1984. Predicting supreme court cases dena, Kyle Martin, Stewart Massie, Ikechukwu Nkisi-
probabilistically: The search and seizure cases, 1962- Orji, Ruvan Weerasinghe, Anne Liret, and Bruno
1981. American Political Science Review, 78(4):891– Fleisch. 2024. Cbr-rag: case-based reasoning for
900. retrieval augmented generation in llms for legal ques-
tion answering. In International Conference on Case-
Dong Shu, Haoran Zhao, Xukun Liu, David Demeter, Based Reasoning. Springer.
Mengnan Du, and Yongfeng Zhang. 2024. Lawllm:
Law large language model for the us legal system. In Shiguang Wu, Zhongkun Liu, Zhen Zhang, Zheng Chen,
Proceedings of the 33rd ACM International Confer- Wentao Deng, Wenhao Zhang, Jiyuan Yang, Zhi-
ence on information and knowledge management. tao Yao, Yougang Lyu, Xin Xin, Shen Gao, Pengjie
Ren, Zhaochun Ren, and Zhumin Chen. 2023a.
Marco Siino, Mariana Falco, Daniele Croce, and Paolo [Link].
Rosso. 2025. Exploring llms applications in law:
A literature review on current legal nlp approaches. Yiquan Wu, Siying Zhou, Yifei Liu, Weiming Lu, Xi-
IEEE Access. aozhong Liu, Yating Zhang, Changlong Sun, Fei Wu,
and Kun Kuang. 2023b. Precedent-enhanced legal
Weihang Su, Baoqing Yue, Qingyao Ai, Yiran Hu, Jiaqi judgment prediction with llm and domain-model col-
Li, Changyue Wang, Kaiyuan Zhang, Yueyue Wu, laboration. arXiv preprint arXiv:2310.09241.
and Yiqun Liu. 2025. Judge: Benchmarking judg-
ment document generation for chinese legal system. Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yi-
In Proceedings of the 48th International ACM SI- wen Ding, Boyang Hong, Ming Zhang, Junzhe Wang,
GIR Conference on Research and Development in Senjie Jin, Enyu Zhou, and 1 others. 2025. The
Information Retrieval. rise and potential of large language model based
agents: A survey. Science China Information Sci-
ences, 68(2):121101.
Octavia-Maria Sulea, Marcos Zampieri, Shervin Mal-
masi, Mihaela Vela, Liviu P Dinu, and Josef Van Gen- Zhishang Xiang, Chuanjie Wu, Qinggang Zhang,
abith. 2017. Exploring the use of text classification in Shengyuan Chen, Zijin Hong, Xiao Huang, and Jin-
the legal domain. arXiv preprint arXiv:1710.09306. song Su. 2025. When to use graphs in rag: A com-
prehensive analysis for graph retrieval-augmented
Vincent A Traag, Ludo Waltman, and Nees Jan Van Eck. generation. arXiv preprint arXiv:2506.05690.
2019. From louvain to leiden: guaranteeing well-
connected communities. Scientific reports, 9(1):1– Zhishang Xiang, Chengyi Yang, Zerui Chen, Zhimin
12. Wei, Yunbo Tang, Zongpei Teng, Zexi Peng, Zongxia
37466
Li, Chengsong Huang, Yicheng He, Chang Yang, Jianqiiu Zhang. 2024. Should we fear large language
Xinrun Wang, Xiao Huang, Qinggang Zhang, and Jin- models? a structural analysis of the human reason-
song Su. 2026. A systematic survey of self-evolving ing system for elucidating llm capabilities and risks
agents: From model-centric to environment-driven through the lens of heidegger’s philosophy. arXiv
co-evolution. TechRxiv, 2026(0227). preprint arXiv:2403.03288.

Chaojun Xiao, Haoxi Zhong, Zhipeng Guo, Cunchao Qinggang Zhang, Shengyuan Chen, Yuanchen Bei,
Tu, Zhiyuan Liu, Maosong Sun, Yansong Feng, Xi- Zheng Yuan, Huachi Zhou, Zijin Hong, Hao Chen,
anpei Han, Zhen Hu, Heng Wang, and 1 others. 2018. Yilin Xiao, Chuang Zhou, Junnan Dong, and 1 others.
Cail2018: A large-scale legal dataset for judgment 2025a. A survey of graph retrieval-augmented gener-
prediction. arXiv preprint arXiv:1807.02478. ation for customized large language models. arXiv
preprint arXiv:2501.13958.
Nuo Xu, Pinghui Wang, Long Chen, Li Pan, Xiaoyan
Wang, and Junzhou Zhao. 2020. Distinguish confus- Qinggang Zhang, Zhishang Xiang, Yilin Xiao, Le Wang,
ing law articles for legal judgment prediction. arXiv Junhui Li, Xinrun Wang, and Jinsong Su. 2025b.
preprint arXiv:2004.02557. Faithfulrag: Fact-level conflict modeling for context-
faithful retrieval-augmented generation. In Proceed-
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, ings of the 63rd Annual Meeting of the Association
Binyuan Hui, Bo Zheng, Bowen Yu, Chang for Computational Linguistics.
Gao, Chengen Huang, Chenxu Lv, and 1 others.
2025a. Qwen3 technical report. arXiv preprint Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang,
arXiv:2505.09388. Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen
Zhang, Junjie Zhang, Zican Dong, and 1 others. 2023.
Chang Yang, Chuang Zhou, Yilin Xiao, Su Dong, A survey of large language models. arXiv preprint
Luyao Zhuang, Yujing Zhang, Zhu Wang, Zijin Hong, arXiv:2303.18223, 1(2).
Zheng Yuan, Zhishang Xiang, and 1 others. 2026.
Graph-based agent memory: Taxonomy, techniques, Chulun Zhou, Chunkang Zhang, Guoxin Yu, Fandong
and applications. arXiv preprint arXiv:2602.05665. Meng, Jie Zhou, Wai Lam, and Mo Yu. 2025. Improv-
ing multi-step rag with hypergraph-based memory
Rui Yang. 2024. Casegpt: a case reasoning framework for long-context complex relational modeling. arXiv
based on language models and retrieval-augmented preprint arXiv:2512.23959.
generation. Preprint, arXiv:2407.07913.
Zhi Zhou, Jiang-Xin Shi, Peng-Xiao Song, Xiao-Wen
Xinyu Yang, Chenlong Deng, and Zhicheng Dou. 2025b. Yang, Yi-Xuan Jin, Lan-Zhe Guo, and Yu-Feng
Glare: Agentic reasoning for legal judgment predic- Li. 2024. Lawgpt: A chinese legal knowledge-
tion. arXiv preprint arXiv:2508.16383. enhanced large language model. arXiv preprint
arXiv:2406.04614.
Fangyi Yu, Lee Quartey, and Frank Schilder. 2022. Le-
gal prompting: Teaching a language model to think
like a lawyer. Preprint, arXiv:2212.01326.

Weikang Yuan, Junjie Cao, Zhuoren Jiang, Yangyang


Kang, Jun Lin, Kaisong Song, Tianqianjin Lin, Peng-
wei Yan, Changlong Sun, and Xiaozhong Liu. 2024.
Can large language models grasp legal theories? en-
hance legal reasoning with insights from multi-agent
collaboration. In Findings of the Association for
Computational Linguistics: EMNLP 2024, Miami,
Florida, USA. Association for Computational Lin-
guistics.

Shengbin Yue, Wei Chen, Siyuan Wang, Bingxuan Li,


Chenchen Shen, Shujun Liu, Yuxuan Zhou, Yao
Xiao, Song Yun, Xuanjing Huang, and 1 others.
2023. Disc-lawllm: Fine-tuning large language mod-
els for intelligent legal services. arXiv preprint
arXiv:2309.11325.

Shengbin Yue, Shujun Liu, Yuxuan Zhou, Chenchen


Shen, Siyuan Wang, Yao Xiao, Bingxuan Li, Yun
Song, Xiaoyu Shen, Wei Chen, and 1 others. 2024.
Lawllm: Intelligent legal system with legal reason-
ing and verifiable retrieval. In International Confer-
ence on Database Systems for Advanced Applications.
Springer.
37467
A Frequently Asked Questions (FAQs) Validating Hierarchical Knowledge Alignment
Introduction Challenge (i) highlights the difficulty
A.1 What are the advantages of of managing heterogeneous knowledge. LJP ex-
LegalGraphRAG? emplifies this struggle by requiring the model to
LegalGraphRAG introduces several key advance- bridge the semantic gap between concrete case
ments over traditional retrieval-augmented genera- facts and abstract statutory rules. Successfully
tion methods and specialized legal LLMs, address- mapping these distinct granularities validates the
ing critical challenges in the legal domain through effectiveness of our hierarchical graph structure in
its hierarchical structure and multi-agent workflow. organizing multi-level domain knowledge.
Superior Retrieval Effectiveness. First, our Evaluating Rigorous Logical Deduction Unlike
framework significantly improves how legal infor- general Question Answering tasks that may rely
mation is retrieved. While traditional flat retrieval on surface-level semantic matching, LJP necessi-
methods often struggle to differentiate between spe- tates strict syllogistic reasoning (Major Premise
cific case facts and abstract statutory rules, our hier- → Minor Premise → Conclusion). This structural
archical graph organizes this complex information dependency provides an ideal setting to stress-test
into distinct levels. This structure ensures that the our Multi-Agent framework, specifically validating
system captures both detailed evidence and high- whether the Auditor can effectively filter irrelevant
level principles, providing a much more compre- distractions and enforce the logical consistency.
hensive context than standard baselines.
Benchmarking High-Stakes Reliability In pro-
Trustworthy and Transparent Reasoning. Sec- fessional domains, plausibility is insufficient; accu-
ond, LegalGraphRAG addresses the “black box” racy is paramount. LJP imposes a zero-tolerance
issue common in standard LLMs. Instead of gen- standard for hallucination, as every judgment must
erating answers directly, which can lead to hallu- be supported by cited articles. By demonstrat-
cinations or correct predictions based on wrong ing that LegalGraphRAG can produce verifiable,
premises, our system employs a multi-agent work- evidence-based judgments in this demanding con-
flow. This process strictly verifies the retrieved evi- text, we establish a strong precedent for its applica-
dence against the facts of the case. Consequently, bility to critical domains like medicine and finance.
it constructs a logical chain of evidence, ensuring
A.3 How was the retrieval performance
that the final judgment is grounded in valid legal
evaluated and compared across different
logic rather than statistical probability.
strategies?
Flexibility and Model Agnosticism. Finally, the To ensure consistent assessment throughout our
framework offers superior flexibility compared to study (spanning both the preliminary investigation
rigid, specialized legal models. Unlike methods and main comparative experiments), we established
that require extensive and costly fine-tuning on le- a standardized evaluation pipeline based on the le-
gal datasets, LegalGraphRAG functions as a mod- gal corpora and CAIL (Xiao et al., 2018) dataset
ular system. It allows users to easily swap the described in Appendix D.2. The evaluation proce-
underlying backbone model. As demonstrated in dure consists of three steps.
our experiments, this capability enables the inte-
Execution on Test Set For every case query in
gration of powerful advanced models to achieve
the test dataset, we executed two representative
state-of-the-art performance without the need for
RAG strategies. The Flat Strategy follows the tra-
additional training.
ditional baseline approach, indexing all legal docu-
A.2 How to Evaluate Legal Reasoning? ments in a unified flat repository. The Hierarchical
Strategy, built upon Naive RAG, adopts a decou-
We utilize Legal Judgment Prediction (LJP) as the
pled approach by separately storing and retrieving
primary experimental testbed because it serves as
legal articles and historical cases. Based on these
a rigorous “cognitive touchstone” for evaluating
strategies, each model retrieved a set of candidate
complex reasoning in specialized domains. While
evidence from the corpus.
our introduction highlights broader challenges in
healthcare, finance, and law, LJP uniquely encapsu- Ground Truth Alignment We utilized the arti-
lates the core difficulties of high-stakes reasoning. cles provided in the dataset (ground truth articles)
37468
as the “Gold Standard,” as all cases in the dataset By mixing these irrelevant documents into the con-
are inherently annotated with their relevant statu- text, we created a challenging environment that
tory articles. Any retrieved node matching these forces the model to discern legal essence from su-
articles was marked as a True Positive. perficial similarity.
Metric Calculation Based on the alignment re- A.5 Reliability Analysis Definitions
sults, we quantified performance using the two key
We analyzed the CAIL test results to categorize
indicators defined in Appendix D.3:
correct predictions based on evidence support. A
• Retrieval Effectiveness: This measures the Re- prediction is classified as Traceable Correct if the
call of gold-standard evidence, indicating the model correctly predicts the charge and success-
system’s ability to locate legal evidence. fully retrieves the ground-truth articles. Conversely,
it is Untraceable Correct if the correct charge is
• Error Rate: This assesses the proportion of predicted despite failing to retrieve the necessary
irrelevant or misleading nodes within the re- articles.
trieved context reflecting the system’s ability
to filter distractions. B Method Details
This allows us to objectively compare how dif- In this section, we provide the comprehensive tech-
ferent structural approaches (flat vs. hierarchical) nical specifications and implementation details of
impact the precision of legal reasoning. the proposed LegalGraphRAG framework. As
outlined in the main text, our approach operates
A.4 How were the “context with irrelevant
in two distinct phases: (i) Hierarchical Knowl-
information” constructed for the
edge Construction (Figure 7), which organizes
Generation Quality investigation?
legal knowledge into a layered graph structure to
To rigorously test the model’s verification capabili- effectively decouple historical cases, relevant arti-
ties, we constructed evaluation contexts containing cles, and judicial interpretations; and (ii) Evidence-
High-Similarity Irrelevant Information. Instead of based Legal Reasoning (Figure 8), which employs
including randomly selected texts, we curated sets a collaborative agent workflow to retrieve relevant
of documents that are semantically similar to the evidence and generate verifiable judgments.
correct evidence but legally inapplicable. This de-
sign mirrors real-world scenarios where documents B.1 Hierarchical Knowledge Construction
share surface-level keywords but differ fundamen- We construct a Hierarchical Legal Graph G com-
tally in domain applicability. The construction pro- posed of three specialized subgraphs, as illustrated
cess involved two steps: in Figure 3. This multi-layered structure explicitly
differentiates between specific precedents, abstract
Ground Truth Context: First, we established
case relationships, and rigorous statutory rules, as
the baseline context using the ground truth pro-
illustrated in Figure 7.
vided in the CAIL (Xiao et al., 2018) dataset. For
Fact Graph (Gf ac ) serves as the repository for
each case, this set consists exclusively of the cor-
ground-truth precedents. It encodes the natural
rect applicable articles required for the judgment.
structure of legal documents by explicitly linking
Injection of Irrelevant Distractors: To simulate Cases (C), Articles (A), and Offenses (O). Edges
the presence of legally plausible but factually irrel- are established to represent citation relationships
evant documents, we utilized the entire Criminal (eca : C → A) and conviction outcomes (eco : C →
Law code as a retrieval corpus. For each correct O). Formally, it is defined as:
article in the Ground Truth Context, we performed  
|Gf ac |
a vector-based similarity search over this corpus to Gf ac = (Vf ac , Ef ac ) = ci , ai , oi i=1 .
identify the top-k most similar articles that were (13)
not part of the ground truth. Ontology Graph (Gont ) abstracts case features to
These retrieved articles serve as High-Similarity model inter-case relationships. To map unstruc-
Distractors: they share significant lexical and se- tured narratives into a structured semantic space,
mantic overlap with the correct laws (e.g., sharing we define a domain-specific ontology along four
keywords like “theft” or “fraud”) but differ in spe- dimensions: Defendant Attributes, Criminal Behav-
cific constitutive elements or sentencing standards. iors, Victim Characteristics, and Subjective Mental
37469
Figure 7: Overview of the Hierarchical Knowledge Construction phase in LegalGraphRAG.

States. Keywords are extracted and aligned with where


these dimensions to form Case Feature Nodes (F). D(ai ) = {d1 , . . . , d|C| }. (16)
Structurally, we utilize the k-Nearest Neighbors
(k-NN) algorithm (Jiang et al., 2022; Cao et al., By integrating these three layers, HierarGraph
2023; Gao et al., 2024) to establish semantic edges G transforms heterogeneous legal corpora into a
between cases. Based on this topology, we ap- structured ecosystem. This architecture directly
ply the Leiden algorithm (Traag et al., 2019) to addresses the limitations of flat retrieval by offering
cluster related cases into Community Nodes (K), multi-granular support for our multi-agent system.
facilitating coarse-to-fine retrieval. The subgraph
is formally defined as: B.2 Evidence-based Legal Reasoning
 
|Gont |
Gont = (Vont , Eont ) = ci , kj i=1,j=1 . (14) We propose a multi-agent framework to emulate
the rigorous workflow of legal professionals, as il-
Rule Graph (Grul ) incorporates fine-grained legal lustrated in Figure 8. This system operates sequen-
knowledge to resolve statutory ambiguities. This tially through three specialized agents (Researcher,
graph consists of Articles (A) and Judicial Inter- Auditor, and Adjudicator) to transform a case
pretations (J ), linked by explicit cross-references. query into a verifiable judgment.
To further enhance precision, each article node ai Researcher Agent: This agent is responsible for
is equipped with a Diagnostic Checklist D(ai ). grounding the unstructured case query in relevant
Generated by parsing statutory texts, this check- legal knowledge. First, it aligns the raw case de-
list decomposes complex legal provisions into scription with the ontology structure in Gont , ex-
atomic boolean queries. For instance, regarding tracting standardized evidentiary features (e.g., de-
Article 266 (Fraud), the checklist validates the logi- fendant characteristics and criminal behaviors).
cal chain of the crime: “Did the defendant fabricate
Based on these features, we formulate the evi-
facts or conceal the truth?”, “Did the victim fall
dence retrieval process R(q) as the union of three
into a mistake due to this act?”, and “Did the victim
parallel strategies, where q is the legal query:
dispose of property based on this mistake?”. This
mechanism forces the model to verify each consti-
R(q) = Rsem (q) ∪ Rcom (q) ∪ Rchg (q) (17)
tutive element step-by-step, rather than relying on
vague semantic overlaps. Formally, this subgraph
(i) Semantic Match Retrieval: We first locate di-
and its associated checklists are defined as:
  rect evidentiary analogues via fine-grained seman-
|Grul |
Grul = (Vrul , Erul ) = ai , ji i=1 , (15) tic similarity. Let ϕ(·) denote the ontology-aligned
37470
Figure 8: The workflow of the Evidence-based Legal Reasoning phase.

embeddings, this process is defined as: (i) Diagnostic Retrieval: For each article va , the
 agent retrieves its specific Diagnostic Checklist
Rsem (q) = Top-k sim ϕ(q), ϕ(c) (18) D(va ) = {d1 , . . . , d|C| } and relevant Judicial In-
c∈Gont
terpretations J from the Rule Graph Grul .
(ii) Community Expansion Retrieval: To cap- (ii) Item-wise Verification: The agent executes a
ture broader structural context, we employ a verification loop for each diagnostic item dk ∈
community-guided strategy. We identify the single D. It evaluates whether the raw case facts
most relevant thematic community K∗ aligned with q satisfy the specific legal condition dk , sup-
the query, and then retrieve the top-k similar cases porting the judgment with the interpretive con-
restricted within this community: text J . This produces a set of boolean veri-
 fication results Vresults = {rk }, where rk ←
K∗ = argmax sim ϕ(q), ϕ(K)
K∈Gont CheckCondition(q, dk , va , J ).
 (19) (iii) Decision and Pruning: Finally, the Audi-
Rcom (q) = Top-k sim ϕ(q), ϕ(c)
c∈K∗ tor synthesizes the verification results Vresults to
determine the overall applicability of the article.
(iii) Charge-Anchored Retrieval: Finally, we an-
If the article fails to meet the necessary criteria
chor the legal basis by retrieving cases linked to
(IsApplicable is False), the Auditor executes a
inferred charges. Here, O(q) denotes the set of pre-
pruning operation:
dicted charges and NGf ac (o) represents the neigh-
boring cases connected to charge o in the Fact Sverif ied ← Prune(Sverif ied , va ) (21)
Graph:
[ This step removes the inapplicable article node
Rchg (q) = NGf ac (o) (20)
va along with its dependent case precedents and
o∈O(q)
charge nodes, ensuring that the final subgraph
These three retrieval strategies organize the can- Sverif ied contains only logically valid and appli-
didate evidence set Scand . cable evidence.
Auditor Agent: Operating on the candidate evi- Adjudicator Agent: The Adjudicator synthesizes
dence set Scand , the Auditor validates the appli- the verified subgraph Sverif ied to render the final
cability of each retrieved article va through a rig- judgment. Specifically, it organizes the valid nodes
orous verify-and-prune mechanism. This process extracted from Sverif ied into sets of confirmed arti-
proceeds in three specific steps: cles (VAf ), case precedents (VCf ), and charge infor-
37471
Algorithm 1 Evidence-based Legal Reasoning
Require: Raw case query q; Ontology Graph Gont ; Fact Graph Gf ac ; Rule Graph Grul .
Ensure: Final Judgment J with citations.
Stage 1: Researcher Agent (Multi-Strategy Retrieval)
1: ϕ(q) ← OntologyAlign(q, Gont ) ▷ Align query to ontology features
Parallel Evidence Retrieval Strategies:
2: Scand ← Rsem ∪ Rcom ∪ Rchg ▷ Union of candidate evidence
Stage 2: Auditor Agent (Verification & Pruning)
3: Sverif ied ← Scand
4: for each article node va ∈ Scand do
5: D ← RetrieveChecklist(va , Grul ) ▷ Get checklist D(va ) = {d1 , . . . , d|C| }
6: J ← RetrieveInterpretations(va , Grul ) ▷ Get Judicial Interpretations
7: Vresults ← ∅
8: for each diagnostic item dk ∈ D do ▷ Item-wise verification loop
9: rk ← CheckCondition(q, dk , va , J ) ▷ Verify if fact q satisfies condition dk
10: Vresults ← Vresults ∪ {rk }
11: end for
12: IsApplicable ← Decide(Vresults ) ▷ Final determination for node va
13: if not IsApplicable then
14: Sverif ied ← Prune(Sverif ied , va ) ▷ Remove article and linked nodes
15: end if
16: end for
f f f f
17: Gsub ← {VA , VC , VO } ← Organize(Sverif ied ) ▷ Structure the verified subgraph
Stage 3: Adjudicator Agent (Synthesis)
f f f
18: Y ← Adjudicator(q ⊕ VA ⊕ VC ⊕ VO ) ▷ Synthesize judgment with citations
19: return Y

mation (VOf ). By integrating these evidence com- CAIL and CMDL datasets, regardless of the back-
ponents with the original query q, it generates a bone model employed. Notably, even with the
response with explicit citations. This process is lighter GPT-4o-mini on the CMDL dataset, our
formulated as: method achieves a remarkable performance gain
(e.g., significantly exceeding the strong baseline
Y = Adjudicator(q ⊕ VAf ⊕ VCf ⊕ VOf ) (22) RAPTOR in Accuracy), while maintaining its lead
with the more powerful DeepSeek-V3.1. This
The output Y ensures that every conclusion is di-
demonstrates that LegalGraphRAG’s structured
rectly traceable to specific nodes in the knowledge
reasoning capabilities effectively complement the
graph, enforcing transparency and evidence-based
generation power of various state-of-the-art LLMs,
reasoning.
enhancing their precision in complex legal applica-
C Additional Experiments tion scenarios independent of the underlying model
architecture.
C.1 Extensions to the Main Experiment (Q4) Obs.8. Exactness in Law Article Prediction. Ta-
In this section, we conduct a series of extended ble 5 illustrates the model’s capability in Law Arti-
experiments to verify the universality of our frame- cle Prediction, a task demanding precise statutory
work across different model architectures and its ro- grounding rather than generative flexibility. Legal-
bustness in specific, high-difficulty legal sub-tasks. GraphRAG achieves a superior overall accuracy
Obs.7. Universality across Advanced Backbones. of 47.9%, establishing a substantial lead over both
To verify the universality of our framework, we ex- the strongest RAG baseline, HippoRAG2 (39.8%),
tended the evaluation to advanced large language and the domain-specific state-of-the-art, ADAPT
models, specifically DeepSeek-V3.1 and GPT-4o- (41.3%). Remarkably, our 8B-parameter frame-
mini. As shown in Table 4, LegalGraphRAG work even surpasses the massive DeepSeek-V3.1
consistently outperforms all baselines across both (44.9%), highlighting that our structured, evidence-
37472
CAIL
Public Safety Economic Social Order Person Rights
Model Size All ∆
ACC F1 ACC F1 ACC F1 ACC F1
GPT-4o-mini
Naive RAG ∼8B 27.5 37.6 18.8 33.6 18.0 28.8 22.1 39.2 22.2 ↑ 18.7
G-Retriever ∼8B 17.5 24.8 20.3 32.8 20.5 31.0 24.1 31.6 21.4 ↑ 19.5
LightRAG ∼8B 25.4 37.4 21.3 36.9 21.7 38.7 23.4 42.8 23.1 ↑ 17.8
RAPTOR (Sarthi et al., 2024) ∼8B 33.1 49.0 29.3 44.2 25.9 39.9 28.3 43.1 30.5 ↑ 10.4
HippoRAG2 (Gutiérrez et al., 2025) ∼8B 33.1 50.1 28.1 46.1 21.7 43.6 37.2 54.2 31.9 ↑ 9.0
LegalGraphRAG (Ours) ∼8B 39.6 54.8 36.3 52.9 37.3 51.2 42.1 62.4 40.9 –
DeepSeek-V3.1
Naive RAG ∼200B 38.0 54.0 32.3 49.9 33.4 47.3 40.7 53.4 37.8 ↑ 12.1
G-Retriever ∼200B 36.5 54.6 35.1 49.8 36.2 47.2 39.5 48.3 37.2 ↑ 12.7
LightRAG ∼200B 36.6 48.5 26.3 50.2 33.5 46.3 39.7 53.1 45.4 ↑ 4.5
RAPTOR (Sarthi et al., 2024) ∼200B 42.2 56.3 37.8 53.0 39.2 50.7 45.5 52.3 44.4 ↑ 5.5
HippoRAG2 (Gutiérrez et al., 2025) ∼200B 41.5 49.1 33.1 46.8 34.3 47.0 38.6 46.4 41.2 ↑ 8.7
LegalGraphRAG (Ours) ∼200B 44.4 58.8 41.9 57.8 41.9 56.8 46.2 65.1 49.9 –

CMDL
Public Safety Economic Social Order Person Rights
Model Size All ∆
ACC F1 ACC F1 ACC F1 ACC F1
GPT-4o-mini
Naive RAG ∼8B 38.7 50.8 35.6 44.0 30.9 44.2 33.3 43.0 32.9 ↑ 17.1
G-Retriever ∼8B 24.6 38.4 29.7 38.8 30.4 40.0 41.1 51.6 28.3 ↑ 21.7
LightRAG ∼8B 36.2 43.1 37.5 48.9 46.9 55.1 34.6 50.8 34.2 ↑ 15.8
RAPTOR (Sarthi et al., 2024) ∼8B 42.0 49.1 48.2 58.0 57.3 60.9 46.0 59.5 48.7 ↑ 11.3
HippoRAG2 (Gutiérrez et al., 2025) ∼8B 45.0 52.5 42.9 60.8 45.3 64.6 59.4 75.0 46.0 ↑ 14.0
LegalGraphRAG (Ours) ∼8B 46.5 57.8 63.7 67.9 58.0 66.2 67.2 75.4 60.0 –
DeepSeek-V3.1
Naive RAG ∼200B 52.0 65.3 66.1 71.6 69.8 70.7 68.8 80.5 62.9 ↑ 12.8
G-Retriever ∼200B 46.2 64.9 58.6 69.6 56.1 67.4 69.7 78.9 55.2 ↑ 23.5
LightRAG ∼200B 47.7 58.7 46.5 67.6 47.6 53.2 52.3 64.2 52.4 ↑ 26.3
RAPTOR (Sarthi et al., 2024) ∼200B 53.4 64.7 68.2 73.0 66.4 75.4 60.3 71.1 56.4 ↑ 12.3
HippoRAG2 (Gutiérrez et al., 2025) ∼200B 62.1 63.5 62.4 65.3 75.9 78.6 76.9 78.2 74.0 ↑ 4.7
LegalGraphRAG (Ours) ∼200B 66.7 69.9 76.0 79.3 72.9 80.4 79.7 85.5 78.7 –

Table 4: Performance comparison on advanced Models. We compared LegalGraphRAG and other baselines
utilizing advanced LLMs as backbones.

based retrieval mechanism is more effective at pin- LegalGraphRAG’s evidence-based retrieval strat-
pointing legal provisions than simply scaling model egy effectively locates relevant sentencing guide-
parameters or employing semantic retrieval. lines and comparable precedents, thereby constrain-
ing the generation to a more precise and legally
Obs.9. Precision in Term of Penalty Predic-
grounded time range.
tion. Table 6 presents the results on the challenging
term of penalty prediction task, which requires fine-
C.2 Hyper-parameter Sensitivity (Q5)
grained quantitative reasoning rather than simple
classification. LegalGraphRAG demonstrates a sig- To evaluate system stability, we investigated the
nificant advantage in minimizing prediction error, sensitivity of the Researcher Agent to the retrieval
consistently achieving the lowest Mean Absolute parameter k, which governs the number of seman-
Error (MAE) across most subdomains compared tic concepts retrieved from the ontology graph Gont .
to other RAG-based methods. For instance, in We varied k over the set {3, 4, 5, 6}. The upper
the Public Safety category, our model achieves an bound is restricted to 6, as empirical evidence sug-
MAE of 20.9, outperforming RAPTOR (21.7) and gests that exceeding this threshold introduces exces-
HippoRAG2 (23.0). This indicates that while ex- sive context noise, which overwhelms the model’s
act term matching remains difficult for all models, effective window and degrades reasoning.
37473
CAIL
Public Safety Economic Social Order Person Rights
Model Size All ∆
ACC F1 ACC F1 ACC F1 ACC F1
Open-Source Models
Qwen-2.5-7B-Instruct 7B-Inst 28.7 53.8 24.1 48.2 26.6 53.5 36.0 56.2 30.1 ↑ 17.8
Qwen-3-8B 8B-Inst 23.9 58.5 27.6 51.2 36.6 62.2 46.2 66.7 35.9 ↑ 12.0
Internlm3-8b-instruct 8B-Inst 27.6 59.9 26.4 52.8 30.8 59.0 33.6 57.6 29.9 ↑ 18.0
Glm-4-9b-chat 9B-Inst 23.9 60.1 25.6 51.0 34.7 59.4 35.1 58.5 30.8 ↑ 17.1
Advanced Models
GPT-4o-mini (Achiam et al., 2023) ∼8B 24.7 57.2 25.7 42.4 24.6 51.8 34.7 55.0 30.9 ↑ 17.0
DeepSeek-V3.1 (Liu et al., 2024) ∼200B 42.3 63.2 37.1 61.3 44.5 68.8 51.9 67.3 44.9 ↑ 3.0
Legal Specific Methods
DISC-LawLLM-7B (Yue et al., 2024) 7B-Inst 36.5 55.9 29.9 48.8 41.5 61.2 39.3 57.8 38.2 ↑ 9.7
ADAPT (Deng et al., 2024b) 7B-Inst 40.1 50.8 31.6 42.4 39.6 54.9 41.0 50.5 41.3 ↑ 6.6
Legal ∆ (Dai et al., 2025) 7B-Inst 33.0 54.9 27.8 51.0 34.6 59.4 44.5 60.6 37.9 ↑ 10.0
RAG Based Methods
Naive RAG 8B-Inst 30.4 45.7 29.2 48.6 37.5 54.1 36.6 51.1 34.8 ↑ 13.1
RAPTOR (Sarthi et al., 2024) 8B-Inst 36.4 59.2 33.4 53.9 38.9 64.1 41.0 61.7 37.2 ↑ 10.7
HippoRAG2 (Gutiérrez et al., 2025) 8B-Inst 35.2 60.3 33.1 53.7 41.7 67.0 42.6 63.8 39.8 ↑ 8.1
LegalGraphRAG (Ours) 8B-Inst 43.0 64.9 37.8 61.0 44.6 69.4 54.5 70.6 47.9 –

Table 5: Extended experiments on Article Prediction. We evaluated the performance of our model and baselines
on the specific sub-task of law article prediction. We visualize the gains of LegalGraphRAG to the each baseline in
the ∆ columns .
CAIL
Public Safety Economic Social Order Person Rights
Model Size All ∆
ACC MAE ACC MAE ACC MAE ACC MAE
Open-Source Models
Qwen-2.5-7B-Instruct 7B-Inst 13.0 23.7 8.1 33.3 6.0 30.3 9.7 29.4 29.5 ↑ 8.4
Qwen-3-8B 8B-Inst 8.3 32.6 11.3 31.6 17.9 29.7 11.4 26.4 27.5 ↑ 7.4
Internlm3-8b-instruct 8B-Inst 7.1 35.2 5.9 37.2 6.0 32.5 7.2 37.3 33.7 ↑ 13.6
Glm-4-9b-chat 9B-Inst 3.6 36.0 3.2 32.9 1.5 35.1 6.8 38.6 33.1 ↑ 13.0
Advanced Models
GPT-4o-mini (Achiam et al., 2023) ∼8B 6.9 38.2 7.3 34.2 8.3 31.7 8.0 34.6 33.6 ↑ 13.5
DeepSeek-V3.1 (Liu et al., 2024) ∼200B 7.1 31.2 8.1 31.5 10.4 25.9 8.4 29.1 29.1 ↑ 8.6
Legal Specific Methods
DISC-LawLLM-7B (Yue et al., 2024) 7B-Inst 15.5 24.5 5.0 33.2 3.0 37.9 8.4 34.5 31.6 ↑ 11.5
ADAPT (Deng et al., 2024b) 7B-Inst 8.3 21.8 10.9 21.9 3.0 24.3 9.7 22.3 20.4 ↑ 0.3
Legal ∆ (Dai et al., 2025) 7B-Inst 11.9 25.0 9.0 29.0 9.0 27.5 8.9 27.1 26.3 ↑ 6.2
RAG Based Methods
Naive RAG 8B-Inst 10.5 29.8 11.4 30.8 18.1 21.7 12.6 24.3 26.5 ↑ 4.4
RAPTOR (Sarthi et al., 2024) 8B-Inst 12.0 21.7 11.7 28.3 15.8 26.1 16.2 21.8 24.3 ↑ 4.2
HippoRAG2 (Gutiérrez et al., 2025) 8B-Inst 13.1 23.0 12.7 25.8 17.9 23.8 13.5 23.4 23.8 ↑ 3.7
LegalGraphRAG (Ours) 8B-Inst 14.0 20.9 13.7 22.1 19.4 23.6 17.1 22.7 20.1 –

Table 6: Extended experiments on Term of Penalty Prediction. We assessed the accuracy and error rates of
imprisonment term predictions compared to baselines. We visualize the gains of LegalGraphRAG to the each
baseline in the ∆ columns .

37474
Token Consumption (< 106 ) Avg Cost
Method Indexing Time (s)
Prompt Completion Time Token
RAPTOR 13696.90 5.64 0.72 5.86 3589
HippoRAG2 4581.60 10.58 2.79 11.2 5199
LegalGraphRAG (Ours) 3687.49 3.97 0.78 46.1 10664

Table 7: Comparison of Computational Efficiency: Offline Indexing vs. Online Inference. We report the total time
and token usage for graph construction (Indexing) and the average cost per query (Online).

offline efficiency with the lowest indexing time


(3687.49s) and token consumption. However, dur-
ing the online phase, it incurs higher latency (46.1s)
and token usage. This overhead is a necessary
trade-off for evidence-based reasoning. Unlike
baseline GraphRAG approaches that often oper-
ate as opaque "black boxes", our method explicitly
constructs credible reasoning chains to support its
judgments. While generating such transparent evi-
dence consumes more resources, it is indispensable
for ensuring the trustworthiness and interpretability
Figure 9: Impact of the retrieval parameter k on charge required in legal domains.
prediction performance (CAIL dataset). The backbone
model is Qwen3-8B. D Implementation Details

Obs.10. Robustness to Retrieval Hyperparam- D.1 Benchmark Dataset


eter Variations. As illustrated in Figure 9, Legal- We evaluate LegalGraphRAG on two benchmark
GraphRAG exhibits strong robustness to variations datasets, including CAIL2018(Xiao et al., 2018)
in k. Although performance peaks at k = 5 (achiev- and CMDL(Huang et al., 2024).
ing 40.9 Accuracy and 52.8 F1), the variance across CAIL2018(Xiao et al., 2018): A large-scale Chi-
the tested range is marginal. This indicates that the nese legal dataset designed for the task of Legal
Researcher Agent reliably captures essential se- Judgment Prediction (LJP), where models predict
mantic information without being hypersensitive court outcomes based on factual case descriptions.
to exact thresholding, provided k remains within a It comprises over 2.6 million criminal cases pub-
reasonable bound. lished by the Supreme People’s Court of China,
making it the largest publicly available dataset of
C.3 Latency and Token Cost (Q6) its kind. Each case includes a detailed fact descrip-
In this section, we shift our focus to the practical tion along with structured judgment annotations,
efficiency and computational overhead of the com- namely applicable law articles (183 categories),
pared frameworks. Beyond accuracy, the latency charges (202 categories), and prison terms. The
and token consumption are crucial factors for real- dataset was created to address the lack of high-
world deployment. We provide a detailed break- quality, large-scale resources in legal AI and to
down and comparison of the computational costs provide a realistic benchmark that reflects the com-
for RAPTOR, HippoRAG2, and LegalGraphRAG plexity and imbalance inherent in real judicial data,
across two primary phases: (i) the offline graph where frequent charges dominate the case distribu-
construction stage, reporting both the time and to- tion. CAIL2018 has since become a foundational
ken cost required to build the knowledge base; and resource for evaluating and advancing automated
(ii) the online query-answering stage, reporting the legal judgment prediction systems.
time and token cost incurred during the retrieval CMDL(Huang et al., 2024): A large-scale, real-
and reasoning process for a given query. world Chinese Multi-Defendant Legal Judgment
Obs.11. Trade-off between Efficiency and In- Prediction dataset designed to address the under-
terpretability. Table 7 presents the computational explored challenge of predicting judicial outcomes
efficiency. LegalGraphRAG demonstrates superior in cases involving multiple defendants. It com-
37475
Dataset CAIL CMDL Dataset Case
# Case Num 568 572 # Num 14049
# Charges 168 239 # Charges* 818
# Average criminal per case 1.25 2.40 # Average defendant per case 3.39
# Average defendant per case 1.72 1.14 # Average length per case 399.73
# Average length per case 654.79 517.13
Table 9: Basic statistics of the cases in corpus. *The
Table 8: Basic statistics of the test datasets. large number of “Charges” is due to inconsistencies in
the descriptions of crimes across different datasets.

prises 393,945 criminal cases with approximately Dataset Article Judicial interpretations
1.2 million defendants, covering 321 distinct
# Num 452 656
charges and 275 legal articles. Notably, CMDL in- # Average length 128.54 243.95
troduces case-level evaluation metrics that account
for case complexity and varying numbers of defen- Table 10: Basic statistics of the legal knowledges in
dants, offering a more holistic assessment of model corpus.
performance in multi-defendant scenarios. For ex-
perimental feasibility, the subset CMDL-small is
a subset of cases from multiple authoritative legal
often utilized, as it preserves the data distribution
datasets: JuDGE(Su et al., 2025), CAIL2018(Xiao
while significantly reducing computational costs,
et al., 2018), CMDL(Huang et al., 2024), and
making it suitable for preliminary benchmarking
LeCaRDv2(Li et al., 2024). The construction fol-
and model validation in resource-constrained re-
lows a procedure similar to that used for the ex-
search settings.
perimental subsets: we first filter cases by fact de-
Dataset Construction Due to considerations re- scription length (under 1,024 characters) and apply
garding the generation speed and token cost of the sampling designed to balance the representation of
GraphRAG method, we constructed focused sub- different charges, thereby creating a broad and di-
sets from both the CAIL2018 and CMDL datasets verse collection of historical precedents and factual
using a uniform procedure aimed at controlling patterns. Furthermore, to ground the system in au-
input length, balancing charge distribution, and thoritative legal provisions, we incorporate the full
elevating task complexity. The construction first fil- text of the “Criminal Law of the People’s Republic
tered cases to retain only those with factual descrip- of China” along with its relevant judicial interpre-
tions under 1,024 characters. To ensure broader tations as a core statutory knowledge library. The
coverage of under-represented charges and increase combination of this curated historical case library
predictive difficulty, the sampling prioritized de- and the official legal provisions library forms the
fendants whose charges included low-frequency of- complete corpus, enabling models to retrieve both
fenses and deliberately retained a higher proportion experiential precedents and statutory knowledge
of multi-charge cases. Consequently, the resulting during [Link], we have carefully ver-
subsets feature a more balanced charge distribution ified that all cases in the corpus are distinct from
with elevated presence of rare charges and greater those in the test subsets, ensuring no data leakage
average case complexity compared to the original between the knowledge base and the evaluation
datasets. While this design provides a more chal- benchmarks. The size and composition statistics of
lenging testbed for evaluating model performance the final corpus are detailed in Table 9 & 10.
on complex and low-frequency legal scenarios, it
may also lead to lower reported performance for D.3 Evaluation Metrics
some methods relative to their results on the orig- To comprehensively evaluate the Legal Judgment
inal, more naturalistic data distribution. Table 8 Prediction (LJP) tasks, we employ specific metrics
presents detailed statistics of the subsets. for different sub-tasks: Charge and Article Pre-
diction are evaluated using Accuracy (ACC) and
D.2 Corpus
Micro-F1, Term of Penalty Prediction is assessed
In this section, we provide detailed descriptions of using ACC and Mean Absolute Error (MAE), and
corpus used in our experiments. To ensure compre- the Retrieval Quality of our RAG system is mea-
hensive coverage of criminal statutes and charges, sured by Retrieval Effectiveness and Error Rate.
we construct this knowledge base by aggregating Accuracy (ACC) measures the exact match ratio.
37476
For classification tasks, it requires the predicted D.4 Baseline Details
label set Oi to be identical to the ground truth Oi′ .
In this section, we provide detailed descriptions of
For Term of Penalty, it measures the exact match
each baseline used in our comparison, as detailed
of the predicted term. It is calculated as:
in figure 10.
N Naive RAG uses the standard RAG paradigm:
1 X a retriever model first retrieves relevant context
ACC = I(Oi = Oi′ )
N from the corpus based on the given question, and
i=1
then the question is concatenated with the retrieved
Where I(·) is the indicator function. context to form a query for the generation model
Micro-F1 is used for multi-label classification to to produce the final answer.
account for class imbalance. It is the harmonic G-retriever(He et al., 2024a) introduces a re-
mean of micro-averaged precision (Pmicro ) and re- trieval augmented generation framework for tex-
call (Rmicro ): tual graphs by formulating subgraph retrieval as
a Prize-Collecting Steiner Tree optimization prob-
2 · Pmicro · Rmicro lem, enabling conversational question answering
Micro-F1 =
Pmicro + Rmicro across diverse domains like scene understanding
and knowledge graphs while mitigating LLM hal-
Mean Absolute Error (MAE) reflects the devia- lucinations and scaling to large graph sizes.
tion in the predicted term of penalty. Let Ti and Ti′ LightRAG(Guo et al., 2024) introduces a graph-
denote the predicted and ground-truth prison terms enhanced retrieval-augmented generation frame-
(in months), respectively. MAE is defined as: work that integrates entity-relationship graphs into
text indexing, combining low-level precise entity
N
1 X retrieval with high-level thematic discovery for ef-
MAE = |Ti − Ti′ | ficient and adaptive knowledge integration.
N
i=1
HippoRAG2(Gutiérrez et al., 2025) builds on
HippoRAG’s Personalized PageRank framework
Retrieval Effectiveness measures how well the
by integrating dense-sparse coding for passages
retrieved content aligns with the question’s intent.
and phrases in the knowledge graph, enabling
Higher values indicate more focused and pertinent
deeper contextualization and recognition memory
information. It is defined as:
for triple filtering. Enhances online retrieval with
1 X query-to-triple matching and optimized seed node
Retrieval Effectiveness = R(c, Q, E) weighting, outperforming standard RAG across fac-
|C|
c∈C
tual, sense-making, and associative memory tasks.
where C denotes the set of retrieved contexts, Q rep- RAPTOR(Sarthi et al., 2024) constructs a hier-
resents the question, E denotes the set of evidence, archical tree by recursively clustering and summa-
and the operator R(·) determines the relevance of rizing embedded text chunks, enabling retrieval
a context c. of information at multiple levels of abstraction to
improve performance on long-document question-
Error Rate quantifies the incompleteness of the
answering tasks.
retrieval process. Instead of measuring recall di-
Disc-LawLLM(Yue et al., 2024) is a retrieval-
rectly, we assess the proportion of reference claims
augmented large language model fine-tuned on Chi-
not supported by the retrieved context:
nese judicial datasets using legal syllogism prompt-
! ing to provide reasoning-capable legal services,
1 X
Error Rate = 1 − I(S(c, C)) (23) including consultation, judgment prediction, and
|R| examination assistance. In this work, we use the
c∈R
officially open-sourced LawLLM-7B, which is fine-
where R is the set of reference claims, S(·) de- tuned from Qwen2.5-Instruct-7B.
termines whether a claim c is supported by the re- Legal ∆(Dai et al., 2025) employs a reinforce-
trieved context C, and I(·) is the indicator function. ment learning framework that enhances legal rea-
A lower Error Rate indicates a more comprehensive soning in LLMs by maximizing chain-of-thought
evidence collection. guided information gain through dual-mode inputs
37477
RAG Configuration RAPTOR Configuration

{ {
embedding_model: bge-m3, embedding_model: bge-m3,
retrieval_topk: 5, chunk_token_size: 1200,
chunk_token_size: 1000, chunk_overlap_token_size: 100,
chunk_overlap_token_size: 200 num_layers: 5,
} max_length_in_cluster: 3500,
threshold: 0.1,
RAG Configuration cluster_metric: cosine,
threshold_cluster_num: 5000
{ }
embedding_model: bge-m3,
retrieval_topk: 5,
chunk_token_size: 1000, Figure 10: Hyperparameter configurations for the base-
chunk_overlap_token_size: 200 line RAG models.
}
and differential Q-value analysis. In this work, we
G-retriever Configuration use the officially open-sourced model, which is
fine-tuned from Qwen2.5-Instruct-7B.
{
ADAPT(Deng et al., 2024b) is a discriminative
embedding_model: bge-m3,
reasoning framework for LLMs in legal judgment
retrieval_topk: 3,
prediction that emulates human judicial processes
chunk_token_size: 1200,
by asking to decompose case facts into key ele-
chunk_overlap_token_size: 100,
ments, discriminating among candidate charges for
entities_max_tokens: 2000,
alignment, and predicting final judgments, further
relationships_max_tokens: 2000
improved via multi-task fine-tuning with synthetic
}
trajectories. In this work, we use the officially open-
sourced model, which is fine-tuned from Qwen2-
LightRAG Configuration 7B.
{ D.5 LegalGraphRAG Setup
embedding_model: bge-m3,
In the experimental setup of LegalGraphRAG, hy-
query_type: hybrid,
perparameters are configured to optimize retrieval
chunk_token_size: 1200,
precision. During graph construction, the Ontol-
retrieval_topk: 20,
ogy Graph utilizes k-nearest neighbors (kNN) to
chunk_overlap_token_size: 100,
select the top-3 case feature nodes based on cosine
max_token_text_unit: 2000,
similarity for direct semantic matching.
max_token_global_context: 2000,
For the Researcher agent, the evidence-based re-
max_token_local_context: 2000
trieval strategy operates with specific thresholds:
}
the retrieval parameter is set to k = 5. This parame-
ter governs both the number of top-ranked semantic
HippoRAG2 Configuration concepts retrieved from the ontology graph and the
{ scope of community expansion, a value selected
embedding_model: bge-m3, based on the sensitivity analysis to balance context
retrieval_top_k: 5, coverage and noise control.
linking_top_k: 5,
E Extended Case Examples
max_qa_steps: 3,
qa_top_k: 5, In this section, we will walk through several cases
graph_type: facts_and_sim_passage to detail the retrieval and reasoning pipeline of
_node_unidirectional LegalGraphRAG. Figure 11 illustrates evidence re-
} trieval for a Dangerous Driving case, while Figure
37478
Query The procuratorial organ alleges that at approximately 18:45 on November 10, 2020, the defendant Wang Moumou, while driving under the
influence of alcohol and without a valid driver's license, was operating a passenger car (License Plate: Liao D×××××; the vehicle's owner was the defendant Liu
Mou) southbound on Changling Street in a certain town of a certain county, Qingyuan. When reaching a traffic signal intersection, he collided with a compact
sport utility vehicle (License Plate: Liao D×××××) driven by Zhao Moumou, who was traveling ahead in the same direction. This resulted in a road traffic
accident causing damage to both [Link] Traffic Police Brigade of a certain county in Qingyuan determined that the defendant Wang Moumou bore full
responsibility for this accident, while Zhao Moumou bore no responsibility. An appraisal conducted by the Fushun Gongzheng Judicial Appraisal Institute detected
the presence of ethanol in the defendant Wang Moumou's venous blood, with a concentration of 242.5 mg/100 ml, constituting drunk driving.

Case Feature D: "Driving under the influence"; O: "Traffic collision"; V: "Personal injury", "Passenger car driver"; M: "Negligence"

Community Expansion Retrieval Community Summary: Cases often involve offenses such as drunk driving and traffic accidents.
Researcher
Fact: At approximately 14:10 on October 18, 2016, the defendant Chen Wenli, while reversing a gray Wuling-brand light...
Case Feature: D: "Adult"; O: "Driving a vehicle", "Failed to observe surroundings"; V: "Death after failed rescue efforts"; M: "Negligence"
Name: Chen Wenli Charges: Crime of Causing Death by Negligence Law Articles: 233...

Semantic Match Retrieval


Fact: The procuratorial organ alleges that at approximately 18:45 on November 10, 2020, the defendant Wang Moumou, while...
Case Feature: D: "Adult"; O: "Causing a traffic accident"; V: "Vehicle damage", "Public road"; M: "Negligence"
Name: Wang Charges: Crime of Dangerous Driving Law Articles: 133

Charge-Anchored Retrieval "Crime of Dangerous Driving", "Crime of Causing a Traffic Accident"... Laws: 133

Article 133 Paragraph 1


Content: Whoever drives a motor vehicle on a road under any of the following circumstances shall be sentenced to criminal detention ...
Related: "Interpretation on ... of Criminal Cases of Dangerous Driving"
Evidence-based Reasoning:
Auditor
Is the act committed while driving a motor vehicle on a road?Yes. The defendant, Li Moujia, was operating an...
Is it driving a motor vehicle while intoxicated?Yes. The forensic appraisal determined that the ethanol content in ...
...
Final: Yes

Article 233 Final: No

Final Context
Related Laws: [Article 133 Paragraph 1] [Crime of Dangerous Driving] Whoever drives a motor vehicle on a road under any of the ...
(Other related laws are omitted below)
Similar Cases: [Charges: Crime of Dangerous Driving] [Name: Wang] The procuratorial organ alleges that at approximately 18:45 on...
Adjudicator (Other similar cases are omitted below)

Final Results
Charge: Crime of Dangerous Driving Articles: 133 Imprisonment: 18 months
Reasoning: Li Moujia drove a motor vehicle on a road while intoxicated (with a blood ethanol content of 155.7 mg/100 ml). His conduct
satisfies the circumstance specified in Article 133-1(1)(2) of the Criminal Law, "driving a motor vehicle while intoxicated," and thus
constitutes the crime of dangerous driving [1]...
[1] corresponds to Article 133
[2] corresponds to Cases1([Charges: Crime of Dangerous Driving] [Name: Wang])
...

Figure 11: Qualitative analysis of a representative case regarding the crime of Dangerous Driving. The visualization
highlights the retrieval of evidence related to specific statutory conditions.

12 demonstrates the legal reasoning process for an G Related Works


Occupational Embezzlement case.
G.1 Legal Large Language Models

The rapid evolution of LLMs has catalyzed the de-


F Prompt Set velopment of domain-specific models tailored for
the legal sphere. For Chinese law, ChatLaw (Cui
et al., 2023), DISC-LawLLM (Yue et al., 2023),
To facilitate reproducibility and provide trans- and InternLM2Law (Fei et al., 2025) leverage
parency into the agent behaviors, we present extensive legal corpora, including judicial interpre-
the specific instruction sets designed for the Re- tations and statutes, to handle diverse legal tasks.
searcher (Figures 13 and 14), Auditor (Figures 15 Other notable models like LawGPT (Zhou et al.,
and 16), and Adjudicator (Figures 17 and 18) 2024) and Fuzi-Mingcha (Wu et al., 2023a) in-
agents. These prompts orchestrate the multi-stage tegrate unsupervised legal texts with supervised
reasoning process described in the main text. fine-tuning to enhance domain understanding. Be-
37479
Query
The defendant Zhang Mou entered into a labor contract with the Yaojiaba Tunnel Project Department of the Xiangjiaba Hydropower Station Project,
originally under Sichuan Road and Bridge Construction Group Co., Ltd., serving as the head chef of the project department. His primary responsibility was the
financial reimbursement for all monthly canteen expenses of the project department. On December 19, 2016, the Yaojiaba Tunnel Project Department transferred
98,160 yuan, which was the reimbursed canteen expenses, into Zhang Mou's Agricultural Bank of China card. Zhang Mou withdrew 95,000 yuan of that amount
on the same day and took out the remaining 3,000 yuan the following day.

Case Feature D: "Employee", "Staff"; O: "Misappropriation of funds"; V: "Company funds", "Significant amount"; M: "Subjective intent"

Community Expansion Retrieval Community Summary: Cases often involve duty-related crimes, bribery crimes, and property...
Researcher
Fact: From 2009 to 2011, while serving as a credit officer at the Qiuji Credit Cooperative in Suining County, the defendant Zhang ...
Case Feature: D: "State Functionary"; O: "Embezzlement of funds", "Taking advantage of one's position"; V: "Unit funds"; M: "Direct intent"
Name: Zhang Zhi'an Charges: Misappropriation of Funds Law Articles: 185, 272

Semantic Match Retrieval


Fact: Since November 2009, defendant Bai has served as a client manager at the Zhushan County Branch of Bank of China Limited as ...
Case Feature: D: "State Functionary"; O: "Embezzlement of public funds"; V: "Public property", "Significant amount"; M: "Direct intent"
Name: Bai Charges: Misappropriation of Public Funds Law Articles: 185, 384

Charge-Anchored Retrieval "Embezzlement", "Misappropriation of funds", "Corruption" Laws: 271, 272

Article 271 Paragraph 1


Content: Employees of companies, enterprises, or other entities who, taking advantage of their positions, illegally take possession of ...
Related: "Interpretation on ... Criminal Cases of Embezzlement and Bribery"
Evidence-based Reasoning:
Auditor
Is the defendant an employee of a company, enterprise, or other entity?Yes. The defendant Zhang signed a labor contract...
Did the defendant abuse his position?Yes. Zhang's main responsibility was handling financial reimbursement for the ...
...
Final: Yes

Article 271 Paragraph 2 Final: No

Article 272 Final: No

Final Context
Related Laws: [Article 271 Paragraph 1] [job-related embezzlement] Employees of companies, enterprises, or other entities who...
(Other related laws are omitted below)
Similar Cases: [Charges: job-related embezzlement] [Name: Xu 3Mou] Between December 2007 and December 2009, the defendant Xu...
Adjudicator (Other similar cases are omitted below)

Final Results
Charge: job-related embezzlement Articles: 271 Imprisonment: 18 months.
Resoning: Zhang served as the head chef of the project department and was responsible for financial reimbursement, which qualifies him
as a "personnel of a company, enterprise, or other unit" [1]. By taking advantage of his duty in handling canteen expenses...
[1] corresponds to Article 271 Paragraph 1
[2] corresponds to Cases1([Charges: job-related embezzlement] [Name: Xu 3Mou])
...

Figure 12: Qualitative analysis of a representative case regarding the crime of Occupational Embezzlement. The
example demonstrates the model’s reasoning in identifying the abuse of professional position.

yond Chinese, SaulLM (Colombo et al., 2024) fo- gal, 1984) relied on artificially designed features,
cuses on English legal texts based on the Mixtral and traditional machine learning methods (Sulea
architecture, while LawLLM (Shu et al., 2024) et al., 2017) were applied to predict legal judg-
addresses US legal tasks such as similar case re- ments. Recent advances in deep learning (Xu
trieval. Additionally, specialized models like In- et al., 2020; Han and Zhicheng, 2023) have mo-
LegalLLaMA (Ghosh et al., 2024) target Indian tivated researchers to leverage neural networks
and French legal domains respectively. These mod- for automated text representation learning. Re-
els provide crucial baselines for downstream tasks cently, LLMs have further promoted the progress of
but often lack the specific reasoning architecture LJP (Deng et al., 2024a). Several studies (Wu et al.,
required for complex judgment prediction. 2023b; Peng and Chen, 2024) employ Retrieval-
Augmented Generation (RAG) to enhance LLMs
G.2 Legal judgment prediction by incorporating external legal knowledge. To re-
fine decision-making, recent works have introduced
Legal judgment prediction (LJP) has experienced
structured reasoning frameworks that systemati-
significant development and become an increas-
cally decompose case facts to distinguish confusing
ingly crucial NLP task. Earlier research (Se-
37480
Criminal Case Keyword Extraction
Task Definitions: Extract legal keywords from a criminal case description and
classify them into four categories: Defendant Attributes, Criminal Behaviors,
Victim Characteristics, and Subjective Mental States. The output must be a
strictly valid JSON object without additional text.
Keyword Definitions:
• Defendant Attributes: Legal traits (e.g., age group, criminal history,
occupation). Avoid specific names or numbers.
• Criminal Behaviors: Legal types of acts and significant methods. Exclude
specific time/location details.
• Victim Characteristics: Nature of the property or location. Generalize
specific amounts (e.g., “large amount”).
• Subjective Mental States: Legal descriptions of intent and remorse.
Output Example: {
“Defendant_Attribute”: [“Adult”, “Prior Criminal Record”],
“Criminal_Behaviors”: [“Theft”, “Burglary”],
“Victim_Characteristics”: [“Private Residence”, “Large Amount”],
“Subjective_Mental_States”: [“Direct Intent”, “Voluntary Surrender”]
}

Figure 13: Prompt for the Researcher agent to extract and classify legal keywords from case descriptions.

Charge Pre-judge(for Charge-Anchored Retrieval)


Task Definitions: Act as a criminal law expert to analyze the provided case
({case_text}).
• Output reasonably possible charges (confidence > 30%) sorted by probability
(descending).
• If a dominant charge exists (confidence > 70%), prioritize it; if it is the
only certain charge, output it exclusively.
• Exclude charges with probability < 10%.
Format Definitions:
• Output strictly as a Python list: [’Charge 1’, ’Charge 2’, ...].
• The output must start with [.
• Return an empty list [] if no charge matches.
• No additional explanations or text allowed.

Figure 14: Prompt for the Researcher agent to pre-judge potential charges for charge-anchored retrieval.

charges (Jiang and Yang, 2023; Deng et al., 2024b; we make full use of external knowledge and prece-
Wang et al., 2024). Furthermore, multi-agent simu- dents within a unified framework.
lation frameworks have been explored to improve
performance by simulating court debates and ana- G.3 Reasoning skills in legal domain
lyzing cases from diverse perspectives (He et al., Recent work has improved LLMs’ reasoning
2024b). However, existing LLM-based methods through better prompting techniques (Sahoo et al.,
still struggle to utilize comprehensive legal knowl- 2024). Chain-of-thought (CoT) (Wei et al., 2022)
edge (Fei et al., 2024) effectively. In this context, prompting can explicitly guide LLMs to reason
37481
Auditor Checklist(item)
Task Definitions: Act as a legal AI assistant to assess if the case facts
strictly satisfy a specific constituent element of the law.
• Analyze the law_item and case facts.
• Focus exclusively on the target element (e.g., “intent”), using related
materials (if provided) for interpretation.
• determine applicability based on facts and logic.
I/O Specifications:
• Input: law_item, related (supplementary materials), element, case.
• Output: Provide reasoning first, then enclose the final result strictly within
tags: <answer>true</answer> or <answer>false</answer>.
Template:
law: {law_item}, related: {related}
element: {element}, case: {case}

Figure 15: Prompt for the Auditor agent to verify if case facts satisfy specific constituent elements of the law.

Auditor Checklist(final)
Task Definitions: Act as a legal analysis assistant to determine if the provided
law article applies to the specific case (i.e., verify violation or crime).
• Identify all relevant constituent elements from the law text.
• Verify critical elements independently; note that the provided true_list and
false_list may be incomplete.
I/O Specifications:
• Input Variables: case, law, true_list (proven elements), false_list (disproven
elements).
• Output: Provide reasoning first, then enclose the final result strictly within
tags: <answer>true</answer> or <answer>false</answer>.
Template:
case: {case}, law: {law}
true_list: {true_list}, false_list: {false_list}

Figure 16: Prompt for the Auditor agent to assess the overall applicability of a law based on verified elements.

step by step. In the legal domain, researchers have CaseGPT (Yang, 2024) combines LLMs with RAG
adapted CoT to legal-specific frameworks. For in- to support semi-structured reasoning and legal ar-
stance, Yu et al. (Yu et al., 2022) demonstrated that gumentation (Westermann, 2024). Additionally,
incorporating the IRAC (Issue, Rule, Application, GLARE (Yang et al., 2025b) leverages an agen-
Conclusion) framework significantly enhances rea- tic framework and web data for legal reasoning.
soning capabilities. LoT (Jiang and Yang, 2023) However, these approaches primarily rely on in-
proposed legal syllogism reasoning to improve trinsic capabilities or noisy external data, which
performance on LJP tasks, and ADAPT (Deng constraints reasoning depth (Zhang, 2024; Ke et al.,
et al., 2024b) established a workflow enabling dis- 2025). Therefore, we propose an agentic frame-
criminative reasoning. Moreover, approaches like work to dynamically acquire key legal knowledge,
MALR (Yuan et al., 2024) utilize parameter-free enhancing both breadth and depth.
learning to decompose complex legal tasks, while
37482
Charge & Sentencing (JSON)
Task Definitions: Act as a legal expert to adjudge the defendant based on
candidate charges.
• 1. Final Charge Application: For concurrence, apply the “heavier penalty”
rule; for multiple acts, apply combined punishment.
• 2. Sentencing: Predict the specific law article and a reasonable sentencing
range based on facts and judicial practice.
Format Definitions:
• {
charge_name: [Charge A, ...],
law_article: [Art. X, ...],
term_of_imprisonment: {
death_penalty: boolean,
imprisonment: integer (months),
life_imprisonment: boolean
}
}

Figure 17: Prompt for the Adjudicator agent to generate structured sentencing predictions and apply legal rules.

Legal Reasoning & Verdict


Task Definitions: Act as a legal consultant to analyze the case using the
provided Context documents.
• Step 1: Fact & Act Analysis: Analyze how many independent criminal acts exist.
Explicitly cite the supporting evidence from the context using [1][2]....
• Step 2: Law Application: Resolve any legal concurrence (e.g., Imaginative
Concurrence vs. Combined Punishment). Explain why specific articles apply
over others.
• Step 3: Sentencing Prediction: comprehensive assessment of sentencing based
on statutory rules.
Output Format:
• Structure: Output in two clear sections: Legal Analysis (reasoning with
citations) and Final Verdict (conclusion).
• Requirement: You must mark the source of your facts or laws using brackets
like [1].
Template:
Context: {context_list}, Case: {case_description}

Figure 18: Prompt for the Adjudicator agent to synthesize legal reasoning and output the final verdict.

37483
G.4 RAG in legal domain
In the legal domain, recent studies have adapted
RAG and graph-based solutions for specific tasks,
particularly Legal Question Answering (LQA) and
retrieval pipelines. For instance, recent works
optimize LLM outputs by incorporating external
case-based information (Wiratunga et al., 2024;
Louis et al., 2024) or utilizing adapt-retrieve-revise
pipelines (Wan et al., 2024) to combine contin-
ual training with evidence revision. Dedicated
benchmarks such as LegalBench-RAG show that
the retrieval stage remains a bottleneck (Pipitone
and Alami, 2024). Works enriching retrieval with
structural information (e.g., graphs of articles)
demonstrate gains in tasks like statutory article re-
trieval (Louis et al., 2023; Hei et al., 2024; Ho
et al., 2025). Recently, the SAT-Graph RAG frame-
work (de Martim, 2025) was proposed to model
the hierarchical structure of legal norms. However,
its sophisticated ontology-driven approach requires
heavily structured input data, limiting its applica-
bility to less curated corpora.

The Usage of LLMs


In this paper, LLMs were used only to polish the
writing and correct grammatical errors for clarity.
In the preparation of this manuscript, we utilized
Large Language Models (LLMs) to assist with the
writing process. Specifically, the model was used to
refine the English text, including correcting gram-
matical errors and improving sentence clarity. Ad-
ditionally, LLMs assisted in the initial formatting
of several tables. The authors reviewed all model
suggestions and retain full responsibility for the sci-
entific accuracy and integrity of the final content.

37484

You might also like