Legal Graph Rag
Legal Graph Rag
Figure 3: The architecture of LegalGraphRAG. The framework consists of two main phases: (1) Hierarchical
Knowledge Construction, which builds a Hierarchical Legal Graph (HierarGraph) comprising an Fact Graph,
Ontology Graph and Rule Graph to organize heterogeneous legal knowledge; and (2) Evidence-based Legal
Reasoning, where a multi-agent system (Researcher, Auditor, and Adjudicator) performs structured retrieval,
validation, and synthesis over the HierarGraph to generate interpretable legal decisions.
based legal reasoning. The detailed construction The specific retrieval algorithms and parameter
procedures are provided in Appendix B.1. settings are detailed in Appendix B.2.
4.3 Evidence-based Legal Reasoning 4.3.2 Evidence Validation
To leverage the multi-granular knowledge encoded Given the candidate evidence retrieved in the Ev-
in our HierarGraph, we propose a multi-agent sys- idence Retrieval, this stage focuses on validating
tem for evidence-based reasoning, in which spe- whether the case facts genuinely satisfy the con-
cialized agents sequentially traverse the graph to ditions required by the law, rather than relying on
perform evidence retrieval, validation, and synthe- surface-level semantic relevance.
sis. Specifically, the workflow consists of three Specifically, for each candidate article, we verify
agents:1) Researcher, 2) Auditor, and 3) Adjudica- its applicability by evaluating the case facts using
tor. Through structured graph traversal and logical the associated Diagnostic Checklist and Judicial
37459
CAIL CMDL Average
Public Safety Economic Social Order Person Rights Public Safety Economic Social Order Person Rights
Model Size ACC / F1 ACC / F1 ACC / F1 ACC / F1 ACC / F1 ACC / F1 ACC / F1 ACC / F1 All ∆
Open-Source Models
Qwen-2.5-7B-Instruct 7B-Inst 24.0 45.8 23.1 42.5 22.9 36.7 27.4 46.0 25.8 32.4 28.7 35.8 27.2 42.1 32.8 49.6 26.7 ↑ 22.8
Qwen-3-8B 8B-Inst 31.7 49.2 25.8 42.7 26.3 39.8 27.6 47.8 44.0 52.3 44.7 53.1 42.7 51.9 53.0 57.7 35.2 ↑ 19.9
Internlm3-8b-instruct 8B-Inst 29.8 49.1 26.7 42.0 25.2 34.3 28.1 47.3 25.4 32.1 35.7 37.0 27.5 36.2 34.1 53.6 26.6 ↑ 22.9
Glm-4-9b-chat 9B-Inst 18.4 33.7 19.7 36.1 15.8 32.1 26.0 44.5 23.5 34.2 23.6 40.8 19.1 37.0 41.5 47.0 21.2 ↑ 28.2
Advanced Models
GPT-4o-mini ∼8B 19.7 35.5 19.6 33.3 15.5 35.2 29.0 46.3 18.0 28.0 22.7 31.9 21.9 32.0 35.9 50.3 28.4 ↑ 21.1
DeepSeek-V3.1 ∼200B 31.0 51.3 29.0 48.4 29.8 50.2 35.2 54.8 35.0 64.0 54.7 62.7 58.2 61.9 62.5 71.6 42.8 ↑ 6.7
Legal Specific Methods
DISC-LawLLM-7B 7B-Inst 40.1 50.9 31.0 51.5 34.8 47.7 34.5 56.0 49.7 53.6 39.6 52.1 30.3 49.5 48.4 63.3 30.3 ↑ 19.1
ADAPT 7B-Inst 38.7 43.7 32.7 43.4 27.6 41.7 35.2 50.7 54.5 58.8 57.1 59.4 40.9 43.4 61.5 62.1 42.8 ↑ 6.7
Legal ∆ 7B-Inst 40.8 50.6 25.1 37.4 32.1 43.7 34.1 53.6 58.3 61.5 51.8 55.8 50.2 54.8 65.8 64.4 42.4 ↑ 7.1
RAG Based Methods
Naive RAG 8B-Inst 31.0 45.7 24.4 38.7 28.1 38.4 34.5 46.8 45.8 57.3 44.8 55.2 46.8 58.5 49.6 57.8 33.3 ↑ 16.1
G-retriever 8B-Inst 33.8 48.0 26.0 39.8 23.8 39.3 32.6 50.1 36.8 40.0 42.5 48.8 45.3 50.7 46.2 52.4 34.4 ↑ 13.2
LightRAG 8B-Inst 20.4 43.6 21.7 42.5 19.0 42.5 26.9 50.6 37.9 50.1 43.2 45.1 44.2 51.3 43.7 46.9 30.5 ↑ 19.0
RAPTOR 8B-Inst 34.6 50.4 31.6 43.9 32.1 45.6 32.4 45.7 53.8 62.6 53.6 60.1 52.5 62.8 52.1 66.9 43.1 ↑ 6.3
HippoRAG2 8B-Inst 34.5 38.2 24.0 33.5 28.8 35.0 31.0 36.3 53.5 56.5 50.6 52.7 53.5 55.0 62.4 62.8 43.1 ↑ 6.3
LegalGraphRAG (Ours) 8B-Inst 42.9 54.3 38.5 53.6 37.6 51.1 37.2 58.3 65.5 66.5 59.8 65.1 58.5 63.7 70.1 72.7 49.5 –
Table 2: Performance comparison on CAIL and CMDL. We employ Qwen3-8B as the default backbone
model. The best results are highlighted in bold, and the second-best are underlined. We visualize the gains of
LegalGraphRAG over each baseline in the ∆ columns .
Interpretations encoded in the Grul . The verifica- and synthesis, the system enforces stepwise verifi-
tion outcomes are then aggregated to produce a cation and ensures that every conclusion is explic-
definitive applicability judgment for each article. itly derived from and supported by verified legal
Based on these judgments, Auditor filters the evidence, resulting in reliable judicial decisions.
retrieval subgraph by pruning inapplicable articles
and their associated case and charge nodes. Finally, 5 Experiment
it organizes the remaining nodes into a legally con- This section presents a comprehensive evaluation
sistent and evidence-supported subgraph, which of LegalGraphRAG on two legal judgment bench-
serves as a validated knowledge basis for subse- marks. Our experiments are designed to answer the
quent decision-making. Further implementation following three questions. Q1 (Generation Accu-
details can be found in Appendix B.2. racy): Does LegalGraphRAG outperform SOTA
4.3.3 Evidence Synthesis GraphRAG methods and leading legal-domain
LLMs in generation quality? Q2 (Case Study):
In the final stage, the validated evidence produced
How does LegalGraphRAG handle specific legal
in the previous steps is synthesized to derive a
cases, and does it provide more interpretable out-
legally grounded judgment. Based on the verified
puts compared to baselines? Q3 (Ablation Study):
subgraph, Adjudicator integrates the confirmed ar-
What is the contribution of each core component to
ticles (Af ), cases (C f ), and offense information
the final performance of LegalGraphRAG? More
(Of ) to determine the applicable charges and their
additional experiments are provided in Appendix C
statutory basis. This process is formulated as:
5.1 Experiment Setup
J = Adjudicator(q ⊕ Af ⊕ C f ⊕ Of ) (12)
Datasets We evaluate on two widely used legal
Crucially, the judgment is not produced as a di- benchmarks: CAIL2018 (Xiao et al., 2018) and
rect verdict. Instead, it is accompanied by explicit CMDL (Huang et al., 2024), covering diverse crim-
citations to the statutory articles and judicial inter- inal sub-fields such as Public Safety, Social Order,
pretations used in the reasoning process, ensuring Economic Offenses, and Person Rights. The re-
that every conclusion is directly traceable to veri- trieval knowledge base is built from a collection of
fied evidence in the HierarGraph. authoritative legal sources, including case datasets
Overall, LegalGraphRAG formulates legal judg- and statutory texts. Further dataset details are pro-
ment as a transparent, evidence-based reasoning vided in Appendix D.1 & D.2.
pipeline rather than a black-box generation process. Baselines To ensure a comprehensive evalua-
Through sequential evidence grounding, validation, tion, we categorize our comparative experiments
37460
Query Please render a judgment against the defendant Zhang based on the following facts: “Defendant Zhang signed a Answer Charge: job-related embezzlement
labor contract ... On December 19, 2016, the Yaojiaba Tunnel Project Department transferred 98,160 yuan into Zhang's account Article: 271 Imprisonment: 18 months
after the canteen expense reimbursement. Zhang withdrew 95,000 yuan ... and took an additional 3,000 yuan the following day.”
Answer Charge:Embezzlement Article: 382 Imprisonment: 36 months Answer Charge: Job-Related Embezzlement Article: 271 Imprisonment: 18 months
Figure 4: A comparative case study illustrating the reasoning trajectories of different methods. While Naive
RAG fails due to missing legal articles and syllogism-based methods struggle with ambiguities, LegalGraphRAG
derives the correct judgment. By leveraging the HierarGraph and Evidence-based Legal Reasoning, our framework
demonstrates transparency and reliability, providing a verifiable reasoning chain grounded in legal evidence.
Figure 6: Reliability Analysis. LegalGraphRAG signif- Table 3: Ablation study of LegalGraphRAG compo-
icantly increases the proportion of Traceable Correct nents on the CAIL dataset. Results underscore the in-
samples, effectively minimizing Untraceable Correct dispensable role of the HierarGraph for knowledge or-
predictions where the answer is correct but lacks sup- ganization and the synergy between the Researcher and
porting evidence in the retrieved context. Auditor agents in ensuring reasoning accuracy.
diction accuracy overall. chain. LegalGraphRAG significantly increases the
Obs.2. LegalGraphRAG substantially surpasses ratio of “Traceable Correct” samples (defined in
existing specialized legal LLMs. Our approach Appendix A.5). By enforcing strict verification, our
outperforms Legal ∆ and ADAPT by an average of system ensures that every statute cited in the judg-
7.1% and 6.7%, respectively. Moreover, as shown ment is explicitly present in the retrieved context,
in Table 4 in Appendix, LegalGraphRAG integrates transforming opaque predictions into transparent,
flexibly with different backbone models, achieving traceable decisions.
a peak performance of 78.7% on CMDL when com-
bined with strong backbones. This demonstrates 5.4 Ablation Study (Q3)
strong adaptability and robust reasoning compared To quantify the impact of each component, we per-
to specialized legal-domain baselines. formed a systematic ablation study by removing
specific modules from the full LegalGraphRAG
5.3 Case Study (Q2)
framework. Results are detailed in Table 3.
To demonstrate the superior interpretability of our Obs.5. Hierarchical structure is the cornerstone
framework, we present a qualitative analysis of of performance. Removing the hierarchical graph
a representative criminal case in Figure 4. More (w/o HierarGraph) causes the sharpest accuracy
cases are provided in Appendix E. drop of 7.2%. This confirms that separating con-
Obs.3. LegalGraphRAG retrieves significantly crete facts from abstract rules into distinct granular
more relevant and comprehensive evidence. As levels is essential, providing structural precision
illustrated in Figure 5, conventional flat graph struc- that flat indexing lacks.
tures (e.g., HippoRAG2) struggle to handle hetero- Obs.6. The multi-agent workflow guarantees
geneous legal documents, often failing to capture reasoning reliability. Excluding the Researcher
essential statutes. This structural limitation leads and Auditor degrades accuracy by 4.0% and
to fragmented context. In contrast, our hierarchical 3.4%, respectively. This validates their synergis-
organization effectively structures legal knowledge, tic roles: the Researcher maximizes evidence cov-
ensuring that the retrieved context is sufficient to erage through diverse retrieval strategies, while
support robust reasoning. the Auditor enforces rigorous verification, ensuring
Obs.4. LegalGraphRAG guarantees decision only validated evidence supports the judgment.
traceability through rigorous evidence ground-
ing. While baseline models often achieve cor- 6 Conclusion
rect predictions, our reliability analysis (Figure In conclusion, we have presented LegalGraphRAG,
6) reveals a critical issue of “unsupported correct- an evidence-based legal reasoning framework that
ness”, where the model predicts the right charge addresses the critical challenges of legal hetero-
but fails to retrieve the necessary supporting evi- geneity and reasoning reliability. By integrating a
dence. This implies that the prediction is not sup- hierarchical knowledge graph with a collaborative
ported by relevant evidence or a valid reasoning multi-agent system, our approach transforms the le-
37462
gal reasoning process into a transparent pipeline of Bias and Fairness We acknowledge that mod-
retrieval, verification, and synthesis. Extensive ex- els trained on historical legal judgment data may
periments on legal judgment benchmarks validate inadvertently capture or amplify inherent biases
that LegalGraphRAG establishes a new state-of- present in the judicial system, such as those related
the-art, significantly advancing accurate and trust- to region or gender. While our work focuses on
worthy AI for reliable and complex legal analysis. improving the logical reasoning and retrieval capa-
bilities of legal LLMs through GraphRAG, where
Limitation the outputs are interpreted with clear evidence.
While LegalGraphRAG demonstrates significant Intended Use and Misuse The proposed Legal-
proficiency in processing textual legal documents GraphRAG is designed as an assistive tool to sup-
and statutes, its current scope is confined to uni- port legal professionals and researchers in retriev-
modal textual inputs. Real-world judicial proceed- ing precedents and analyzing case facts. It is not
ings, however, often rely on a heterogeneity of intended to replace human judges or lawyers, nor
evidence types, including crime scene photogra- should it be deployed as a fully automated decision-
phy, surveillance footage, scanned handwritten doc- making system in real-world judicial scenarios.
uments, and audio recordings of court hearings. The “prison term” and “judgment” predictions gen-
Currently, our framework requires all non-textual erated by the model should be viewed as reference
evidence to be transcribed or described textually probabilities rather than enforceable verdicts.
before processing, which may result in the loss
of critical visual or auditory nuances essential for Acknowledgements
fact verification. For instance, distinguishing be-
tween “inten” and “negligence” might sometimes The project was supported by National Key
rely on visual cues in surveillance video that tex- R&D Program of China (No. 2022ZD0160501),
tual descriptions fail to capture fully. Extending Natural Science Foundation of Fujian Province
the Hierarchical Legal Knowledge Graph to incor- of China (No. 2024J011001), and the Public
porate multimodal nodes (e.g., embedding visual Technology Service Platform Project of Xiamen
evidence into the Fact Graph) represents a promis- (No.3502Z20231043). We also thank the reviewers
ing avenue for future research. Such an extension for their insightful comments.
would enable the model to perform cross-modal
reasoning, verifying textual testimony against vi-
References
sual evidence, thereby moving closer to a holistic
and robust “Smart Court” system. Josh Achiam, Steven Adler, Sandhini Agarwal, Lama
Ahmad, Ilge Akkaya, Florencia Leoni Aleman,
Ethics Statement Diogo Almeida, Janko Altenschmidt, Sam Altman,
Shyamal Anadkat, and 1 others. 2023. Gpt-4 techni-
We confirm that this study fully complies with the cal report. arXiv preprint arXiv:2303.08774.
ACL Ethics Policy. Below, we address specific eth- Sebastian Borgeaud, Arthur Mensch, Jordan Hoff-
ical considerations regarding the data and the ap- mann, Trevor Cai, Eliza Rutherford, Katie Milli-
plication of our proposed model, LegalGraphRAG. can, George Bm Van Den Driessche, Jean-Baptiste
Lespiau, Bogdan Damoc, Aidan Clark, and 1 others.
Data Privacy and Compliance Our experiments 2022. Improving language models by retrieving from
involve four publicly available datasets (CAIL2018, trillions of tokens. In International conference on
CMDL, JuDGE, and LeCaRDv2) and statutory machine learning. PMLR.
texts. These resources are established benchmarks Zhiwei Cao, Qian Cao, Yu Lu, Ningxin Peng, Luyang
in the legal NLP community. We emphasize that Huang, Shanbo Cheng, and Jinsong Su. 2024. Retain-
all court judgments utilized in this work have been ing key information under high compression ratios:
Query-guided compressor for llms. In Proceedings
pre-processed and anonymized by the original data
of the 62nd Annual Meeting of the Association for
providers. Private details, including the real names Computational Linguistics (Volume 1: Long Papers),
of defendants and victims, have been removed or pages 12685–12695.
masked to ensure no personally identifiable infor-
Zhiwei Cao, Baosong Yang, Huan Lin, Suhang Wu, Xi-
mation (PII) is exposed. We strictly use this data angpeng Wei, Dayiheng Liu, Jun Xie, Min Zhang,
for academic research purposes and adhere to their and Jinsong Su. 2023. Bridging the domain gaps in
respective data usage licenses. context representations for k-nearest neighbor neural
37463
machine translation. In Proceedings of the 61st An- Zhiwei Fei, Songyang Zhang, Xiaoyu Shen, Dawei
nual Meeting of the Association for Computational Zhu, Xiao Wang, Jidong Ge, and Vincent Ng. 2025.
Linguistics (Volume 1: Long Papers), pages 5841– Internlm-law: An open-sourced chinese legal large
5853. language model. In Proceedings of the 31st Interna-
tional Conference on Computational Linguistics.
Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu
Lian, and Zheng Liu. 2024. Bge m3-embedding: Yan Gao, Zhiwei Cao, Zhongjian Miao, Baosong Yang,
Multi-lingual, multi-functionality, multi-granularity Shiyu Liu, Min Zhang, and Jinsong Su. 2024. Ef-
text embeddings through self-knowledge distillation. ficient k-nearest-neighbor machine translation with
arXiv preprint arXiv:2402.03216. dynamic retrieval. In Findings of the Association for
Computational Linguistics: ACL 2024, pages 7990–
Pierre Colombo, Telmo Pessoa Pires, Malik Boudiaf,
8001.
Dominic Culver, Rui Melo, Caio Corro, Andre FT
Martins, Fabrizio Esposito, Vera Lúcia Raposo, Sofia Sudipto Ghosh, Devanshu Verma, Balaji Ganesan, Purn-
Morgado, and 1 others. 2024. Saullm-7b: A pioneer- ima Bindal, Vikas Kumar, and Vasudha Bhatnagar.
ing large language model for law. arXiv preprint 2024. Inlegalllama: Indian legal knowledge en-
arXiv:2403.03883. hanced large language model. In International Joint
Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Conference on Artificial Intelligence.
Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Mar-
cel Blistein, Ori Ram, Dan Zhang, Evan Rosen, and Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, Chen-
1 others. 2025. Gemini 2.5: Pushing the frontier with hui Zhang, Da Yin, Dan Zhang, Diego Rojas, Guanyu
advanced reasoning, multimodality, long context, and Feng, Hanlin Zhao, and 1 others. 2024. Chatglm: A
next generation agentic capabilities. arXiv preprint family of large language models from glm-130b to
arXiv:2507.06261. glm-4 all tools. arXiv preprint arXiv:2406.12793.
Jiaxi Cui, Munan Ning, Zongjian Li, Bohua Chen, Yang Zirui Guo, Lianghao Xia, Yanhua Yu, Tu Ao, and
Yan, Hao Li, Bin Ling, Yonghong Tian, and Li Yuan. Chao Huang. 2024. Lightrag: Simple and fast
2023. Chatlaw: A multi-agent collaborative legal retrieval-augmented generation. arXiv preprint
assistant with knowledge graph enhanced mixture- arXiv:2410.05779.
of-experts large language model. arXiv preprint
arXiv:2306.16092. Bernal Jiménez Gutiérrez, Yiheng Shu, Weijian Qi,
Sizhe Zhou, and Yu Su. 2025. From rag to memory:
Xin Dai, Buqiang Xu, Zhenghao Liu, Yukun Yan, Non-parametric continual learning for large language
Huiyuan Xie, Xiaoyuan Yi, Shuo Wang, and Ge Yu. models. arXiv preprint arXiv:2502.14802.
2025. Legal δ: Enhancing legal reasoning in llms via
reinforcement learning with chain-of-thought guided Zhang Han and Dou Zhicheng. 2023. Case retrieval
information gain. arXiv preprint arXiv:2508.12281. for legal judgment prediction in legal artificial intelli-
gence. In Proceedings of the 22nd Chinese National
Hudson de Martim. 2025. Graph rag for legal norms: A Conference on Computational Linguistics.
hierarchical and temporal approach. arXiv preprint
arXiv:2505.00039. Xiaoxin He, Yijun Tian, Yifei Sun, Nitesh Chawla,
Thomas Laurent, Yann LeCun, Xavier Bresson,
Chenlong Deng, Kelong Mao, and Zhicheng Dou. and Bryan Hooi. 2024a. G-retriever: Retrieval-
2024a. Learning interpretable legal case retrieval augmented generation for textual graph understand-
via knowledge-guided case reformulation. arXiv ing and question answering. Advances in Neural
preprint arXiv:2406.19760. Information Processing Systems, 37:132876–132907.
Chenlong Deng, Kelong Mao, Yuyao Zhang, and
Zhicheng Dou. 2024b. Enabling discriminative rea- Zhitao He, Pengfei Cao, Chenhao Wang, Zhuoran Jin,
soning in llms for legal judgment prediction. arXiv Yubo Chen, Jiexin Xu, Huaijun Li, Kang Liu, and
preprint arXiv:2407.01964. Jun Zhao. 2024b. Agentscourt: Building judicial
decision-making agents with court debate simula-
Darren Edge, Ha Trinh, Newman Cheng, Joshua tion and legal knowledge augmentation. In Find-
Bradley, Alex Chao, Apurva Mody, Steven Truitt, ings of the Association for Computational Linguis-
Dasha Metropolitansky, Robert Osazuwa Ness, and tics: EMNLP 2024.
Jonathan Larson. 2024. From local to global: A
graph rag approach to query-focused summarization. Mengzhe Hei, Qingbao Liu, Sheng Zhang, Honglin
arXiv preprint arXiv:2404.16130. Shi, Jiashun Duan, and Xin Zhang. 2024. A het-
erogeneous graph based on legal documents and le-
Zhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou, gal statute hierarchy for chinese legal case retrieval.
Zhuo Han, Alan Huang, Songyang Zhang, Kai Chen, IEEE Access, 12:93502–93516.
Zhixin Yin, Zongwen Shen, and 1 others. 2024. Law-
bench: Benchmarking legal knowledge of large lan- Justin Ho, Alexandra Colby, and William Fisher. 2025.
guage models. In Proceedings of the 2024 conference Incorporating legal structure in retrieval-augmented
on empirical methods in natural language process- generation: A case study on copyright fair use. arXiv
ing. preprint arXiv:2505.02164.
37464
Zhitian Hou, Zihan Ye, Nanli Zeng, Tianyong Hao, and legal consultation conversation. In Proceedings of
Kun Zeng. 2025. Large language models meet le- the 48th International ACM SIGIR Conference on
gal artificial intelligence: A survey. arXiv preprint Research and Development in Information Retrieval.
arXiv:2509.09969.
Haitao Li, Yunqiu Shao, Yueyue Wu, Qingyao Ai, Yix-
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan iao Ma, and Yiqun Liu. 2024. Lecardv2: A large-
Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, scale chinese legal case retrieval dataset. In Proceed-
Weizhu Chen, and 1 others. 2022. Lora: Low-rank ings of the 47th International ACM SIGIR Confer-
adaptation of large language models. ICLR, 1(2):3. ence on Research and Development in Information
Retrieval.
Wanhong Huang, Yi Feng, Chuanyi Li, Honghan Wu, Ji-
dong Ge, and Vincent Ng. 2024. Cmdl: A large-scale Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang,
chinese multi-defendant legal judgment prediction Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi
dataset. In Findings of the Association for Computa- Deng, Chenyu Zhang, Chong Ruan, and 1 others.
tional Linguistics ACL 2024. 2024. Deepseek-v3 technical report. arXiv preprint
arXiv:2412.19437.
Cong Jiang and Xiaolei Yang. 2023. Legal syllogism
prompting: Teaching large language models for legal Qichuan Liu, Chentao Zhang, Yuxuan Hu, Chenfeng
judgment prediction. In Proceedings of the nine- Zheng, Qinggang Zhang, and Zhihong Zhang. 2026.
teenth international conference on artificial intelli- Facilitating generative retrieval with logical denois-
gence and law. ing for interpretable conversational search. In Pro-
ceedings of the ACM Web Conference 2026, page
Hui Jiang, Ziyao Lu, Fandong Meng, Chulun Zhou, 2296–2307.
Jie Zhou, Degen Huang, and Jinsong Su. 2022. To-
wards robust k-nearest-neighbor machine translation. Qichuan Liu, Chentao Zhang, Chenfeng Zheng, Gu-
In Proceedings of the 2022 Conference on Empiri- osheng Hu, Xiaodong Li, and Zhihong Zhang. 2025.
cal Methods in Natural Language Processing, pages Beyond the answer: Advancing multi-hop QA with
5468–5477. fine-grained graph reasoning and evaluation. In Pro-
ceedings of the 63rd Annual Meeting of the Associa-
Vladimir Karpukhin, Barlas Oguz, Sewon Min, tion for Computational Linguistics (Volume 1: Long
Patrick SH Lewis, Ledell Wu, Sergey Edunov, Danqi Papers), pages 23433–23456.
Chen, and Wen-tau Yih. 2020. Dense passage re-
trieval for open-domain question answering. In Antoine Louis, Gijs Van Dijck, and Gerasimos Spanakis.
EMNLP (1), pages 6769–6781. 2023. Finding the law: Enhancing statutory article
retrieval via graph neural networks. arXiv preprint
Zixuan Ke, Fangkai Jiao, Yifei Ming, Xuan-Phi Nguyen, arXiv:2301.12847.
Austin Xu, Do Xuan Long, Minzhi Li, Chengwei Qin,
Peifeng Wang, Silvio Savarese, and 1 others. 2025. Antoine Louis, Gijs van Dijck, and Gerasimos Spanakis.
A survey of frontiers in llm reasoning: Inference scal- 2024. Interpretable long-form legal question answer-
ing, learning to reason, and agentic systems. arXiv ing with retrieval-augmented large language models.
preprint arXiv:2504.09037. In Proceedings of the AAAI Conference on Artificial
Intelligence, volume 38.
Hyunjae Kim, Jiwoong Sohn, Aidan Gilson, Nicholas
Cochran-Caggiano, Serina Applebaum, Heeju Jin, Yun Luo, Zhen Yang, Fandong Meng, Yafu Li, Jie Zhou,
Seihee Park, Yujin Park, Jiyeong Park, Seoyoung and Yue Zhang. 2025. An empirical study of catas-
Choi, and 1 others. 2025. Rethinking retrieval- trophic forgetting in large language models during
augmented generation for medicine: A large-scale, continual fine-tuning. IEEE Transactions on Audio,
systematic expert evaluation and practical insights. Speech and Language Processing.
arXiv preprint arXiv:2511.06738.
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das,
Jinqi Lai, Wensheng Gan, Jiayang Wu, Zhenlian Qi, and Daniel Khashabi, and Hannaneh Hajishirzi. 2023.
Philip S Yu. 2024. Large language models in law: A When not to trust language models: Investigating
survey. AI Open, 5:181–196. effectiveness of parametric and non-parametric mem-
ories. In Proceedings of the 61st Annual Meeting of
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio the Association for Computational Linguistics (Vol-
Petroni, Vladimir Karpukhin, Naman Goyal, Hein- ume 1: Long Papers).
rich Küttler, Mike Lewis, Wen-tau Yih, Tim Rock-
täschel, and 1 others. 2020. Retrieval-augmented gen- Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad
eration for knowledge-intensive nlp tasks. Advances Saqib, Saeed Anwar, Muhammad Usman, Naveed
in neural information processing systems, 33:9459– Akhtar, Nick Barnes, and Ajmal Mian. 2025. A com-
9474. prehensive overview of large language models. ACM
Transactions on Intelligent Systems and Technology,
Haitao Li, Yifan Chen, Hu YiRan, Qingyao Ai, Jun- 16(5):1–72.
jie Chen, Xiaoyu Yang, Jianhui Yang, Yueyue Wu,
Zeyang Liu, and Yiqun Liu. 2025. Lexrag: Bench- Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida,
marking retrieval-augmented generation in multi-turn Carroll Wainwright, Pamela Mishkin, Chong Zhang,
37465
Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 Zhen Wan, Yating Zhang, Yexiang Wang, Fei Cheng,
others. 2022. Training language models to follow in- and Sadao Kurohashi. 2024. Reformulating domain
structions with human feedback. Advances in neural adaptation of large language models as adapt-retrieve-
information processing systems, 35:27730–27744. revise: A case study on chinese legal domain. In
Findings of the Association for Computational Lin-
Xiao Peng and Liang Chen. 2024. Athena: Retrieval- guistics: ACL 2024.
augmented legal judgment prediction with large lan-
guage models. arXiv preprint arXiv:2410.11195. Cunxiang Wang, Xiaoze Liu, Yuanhao Yue, Xiangru
Tang, Tianhang Zhang, Cheng Jiayang, Yunzhi Yao,
Nicholas Pipitone and Ghita Houir Alami. 2024. Wenyang Gao, Xuming Hu, Zehan Qi, and 1 others.
Legalbench-rag: A benchmark for retrieval- 2023. Survey on factuality in large language models:
augmented generation in the legal domain. arXiv Knowledge, retrieval and domain-specificity. arXiv
preprint arXiv:2408.10343. preprint arXiv:2310.07521.
Xuran Wang, Xinguang Zhang, Vanessa Hoo, Zhouhang
B. Rüthers, C. Fischer, and A. Birk. 2013. Rechtstheo- Shao, and Xuguang Zhang. 2024. Legalreasoner: A
rie mit juristischer Methodenlehre. Grundrisse des multi-stage framework for legal judgment prediction
Rechts. C.H. Beck. via large language models and knowledge integration.
IEEE Access.
Pranab Sahoo, Ayush Kumar Singh, Sriparna Saha,
Vinija Jain, Samrat Mondal, and Aman Chadha. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten
2024. A systematic survey of prompt engineering in Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou,
large language models: Techniques and applications. and 1 others. 2022. Chain-of-thought prompting elic-
arXiv preprint arXiv:2402.07927. its reasoning in large language models. Advances
in neural information processing systems, 35:24824–
Parth Sarthi, Salman Abdullah, Aditi Tuli, Shubh 24837.
Khanna, Anna Goldie, and Christopher D Manning.
2024. Raptor: Recursive abstractive processing for Hannes Westermann. 2024. Dallma: Semi-structured
tree-organized retrieval. In The Twelfth International legal reasoning and drafting with large language mod-
Conference on Learning Representations. els. In 2nd Workshop on Generative AI and Law.
Nirmalie Wiratunga, Ramitha Abeyratne, Lasal Jayawar-
Jeffrey A Segal. 1984. Predicting supreme court cases dena, Kyle Martin, Stewart Massie, Ikechukwu Nkisi-
probabilistically: The search and seizure cases, 1962- Orji, Ruvan Weerasinghe, Anne Liret, and Bruno
1981. American Political Science Review, 78(4):891– Fleisch. 2024. Cbr-rag: case-based reasoning for
900. retrieval augmented generation in llms for legal ques-
tion answering. In International Conference on Case-
Dong Shu, Haoran Zhao, Xukun Liu, David Demeter, Based Reasoning. Springer.
Mengnan Du, and Yongfeng Zhang. 2024. Lawllm:
Law large language model for the us legal system. In Shiguang Wu, Zhongkun Liu, Zhen Zhang, Zheng Chen,
Proceedings of the 33rd ACM International Confer- Wentao Deng, Wenhao Zhang, Jiyuan Yang, Zhi-
ence on information and knowledge management. tao Yao, Yougang Lyu, Xin Xin, Shen Gao, Pengjie
Ren, Zhaochun Ren, and Zhumin Chen. 2023a.
Marco Siino, Mariana Falco, Daniele Croce, and Paolo [Link].
Rosso. 2025. Exploring llms applications in law:
A literature review on current legal nlp approaches. Yiquan Wu, Siying Zhou, Yifei Liu, Weiming Lu, Xi-
IEEE Access. aozhong Liu, Yating Zhang, Changlong Sun, Fei Wu,
and Kun Kuang. 2023b. Precedent-enhanced legal
Weihang Su, Baoqing Yue, Qingyao Ai, Yiran Hu, Jiaqi judgment prediction with llm and domain-model col-
Li, Changyue Wang, Kaiyuan Zhang, Yueyue Wu, laboration. arXiv preprint arXiv:2310.09241.
and Yiqun Liu. 2025. Judge: Benchmarking judg-
ment document generation for chinese legal system. Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yi-
In Proceedings of the 48th International ACM SI- wen Ding, Boyang Hong, Ming Zhang, Junzhe Wang,
GIR Conference on Research and Development in Senjie Jin, Enyu Zhou, and 1 others. 2025. The
Information Retrieval. rise and potential of large language model based
agents: A survey. Science China Information Sci-
ences, 68(2):121101.
Octavia-Maria Sulea, Marcos Zampieri, Shervin Mal-
masi, Mihaela Vela, Liviu P Dinu, and Josef Van Gen- Zhishang Xiang, Chuanjie Wu, Qinggang Zhang,
abith. 2017. Exploring the use of text classification in Shengyuan Chen, Zijin Hong, Xiao Huang, and Jin-
the legal domain. arXiv preprint arXiv:1710.09306. song Su. 2025. When to use graphs in rag: A com-
prehensive analysis for graph retrieval-augmented
Vincent A Traag, Ludo Waltman, and Nees Jan Van Eck. generation. arXiv preprint arXiv:2506.05690.
2019. From louvain to leiden: guaranteeing well-
connected communities. Scientific reports, 9(1):1– Zhishang Xiang, Chengyi Yang, Zerui Chen, Zhimin
12. Wei, Yunbo Tang, Zongpei Teng, Zexi Peng, Zongxia
37466
Li, Chengsong Huang, Yicheng He, Chang Yang, Jianqiiu Zhang. 2024. Should we fear large language
Xinrun Wang, Xiao Huang, Qinggang Zhang, and Jin- models? a structural analysis of the human reason-
song Su. 2026. A systematic survey of self-evolving ing system for elucidating llm capabilities and risks
agents: From model-centric to environment-driven through the lens of heidegger’s philosophy. arXiv
co-evolution. TechRxiv, 2026(0227). preprint arXiv:2403.03288.
Chaojun Xiao, Haoxi Zhong, Zhipeng Guo, Cunchao Qinggang Zhang, Shengyuan Chen, Yuanchen Bei,
Tu, Zhiyuan Liu, Maosong Sun, Yansong Feng, Xi- Zheng Yuan, Huachi Zhou, Zijin Hong, Hao Chen,
anpei Han, Zhen Hu, Heng Wang, and 1 others. 2018. Yilin Xiao, Chuang Zhou, Junnan Dong, and 1 others.
Cail2018: A large-scale legal dataset for judgment 2025a. A survey of graph retrieval-augmented gener-
prediction. arXiv preprint arXiv:1807.02478. ation for customized large language models. arXiv
preprint arXiv:2501.13958.
Nuo Xu, Pinghui Wang, Long Chen, Li Pan, Xiaoyan
Wang, and Junzhou Zhao. 2020. Distinguish confus- Qinggang Zhang, Zhishang Xiang, Yilin Xiao, Le Wang,
ing law articles for legal judgment prediction. arXiv Junhui Li, Xinrun Wang, and Jinsong Su. 2025b.
preprint arXiv:2004.02557. Faithfulrag: Fact-level conflict modeling for context-
faithful retrieval-augmented generation. In Proceed-
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, ings of the 63rd Annual Meeting of the Association
Binyuan Hui, Bo Zheng, Bowen Yu, Chang for Computational Linguistics.
Gao, Chengen Huang, Chenxu Lv, and 1 others.
2025a. Qwen3 technical report. arXiv preprint Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang,
arXiv:2505.09388. Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen
Zhang, Junjie Zhang, Zican Dong, and 1 others. 2023.
Chang Yang, Chuang Zhou, Yilin Xiao, Su Dong, A survey of large language models. arXiv preprint
Luyao Zhuang, Yujing Zhang, Zhu Wang, Zijin Hong, arXiv:2303.18223, 1(2).
Zheng Yuan, Zhishang Xiang, and 1 others. 2026.
Graph-based agent memory: Taxonomy, techniques, Chulun Zhou, Chunkang Zhang, Guoxin Yu, Fandong
and applications. arXiv preprint arXiv:2602.05665. Meng, Jie Zhou, Wai Lam, and Mo Yu. 2025. Improv-
ing multi-step rag with hypergraph-based memory
Rui Yang. 2024. Casegpt: a case reasoning framework for long-context complex relational modeling. arXiv
based on language models and retrieval-augmented preprint arXiv:2512.23959.
generation. Preprint, arXiv:2407.07913.
Zhi Zhou, Jiang-Xin Shi, Peng-Xiao Song, Xiao-Wen
Xinyu Yang, Chenlong Deng, and Zhicheng Dou. 2025b. Yang, Yi-Xuan Jin, Lan-Zhe Guo, and Yu-Feng
Glare: Agentic reasoning for legal judgment predic- Li. 2024. Lawgpt: A chinese legal knowledge-
tion. arXiv preprint arXiv:2508.16383. enhanced large language model. arXiv preprint
arXiv:2406.04614.
Fangyi Yu, Lee Quartey, and Frank Schilder. 2022. Le-
gal prompting: Teaching a language model to think
like a lawyer. Preprint, arXiv:2212.01326.
embeddings, this process is defined as: (i) Diagnostic Retrieval: For each article va , the
agent retrieves its specific Diagnostic Checklist
Rsem (q) = Top-k sim ϕ(q), ϕ(c) (18) D(va ) = {d1 , . . . , d|C| } and relevant Judicial In-
c∈Gont
terpretations J from the Rule Graph Grul .
(ii) Community Expansion Retrieval: To cap- (ii) Item-wise Verification: The agent executes a
ture broader structural context, we employ a verification loop for each diagnostic item dk ∈
community-guided strategy. We identify the single D. It evaluates whether the raw case facts
most relevant thematic community K∗ aligned with q satisfy the specific legal condition dk , sup-
the query, and then retrieve the top-k similar cases porting the judgment with the interpretive con-
restricted within this community: text J . This produces a set of boolean veri-
fication results Vresults = {rk }, where rk ←
K∗ = argmax sim ϕ(q), ϕ(K)
K∈Gont CheckCondition(q, dk , va , J ).
(19) (iii) Decision and Pruning: Finally, the Audi-
Rcom (q) = Top-k sim ϕ(q), ϕ(c)
c∈K∗ tor synthesizes the verification results Vresults to
determine the overall applicability of the article.
(iii) Charge-Anchored Retrieval: Finally, we an-
If the article fails to meet the necessary criteria
chor the legal basis by retrieving cases linked to
(IsApplicable is False), the Auditor executes a
inferred charges. Here, O(q) denotes the set of pre-
pruning operation:
dicted charges and NGf ac (o) represents the neigh-
boring cases connected to charge o in the Fact Sverif ied ← Prune(Sverif ied , va ) (21)
Graph:
[ This step removes the inapplicable article node
Rchg (q) = NGf ac (o) (20)
va along with its dependent case precedents and
o∈O(q)
charge nodes, ensuring that the final subgraph
These three retrieval strategies organize the can- Sverif ied contains only logically valid and appli-
didate evidence set Scand . cable evidence.
Auditor Agent: Operating on the candidate evi- Adjudicator Agent: The Adjudicator synthesizes
dence set Scand , the Auditor validates the appli- the verified subgraph Sverif ied to render the final
cability of each retrieved article va through a rig- judgment. Specifically, it organizes the valid nodes
orous verify-and-prune mechanism. This process extracted from Sverif ied into sets of confirmed arti-
proceeds in three specific steps: cles (VAf ), case precedents (VCf ), and charge infor-
37471
Algorithm 1 Evidence-based Legal Reasoning
Require: Raw case query q; Ontology Graph Gont ; Fact Graph Gf ac ; Rule Graph Grul .
Ensure: Final Judgment J with citations.
Stage 1: Researcher Agent (Multi-Strategy Retrieval)
1: ϕ(q) ← OntologyAlign(q, Gont ) ▷ Align query to ontology features
Parallel Evidence Retrieval Strategies:
2: Scand ← Rsem ∪ Rcom ∪ Rchg ▷ Union of candidate evidence
Stage 2: Auditor Agent (Verification & Pruning)
3: Sverif ied ← Scand
4: for each article node va ∈ Scand do
5: D ← RetrieveChecklist(va , Grul ) ▷ Get checklist D(va ) = {d1 , . . . , d|C| }
6: J ← RetrieveInterpretations(va , Grul ) ▷ Get Judicial Interpretations
7: Vresults ← ∅
8: for each diagnostic item dk ∈ D do ▷ Item-wise verification loop
9: rk ← CheckCondition(q, dk , va , J ) ▷ Verify if fact q satisfies condition dk
10: Vresults ← Vresults ∪ {rk }
11: end for
12: IsApplicable ← Decide(Vresults ) ▷ Final determination for node va
13: if not IsApplicable then
14: Sverif ied ← Prune(Sverif ied , va ) ▷ Remove article and linked nodes
15: end if
16: end for
f f f f
17: Gsub ← {VA , VC , VO } ← Organize(Sverif ied ) ▷ Structure the verified subgraph
Stage 3: Adjudicator Agent (Synthesis)
f f f
18: Y ← Adjudicator(q ⊕ VA ⊕ VC ⊕ VO ) ▷ Synthesize judgment with citations
19: return Y
mation (VOf ). By integrating these evidence com- CAIL and CMDL datasets, regardless of the back-
ponents with the original query q, it generates a bone model employed. Notably, even with the
response with explicit citations. This process is lighter GPT-4o-mini on the CMDL dataset, our
formulated as: method achieves a remarkable performance gain
(e.g., significantly exceeding the strong baseline
Y = Adjudicator(q ⊕ VAf ⊕ VCf ⊕ VOf ) (22) RAPTOR in Accuracy), while maintaining its lead
with the more powerful DeepSeek-V3.1. This
The output Y ensures that every conclusion is di-
demonstrates that LegalGraphRAG’s structured
rectly traceable to specific nodes in the knowledge
reasoning capabilities effectively complement the
graph, enforcing transparency and evidence-based
generation power of various state-of-the-art LLMs,
reasoning.
enhancing their precision in complex legal applica-
C Additional Experiments tion scenarios independent of the underlying model
architecture.
C.1 Extensions to the Main Experiment (Q4) Obs.8. Exactness in Law Article Prediction. Ta-
In this section, we conduct a series of extended ble 5 illustrates the model’s capability in Law Arti-
experiments to verify the universality of our frame- cle Prediction, a task demanding precise statutory
work across different model architectures and its ro- grounding rather than generative flexibility. Legal-
bustness in specific, high-difficulty legal sub-tasks. GraphRAG achieves a superior overall accuracy
Obs.7. Universality across Advanced Backbones. of 47.9%, establishing a substantial lead over both
To verify the universality of our framework, we ex- the strongest RAG baseline, HippoRAG2 (39.8%),
tended the evaluation to advanced large language and the domain-specific state-of-the-art, ADAPT
models, specifically DeepSeek-V3.1 and GPT-4o- (41.3%). Remarkably, our 8B-parameter frame-
mini. As shown in Table 4, LegalGraphRAG work even surpasses the massive DeepSeek-V3.1
consistently outperforms all baselines across both (44.9%), highlighting that our structured, evidence-
37472
CAIL
Public Safety Economic Social Order Person Rights
Model Size All ∆
ACC F1 ACC F1 ACC F1 ACC F1
GPT-4o-mini
Naive RAG ∼8B 27.5 37.6 18.8 33.6 18.0 28.8 22.1 39.2 22.2 ↑ 18.7
G-Retriever ∼8B 17.5 24.8 20.3 32.8 20.5 31.0 24.1 31.6 21.4 ↑ 19.5
LightRAG ∼8B 25.4 37.4 21.3 36.9 21.7 38.7 23.4 42.8 23.1 ↑ 17.8
RAPTOR (Sarthi et al., 2024) ∼8B 33.1 49.0 29.3 44.2 25.9 39.9 28.3 43.1 30.5 ↑ 10.4
HippoRAG2 (Gutiérrez et al., 2025) ∼8B 33.1 50.1 28.1 46.1 21.7 43.6 37.2 54.2 31.9 ↑ 9.0
LegalGraphRAG (Ours) ∼8B 39.6 54.8 36.3 52.9 37.3 51.2 42.1 62.4 40.9 –
DeepSeek-V3.1
Naive RAG ∼200B 38.0 54.0 32.3 49.9 33.4 47.3 40.7 53.4 37.8 ↑ 12.1
G-Retriever ∼200B 36.5 54.6 35.1 49.8 36.2 47.2 39.5 48.3 37.2 ↑ 12.7
LightRAG ∼200B 36.6 48.5 26.3 50.2 33.5 46.3 39.7 53.1 45.4 ↑ 4.5
RAPTOR (Sarthi et al., 2024) ∼200B 42.2 56.3 37.8 53.0 39.2 50.7 45.5 52.3 44.4 ↑ 5.5
HippoRAG2 (Gutiérrez et al., 2025) ∼200B 41.5 49.1 33.1 46.8 34.3 47.0 38.6 46.4 41.2 ↑ 8.7
LegalGraphRAG (Ours) ∼200B 44.4 58.8 41.9 57.8 41.9 56.8 46.2 65.1 49.9 –
CMDL
Public Safety Economic Social Order Person Rights
Model Size All ∆
ACC F1 ACC F1 ACC F1 ACC F1
GPT-4o-mini
Naive RAG ∼8B 38.7 50.8 35.6 44.0 30.9 44.2 33.3 43.0 32.9 ↑ 17.1
G-Retriever ∼8B 24.6 38.4 29.7 38.8 30.4 40.0 41.1 51.6 28.3 ↑ 21.7
LightRAG ∼8B 36.2 43.1 37.5 48.9 46.9 55.1 34.6 50.8 34.2 ↑ 15.8
RAPTOR (Sarthi et al., 2024) ∼8B 42.0 49.1 48.2 58.0 57.3 60.9 46.0 59.5 48.7 ↑ 11.3
HippoRAG2 (Gutiérrez et al., 2025) ∼8B 45.0 52.5 42.9 60.8 45.3 64.6 59.4 75.0 46.0 ↑ 14.0
LegalGraphRAG (Ours) ∼8B 46.5 57.8 63.7 67.9 58.0 66.2 67.2 75.4 60.0 –
DeepSeek-V3.1
Naive RAG ∼200B 52.0 65.3 66.1 71.6 69.8 70.7 68.8 80.5 62.9 ↑ 12.8
G-Retriever ∼200B 46.2 64.9 58.6 69.6 56.1 67.4 69.7 78.9 55.2 ↑ 23.5
LightRAG ∼200B 47.7 58.7 46.5 67.6 47.6 53.2 52.3 64.2 52.4 ↑ 26.3
RAPTOR (Sarthi et al., 2024) ∼200B 53.4 64.7 68.2 73.0 66.4 75.4 60.3 71.1 56.4 ↑ 12.3
HippoRAG2 (Gutiérrez et al., 2025) ∼200B 62.1 63.5 62.4 65.3 75.9 78.6 76.9 78.2 74.0 ↑ 4.7
LegalGraphRAG (Ours) ∼200B 66.7 69.9 76.0 79.3 72.9 80.4 79.7 85.5 78.7 –
Table 4: Performance comparison on advanced Models. We compared LegalGraphRAG and other baselines
utilizing advanced LLMs as backbones.
based retrieval mechanism is more effective at pin- LegalGraphRAG’s evidence-based retrieval strat-
pointing legal provisions than simply scaling model egy effectively locates relevant sentencing guide-
parameters or employing semantic retrieval. lines and comparable precedents, thereby constrain-
ing the generation to a more precise and legally
Obs.9. Precision in Term of Penalty Predic-
grounded time range.
tion. Table 6 presents the results on the challenging
term of penalty prediction task, which requires fine-
C.2 Hyper-parameter Sensitivity (Q5)
grained quantitative reasoning rather than simple
classification. LegalGraphRAG demonstrates a sig- To evaluate system stability, we investigated the
nificant advantage in minimizing prediction error, sensitivity of the Researcher Agent to the retrieval
consistently achieving the lowest Mean Absolute parameter k, which governs the number of seman-
Error (MAE) across most subdomains compared tic concepts retrieved from the ontology graph Gont .
to other RAG-based methods. For instance, in We varied k over the set {3, 4, 5, 6}. The upper
the Public Safety category, our model achieves an bound is restricted to 6, as empirical evidence sug-
MAE of 20.9, outperforming RAPTOR (21.7) and gests that exceeding this threshold introduces exces-
HippoRAG2 (23.0). This indicates that while ex- sive context noise, which overwhelms the model’s
act term matching remains difficult for all models, effective window and degrades reasoning.
37473
CAIL
Public Safety Economic Social Order Person Rights
Model Size All ∆
ACC F1 ACC F1 ACC F1 ACC F1
Open-Source Models
Qwen-2.5-7B-Instruct 7B-Inst 28.7 53.8 24.1 48.2 26.6 53.5 36.0 56.2 30.1 ↑ 17.8
Qwen-3-8B 8B-Inst 23.9 58.5 27.6 51.2 36.6 62.2 46.2 66.7 35.9 ↑ 12.0
Internlm3-8b-instruct 8B-Inst 27.6 59.9 26.4 52.8 30.8 59.0 33.6 57.6 29.9 ↑ 18.0
Glm-4-9b-chat 9B-Inst 23.9 60.1 25.6 51.0 34.7 59.4 35.1 58.5 30.8 ↑ 17.1
Advanced Models
GPT-4o-mini (Achiam et al., 2023) ∼8B 24.7 57.2 25.7 42.4 24.6 51.8 34.7 55.0 30.9 ↑ 17.0
DeepSeek-V3.1 (Liu et al., 2024) ∼200B 42.3 63.2 37.1 61.3 44.5 68.8 51.9 67.3 44.9 ↑ 3.0
Legal Specific Methods
DISC-LawLLM-7B (Yue et al., 2024) 7B-Inst 36.5 55.9 29.9 48.8 41.5 61.2 39.3 57.8 38.2 ↑ 9.7
ADAPT (Deng et al., 2024b) 7B-Inst 40.1 50.8 31.6 42.4 39.6 54.9 41.0 50.5 41.3 ↑ 6.6
Legal ∆ (Dai et al., 2025) 7B-Inst 33.0 54.9 27.8 51.0 34.6 59.4 44.5 60.6 37.9 ↑ 10.0
RAG Based Methods
Naive RAG 8B-Inst 30.4 45.7 29.2 48.6 37.5 54.1 36.6 51.1 34.8 ↑ 13.1
RAPTOR (Sarthi et al., 2024) 8B-Inst 36.4 59.2 33.4 53.9 38.9 64.1 41.0 61.7 37.2 ↑ 10.7
HippoRAG2 (Gutiérrez et al., 2025) 8B-Inst 35.2 60.3 33.1 53.7 41.7 67.0 42.6 63.8 39.8 ↑ 8.1
LegalGraphRAG (Ours) 8B-Inst 43.0 64.9 37.8 61.0 44.6 69.4 54.5 70.6 47.9 –
Table 5: Extended experiments on Article Prediction. We evaluated the performance of our model and baselines
on the specific sub-task of law article prediction. We visualize the gains of LegalGraphRAG to the each baseline in
the ∆ columns .
CAIL
Public Safety Economic Social Order Person Rights
Model Size All ∆
ACC MAE ACC MAE ACC MAE ACC MAE
Open-Source Models
Qwen-2.5-7B-Instruct 7B-Inst 13.0 23.7 8.1 33.3 6.0 30.3 9.7 29.4 29.5 ↑ 8.4
Qwen-3-8B 8B-Inst 8.3 32.6 11.3 31.6 17.9 29.7 11.4 26.4 27.5 ↑ 7.4
Internlm3-8b-instruct 8B-Inst 7.1 35.2 5.9 37.2 6.0 32.5 7.2 37.3 33.7 ↑ 13.6
Glm-4-9b-chat 9B-Inst 3.6 36.0 3.2 32.9 1.5 35.1 6.8 38.6 33.1 ↑ 13.0
Advanced Models
GPT-4o-mini (Achiam et al., 2023) ∼8B 6.9 38.2 7.3 34.2 8.3 31.7 8.0 34.6 33.6 ↑ 13.5
DeepSeek-V3.1 (Liu et al., 2024) ∼200B 7.1 31.2 8.1 31.5 10.4 25.9 8.4 29.1 29.1 ↑ 8.6
Legal Specific Methods
DISC-LawLLM-7B (Yue et al., 2024) 7B-Inst 15.5 24.5 5.0 33.2 3.0 37.9 8.4 34.5 31.6 ↑ 11.5
ADAPT (Deng et al., 2024b) 7B-Inst 8.3 21.8 10.9 21.9 3.0 24.3 9.7 22.3 20.4 ↑ 0.3
Legal ∆ (Dai et al., 2025) 7B-Inst 11.9 25.0 9.0 29.0 9.0 27.5 8.9 27.1 26.3 ↑ 6.2
RAG Based Methods
Naive RAG 8B-Inst 10.5 29.8 11.4 30.8 18.1 21.7 12.6 24.3 26.5 ↑ 4.4
RAPTOR (Sarthi et al., 2024) 8B-Inst 12.0 21.7 11.7 28.3 15.8 26.1 16.2 21.8 24.3 ↑ 4.2
HippoRAG2 (Gutiérrez et al., 2025) 8B-Inst 13.1 23.0 12.7 25.8 17.9 23.8 13.5 23.4 23.8 ↑ 3.7
LegalGraphRAG (Ours) 8B-Inst 14.0 20.9 13.7 22.1 19.4 23.6 17.1 22.7 20.1 –
Table 6: Extended experiments on Term of Penalty Prediction. We assessed the accuracy and error rates of
imprisonment term predictions compared to baselines. We visualize the gains of LegalGraphRAG to the each
baseline in the ∆ columns .
37474
Token Consumption (< 106 ) Avg Cost
Method Indexing Time (s)
Prompt Completion Time Token
RAPTOR 13696.90 5.64 0.72 5.86 3589
HippoRAG2 4581.60 10.58 2.79 11.2 5199
LegalGraphRAG (Ours) 3687.49 3.97 0.78 46.1 10664
Table 7: Comparison of Computational Efficiency: Offline Indexing vs. Online Inference. We report the total time
and token usage for graph construction (Indexing) and the average cost per query (Online).
prises 393,945 criminal cases with approximately Dataset Article Judicial interpretations
1.2 million defendants, covering 321 distinct
# Num 452 656
charges and 275 legal articles. Notably, CMDL in- # Average length 128.54 243.95
troduces case-level evaluation metrics that account
for case complexity and varying numbers of defen- Table 10: Basic statistics of the legal knowledges in
dants, offering a more holistic assessment of model corpus.
performance in multi-defendant scenarios. For ex-
perimental feasibility, the subset CMDL-small is
a subset of cases from multiple authoritative legal
often utilized, as it preserves the data distribution
datasets: JuDGE(Su et al., 2025), CAIL2018(Xiao
while significantly reducing computational costs,
et al., 2018), CMDL(Huang et al., 2024), and
making it suitable for preliminary benchmarking
LeCaRDv2(Li et al., 2024). The construction fol-
and model validation in resource-constrained re-
lows a procedure similar to that used for the ex-
search settings.
perimental subsets: we first filter cases by fact de-
Dataset Construction Due to considerations re- scription length (under 1,024 characters) and apply
garding the generation speed and token cost of the sampling designed to balance the representation of
GraphRAG method, we constructed focused sub- different charges, thereby creating a broad and di-
sets from both the CAIL2018 and CMDL datasets verse collection of historical precedents and factual
using a uniform procedure aimed at controlling patterns. Furthermore, to ground the system in au-
input length, balancing charge distribution, and thoritative legal provisions, we incorporate the full
elevating task complexity. The construction first fil- text of the “Criminal Law of the People’s Republic
tered cases to retain only those with factual descrip- of China” along with its relevant judicial interpre-
tions under 1,024 characters. To ensure broader tations as a core statutory knowledge library. The
coverage of under-represented charges and increase combination of this curated historical case library
predictive difficulty, the sampling prioritized de- and the official legal provisions library forms the
fendants whose charges included low-frequency of- complete corpus, enabling models to retrieve both
fenses and deliberately retained a higher proportion experiential precedents and statutory knowledge
of multi-charge cases. Consequently, the resulting during [Link], we have carefully ver-
subsets feature a more balanced charge distribution ified that all cases in the corpus are distinct from
with elevated presence of rare charges and greater those in the test subsets, ensuring no data leakage
average case complexity compared to the original between the knowledge base and the evaluation
datasets. While this design provides a more chal- benchmarks. The size and composition statistics of
lenging testbed for evaluating model performance the final corpus are detailed in Table 9 & 10.
on complex and low-frequency legal scenarios, it
may also lead to lower reported performance for D.3 Evaluation Metrics
some methods relative to their results on the orig- To comprehensively evaluate the Legal Judgment
inal, more naturalistic data distribution. Table 8 Prediction (LJP) tasks, we employ specific metrics
presents detailed statistics of the subsets. for different sub-tasks: Charge and Article Pre-
diction are evaluated using Accuracy (ACC) and
D.2 Corpus
Micro-F1, Term of Penalty Prediction is assessed
In this section, we provide detailed descriptions of using ACC and Mean Absolute Error (MAE), and
corpus used in our experiments. To ensure compre- the Retrieval Quality of our RAG system is mea-
hensive coverage of criminal statutes and charges, sured by Retrieval Effectiveness and Error Rate.
we construct this knowledge base by aggregating Accuracy (ACC) measures the exact match ratio.
37476
For classification tasks, it requires the predicted D.4 Baseline Details
label set Oi to be identical to the ground truth Oi′ .
In this section, we provide detailed descriptions of
For Term of Penalty, it measures the exact match
each baseline used in our comparison, as detailed
of the predicted term. It is calculated as:
in figure 10.
N Naive RAG uses the standard RAG paradigm:
1 X a retriever model first retrieves relevant context
ACC = I(Oi = Oi′ )
N from the corpus based on the given question, and
i=1
then the question is concatenated with the retrieved
Where I(·) is the indicator function. context to form a query for the generation model
Micro-F1 is used for multi-label classification to to produce the final answer.
account for class imbalance. It is the harmonic G-retriever(He et al., 2024a) introduces a re-
mean of micro-averaged precision (Pmicro ) and re- trieval augmented generation framework for tex-
call (Rmicro ): tual graphs by formulating subgraph retrieval as
a Prize-Collecting Steiner Tree optimization prob-
2 · Pmicro · Rmicro lem, enabling conversational question answering
Micro-F1 =
Pmicro + Rmicro across diverse domains like scene understanding
and knowledge graphs while mitigating LLM hal-
Mean Absolute Error (MAE) reflects the devia- lucinations and scaling to large graph sizes.
tion in the predicted term of penalty. Let Ti and Ti′ LightRAG(Guo et al., 2024) introduces a graph-
denote the predicted and ground-truth prison terms enhanced retrieval-augmented generation frame-
(in months), respectively. MAE is defined as: work that integrates entity-relationship graphs into
text indexing, combining low-level precise entity
N
1 X retrieval with high-level thematic discovery for ef-
MAE = |Ti − Ti′ | ficient and adaptive knowledge integration.
N
i=1
HippoRAG2(Gutiérrez et al., 2025) builds on
HippoRAG’s Personalized PageRank framework
Retrieval Effectiveness measures how well the
by integrating dense-sparse coding for passages
retrieved content aligns with the question’s intent.
and phrases in the knowledge graph, enabling
Higher values indicate more focused and pertinent
deeper contextualization and recognition memory
information. It is defined as:
for triple filtering. Enhances online retrieval with
1 X query-to-triple matching and optimized seed node
Retrieval Effectiveness = R(c, Q, E) weighting, outperforming standard RAG across fac-
|C|
c∈C
tual, sense-making, and associative memory tasks.
where C denotes the set of retrieved contexts, Q rep- RAPTOR(Sarthi et al., 2024) constructs a hier-
resents the question, E denotes the set of evidence, archical tree by recursively clustering and summa-
and the operator R(·) determines the relevance of rizing embedded text chunks, enabling retrieval
a context c. of information at multiple levels of abstraction to
improve performance on long-document question-
Error Rate quantifies the incompleteness of the
answering tasks.
retrieval process. Instead of measuring recall di-
Disc-LawLLM(Yue et al., 2024) is a retrieval-
rectly, we assess the proportion of reference claims
augmented large language model fine-tuned on Chi-
not supported by the retrieved context:
nese judicial datasets using legal syllogism prompt-
! ing to provide reasoning-capable legal services,
1 X
Error Rate = 1 − I(S(c, C)) (23) including consultation, judgment prediction, and
|R| examination assistance. In this work, we use the
c∈R
officially open-sourced LawLLM-7B, which is fine-
where R is the set of reference claims, S(·) de- tuned from Qwen2.5-Instruct-7B.
termines whether a claim c is supported by the re- Legal ∆(Dai et al., 2025) employs a reinforce-
trieved context C, and I(·) is the indicator function. ment learning framework that enhances legal rea-
A lower Error Rate indicates a more comprehensive soning in LLMs by maximizing chain-of-thought
evidence collection. guided information gain through dual-mode inputs
37477
RAG Configuration RAPTOR Configuration
{ {
embedding_model: bge-m3, embedding_model: bge-m3,
retrieval_topk: 5, chunk_token_size: 1200,
chunk_token_size: 1000, chunk_overlap_token_size: 100,
chunk_overlap_token_size: 200 num_layers: 5,
} max_length_in_cluster: 3500,
threshold: 0.1,
RAG Configuration cluster_metric: cosine,
threshold_cluster_num: 5000
{ }
embedding_model: bge-m3,
retrieval_topk: 5,
chunk_token_size: 1000, Figure 10: Hyperparameter configurations for the base-
chunk_overlap_token_size: 200 line RAG models.
}
and differential Q-value analysis. In this work, we
G-retriever Configuration use the officially open-sourced model, which is
fine-tuned from Qwen2.5-Instruct-7B.
{
ADAPT(Deng et al., 2024b) is a discriminative
embedding_model: bge-m3,
reasoning framework for LLMs in legal judgment
retrieval_topk: 3,
prediction that emulates human judicial processes
chunk_token_size: 1200,
by asking to decompose case facts into key ele-
chunk_overlap_token_size: 100,
ments, discriminating among candidate charges for
entities_max_tokens: 2000,
alignment, and predicting final judgments, further
relationships_max_tokens: 2000
improved via multi-task fine-tuning with synthetic
}
trajectories. In this work, we use the officially open-
sourced model, which is fine-tuned from Qwen2-
LightRAG Configuration 7B.
{ D.5 LegalGraphRAG Setup
embedding_model: bge-m3,
In the experimental setup of LegalGraphRAG, hy-
query_type: hybrid,
perparameters are configured to optimize retrieval
chunk_token_size: 1200,
precision. During graph construction, the Ontol-
retrieval_topk: 20,
ogy Graph utilizes k-nearest neighbors (kNN) to
chunk_overlap_token_size: 100,
select the top-3 case feature nodes based on cosine
max_token_text_unit: 2000,
similarity for direct semantic matching.
max_token_global_context: 2000,
For the Researcher agent, the evidence-based re-
max_token_local_context: 2000
trieval strategy operates with specific thresholds:
}
the retrieval parameter is set to k = 5. This parame-
ter governs both the number of top-ranked semantic
HippoRAG2 Configuration concepts retrieved from the ontology graph and the
{ scope of community expansion, a value selected
embedding_model: bge-m3, based on the sensitivity analysis to balance context
retrieval_top_k: 5, coverage and noise control.
linking_top_k: 5,
E Extended Case Examples
max_qa_steps: 3,
qa_top_k: 5, In this section, we will walk through several cases
graph_type: facts_and_sim_passage to detail the retrieval and reasoning pipeline of
_node_unidirectional LegalGraphRAG. Figure 11 illustrates evidence re-
} trieval for a Dangerous Driving case, while Figure
37478
Query The procuratorial organ alleges that at approximately 18:45 on November 10, 2020, the defendant Wang Moumou, while driving under the
influence of alcohol and without a valid driver's license, was operating a passenger car (License Plate: Liao D×××××; the vehicle's owner was the defendant Liu
Mou) southbound on Changling Street in a certain town of a certain county, Qingyuan. When reaching a traffic signal intersection, he collided with a compact
sport utility vehicle (License Plate: Liao D×××××) driven by Zhao Moumou, who was traveling ahead in the same direction. This resulted in a road traffic
accident causing damage to both [Link] Traffic Police Brigade of a certain county in Qingyuan determined that the defendant Wang Moumou bore full
responsibility for this accident, while Zhao Moumou bore no responsibility. An appraisal conducted by the Fushun Gongzheng Judicial Appraisal Institute detected
the presence of ethanol in the defendant Wang Moumou's venous blood, with a concentration of 242.5 mg/100 ml, constituting drunk driving.
Case Feature D: "Driving under the influence"; O: "Traffic collision"; V: "Personal injury", "Passenger car driver"; M: "Negligence"
Community Expansion Retrieval Community Summary: Cases often involve offenses such as drunk driving and traffic accidents.
Researcher
Fact: At approximately 14:10 on October 18, 2016, the defendant Chen Wenli, while reversing a gray Wuling-brand light...
Case Feature: D: "Adult"; O: "Driving a vehicle", "Failed to observe surroundings"; V: "Death after failed rescue efforts"; M: "Negligence"
Name: Chen Wenli Charges: Crime of Causing Death by Negligence Law Articles: 233...
Charge-Anchored Retrieval "Crime of Dangerous Driving", "Crime of Causing a Traffic Accident"... Laws: 133
Final Context
Related Laws: [Article 133 Paragraph 1] [Crime of Dangerous Driving] Whoever drives a motor vehicle on a road under any of the ...
(Other related laws are omitted below)
Similar Cases: [Charges: Crime of Dangerous Driving] [Name: Wang] The procuratorial organ alleges that at approximately 18:45 on...
Adjudicator (Other similar cases are omitted below)
Final Results
Charge: Crime of Dangerous Driving Articles: 133 Imprisonment: 18 months
Reasoning: Li Moujia drove a motor vehicle on a road while intoxicated (with a blood ethanol content of 155.7 mg/100 ml). His conduct
satisfies the circumstance specified in Article 133-1(1)(2) of the Criminal Law, "driving a motor vehicle while intoxicated," and thus
constitutes the crime of dangerous driving [1]...
[1] corresponds to Article 133
[2] corresponds to Cases1([Charges: Crime of Dangerous Driving] [Name: Wang])
...
Figure 11: Qualitative analysis of a representative case regarding the crime of Dangerous Driving. The visualization
highlights the retrieval of evidence related to specific statutory conditions.
Case Feature D: "Employee", "Staff"; O: "Misappropriation of funds"; V: "Company funds", "Significant amount"; M: "Subjective intent"
Community Expansion Retrieval Community Summary: Cases often involve duty-related crimes, bribery crimes, and property...
Researcher
Fact: From 2009 to 2011, while serving as a credit officer at the Qiuji Credit Cooperative in Suining County, the defendant Zhang ...
Case Feature: D: "State Functionary"; O: "Embezzlement of funds", "Taking advantage of one's position"; V: "Unit funds"; M: "Direct intent"
Name: Zhang Zhi'an Charges: Misappropriation of Funds Law Articles: 185, 272
Final Context
Related Laws: [Article 271 Paragraph 1] [job-related embezzlement] Employees of companies, enterprises, or other entities who...
(Other related laws are omitted below)
Similar Cases: [Charges: job-related embezzlement] [Name: Xu 3Mou] Between December 2007 and December 2009, the defendant Xu...
Adjudicator (Other similar cases are omitted below)
Final Results
Charge: job-related embezzlement Articles: 271 Imprisonment: 18 months.
Resoning: Zhang served as the head chef of the project department and was responsible for financial reimbursement, which qualifies him
as a "personnel of a company, enterprise, or other unit" [1]. By taking advantage of his duty in handling canteen expenses...
[1] corresponds to Article 271 Paragraph 1
[2] corresponds to Cases1([Charges: job-related embezzlement] [Name: Xu 3Mou])
...
Figure 12: Qualitative analysis of a representative case regarding the crime of Occupational Embezzlement. The
example demonstrates the model’s reasoning in identifying the abuse of professional position.
yond Chinese, SaulLM (Colombo et al., 2024) fo- gal, 1984) relied on artificially designed features,
cuses on English legal texts based on the Mixtral and traditional machine learning methods (Sulea
architecture, while LawLLM (Shu et al., 2024) et al., 2017) were applied to predict legal judg-
addresses US legal tasks such as similar case re- ments. Recent advances in deep learning (Xu
trieval. Additionally, specialized models like In- et al., 2020; Han and Zhicheng, 2023) have mo-
LegalLLaMA (Ghosh et al., 2024) target Indian tivated researchers to leverage neural networks
and French legal domains respectively. These mod- for automated text representation learning. Re-
els provide crucial baselines for downstream tasks cently, LLMs have further promoted the progress of
but often lack the specific reasoning architecture LJP (Deng et al., 2024a). Several studies (Wu et al.,
required for complex judgment prediction. 2023b; Peng and Chen, 2024) employ Retrieval-
Augmented Generation (RAG) to enhance LLMs
G.2 Legal judgment prediction by incorporating external legal knowledge. To re-
fine decision-making, recent works have introduced
Legal judgment prediction (LJP) has experienced
structured reasoning frameworks that systemati-
significant development and become an increas-
cally decompose case facts to distinguish confusing
ingly crucial NLP task. Earlier research (Se-
37480
Criminal Case Keyword Extraction
Task Definitions: Extract legal keywords from a criminal case description and
classify them into four categories: Defendant Attributes, Criminal Behaviors,
Victim Characteristics, and Subjective Mental States. The output must be a
strictly valid JSON object without additional text.
Keyword Definitions:
• Defendant Attributes: Legal traits (e.g., age group, criminal history,
occupation). Avoid specific names or numbers.
• Criminal Behaviors: Legal types of acts and significant methods. Exclude
specific time/location details.
• Victim Characteristics: Nature of the property or location. Generalize
specific amounts (e.g., “large amount”).
• Subjective Mental States: Legal descriptions of intent and remorse.
Output Example: {
“Defendant_Attribute”: [“Adult”, “Prior Criminal Record”],
“Criminal_Behaviors”: [“Theft”, “Burglary”],
“Victim_Characteristics”: [“Private Residence”, “Large Amount”],
“Subjective_Mental_States”: [“Direct Intent”, “Voluntary Surrender”]
}
Figure 13: Prompt for the Researcher agent to extract and classify legal keywords from case descriptions.
Figure 14: Prompt for the Researcher agent to pre-judge potential charges for charge-anchored retrieval.
charges (Jiang and Yang, 2023; Deng et al., 2024b; we make full use of external knowledge and prece-
Wang et al., 2024). Furthermore, multi-agent simu- dents within a unified framework.
lation frameworks have been explored to improve
performance by simulating court debates and ana- G.3 Reasoning skills in legal domain
lyzing cases from diverse perspectives (He et al., Recent work has improved LLMs’ reasoning
2024b). However, existing LLM-based methods through better prompting techniques (Sahoo et al.,
still struggle to utilize comprehensive legal knowl- 2024). Chain-of-thought (CoT) (Wei et al., 2022)
edge (Fei et al., 2024) effectively. In this context, prompting can explicitly guide LLMs to reason
37481
Auditor Checklist(item)
Task Definitions: Act as a legal AI assistant to assess if the case facts
strictly satisfy a specific constituent element of the law.
• Analyze the law_item and case facts.
• Focus exclusively on the target element (e.g., “intent”), using related
materials (if provided) for interpretation.
• determine applicability based on facts and logic.
I/O Specifications:
• Input: law_item, related (supplementary materials), element, case.
• Output: Provide reasoning first, then enclose the final result strictly within
tags: <answer>true</answer> or <answer>false</answer>.
Template:
law: {law_item}, related: {related}
element: {element}, case: {case}
Figure 15: Prompt for the Auditor agent to verify if case facts satisfy specific constituent elements of the law.
Auditor Checklist(final)
Task Definitions: Act as a legal analysis assistant to determine if the provided
law article applies to the specific case (i.e., verify violation or crime).
• Identify all relevant constituent elements from the law text.
• Verify critical elements independently; note that the provided true_list and
false_list may be incomplete.
I/O Specifications:
• Input Variables: case, law, true_list (proven elements), false_list (disproven
elements).
• Output: Provide reasoning first, then enclose the final result strictly within
tags: <answer>true</answer> or <answer>false</answer>.
Template:
case: {case}, law: {law}
true_list: {true_list}, false_list: {false_list}
Figure 16: Prompt for the Auditor agent to assess the overall applicability of a law based on verified elements.
step by step. In the legal domain, researchers have CaseGPT (Yang, 2024) combines LLMs with RAG
adapted CoT to legal-specific frameworks. For in- to support semi-structured reasoning and legal ar-
stance, Yu et al. (Yu et al., 2022) demonstrated that gumentation (Westermann, 2024). Additionally,
incorporating the IRAC (Issue, Rule, Application, GLARE (Yang et al., 2025b) leverages an agen-
Conclusion) framework significantly enhances rea- tic framework and web data for legal reasoning.
soning capabilities. LoT (Jiang and Yang, 2023) However, these approaches primarily rely on in-
proposed legal syllogism reasoning to improve trinsic capabilities or noisy external data, which
performance on LJP tasks, and ADAPT (Deng constraints reasoning depth (Zhang, 2024; Ke et al.,
et al., 2024b) established a workflow enabling dis- 2025). Therefore, we propose an agentic frame-
criminative reasoning. Moreover, approaches like work to dynamically acquire key legal knowledge,
MALR (Yuan et al., 2024) utilize parameter-free enhancing both breadth and depth.
learning to decompose complex legal tasks, while
37482
Charge & Sentencing (JSON)
Task Definitions: Act as a legal expert to adjudge the defendant based on
candidate charges.
• 1. Final Charge Application: For concurrence, apply the “heavier penalty”
rule; for multiple acts, apply combined punishment.
• 2. Sentencing: Predict the specific law article and a reasonable sentencing
range based on facts and judicial practice.
Format Definitions:
• {
charge_name: [Charge A, ...],
law_article: [Art. X, ...],
term_of_imprisonment: {
death_penalty: boolean,
imprisonment: integer (months),
life_imprisonment: boolean
}
}
Figure 17: Prompt for the Adjudicator agent to generate structured sentencing predictions and apply legal rules.
Figure 18: Prompt for the Adjudicator agent to synthesize legal reasoning and output the final verdict.
37483
G.4 RAG in legal domain
In the legal domain, recent studies have adapted
RAG and graph-based solutions for specific tasks,
particularly Legal Question Answering (LQA) and
retrieval pipelines. For instance, recent works
optimize LLM outputs by incorporating external
case-based information (Wiratunga et al., 2024;
Louis et al., 2024) or utilizing adapt-retrieve-revise
pipelines (Wan et al., 2024) to combine contin-
ual training with evidence revision. Dedicated
benchmarks such as LegalBench-RAG show that
the retrieval stage remains a bottleneck (Pipitone
and Alami, 2024). Works enriching retrieval with
structural information (e.g., graphs of articles)
demonstrate gains in tasks like statutory article re-
trieval (Louis et al., 2023; Hei et al., 2024; Ho
et al., 2025). Recently, the SAT-Graph RAG frame-
work (de Martim, 2025) was proposed to model
the hierarchical structure of legal norms. However,
its sophisticated ontology-driven approach requires
heavily structured input data, limiting its applica-
bility to less curated corpora.
37484