Legal judgment generation is a critical task in legal intelligence. However, existing research in legal judgment generation has predominantly focused on first-instance trials, relying on static fact-to-verdict mappings while neglecting the dialectical nature of appellate (second-instance) review. To address this, we introduce AppellateGen, a benchmark for second-instance legal judgment generation comprising 7,351 case pairs. The task requires models to draft legally binding judgments by reasoning over the initial verdict and evidentiary updates, thereby modeling the causal dependency between trial stages. We further propose a judicial Standard Operating Procedure (SOP)-based Legal Multi-Agent System (SLMAS) to simulate judicial workflows, which decomposes the generation process into discrete stages of issue identification, retrieval, and drafting. Experimental results indicate that while SLMAS improves logical consistency, the complexity of appellate reasoning remains a substantial challenge for current LLMs. The dataset and code are publicly available at: https://anonymous.4open.science/r/AppellateGen-5763.
@article{arxiv.2601.01331,
title = {AppellateGen: A Benchmark for Appellate Legal Judgment Generation},
author = {Hongkun Yang and Lionel Z. Wang and Wei Fan and Yiran Hu and Lixu Wang and Chenyu Liu and Yu Zeng and Shenghong Fu and Lei Gong and Zhengxin Zhang and Haoyang Li and Jiexin Zheng and Xin Xu},
journal= {arXiv preprint arXiv:2601.01331},
year = {2026}
}