Large language models decode narrative pathology reports to define clinically relevant subtypes in IgA nephropathy
Immunoglobulin A nephropathy (IgAN) is the most common primary glomerular disease worldwide, with a highly heterogeneous clinical course. Current biopsy-based risk assessment relies largely on structured histological scores, such as the Oxford MEST-C score, which summarize selected lesions but may omit information contained in routine narrative pathology reports. Here we show that large language model-based analysis of routine biopsy reports identifies reproducible pathological subtypes with prognostic relevance in IgAN. We studied 3,078 adults with primary IgAN from two hospitals in Wenzhou, China, using the First Affiliated Hospital cohort for subtype discovery and model development and the Second Affiliated Hospital cohort for independent external validation. Unsupervised clustering identified two subtypes, low-injury and high-injury, that differed in clinical presentation, histological severity and kidney outcomes. The high-injury subtype was independently associated with a higher risk of the composite kidney endpoint after adjustment for baseline clinical factors and the complete Oxford MEST-C score set (hazard ratio, 2.57; 95% confidence interval, 1.25-5.29). Interpretability analyses indicated that subtype separation was driven mainly by chronic tubulointerstitial injury and inflammatory burden. Exploratory analyses suggested heterogeneity in corticosteroid-associated outcome patterns across subtypes. These findings indicate that routine narrative biopsy reports contain prognostic information that complements structured pathological scoring. Kidney pathology reports contain rich information that structured assessment does not fully capture. Here, the authors show that large language models extract and standardize it to define reproducible IgA nephropathy subtypes that improve risk prediction and reveal differential treatment responses.