Transforming Liver Transplant Care with Artificial Intelligence: A Narrative Review
Kathryn L White 1
, Carol Crochet 1
, Maheswaran Pitchaimuthu 2,*
![]()
-
Department of Surgery, The University of Oklahoma College of Medicine, Oklahoma, OK, USA
-
Division of Transplant Surgery, Department of Surgery, The University of Oklahoma College of Medicine, Oklahoma, OK, USA
* Correspondence: Maheswaran Pitchaimuthu
![]()
Academic Editor: Massimo Pinzani
Received: April 09, 2026 | Accepted: September 01, 2026 | Published: September 11, 2026
OBM Transplantation 2026, Volume 10, Issue 3, doi:10.21926/obm.transplant.2603275
Recommended citation: White KL, Crochet C, Pitchaimuthu M. Transforming Liver Transplant Care with Artificial Intelligence: A Narrative Review. OBM Transplantation 2026; 10(3): 275; doi:10.21926/obm.transplant.2603275.
© 2026 by the authors. This is an open access article distributed under the conditions of the Creative Commons by Attribution License, which permits unrestricted use, distribution, and reproduction in any medium or format, provided the original work is correctly cited.
Abstract
Outcomes in liver transplantation remain constrained by challenges in donor selection, organ allocation, and long-term graft survival, and traditional scoring systems such as MELD, SOFT, and the Donor Risk Index have well-recognized limitations in capturing the nonlinear complexity of transplant data. Artificial intelligence (AI), including machine learning (ML) and deep learning (DL), offers analytical methods capable of processing high-dimensional data beyond conventional regression, but the maturity and clinical readiness of this evidence base across the liver transplant continuum have not been comprehensively characterized. A narrative review was conducted via systematic search of PubMed, MEDLINE, and the Cochrane Library (January 1991-June 2025) using combinations of terms including “artificial intelligence,” “machine learning,” “deep learning,” “liver transplant,” “organ allocation,” and “graft failure.” Of 754 records identified, 45 studies met inclusion criteria following independent dual-reviewer screening. Each study was classified by evidence level as exploratory (internal validation only), externally validated, or implementation-oriented, and organized according to the pre-transplant and post-transplant continuum. AI applications span transplant candidacy assessment, organ allocation and waitlist prioritization, donor-recipient matching, graft quality assessment, and post-transplant prediction of acute kidney injury, sepsis, cardiovascular events, graft dysfunction, biliary complications, hepatocellular carcinoma (HCC) recurrence, immunosuppression dosing, and mortality. Performance varied widely (AUC/C-index 0.63-0.99), with the most methodologically mature evidence found for HCC recurrence prediction (C-index 0.75-0.839, including one internationally validated model) and waitlist mortality modeling. Notably, logistic regression outperformed several ML algorithms in the largest donor-recipient matching registry analyzed (n > 39,000), and cross-national validation of mortality models revealed substantial performance degradation (AUROC 0.70-0.74) when applied across healthcare systems. Fewer than half of included studies used explainability methods such as SHAP, and only two models explicitly incorporated equity considerations into their design. AI is increasingly applied across the liver transplant continuum and shows particular promise for HCC recurrence prediction and waitlist risk stratification. However, the field is limited by a predominance of single-center exploratory studies with internal validation only, inadequate attention to algorithmic fairness and bias, and an absence of prospective evaluation or regulatory engagement-no AI model in liver transplantation has yet received regulatory approval for clinical use. Future progress will require standardized data frameworks, multicenter prospective validation, formal reporting guidelines, and equity auditing before AI tools can be considered genuinely practice-changing.
Keywords
Artificial intelligence; liver transplantation; machine learning; donor-recipient matching; post-transplant outcomes
1. Introduction
Outcomes in liver transplantation remain constrained by challenges in donor selection, organ allocation, perioperative risk stratification, and long-term graft survival. Increasing recipient acuity and reliance on marginal donors have amplified the need for more precise, data-driven decision-making throughout the transplant process [1,2]. Traditional scoring systems-including the Model for End-Stage Liver Disease (MELD), Survival Following Liver Transplantation (SOFT), and the Donor Risk Index (DRI)-have provided foundational prognostic frameworks, yet each carries well-recognized limitations in capturing the nonlinear complexity of transplant outcomes. Approximately 10% of patients with low MELD-Na scores still experience high waitlist mortality, illustrating the ceiling of these tools.
Artificial intelligence (AI), including machine learning (ML) and deep learning (DL), offers analytical methods capable of processing high-dimensional, nonlinear clinical data beyond the reach of conventional regression [3]. Machine learning encompasses algorithms that learn patterns from data without being explicitly programmed-including random forests, gradient boosting, support vector machines, and logistic regression variants. Deep learning is a subset of ML that uses multilayer neural networks and is particularly suited to image and sequence data. These are not interchangeable terms: the choice of method has direct implications for interpretability, data requirements, and validation strategy. In liver transplantation, AI has been applied to donor and recipient risk prediction, waitlist mortality, allocation modeling, imaging and histopathologic assessment, perioperative outcome prediction, and graft survival forecasting [4]. While early studies demonstrate promising performance, clinical adoption remains limited by concerns regarding interpretability, bias, generalizability, and integration into existing workflows [4,5].
A key question remains: what additional value does AI provide beyond established clinical assessment, pathology review, imaging interpretation, and multidisciplinary decision-making? While early studies are promising, the evidence is still evolving, and further validation is needed before AI can be routinely integrated into clinical practice.AI may offer advantages in settings where: (i) the relevant signal is genuinely nonlinear or involves high-dimensional interactions undetectable by conventional methods; (ii) image analysis requires scale, speed, or consistency beyond human capacity; or (iii) longitudinal time-series data can be monitored continuously in ways impossible in clinical practice. This review critically evaluates where the evidence currently stands for each of these claims across the transplant continuum.
This narrative review provides a comprehensive overview of current applications of artificial intelligence (AI) in liver transplantation, spanning donor selection, graft assessment, recipient evaluation, immunosuppression optimization, and equity-focused organ allocation. It critically appraises the existing evidence, highlighting the quality and strength of published studies while identifying key limitations, including study heterogeneity, risk of bias, limited external validation, and challenges to clinical implementation. Finally, the review discusses the essential requirements for the responsible integration of AI into routine liver transplant practice and outlines future directions for research and clinical adoption.
2. Methods
This paper is presented as a narrative review, reflecting the breadth and heterogeneity of the literature, the absence of standardized outcome definitions across AI studies in liver transplantation, and the impossibility of meaningful quantitative synthesis across studies using incompatible models, endpoints, and validation strategies. A formal meta-analysis was not performed because the included studies differ substantially in AI methodology, patient populations, outcome definitions, and performance metrics; pooling AUC values across such studies would produce misleading estimates.
A systematic literature search was conducted in PubMed, MEDLINE, and the Cochrane Library to identify studies evaluating AI applications in liver transplantation. The search incorporated combinations of the following terms: “artificial intelligence,” “machine learning,” “deep learning,” “neural network,” “explainable AI,” “liver transplant,” “liver transplantation,” “organ allocation,” “waitlist mortality,” “graft failure,” “donor-recipient matching,” and “hepatocellular carcinoma recurrence.” The search covered publications from January 1991 through June 2025.
A total of 754 records were identified. After removal of duplicates, titles and abstracts were independently screened by two reviewers. Studies were excluded if they were non-English publications, review articles, case reports, conference abstracts, or if they lacked sufficient methodological detail to permit appraisal of model development and validation strategy. Full-text articles were subsequently assessed for eligibility, with discrepancies resolved by discussion and consensus. Following this process, 45 studies were included in the final qualitative synthesis (Figure 1).
Figure 1 Flow diagram of study selection process.
For each included study, data were extracted on study design, patient population, data source, AI model type, outcome definition, performance metrics (AUC, AUROC, C-index, MDAPE), validation approach (internal only, external, or cross-national), and interpretability method (e.g., SHAP, inherently interpretable model, or uninterpreted). Studies were classified by evidence level as: (i) exploratory model development with internal validation only; (ii) externally validated models; or (iii) implementation-oriented or prospectively evaluated models. Results were organized according to the liver transplant continuum, from candidacy and allocation through perioperative complications to long-term outcomes. No formal quality appraisal tool was applied, consistent with narrative review methodology; however, key methodological strengths and limitations are explicitly described for each study and summarized in Table 1.
Table 1 Summary of Artificial Intelligence Applications in Liver Transplantation: Evidence Level, Key Findings, and Methodological Considerations.

3. Pre-Transplant Applications
The persistent imbalance between demand for liver transplantation and the limited availability of donor organs has prompted increasing interest in AI to improve all parts of the transplant process. Among these, (AI) techniques are being actively explored to enhance organ allocation systems and refine recipient risk assessment prior to transplantation [5]. Traditional statistical approaches often struggle to incorporate the high-dimensional and nonlinear relationships inherent in transplant datasets. In contrast, machine learning algorithms can process large volumes of complex data containing numerous donors, recipient, and procedural variables, allowing them to identify intricate interactions that may not be readily apparent using conventional methods. As a result, these tools hold significant promise in supporting clinical decision-making across the liver transplantation continuum, particularly during the pre-transplant phase.
3.1 Transplant Candidacy
Determining candidacy for liver transplantation requires comprehensive multidisciplinary evaluation, integrating psychosocial stability, adherence potential, and medical readiness. AI models are positioned to improve objectivity, but their use in candidacy decisions carries significant equity implications, since errors that exclude patients from transplantation may be irreversible [5]. AI tools should complement, not replace, expert clinical judgment.
Determining candidacy for liver transplantation requires a comprehensive and multidisciplinary evaluation aimed at assessing whether a patient is likely to derive meaningful benefit from transplantation. This process extends beyond medical urgency to include psychosocial stability, adherence potential, and economic readiness. Increasingly, artificial intelligence has emerged as a novel adjunct to traditional evaluation frameworks, with the potential to improve objectivity and consistency in transplant candidacy decisions [1,2]. AI-driven models are uniquely positioned to integrate and analyze complex, multidimensional datasets that often exceed the capacity of standard risk scores or clinician judgment alone [1].
Despite its promise, the use of AI in transplant candidacy assessment remains an evolving area of research. Successful clinical implementation requires transparency in model design, rigorous validation across diverse populations, and careful oversight to avoid perpetuating or amplifying existing healthcare disparities and biases [5]. Accordingly, AI tools should complement, rather than replace, expert clinical judgment during candidate evaluation.
One area in which AI has gained particular attention is psychosocial risk assessment prior to liver transplantation. Psychosocial factors, including substance use history, social support, and psychiatric comorbidities, are known to influence post-transplant outcomes but are often difficult to quantify consistently. Lee et al. applied XGBoost to predict post-transplant alcohol relapse in 116 recipients across 10 US centers for early liver transplantation in alcohol-associated hepatitis, after 6 months abstinence period has been removed from mandatory requirement [6]. The model achieved an internal AUC of 0.93 (PPV 89.1%), which declined to 0.69 on external validation (PPV 82%). This performance drop at external validation-a 25% relative decrease in AUC-is a critical finding that the internal performance alone would obscure. It illustrates that psychosocial predictors are highly context-dependent, limiting generalizability across centers with different evaluation protocols. The authors appropriately recommended this tool for guiding targeted relapse prevention rather than excluding patients from transplantation.
AI has also been investigated as a means of refining pre-transplant cardiac risk assessment. Traditional cardiac screening for liver transplant candidates is resource-intensive and frequently involves invasive testing, even though many patients ultimately do not experience post-transplant cardiac complications. Zaver et al. evaluated an AI-enhanced electrocardiogram (AI-ECG) using a convolutional neural network to detect left ventricular dysfunction (LVEF < 50%) and predict post-transplant atrial fibrillation in 712 liver transplant candidates [7]. AUROC values ranged from 0.63 to 0.70, representing modest discriminative performance. This is an exploratory study with internal validation only; its clinical value lies in demonstrating proof-of-concept feasibility of AI-ECG as an adjunct screening tool, not in replacing echocardiography.
Soldera et al. also predicted mortality by oesophageal bleeding with the use of ANN [32]. Although not a transplant study, this literature illustrates a recurring and important theme: ML models can achieve high internal performance in complex hepatology datasets, yet the clinical implications of that performance depend critically on calibration, external validity, and endpoint definition.
3.2 Organ Allocation
Organ allocation in liver transplantation is a complex process that must balance medical urgency, wait time, donor-recipient compatibility, blood type, organ size, and geographic considerations. The overarching goal is to prioritize the sickest patients while minimizing organ wastage and maximizing survival benefit [2]. Given the life-or-death consequences of allocation decisions, standardization and fairness are paramount. Artificial intelligence has demonstrated potential utility in this domain by improving risk stratification and allocation precision beyond traditional scoring systems.
3.3 Donor-Recipient Matching
Multiple scoring systems have been developed to support this process, including the Model for End-Stage Liver Disease (MELD), Survival Following Liver Transplantation (SOFT), Balance of Risk (BAR), and donor-based scores such as D-MELD and the Donor Risk Index (DRI)-none of which has proven universally optimal [1,2,3]. Balanza et al. evaluated artificial neural networks (ANNs) for D-R matching across 1,003 liver transplants from 11 Spanish centers with 64 donor and recipient variables, achieving 90.79% accuracy for 3-month graft survival and 71.42% for graft loss, outperforming regression approaches [33]. These findings suggest that ANNs may serve as powerful decision support tools for donor-recipient matching within appropriately curated datasets, however, this is a single-registry analysis, and the performance advantage of ANNs over logistic regression in a sufficiently large, well-curated European dataset has not been independently replicated.
A critical counter-example is the large-scale analysis by Guijo-Rubio et al. using over 39,000 UNOS transplants, in which logistic regression (LR) consistently outperformed advanced ML models including random forests, gradient boosting, support vector machines, and neural networks [3]. This finding is not a failure of ML-it is a meaningful insight: in large, well-structured registries with relatively low signal-to-noise ratios, the regularization and calibration properties of well-implemented logistic regression may match or exceed more complex algorithms. The real lesson from this tension is that AI performance depends fundamentally on data quality, endpoint definition, cohort size, and whether nonlinear complexity genuinely exists in the signal. The assumption that more complex models are inherently better is not supported by the liver transplant evidence base.
3.4 Risk Prediction and Outcome Matching
Early applications of machine learning (ML) in liver transplantation focused primarily on improving pre-transplant risk stratification. Molinari et al. developed an ML-based model to predict perioperative mortality using recipient characteristics available during transplant evaluation. Using a national United Network for Organ Sharing (UNOS) cohort of more than 30,000 first-time deceased donor liver transplant recipients, the investigators applied classification tree analysis and artificial neural networks to identify five key predictors of 90-day mortality: recipient age, Model for End-Stage Liver Disease (MELD) score, body mass index, diabetes, and pre-transplant dialysis [34]. These variables were incorporated into a simplified point-based risk score that demonstrated excellent discrimination for identifying high-risk recipients, with an area under the receiver operating characteristic curve (AUC) exceeding 0.90. The model also predicted one-year mortality and five-year survival, highlighting the ability of ML to translate complex nonlinear relationships into clinically interpretable tools for candidate selection and counseling. Although limited by the lack of external validation and omission of frailty-related variables, this study established an important proof of concept for AI-assisted pre-transplant risk assessment.
Subsequent studies have consistently shown that ML models outperform conventional prognostic scores for predicting transplant outcomes. In a recent systematic review of 23 observational studies, Chongo and Soldera reported that ML approaches-including random forests, gradient boosting, artificial neural networks, and deep learning-demonstrated superior predictive performance compared with established scores across multiple clinical outcomes [35]. However, the review also identified major limitations, including retrospective single-center study designs, heterogeneous model development, inconsistent validation, and inadequate reporting standards. Future research should therefore prioritize prospective multicenter validation, external calibration, transparent reporting, and demonstration that AI-assisted decision-making improves patient outcomes beyond existing clinical risk assessment tools [35].
3.5 Advancing Beyond MELD: Waitlist Prioritization
Liver transplant allocation from deceased donors is largely guided by the MELD score, but has inherent limitations including its inability to account for complications of portal hypertension and other features that contribute to mortality risk. Nagai et al. developed a neural network using over 105,000 UNOS/OPTN waitlisted candidates and 28 variables, achieving an AUC-ROC of 0.936 for waitlist mortality prediction, outperforming MELD-Na [8]. However, the exclusion of patients transplanted within 90 days may have biased the cohort toward lower-risk individuals-a selection effect that could inflate predictive performance. Bertsimas et al. proposed OPOM (Optimized Prediction of Mortality) using optimal classification trees, demonstrating a simulated 17.6% reduction in waitlist mortality and improved equity for women and non-HCC candidates [9]. Importantly, OPOM was validated using the Liver Simulated Allocation Model (LSAM), not prospective implementation, and the gap between simulation and real-world performance remains an open question.
Kwong et al. developed an ML model to predict HCC waitlist dropout using OPTN data from nearly 19,000 patients, achieving a C-statistic of 0.74 by incorporating AFP, tumor size, bilirubin, INR, and ascites [10]. Gómez-Orellana et al. introduced GEMA-AI, an explainable AI model using data from over 9,300 candidates that specifically addressed gender disparities in liver allocation, outperforming MELD-based models for women and critically ill patients [11]. The explicit focus on equity and interpretability in GEMA-AI represents the kind of implementation-oriented design that is rarely seen in this literature.
3.6 Graft Quality Assessment
AI has been applied to standardize the subjective assessment of donor liver quality. The 2022 Banff consensus recommendations guide steatosis assessment, and Gambella et al. developed a Banff-aligned deep learning algorithm demonstrating excellent agreement with expert pathologists, outperforming pre-Banff approaches [12]. Narayan et al.’s CVAI model (U-Net architecture) objectively assessed steatosis in 90 biopsies and better predicted early allograft dysfunction than manual pathologist review [13]. Cesaretti et al. applied a semi-supervised SVMSIL model to smartphone images of 117 grafts, achieving 89% accuracy for steatosis ≥30% [14].
These studies are collectively exploratory-they involve small datasets (90-292 samples), lack external validation, and have not been integrated into clinical workflows. The promise of AI-enabled real-time intraoperative graft assessment is real, but these studies represent hypothesis-generating proof-of-concept work rather than practice-changing evidence. As demonstrated by the broader deep learning literature in HCC imaging, technically impressive internal performance can mask serious limitations in annotation quality, cross-center generalizability, and workflow integration [36,37]. Finally, Oh et al. demonstrated high-accuracy automated 3D liver segmentation for living donor CT angiography with improved correlation with actual graft weights, highlighting AI’s growing role in preoperative planning and donor safety [38].
4. Post-Transplant Applications
AI has emerged as a critical tool in the post-transplant period to predict morbidity and mortality. It enables early identification of complications by analyzing complex, high-dimensional clinical data that evolve over time, which helps clinicians to intervene earlier, tailor postoperative management and potentially improve short- and long-term outcomes. The evidence below is organized by complication type and classified by validation robustness.
4.1 Surgical Complications: AKI, Pneumonia, and Sepsis
Post-transplant AKI affects up to 50% of recipients and is strongly associated with mortality and development of CKD. Zhang et al. developed a gradient boosting machine (GBM) model using data from 780 transplants (2015-2019) with external validation on 195 cases (2019-2021), achieving an AUC of 0.76 internally and 0.75 externally-one of the few studies in this review with temporally separated external validation [15]. Shapley Additive Explanations (SHAP) analysis identified high preoperative indirect bilirubin, low intraoperative urine output, prolonged anesthesia time, low platelet count, and graft steatosis as key predictors. The consistency across validation cohorts supports real clinical utility, though the modest AUC (0.75) means approximately one in four high-risk patients would be missed.
Chen et al. (2021) developed XGBoost models predicting postoperative pneumonia in 591 recipients (incidence 42.8%), achieving an AUC of 0.794 using routinely available perioperative variables, including included INR, hematocrit, platelet count, albumin, ALT, fibrinogen, white blood cell count, prothrombin time, serum sodium, total bilirubin, anesthesia duration, preoperative hospital stay, transfusion volume, and operative time [16]. This is an internal validation study without external replication.
Kamaleswaran et al. applied continuous physiological monitoring to predict early sepsis across 5,748 ICU admissions, including 92 post-transplant recipients, achieving an impressive AUC of 0.97 [17]. However, the transplant subset was very small, and performance in a dedicated transplant cohort may differ substantially. Chen et al. (2023) specifically targeted post-transplant sepsis within seven days in 677 recipients (incidence 31.9%), with random forest achieving an AUC of 0.731 internally and 0.755 on external validation [39]-a modest but reproducible signal with genuine clinical relevance.
4.2 Cardiovascular Complications
As the prevalence of MASH-related cirrhosis increases as an indication for transplantation, recipients carry progressively greater cardiovascular comorbidities. Abdelhameed et al. evaluated a BiGRU deep learning model using pre-transplant claims data from over 18,000 patients to predict MACE, achieving an AUC-ROC of 0.841 for 30-day events [4]. Claims-based models capture billing codes rather than granular clinical measurements, and the generalizability of this approach to clinical decision-making is uncertain.
Jain et al. applied XGBoost across 1,459 liver transplant recipients to predict MACE, all-cause mortality, and cardiovascular mortality, with MACE AUC of 0.71-only marginally better than logistic regression in the same dataset [18]. Importantly, SHAP analysis identified age, diabetes, serum creatinine, NASH-cirrhosis, right ventricular systolic pressure, and LVEF as the dominant predictors, providing clinically interpretable insights consistent with established cardiovascular risk factors in this population.
Soldera et al. (2024), who applied XGBoost with SHAP to predict post-LT MACE in 575 transplant recipients from a Southern Brazilian academic center, achieving an AUROC of 0.89 with excellent calibration as assessed by Brier score [19]. This study is notable for following Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD) reporting standards-a level of methodological rigor rarely observed in the transplant AI literature-and for providing an online prediction calculator, representing a step toward clinical implementation. The authors identified noninvasive cardiac stress test outcomes, nonselective beta-blocker use, direct bilirubin, blood type O, and dynamic changes on myocardial perfusion scintigraphy as key predictors, demonstrating the value of integrating hepatic and cardiovascular variables into a unified pre-transplant risk model. However, despite strong internal performance, the model remains limited by its retrospective single-center design, low event rate, and absence of external validation.
Future studies should prioritize prospective multicenter validation, continual model recalibration, interoperability with electronic health record systems, and assessment of clinical impact before AI models can be routinely incorporated into transplant decision-making.
4.3 Graft Dysfunction
Early allograft dysfunction (EAD) affects 20-40% of liver transplant recipients [20]. Meng et al. developed a deep learning ultrasound and flow spectrogram fusion network for EAD diagnosis using 117 transplant patients, achieving an AUC of 0.968 in relation to abnormal hepatic blood flow [20]. While technically impressive, this is a proof-of-concept study in a small, single-center cohort and requires substantially larger multicenter validation before clinical deployment.
Giglio et al. developed an ANN predicting early graft failure within three months after adult-to-adult living donor LT, trained on 2,073 transplants and externally validated with UNOS data (AUC 0.68-0.70) [21]. This is one of the few studies in this review with independent external validation using a geographically and institutionally distinct registry. The modest but consistent AUC across validation datasets suggests genuine predictive signal rather than overfitting. Lau et al. and Cooper et al. demonstrated that ML models can predict early graft failure (AUC up to 0.82) and acute graft-versus-host disease (AUROC 0.93-0.96), respectively, with Cooper et al. validating in a second institutional cohort [22,23].
4.4 Biliary Complications
Biliary complications affect up to 20% of liver transplant recipients. Hu et al. developed support vector machine models to predict biliary complications at 3, 6, and 12 months in 517 patients, achieving AUCs exceeding 0.88 across all time points [40]. Fodor et al. combined hyperspectral imaging with deep learning during normothermic machine perfusion, improving predictive accuracy to 94% for biliary complications-a technically novel approach that remains at proof-of-concept stage and requires specialized intraoperative equipment not yet widely available [41]. Andishgar et al. applied time-to-event random survival forest models, achieving C-indices of 0.699-0.784 for biliary complications and mortality, representing one of the more methodologically rigorous survival modeling approaches in this literature [31].
4.5 MASH Diagnosis and Fibrosis Prediction
MASH is a leading and growing indication for liver transplantation, with high rates of recurrent and de novo disease post-transplant. Hajek et al. developed a noninvasive ML diagnostic model using proton MR spectroscopy, achieving AUC > 0.96 for MASH identification in transplant recipients, suggesting MR spectroscopy may reduce biopsy dependence [42]. Rabindranath et al. demonstrated that multimodal models integrating clinical, laboratory, and ultrasound data achieved AUCs of 0.77-0.81 for fibrosis prediction, while deep learning using ultrasound images alone performed poorly-an important finding underscoring that model architecture must match data modality, and that multimodal integration is typically necessary [43].
4.6 HCC Recurrence
Approximately 20% of liver transplant recipients for HCC develop recurrence, which carries poor prognosis under immunosuppression [24]. Traditional criteria such as Milan may insufficiently identify high-risk patients, and integrating imaging, biomarkers, and molecular profiles using ML may improve selection. Lai et al. developed TRAIN-AI, a deep learning model using eight tumor and patient factors in a large international multicenter cohort, achieving a concordance of 0.77 and outperforming existing criteria [24]-this is one of the few studies with genuine international multicenter validation. Cao et al.’s DeepSurv models provided highly accurate pre- and post-transplant recurrence predictions with external validation C-indices up to 0.839 [25]. Ivanics et al. used CoxNet on 739 recipients, outperforming AFP and MORAL scores with a concordance of 0.75 [1]. Collectively, these studies represent the most mature AI evidence in liver transplantation, reflecting consistent multicenter validation and clinically meaningful performance improvements over established benchmarks.
4.7 Immunosuppression Optimization
Yoon et al. developed an LSTM neural network to predict tacrolimus blood levels in 443 South Korean and 106 US liver transplant recipients by using 6264 samples [26]. The LSTM model outperformed other approaches (MDAPE 22.3%, RMSE 1.7 ng/mL) and was associated with improved achievement of therapeutic levels and reduced ICU stay (5.5 vs 8.0 days; P = 0.042). The cross-national validation component and the clinically meaningful outcome (ICU stay reduction) distinguish this from purely methodological studies. These results demonstrate that AI can enhance immunosuppressive management and transplant outcomes.
4.8 Mortality and Survival Prediction
Post-transplant survival prediction has attracted the largest number of AI studies, yet it also illustrates the field’s core tension most clearly. Deep learning approaches including BiGRU (Abdelhameed et al., AUC 0.84) and multilayer perceptrons (Raji et al., >99% accuracy-a figure that warrants caution without extensive external validation) have demonstrated improved personalized risk prediction [4,28]. DNNs (Ershoff et al., AUC ~0.70) achieved modest improvement over the SOFT and BAR scores despite using 202 features, suggesting that additional model complexity does not guarantee improved discrimination [29]. Random forest models highlighted nonlinear donor-recipient interactions (Yu et al., AUC 0.80-0.85) [44]. Kazemi et al. demonstrated SVM-based models outperforming Cox regression in 902 recipients, with post-transplant complications-particularly graft failure and Aspergillus infection-as the dominant predictors [27].
ML has also been applied to identify modifiable risk factors in specific recipient subgroups. Yasodhara et al. used machine learning on liver transplant recipients with diabetes mellitus, achieving AUROC values of 0.60-0.70 for long-term survival prediction, while identifying modifiable predictors including immunosuppression regimen, post-transplant diabetes control, and renal function trajectories that may guide targeted post-transplant management [45].
Börner et al. developed a novel deep learning model specifically designed as a donor-recipient matching tool to predict survival after liver transplantation, achieving a remarkably high AUC of 0.94 in internal validation [30]. While this result is technically impressive, it derives from a single-center dataset and has not been externally validated; such high performance in small internal datasets warrants cautious interpretation pending replication in independent cohorts.
The most important study for calibrating expectations about cross-country generalizability is Ivanics et al., who evaluated ML models across three national registries (Canada, UK, US) and found that performance declined substantially from individualized (AUROC 0.70-0.74) to harmonized cross-country models, with external validity characterized as poor overall [1]. Chongo and Soldera et al. study underscores the potential of ML models in guiding decisions related to allograft allocation and LT, marking a significant evolution in the field of prognostication [19], This ML models perform satisfactorily in individual registries but show limited external validity, especially when applied across countries with different registry structures and patient populations. Time-to-event models, including RSF (Andishgar et al., C-index 0.699-0.784), enable dynamic mortality and biliary complication prediction incorporating early post-transplant variables [31].
Soldera et al. discovered that machine learning tools were more accurate than other widely used tools at predicting 30 and 365 days post liver transplant mortality. This endorses the idea that machine learning algorithms could improve decision making processes for better organ allocation and short-term mortality prediction especially in the perioperative window where early intervention can alter outcomes [46].
5. Ethical Considerations, Fairness, and Explainability
The ethical dimensions of AI in liver transplantation deserve dedicated analysis rather than a passing caution, given that allocation and candidacy decisions are precisely the contexts where algorithmic unfairness has the greatest consequences. A model that systematically underestimates waitlist mortality for women, minorities, or patients at non-academic centers does not merely produce statistical error-it perpetuates inequity in access to a life-saving intervention.
GEMA-AI (Gómez-Orellana et al.) was explicitly designed to address gender disparities in liver allocation, and OPOM (Bertsimas et al.) demonstrated improved equity for women and non-HCC candidates compared with MELD-Na. These represent the current state of the art for equity-aware AI in liver transplantation [9,11]. However, systematic evaluation of algorithmic fairness-including formal disparity metrics across race, sex, insurance status, and transplant center volume-is absent from the majority of studies included in this review.
Explainability is equally critical, as only a minority explicitly incorporated interpretability methods. SHAP (Shapley Additive Explanations) was used in Zhang et al. (AKI prediction) [15], Jain et al. (MACE) [18], and Soldera et al. (MACE) [46] to identify individual feature contributions, making model predictions more clinically interpretable. Inherently interpretable models, including logistic regression and optimal classification trees (OPOM), were used by Guijo-Rubio et al. [3]. and Bertsimas et al. [9]. The majority of remaining studies-particularly those using deep neural networks-function as black-box models where individual predictions cannot be readily attributed to specific clinical variables. For transplant allocation decisions, where clinicians and patients have a legitimate interest in understanding why a decision was made, black-box tools require additional interpretability safeguards before deployment.
Future AI systems in liver transplantation must incorporate prospective bias auditing, diverse training datasets, standardized fairness metrics, and regulatory oversight aligned with emerging frameworks such as the EU AI Act and FDA guidance on AI/ML-based software as a medical device.
6. Practical Implementation, Clinical Integration, and Infrastructure
The gap between model development and clinical deployment in liver transplantation is substantial and deserves explicit discussion. Most studies reviewed here report AUC values but do not address the conditions required for safe clinical use, including prospective evaluation, user interface design, integration with electronic health record systems, or real-world impact on transplant decisions. From a regulatory perspective, AI tools used to inform allocation or candidacy decisions would be classified as Software as a Medical Device (SaMD) by the FDA, requiring submission under the De Novo or 510(k) pathways depending on risk classification. No AI model in liver transplantation has yet received FDA approval for this purpose.
Infrastructure requirements for AI implementation in transplant centers include: (i) interoperable electronic health record systems with structured, standardized data fields; (ii) prospective data collection frameworks aligned with registry definitions (UNOS, OPTN, ELTR); (iii) dedicated data science support for model maintenance and performance monitoring; and (iv) institutional review and governance structures for AI oversight. Few transplant programs in the US or internationally currently meet all of these requirements, and implementation without robust governance risks deploying models that perform well in trials but fail silently in practice. Strauss et al. specifically examined clinician attitudes toward AI-based clinical decision support in liver transplantation evaluation and identified concerns about transparency, accountability, and the risk of perpetuating existing disparities [5]. These qualitative findings should guide implementation strategies alongside quantitative performance metrics.
7. Conclusion
Artificial intelligence, including machine learning and deep learning, is increasingly influencing liver transplantation by enabling more accurate, individualized risk prediction beyond traditional scoring systems. AI applications span the pretransplant and post-transplant continuum, with the most mature evidence for HCC recurrence prediction, waitlist mortality modeling, donor-recipient matching, perioperative risks including cardiac events, immunosuppression optimization and long-term outcome. Explainable AI-particularly SHAP-based models-has improved clinical interpretability in a subset of studies, but most published models remain opaque and unvalidated externally.
This review highlights three structural weaknesses that limit the clinical readiness of the field: (i) a predominance of exploratory single-center studies with internal validation only, in contrast to a small number of externally validated and implementation-oriented studies; (ii) inadequate attention to algorithmic fairness, bias, and equity in a domain where such errors have life-or-death consequences; and (iii) absence of prospective evaluation and regulatory pathway engagement for any AI tool in liver transplantation to date. Future progress will require standardized data frameworks, multicenter prospective validation, adoption of formal reporting guidelines, equity auditing, and engagement with regulatory pathways to ensure that AI tools are safe, equitable, and genuinely practice-changing when adopted.
Abbreviations

Author Contributions
Conception of the work: MP. Drafting the manuscript, Review and editing the manuscript, Critical review of the manuscript: KW, CC, MP. Final revision and approval for submission: MP.
Competing Interests
The authors have declared that no competing interests exist.
AI-Assisted Technologies Statement
During the preparation of this work the author(s) used Copilot AI for formatting. After using this tool/service, the author(s) reviewed and edited the content as needed and take(s) full responsibility for the content of the published article.
References
- Ivanics T, So D, Claasen MP, Wallace D, Patel MS, Gravely A, et al. Machine learning-based mortality prediction models using national liver transplantation registries are feasible but have limited utility across countries. Am J Transplant. 2023; 23: 64-71. [CrossRef] [Google scholar]
- Liu CL, Soong RS, Lee WC, Jiang GW, Lin YC. Predicting short-term survival after liver transplantation using machine learning. Sci Rep. 2020; 10: 5654. [CrossRef] [Google scholar]
- Guijo-Rubio D, Briceño J, Gutiérrez PA, Ayllón MD, Ciria R, Hervás-Martínez C. Statistical methods versus machine learning techniques for donor-recipient matching in liver transplantation. PLoS One. 2021; 16: e0252068. [CrossRef] [Google scholar]
- Abdelhameed A, Bhangu H, Feng J, Li F, Hu X, Patel P, et al. Deep learning-based prediction modeling of major adverse cardiovascular events after liver transplantation. Mayo Clin Proc Digit Health. 2024; 2: 221-230. [CrossRef] [Google scholar]
- Strauss AT, Sidoti CN, Sung HC, Jain VS, Lehmann H, Purnell TS, et al. Artificial intelligence-based clinical decision support for liver transplant evaluation and considerations about fairness: A qualitative study. Hepatol Commun. 2023; 7: e0239. [CrossRef] [Google scholar]
- Lee BP, Roth N, Rao P, Im GY, Vogel AS, Hasbun J, et al. Artificial intelligence to identify harmful alcohol use after early liver transplant for alcohol‐associated hepatitis. Am J Transplant. 2022; 22: 1834-1841. [CrossRef] [Google scholar]
- Zaver HB, Mzaik O, Thomas J, Roopkumar J, Adedinsewo D, Keaveny AP, et al. Utility of an artificial intelligence enabled electrocardiogram for risk assessment in liver transplant candidates. Dig Dis Sci. 2023; 68: 2379-2388. [CrossRef] [Google scholar]
- Nagai S, Nallabasannagari AR, Moonka D, Reddiboina M, Yeddula S, Kitajima T, et al. Use of neural network models to predict liver transplantation waitlist mortality. Liver Transpl. 2022; 28: 1133-1143. [CrossRef] [Google scholar]
- Bertsimas D, Kung J, Trichakis N, Wang Y, Hirose R, Vagefi PA. Development and validation of an optimized prediction of mortality for candidates awaiting liver transplantation. Am J Transplant. 2019; 19: 1109-1118. [CrossRef] [Google scholar]
- Kwong A, Hameed B, Syed S, Ho R, Mard H, Arshad S, et al. Machine learning to predict waitlist dropout among liver transplant candidates with hepatocellular carcinoma. Cancer Med. 2022; 11: 1535-1541. [CrossRef] [Google scholar]
- Gómez-Orellana AM, Rodríguez-Perálvarez ML, Guijo-Rubio D, Gutiérrez PA, Majumdar A, McCaughan GW, et al. Gender-equity model for liver allocation using artificial intelligence (GEMA-AI) for waiting list liver transplant prioritization. Clin Gastroenterol Hepatol. 2025; 23: 2187-2196. [CrossRef] [Google scholar]
- Gambella A, Salvi M, Molinaro L, Patrono D, Cassoni P, Papotti M, et al. Improved assessment of donor liver steatosis using Banff consensus recommendations and deep learning algorithms. J Hepatol. 2024; 80: 495-504. [CrossRef] [Google scholar]
- Narayan RR, Abadilla N, Yang L, Chen SB, Klinkachorn M, Eddington HS, et al. Artificial intelligence for prediction of donor liver allograft steatosis and early post-transplantation graft failure. HPB. 2022; 24: 764-771. [CrossRef] [Google scholar]
- Cesaretti M, Brustia R, Goumard C, Cauchy F, Poté N, Dondero F, et al. Use of artificial intelligence as an innovative method for liver graft macrosteatosis assessment. Liver Transpl. 2020; 26: 1224-1232. [CrossRef] [Google scholar]
- Zhang Y, Yang D, Liu Z, Chen C, Ge M, Li X, et al. An explainable supervised machine learning predictor of acute kidney injury after adult deceased donor liver transplantation. J Transl Med. 2021; 19: 321. [CrossRef] [Google scholar]
- Chen C, Yang D, Gao S, Zhang Y, Chen L, Wang B, et al. Development and performance assessment of novel machine learning models to predict pneumonia after liver transplantation. Respir Res. 2021; 22: 94. [CrossRef] [Google scholar]
- Kamaleswaran R, Sataphaty SK, Mas VR, Eason JD, Maluf DG. Artificial intelligence may predict early sepsis after liver transplantation. Front Physiol. 2021; 12: 692667. [CrossRef] [Google scholar]
- Jain V, Bansal A, Radakovich N, Sharma V, Khan MZ, Harris K, et al. Machine learning models to predict major adverse cardiovascular events after orthotopic liver transplantation: A cohort study. J Cardiothorac Vasc Anesth. 2021; 35: 2063-2069. [CrossRef] [Google scholar]
- Soldera J, Corso LL, Rech MM, Ballotin VR, Bigarella LG, Tomé F, et al. Predicting major adverse cardiovascular events after orthotopic liver transplantation using a supervised machine learning model: A cohort study. World J Hepatol. 2024; 16: 193-210. [CrossRef] [Google scholar]
- Meng Y, Wang M, Niu N, Zhang H, Yang J, Zhang G, et al. Artificial intelligence-assisted diagnosis of early allograft dysfunction based on ultrasound image and data. Visual Comput Ind Biomed Art. 2025; 8: 13. [CrossRef] [Google scholar]
- Giglio MC, Dolce P, Yilmaz S, Tokat Y, Acarli K, Kilic M, et al. Development of a model to predict the risk of early graft failure after adult-to-adult living donor liver transplantation: An ELTR study. Liver Transpl. 2024; 30: 835-847. [CrossRef] [Google scholar]
- Lau L, Kankanige Y, Rubinstein B, Jones R, Christophi C, Muralidharan V, et al. Machine-learning algorithms predict graft failure after liver transplantation. Transplantation. 2017; 101: e125-e132. [CrossRef] [Google scholar]
- Cooper JP, Perkins JD, Warner PR, Shingina A, Biggins SW, Abkowitz JL, et al. Acute graft‐versus‐host disease after orthotopic liver transplantation: Predicting this rare complication using machine learning. Liver Transpl. 2022; 28: 407-421. [CrossRef] [Google scholar]
- Lai Q, De Stefano C, Emond J, Bhangui P, Ikegami T, Schaefer B, et al. Development and validation of an artificial intelligence model for predicting post‐transplant hepatocellular cancer recurrence. Cancer Commun. 2023; 43: 1381-1385. [CrossRef] [Google scholar]
- Cao S, Yu S, Huang L, Seery S, Xia Y, Zhao Y, et al. Deep learning for hepatocellular carcinoma recurrence before and after liver transplantation: A multicenter cohort study. Sci Rep. 2025; 15: 7730. [CrossRef] [Google scholar]
- Yoon SB, Lee JM, Jung CW, Suh KS, Lee KW, Yi NJ, et al. Machine-learning model to predict the tacrolimus concentration and suggest optimal dose in liver transplantation recipients: A multicenter retrospective cohort study. Sci Rep. 2024; 14: 19996. [CrossRef] [Google scholar]
- Kazemi A, Kazemi K, Sami A, Sharifian R. Identifying factors that affect patient survival after orthotopic liver transplant using machine-learning techniques. Exp Clin Transplant. 2019; 17: 775-783. [CrossRef] [Google scholar]
- Raji CG, Chandra SV, Gracious N, Pillai YR, Sasidharan A. Advanced prognostic modeling with deep learning: Assessing long-term outcomes in liver transplant recipients from deceased and living donors. J Transl Med. 2025; 23: 188. [CrossRef] [Google scholar]
- Ershoff BD, Lee CK, Wray CL, Agopian VG, Urban G, Baldi P, et al. Training and validation of deep neural networks for the prediction of 90-day post-liver transplant mortality using UNOS registry data. Transplant Proc. 2020; 52: 246-258. [CrossRef] [Google scholar]
- Börner N, Schoenberg MB, Pöschke P, Heiliger C, Jacob S, Koch D, et al. A novel deep learning model as a donor-recipient matching tool to predict survival after liver transplantation. J Clin Med. 2022; 11: 6422. [CrossRef] [Google scholar]
- Andishgar A, Bazmi S, Lankarani KB, Taghavi SA, Imanieh MH, Sivandzadeh G, et al. Comparison of time-to-event machine learning models in predicting biliary complication and mortality rate in liver transplant patients. Sci Rep. 2025; 15: 4768. [CrossRef] [Google scholar]
- Soldera J, Tomé F, Corso LL, Rech MM, Ferrazza AD, Terres AZ, et al. Use of a machine learning algorithm to predict rebleeding and mortality for oesophageal variceal bleeding in cirrhotic patients. EMJ Gastroenterol. 2020; 9: 46-48. [Google scholar]
- Pontes Balanza B, Castillo Tuñón JM, Mateos García D, Padillo Ruiz J, Riquelme Santos JC, Álamo Martinez JM, et al. Development of a liver graft assessment expert machine-learning system: When the artificial intelligence helps liver transplant surgeons. Front Surg. 2023; 10: 1048451. [CrossRef] [Google scholar]
- Molinari M, Ayloo S, Tsung A, Jorgensen D, Tevar A, Rahman SH, et al. Prediction of perioperative mortality of cadaveric liver transplant recipients during their evaluations. Transplantation. 2019; 103: e297-e307. [CrossRef] [Google scholar]
- Chongo G, Soldera J. Use of machine learning models for the prognostication of liver transplantation: A systematic review. World J Transplant. 2024; 14: 88891. [CrossRef] [Google scholar]
- Ballotin VR, Bigarella LG, Soldera J, Soldera J. Deep learning applied to the imaging diagnosis of hepatocellular carcinoma. Artif Intell Gastrointest Endosc. 2021; 2: 127-135. [CrossRef] [Google scholar]
- Wang JC, Bai JJ, Wang YY, Wang C, Deng HP. Deep learning and machine learning in image-based hepatocellular carcinoma detection: A systemic review and meta-analysis. Abdom Radiol. 2026. doi: 10.1007/s00261-026-05627-6. [CrossRef] [Google scholar]
- Oh N, Kim JH, Rhu J, Jeong WK, Choi GS, Kim J, et al. Comprehensive deep learning-based assessment of living liver donor CT angiography: From vascular segmentation to volumetric analysis. Int J Surg. 2024; 110: 6551-6657. [CrossRef] [Google scholar]
- Chen C, Chen B, Yang J, Li X, Peng X, Feng Y, et al. Development and validation of a practical machine learning model to predict sepsis after liver transplantation. Ann Med. 2023; 55: 624-633. [CrossRef] [Google scholar]
- Hu F, Li Y, Zeng H, Ju R, Jiang D, Zhang L, et al. Machine learning model for predicting biliary complications after liver transplantation. Clin Transl Gastroenterol. 2025; 16: e00843. [CrossRef] [Google scholar]
- Fodor M, Lanser L, Hofmann J, Otarashvili G, Pühringer M, Cardini B, et al. Hyperspectral imaging as a tool for viability assessment during normothermic machine perfusion of human livers: A proof of concept pilot study. Transpl Int. 2022; 35: 10355. [CrossRef] [Google scholar]
- Hajek M, Sedivy P, Burian M, Mikova I, Trunecka P, Pajuelo D, et al. Liver fat fraction and machine learning improve steatohepatitis diagnosis in liver transplant patients. NMR Biomed. 2025; 38: e70077. [CrossRef] [Google scholar]
- Rabindranath M, Sun Y, Khalili K, Bhat M. Utilizing machine learning to predict liver allograft fibrosis by leveraging clinical and imaging data. Clin Transplant. 2025; 39: e70148. [CrossRef] [Google scholar]
- Yu YD, Lee KS, Kim JM, Ryu JH, Lee JG, Lee KW, et al. Artificial intelligence for predicting survival following deceased donor liver transplantation: Retrospective multi-center study. Int J Surg. 2022; 105: 106838. [CrossRef] [Google scholar]
- Yasodhara A, Dong V, Azhie A, Goldenberg A, Bhat M. Identifying modifiable predictors of long‐term survival in liver transplant recipients with diabetes mellitus using machine learning. Liver Transpl. 2021; 27: 536-547. [CrossRef] [Google scholar]
- Soldera J, Tomé F, Corso LL, Ballotin VR, Bigarella LG, Balbinot RS, et al. 590 Predicting 30 and 365-day mortality after liver transplantation using a machine learning algorithm. Gastroenterology. 2021; 160: S-789-S-790. [CrossRef] [Google scholar]



