Abstract
Anastomotic leak is common enough to shape postoperative surveillance and uncommon enough to make model transport difficult. Most prediction studies report a single pooled performance value, although the meaning of that value depends on where the model is deployed. A model trained in an elective colorectal pathway is not necessarily expected to behave similarly in emergency bowel surgery or transfer-heavy upper gastrointestinal care, even when the same structured variables are available. This study examined how much predictive performance, calibration, and decision thresholds change when leak models are moved across surgery-type and care-setting domains. We assembled a retrospective multicenter cohort of adult gastrointestinal anastomoses from seventeen hospitals and defined twenty-five clinical domains by crossing five procedure families with five episode-level care settings. Models were trained using structured variables available by 18 hours after skin closure. The primary endpoint was clinically managed anastomotic leak within 30 days. The final cohort included 54,910 procedures and 3,625 leaks, yielding an overall incidence of 6.6\%. Random-split validation overstated performance for all methods. A pooled gradient boosting model achieved an apparent area under the receiver operating characteristic curve of 0.812 but declined to 0.736 under leave-domain-out testing. The proposed domain-mixture model with invariant representation learning and calibration regularization achieved 0.781 under the same transport evaluation and reduced expected calibration error from 0.083 to 0.034. Transport failure was concentrated in low pelvic colorectal procedures performed in emergency overnight settings and in upper gastrointestinal reconstructions managed through transfer-in complex care. A global alert threshold produced wide variation in positive predictive value across domains, whereas domain-cluster calibration narrowed that spread substantially. These findings indicate that heterogeneity of leak risk across surgery types and care settings is also a heterogeneity of model validity. In this problem, transportability is itself a clinical outcome.