Abstract
Accurate pancreatic tumor segmentation on contrast-enhanced computed tomography (CECT) is important for staging, treatment planning, and response assessment in pancreatic ductal adenocarcinoma (PDAC). Although foundation models have shown promise for medical image segmentation, their effectiveness for disease-specific tumor delineation remains uncertain. This study evaluated whether task-specific adaptation of a Segment Anything Model (SAM)-based framework improves pancreatic tumor segmentation compared with directly applied foundation models.
In this retrospective multicenter study, CECT examinations from patients with pathologically confirmed PDAC acquired between 2015 and 2025 were included. Two foundation-model baselines, nnInteractive and MedSAM2, were compared with three TAGS-based configurations representing increasing levels of adaptation: TAGS (Zero-Shot), SAM-TAGS (fine-tuned from SAM weights), and MSD-TAGS (fine-tuned from a pancreas-specific checkpoint). Performance was assessed using five-fold cross-validation with the Dice Similarity Coefficient (DSC) and Normalized Surface Dice (NSD).
Across 500 patients from four independent centers, MSD-TAGS achieved the highest overall performance (mean DSC 0.703, mean NSD 0.835), followed by SAM-TAGS (0.691 / 0.822). By comparison, nnInteractive reached 0.632 / 0.782, MedSAM2 0.585 / 0.792, and TAGS (Zero-Shot) 0.496 / 0.639. MSD-TAGS achieved significantly higher DSC and NSD than every competing method (all Holm-adjusted p<0.001), the highest DSC across all four centers, and the highest NSD in three of four. Under leave-one-center-out evaluation, MSD-TAGS was more robust than SAM-TAGS (DSC decline of 0.025 versus 0.089). These findings support the importance of organ-specific pretraining and target-cohort fine-tuning for disease-specific segmentation. The multicenter dataset, with de-identified images and expert annotations, is publicly released.
Key Contributions
The Five Compared Models
Five approaches were evaluated along a spectrum of task-specific adaptation, from zero-shot foundation-model inference to pancreas-specific fine-tuning. All methods received a single foreground point within the tumor, except MedSAM2, which uses its native bounding-box prompt.
- nnInteractive — an interactive 3D foundation model built on nnU-Net, using an AutoZoom refinement mechanism; used unmodified with a single foreground point.
- MedSAM2 — adapts SAM 2 to volumes by propagating a key-slice bounding box bidirectionally through memory-attention; used with its default configuration.
- TAGS (Zero-Shot) — the published TAGS weights trained on MSD-Pancreas, applied directly to the four target cohorts with no further adaptation.
- SAM-TAGS — the TAGS architecture fine-tuned from original SAM-B weights on the combined target cohorts (encoder frozen; adapters, prompt encoder, and decoder updated).
- MSD-TAGS — same recipe as SAM-TAGS but initialized from the pancreas-specific MSD-Pancreas TAGS checkpoint, isolating the benefit of organ-specific pretraining.
Segmentation Performance
Dice Similarity Coefficient (DSC) and Normalized Surface Dice (NSD, 5 mm tolerance) across four centers and overall. Best per column in blue.
| Method | Center A | Center B | Center C | Center D | All-Center Mean | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| DSC | NSD | DSC | NSD | DSC | NSD | DSC | NSD | DSC | NSD | |
| nnInteractive | 0.622 | 0.753 | 0.648 | 0.807 | 0.667 | 0.808 | 0.549 | 0.731 | 0.632 | 0.782 |
| MedSAM2 | 0.495 | 0.724 | 0.586 | 0.795 | 0.683 | 0.853 | 0.466 | 0.726 | 0.585 | 0.792 |
| TAGS (Zero-Shot) | 0.460 | 0.601 | 0.519 | 0.657 | 0.498 | 0.623 | 0.507 | 0.697 | 0.496 | 0.639 |
| SAM-TAGS | 0.658 | 0.786 | 0.666 | 0.776 | 0.745 | 0.868 | 0.633 | 0.805 | 0.691 | 0.822 |
| MSD-TAGS | 0.682 | 0.814 | 0.679 | 0.794 | 0.753 | 0.875 | 0.638 | 0.811 | 0.703 | 0.835 |
MSD-TAGS achieved significantly higher DSC and NSD than every competing method (two-sided Wilcoxon signed-rank, all Holm-adjusted p<0.001), the best DSC at all four centers, and the best NSD at three of four.
Overall Comparison
Figure 2. Mean DSC and NSD for nnInteractive, MedSAM2, TAGS (Zero-Shot), SAM-TAGS, and MSD-TAGS, averaged across all participating centers.
Qualitative Results
Figure 3. Representative pancreatic tumor segmentations across Centers A-D. Columns show the original CT, reference manual segmentation, and predictions from each method. Tumor masks are shown as red overlays; green boxes mark the zoomed insets.
Center-wise Performance
Figure 4. Center-wise segmentation performance across the four institutions: (A) DSC and (B) NSD for each method.
Computational Efficiency
| Method | Parameters (M) | TFLOPs | Inference Time (s) |
|---|---|---|---|
| nnInteractive | 102.35 | 17.996 | 0.255 |
| MedSAM2 | 38.96 | 9.237 | 4.968 |
| TAGS (Zero-Shot) | 126.16 | 7.455 | 0.682 |
| SAM-TAGS | 126.16 | 7.455 | 0.682 |
| MSD-TAGS | 126.16 | 7.455 | 0.682 |
The TAGS variants share the same complexity, require fewer FLOPs than both foundation-model baselines, and maintain feasible inference times — with MSD-TAGS delivering the best accuracy at no extra cost over SAM-TAGS.
Public Multicenter Dataset
We curate and publicly release a multicenter pancreatic tumor CT segmentation dataset of 500 CECT examinations with corresponding expert pancreas and tumor annotations, spanning four independent centers, multiple scanner vendors, and a wide range of acquisition protocols. Publicly available multicenter pancreatic tumor CT datasets remain limited at this scale, and the release aims to support reproducible benchmarking, studies of foundation-model adaptation, and further development of automated pancreatic cancer imaging tools.
Conclusion
This multicenter feasibility study shows that task-specific adaptation is central to reliable pancreatic tumor segmentation with foundation models. General-purpose interactive and medical foundation models provided useful baselines, but fine-tuned TAGS variants consistently achieved superior performance across the cohort. The progressive gains with increasing adaptation suggest that organ-specific pretraining and target-cohort fine-tuning improve disease-specific segmentation, supporting the growing role of task-adapted foundation models in pancreatic cancer imaging and laying groundwork for future clinical translation.
BibTeX
@Article{cancers18172836,
AUTHOR = {Aktas, Halil Ertugrul and Sen Tasci, Eminenur and Peng, Linkai and Tasci, Muhammed Enes and Taktak, Yavuz B. and Taflan, Sitki Safa and Tutun, Baver and Bol, Fergan and Iren, Murat and Ekici, Mucahit and Bejar, Andrea Mia and Keles, Elif and Pan, Hongyi and Dou, Wanying and Gultekin, Burak and Akin, Alper and Ikizgul, Oyku and Cetin, Okan and Uysal, Emre and Mureva, Maide and Nalbant, Mustafa Orhan and Kaya, Nurullah and Medetalibeyoglu, Alpay and Atakir, Kadir and Akkus Yildirim, Berna and Dagoglu Kartal, Gulbiz and Zhou, Zongwei and Erturk, Sukru Mehmet and Miller, Frank H. and Durak, Gorkem and Bagci, Ulas},
TITLE = {Multicenter Validation of Foundation Model Adaptation for Automated Pancreatic Tumor Delineation on CT Scans},
JOURNAL = {Cancers},
VOLUME = {18},
YEAR = {2026},
NUMBER = {17},
ARTICLE-NUMBER = {2836},
URL = {https://www.mdpi.com/2072-6694/18/17/2836},
ISSN = {2072-6694},
DOI = {10.3390/cancers18172836}
}
}