[Nature Medicine] A Clinically-Oriented Foundation Model for Intraoperative Pathology

CRISP article published in Nature Medicine

Recently, Professor Hao Chen’s SmartX Lab team at the Hong Kong University of Science and Technology (HKUST) reported new work in AI-assisted intraoperative pathology. The study was conducted with Professor Muyan Cai and a multidisciplinary team from Sun Yat-sen University Cancer Center, Fudan University Shanghai Cancer Center, Tongji Hospital and other medical institutions.

Published in Nature Medicine, the study introduces CRISP (Clinical Real-time Intraoperative Support for Pathology), a foundation model developed for frozen-section whole-slide images. Frozen-section diagnosis guides time-critical decisions while a patient remains in surgery. Yet these images differ from routine paraffin sections and often contain freezing artefacts, folds, tears and staining variation.

The team pretrained CRISP on 103,797 slides from 10 medical centres and evaluated it across 98 retrospective tasks involving more than 15,000 cases. These tasks covered common cancers, unseen anatomical sites, rare tumours and questions relevant to surgical decisions. A separate registered prospective observational study then assessed CRISP in 3,064 consecutive patients across three centres.

Within the prospective workflows, CRISP supported surgical decisions in 92.6% of cases. Model-assisted triage was associated with an approximately 35% reduction in pathologists’ diagnostic workload, while retaining 79.0% of positive cases for review. The study also recorded 105 avoided ancillary tests and 87.5% accuracy in micrometastasis assessment. CRISP remained a clinical decision-support system under pathologist oversight, not an autonomous diagnostic system.

CRISP multicentre cohort, model development, evaluation, and prospective clinical workflow

Figure 1 | Study cohorts and clinical workflow of CRISP. a to c, The multicentre frozen-section dataset comprises 103,797 slides from 10 medical centres and 25 anatomical sites. d and e, CRISP is pretrained on frozen-section whole-slide images and compared with pathology foundation models and specialist systems across retrospective tasks. f, Prospective observational validation examines real-time intraoperative diagnosis, human-AI collaboration, workload and clinical utility.

Introduction

Frozen-section pathology provides diagnoses during surgery, when findings can alter the operative plan within minutes. Pathologists may need to determine malignancy, assess surgical margins, identify lymph-node metastases or advise whether more tissue should be removed.

This speed creates a demanding diagnostic setting. Rapid preparation can introduce ice-crystal artefacts, tissue deformation, folds, tears, uneven staining and incomplete sampling. Pathologists must interpret these slides under time pressure and communicate their findings to the operating team.

Most computational pathology foundation models are pretrained on formalin-fixed, paraffin-embedded tissue. Their transfer to frozen sections is uncertain because the image distributions and preparation artefacts differ. Retrospective benchmarks alone also cannot show how a model behaves within an active intraoperative workflow.

CRISP addresses both gaps through frozen-section-specific pretraining and staged evaluation. The study progresses from broad retrospective testing to a distinct prospective observational assessment in real-time clinical workflows.

Background

Three barriers have constrained AI support for intraoperative pathology:

  • Frozen sections differ from routine pathology. Rapid preparation introduces artefacts and visual distributions that are poorly represented in paraffin-section training data.
  • The clinical questions are diverse. Intraoperative pathology includes benign-malignant differentiation, margin assessment, lymph-node evaluation and tumour subtyping across many anatomical sites.
  • Retrospective performance does not establish prospective utility. Curated benchmarks cannot reproduce every time, preparation and workflow constraint encountered during surgery.

The study therefore combines domain-specific pretraining, heterogeneous retrospective evaluation and prospective observational validation under clinical conditions.

Method

A multicentre frozen-section foundation dataset

The team collected 103,797 frozen-section whole-slide images from 10 medical centres, covering 25 anatomical sites. The corpus spans variation in institutions, scanners, staining protocols, patient populations and tissue preparation.

Data for pretraining, retrospective evaluation and prospective validation were separated. This staged design distinguishes representation learning and retrospective transfer from assessment in ongoing clinical workflows.

Domain-specific self-supervised pretraining

CRISP learns visual representations from unlabelled frozen-section whole-slide images. Slides are divided into image patches so the model can capture morphological patterns across organs, diseases and institutions. The frozen encoder can then be paired with lightweight task-specific heads or multiple-instance learning pipelines.

This design reduces the amount of labelled data needed for each application. It also allows the study to test whether one frozen-section representation transfers across common tasks and settings absent from pretraining.

Evaluation from retrospective benchmarks to prospective care

The evaluation strategy included three complementary levels:

  1. Retrospective task evaluation across common benign-malignant differentiation, intraoperative decision-making and other clinically relevant endpoints.
  2. Generalization analysis across independent centres, previously unseen anatomical sites and rare cancers.
  3. Prospective observational validation in consecutive intraoperative cases, assessing diagnostic performance, decision support, workload and ancillary testing.

Broad Retrospective Validation

CRISP was evaluated across 98 retrospective downstream tasks involving more than 15,000 frozen-section cases. The benchmark included common diagnostic questions and distribution shifts across centres and anatomical sites. Comparators included specialist systems and pathology foundation models pretrained mainly on routine histology.

Across common benign-malignant differentiation tasks, CRISP showed strong discrimination with limited task-specific adaptation. Analyses across institutions, tissue types and diagnostic categories also examined variation beyond aggregate performance. These retrospective results established broad transfer, while leaving prospective workflow performance as a separate question.

CRISP performance on common benign-malignant differentiation tasks

Figure 2 | Performance on common benign-malignant differentiation tasks. CRISP is compared with specialist systems and pathology foundation models across internal, external and multicentre frozen-section cohorts. The evaluation examines discrimination under institutional and anatomical variation.

Generalization Across Anatomical Sites and Rare Cancers

The next retrospective tests asked whether CRISP could transfer beyond the organs and disease distributions represented during pretraining. The evaluation included previously unseen anatomical sites and rare tumour categories, where task-specific datasets are often limited.

CRISP transferred through both zero-shot-style feature use and supervised downstream adaptation. These findings support the reuse of frozen-section representations across organs and uncommon diagnoses within the evaluated datasets.

Rare cancers nonetheless remain difficult because individual categories contain few cases and may require highly specialised criteria. The results don’t imply that one general model resolves this scarcity. They indicate that domain-specific pretraining can provide a useful starting representation when labelled data are limited.

CRISP generalization across anatomical sites and rare cancers

Figure 3 | Generalization to unseen anatomical sites and rare cancers. The analyses assess transfer across organs, centres and low-frequency diagnostic categories, including settings outside the pretraining taxonomy.

Prospective Clinical Validation

Retrospective evaluation established broad performance, but not behaviour during live surgical workflows. The researchers therefore conducted a registered prospective observational study of 3,064 consecutive patients at three medical centres.

The cohorts covered intraoperative thyroid, breast and central nervous system pathology. CRISP generated predictions while pathologists assessed the cases, and its output supported surgical decisions in 92.6% of cases.

In a model-assisted triage workflow, the reviewed subset retained 79.0% of positive cases while pathologists’ diagnostic workload fell by approximately 35%. The workflow also avoided 105 ancillary tests. These are observational workflow findings and don’t establish effects on patient outcomes.

For lymph-node assessment, CRISP achieved 87.5% accuracy for micrometastases. The model’s role was to screen tissue and direct attention to areas requiring expert review, while pathologists retained diagnostic responsibility.

Prospective clinical validation and human-AI collaboration with CRISP

Figure 4 | Prospective observational validation of CRISP. The multicentre study examines diagnostic performance, surgical decision support, model-assisted triage, workload, ancillary testing and lymph-node micrometastasis assessment in intraoperative workflows.

Translational Potential

CRISP brings together four elements relevant to clinical translation:

  • Frozen-section-specific pretraining. The training corpus represents preparation artefacts and image distributions that differ from routine paraffin histology.
  • Broad retrospective evaluation. One representation was tested across 98 tasks, multiple institutions, unseen anatomical sites and rare tumours.
  • Prospective observational validation. Consecutive patients were assessed under clinical timing, preparation and workflow constraints.
  • Human-AI collaboration. The study evaluated decision support and triage while preserving pathologist oversight.

The evidence has clear limits. The prospective study was observational rather than interventional, so it doesn’t establish improved patient outcomes. Evaluation in further regions, scanners and institutional workflows is needed to assess transportability. Rare cancers and subtle lesions remain difficult, while artefacts, sampling limitations and unseen distributions can still cause errors. Deployment also requires digital slide preparation and computational infrastructure that may not be available in every operating theatre.

CRISP should therefore remain a decision-support and triage tool under pathologist oversight. It is intended to assist time-critical review, not to issue autonomous diagnoses or replace clinical responsibility.


Resources

Paper | A clinically-oriented foundation model for intraoperative pathology

Journal | Nature Medicine (2026)

DOI | https://doi.org/10.1038/s41591-026-04703-0

Code | https://github.com/FT-ZHOU-ZZZ/CRISP

Model weights | https://huggingface.co/FTZhou/CRISP

Public dataset | https://huggingface.co/datasets/FTZhou/CRISP

Co-first authors | Zihan Zhao, Fengtao Zhou, and Ronggang Li

Corresponding authors | Hao Chen, The Hong Kong University of Science and Technology; Muyan Cai, Sun Yat-sen University Cancer Center