×

Poster no: 146

146

Medical imaging foundation model enhances pre-operative prediction of difficult laryngoscopy in cervical spondylosis using MRI

This study developed and validated a Vision Transformer (ViT)-based deep learning model using preoperative cervical MRI images to predict difficult laryngoscopy in patients with cervical spondylosis.

Methods

This study involves 14,407 patients who underwent cervical spine surgery under general anesthesia. A Vision Transformer (ViT) model was pre-trained via self-supervised learning on medical imaging datasets (MRI and chest X-rays) and fine-tuned using T1-weighted MRI sequences. Five-fold cross-validation with data augmentation was employed to evaluate performance against traditional clinical metrics and baseline models.

Results

According to the inclusion criteria, a total of 1,137 patients (233 difficult laryngoscopy cases) were ultimately included in the study. The medical foundation model achieved superior performance (AUC: 0.870), significantly outperforming conventional indicators (combined AUC: 0.734) and baseline architectures (ViT: 0.791; Residual Network: 0.781; Visual Geometry Group: 0.781; Multilayer Perceptron:0.727). Grad-CAM visualizations confirmed clinical relevance, highlighting critical regions, including pharynx, cervical soft tissues, tongue body and cervical spine.

Discussion

The MRI-based ViT model achieved an AUC of 0.870, outperforming conventional clinical indicators (AUC = 0.734). Its superiority stems from domain-specific pre-training on large-scale medical imaging datasets, enabling capture of subtle anatomical nuances such as retrolingual space variations and cervical soft tissue dynamics—features that static metrics like thyromental distance fail to reflect. Grad-CAM visualizations confirmed clinical relevance by highlighting key regions including the pharynx, tongue body, and cervical soft tissues [1]. Unlike prior AI approaches relying on facial images or cervical X-rays [2], this MRI-based framework offers comprehensive three-dimensional spatial analysis. Limitations include reliance on static MRI and restricted generalizability beyond cervical spondylosis populations.

Acknowledgements

This study was funded by Clinical Key Project, Peking University Third Hospital (BYSYZD2025052).

Figure 1. Comparison of model performance using (a) Brier Score and (b) Mean Absolute Error (MAE). Lower values indicate better probabilistic predictions. Median scores are shown. Our Model achieves the lowest errors, outperforming ViT, ResNet, VGG, and MLP. Dashed lines indicate reference levels: blue for perfect prediction, red for random guess.

Contact Author