Knowledge Distillation from Medical Vision Foundation Models to Lightweight Networks for Edge Deployment: A Comparative Study
Abstract
Medical vision foundation models pretrained on large-scale domain-specific data achieve strong clinical performance, but their computational cost precludes deployment on resource-constrained edge devices. We investigate whether the domain-specific representations of a medical foundation model transfer more effectively through knowledge distillation (KD) than generic ImageNet-pretrained features. We compare a domain-specific foundation model (RETFound, ViT-Large, 307M parameters) against a general-purpose pretrained model (ConvNeXt-Base, 89M) as KD teachers for two lightweight students (MobileNetV3-Small, 2.5M; EfficientNet-Lite0, 4.7M) on two clinical benchmarks – HAM10000 (dermoscopy, 7 classes) and APTOS-2019 (diabetic retinopathy, 5 classes) – over a grid of distillation temperature and loss weight, with INT8 quantization and Grad-CAM analysis. Because RETFound is pretrained on retinal images, APTOS- 2019 is in-domain for it whereas HAM10000 (dermoscopy) constitutes an out-of-domain test. On balanced accuracy, the appropriate metric for these class-imbalanced datasets, ConvNeXt-Base yields the stronger student in all four settings (e.g., HAM10000 MobileNetV3: 0.826 vs. 0.763) – including on APTOS-2019, where RETFound’s retinal pretraining is in-domain – despite fewer parameters and no domain-specific pretraining; RETFound distillation can even fall below the no-distillation baseline. RETFound requires higher temperatures for effective transfer, whereas ConvNeXt is robust across a broad range. INT8 quan- tization preserves accuracy within 0.3% while compressing MobileNetV3-Small to 4.2 MB. These results indicate that a teacher’s fine-tuned quality, rather than domain-specific pretraining, predicts distillation success, and offer practical guidance for edge deployment of medical image classifiers.References
M. A. Al-Masud, J. M. Lopez Alcaraz, N. Strodthoff (2025) Benchmarking ECG foundational models: A reality check across clinical tasks, arXiv preprint arXiv:2509.25095. https://doi.org/10.48550/arXiv.2509.25095
Asia Pacific Tele-Ophthalmology Society (2019) APTOS 2019 blindness detection, Kaggle Competition. https://www.kaggle.com/competitions/aptos2019-blindness-detection
L. Beyer, X. Zhai, A. Royer, L. Markeeva, R. Anil, A. Kolesnikov (2022) Knowledge distillation: A good teacher is patient and consistent, Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 10915-10924. https://doi.org/10.1109/CVPR52688.2022.01065
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, et al. (2021) On the opportunities and risks of foundation models, arXiv preprint arXiv:2108.07258. https://doi.org/10.48550/arXiv.2108.07258
R. J. Chen, T. Ding, M. Y. Lu, D. F. K. Williamson, G. Jaume, et al. (2024) Towards a general-purpose foundation model for computational pathology, Nature Medicine, 30, pp. 850-862. https://doi.org/10.1038/s41591-024-02857-3
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei (2009) ImageNet: A large-scale hierarchical image database, Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 248-255. https://doi.org/10.1109/CVPR.2009.5206848
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, et al. (2021) An image is worth 16x16 words: Transformers for image recognition at scale, Proc. Int. Conf. on Learning Representations (ICLR).
A. Esteva, B. Kuprel, R. A. Novoa, J. Ko, S. M. Swetter, H. M. Blau, S. Thrun (2017) Dermatologist-level classification of skin cancer with deep neural networks, Nature, 542(7639), pp. 115-118. https://doi.org/10.1038/nature21056
A. Gholami, S. Kim, Z. Dong, Z. Yao, M. W. Mahoney, K. Keutzer (2022) A survey of quantization methods for efficient neural network inference, Low-Power Computer Vision, Chapman and Hall/CRC, pp. 291-326. https://doi.org/10.1201/9781003162810-13
J. Gou, B. Yu, S. J. Maybank, D. Tao (2021) Knowledge distillation: A survey, International Journal of Computer Vision, 129, pp. 1789-1819. https://doi.org/10.1007/s11263-021-01453-z
V. Gulshan, L. Peng, M. Coram, M. C. Stumpe, D. Wu, et al. (2016) Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs, JAMA, 316(22), pp. 2402-2410. https://doi.org/10.1001/jama.2016.17216
K. He, X. Chen, S. Xie, Y. Li, P. Dollar, R. Girshick (2022) Masked autoencoders are scalable vision learners, Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 15979-15988. https://doi.org/10.1109/CVPR52688.2022.01553
G. Hinton, O. Vinyals, J. Dean (2015) Distilling the knowledge in a neural network, arXiv preprint arXiv:1503.02531. https://doi.org/10.48550/arXiv.1503.02531
A. Howard, M. Sandler, B. Chen, W. Wang, L.-C. Chen, et al. (2019) Searching for MobileNetV3, Proc. IEEE/CVF Int. Conf. on Computer Vision (ICCV), pp. 1314-1324. https://doi.org/10.1109/ICCV.2019.00140
N. Islam, K. M. Hasib, F. A. Joti, A. Karim, S. Azam (2024) Leveraging knowledge distillation for lightweight skin cancer classification: Balancing accuracy and computational efficiency, arXiv preprint arXiv:2406.17051. https://doi.org/10.48550/arXiv.2406.17051
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, et al. (2023) Segment anything, Proc. IEEE/CVF Int. Conf. on Computer Vision (ICCV), pp. 3992-4003. https://doi.org/10.1109/ICCV51070.2023.00371
G. Litjens, T. Kooi, B. E. Bejnordi, A. A. A. Setio, F. Ciompi, et al. (2017) A survey on deep learning in medical image analysis, Medical Image Analysis, 42, pp. 60-88. https://doi.org/10.1016/j.media.2017.07.005
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, S. Xie (2022) A ConvNet for the 2020s, Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 11966-11976. https://doi.org/10.1109/CVPR52688.2022.01167
J. Ma, Y. He, F. Li, L. Han, C. You, B. Wang (2024) Segment anything in medical images, Nature Communications, 15, 654. https://doi.org/10.1038/s41467-024-44824-z
W. Park, D. Kim, Y. Lu, M. Cho (2019) Relational knowledge distillation, Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 3962-3971. https://doi.org/10.1109/CVPR.2019.00409
A. Rancea, I. Anghel, T. Cioara (2024) Edge computing in healthcare: Innovations, opportunities, and challenges, Future Internet, 16(9), 329. https://doi.org/10.3390/fi16090329
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, Y. Bengio (2015) FitNets: Hints for thin deep nets, Proc. Int. Conf. on Learning Representations (ICLR).
Z. Salahuddin, H. C. Woodruff, A. Chatterjee, P. Lambin (2022) Transparency of deep neural networks for medical image analysis: A review of interpretability methods, Computers in Biology and Medicine, 140, 105111. https://doi.org/10.1016/j.compbiomed.2021.105111
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, D. Batra (2017) Grad-CAM: Visual explanations from deep networks via gradient-based localization, Proc. IEEE Int. Conf. on Computer Vision (ICCV), pp. 618-626. https://doi.org/10.1109/ICCV.2017.74
M. Tan, Q. V. Le (2019) EfficientNet: Rethinking model scaling for convolutional neural networks, Proc. 36th Int. Conf. on Machine Learning (ICML), PMLR 97, pp. 6105-6114.
Y. Tian, D. Krishnan, P. Isola (2020) Contrastive representation distillation, Proc. Int. Conf. on Learning Representations (ICLR).
P. Tschandl, C. Rosendahl, H. Kittler (2018) The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions, Scientific Data, 5, 180161. https://doi.org/10.1038/sdata.2018.161
Y. Xu, T. M. Khan, Y. Song, E. Meijerink (2025) Edge deep learning in computer vision and medical diagnostics: A comprehensive survey, Artificial Intelligence Review, 58(3), 93. https://doi.org/10.1007/s10462-024-11033-5
S. Zagoruyko, N. Komodakis (2017) Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer, Proc. Int. Conf. on Learning Representations (ICLR).
S. Zhang, Y. Xu, N. Usuyama, H. Xu, J. Bagga, et al. (2025) A multimodal biomedical foundation model trained from fifteen million image-text pairs, NEJM AI, 2(1). https://doi.org/10.1056/AIoa2400640
J. Zheng, K. Xie, D. Zhang, Z. Lv, X. Yu (2025) Stepwise self-knowledge distillation for skin lesion image classification, Scientific Reports, 15. https://doi.org/10.1038/s41598-025-10717-4
Y. Zhou, M. A. Chia, S. K. Wagner, M. S. Ayhan, D. J. Williamson, et al. (2023) A foundation model for generalizable disease detection from retinal images, Nature, 622(7981), pp. 156-163. https://doi.org/10.1038/s41586-023-06555-x
DOI:
https://doi.org/10.31449/inf.v50i15.15303Keywords:
knowledge distillation, medical foundation models, edge deployment, model compression, quantization, medical image classificationDownloads
Published
Issue
Section
License
Authors retain copyright in their work. By submitting to and publishing with Informatica, authors grant the publisher (Slovene Society Informatika) the non-exclusive right to publish, reproduce, and distribute the article and to identify itself as the original publisher.
All articles are published under the Creative Commons Attribution license CC BY 3.0. Under this license, others may share and adapt the work for any purpose, provided appropriate credit is given and changes (if any) are indicated.
Authors may deposit and share the submitted version, accepted manuscript, and published version, provided the original publication in Informatica is properly cited.







