Skip to main navigation Skip to search Skip to main content

Self-supervised ultrasound-video segmentation with feature prediction and 3D localised loss

  • Edward Ellis
  • , Robert Mendel
  • , Andrew Bulpitt
  • , Nasim Parsa
  • , Michael F. Byrne
  • , Sharib Ali

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

Acquiring and annotating large datasets in ultrasound imaging is challenging due to low contrast, high noise, and susceptibility to artefacts. This process requires significant time and clinical expertise. Self-supervised learning (SSL) offers a promising solution by leveraging unlabelled data to learn useful representations, enabling improved segmentation performance when annotated data is limited. Recent state-of-the-art developments in SSL for video data include V-JEPA, a framework solely based on feature prediction, avoiding pixel level reconstruction or negative samples. We hypothesise that V-JEPA is well-suited to ultrasound imaging, as it is less sensitive to noisy pixel-level detail while effectively leveraging temporal information. To the best of our knowledge, this is the first study to adopt V-JEPA for ultrasound video data. Similar to other patch-based masking SSL techniques such as VideoMAE, V-JEPA is well-suited to ViT-based models. However, ViTs can underperform on small medical datasets due to lack of inductive biases, limited spatial locality and absence of hierarchical feature learning. To improve locality understanding, we propose a novel 3D localisation auxiliary task to improve locality in ViT representations during V-JEPA pre-training. Our results show V-JEPA with our auxiliary task improves segmentation performance significantly across various frozen encoder configurations, with gains up to 3.4% using 100% and up to 8.35% using only 10% of the training data.

Original languageEnglish (US)
Title of host publicationMedical Imaging 2026
Subtitle of host publicationImage Processing
EditorsJhimli Mitra, Yu Gan
PublisherSPIE
ISBN (Electronic)9781510697874
DOIs
StatePublished - Apr 3 2026
Externally publishedYes
EventMedical Imaging 2026: Image Processing - Vancouver, Canada
Duration: Feb 15 2026Feb 19 2026

Publication series

NameProgress in Biomedical Optics and Imaging - Proceedings of SPIE
Volume13925
ISSN (Print)1605-7422
ISSN (Electronic)2410-9045

Conference

ConferenceMedical Imaging 2026: Image Processing
Country/TerritoryCanada
CityVancouver
Period2/15/262/19/26

Bibliographical note

Publisher Copyright:
© COPYRIGHT SPIE. Downloading of the abstract is permitted for personal use only.

Keywords

  • Segmentation
  • Self-Supervised Learning
  • Transformers
  • Ultrasound

Fingerprint

Dive into the research topics of 'Self-supervised ultrasound-video segmentation with feature prediction and 3D localised loss'. Together they form a unique fingerprint.

Cite this