Abstract
Ultrasound images can vary widely across scanners, operators, and anatomical targets, so models trained in one setting often generalize poorly to new hospitals and clinical conditions. The Foundation Model Challenge for Ultrasound Image Analysis (FMC-UIA) reflects this scenario by requiring a single model to handle multiple tasks like segmentation, detection, classification, and landmark regression across diverse organs and datasets. We propose a unified multi-task framework based on a transformer visual encoder from the Qwen3-VL family. Intermediate token features are projected into spatial feature maps and fused using a lightweight multiscale feature pyramid, enabling both pixel-level predictions and global reasoning within a shared representation. Each task is addressed with a small task-specific prediction head, while training uses task-aware sampling and selective loss balancing to manage heterogeneous supervision and reduce task imbalance. Our method is designed to be straightforward to optimize and adaptable across a wide range of ultrasound analysis tasks. The score improved from 67% to 85% on validation and reached 81.84% average on the official test set across all tasks. The code is publicly available at Github.
| Original language | English |
|---|---|
| Title of host publication | 2026 IEEE 23rd International Symposium on Biomedical Imaging (ISBI) |
| DOIs | |
| Publication status | Published (VoR) - 20 May 2026 |
Fingerprint
Dive into the research topics of 'AURORA: Adaptive Unified Representation for Robust Ultrasound Analysis'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver