Skip to main navigation Skip to search Skip to main content

Keep It Frozen: Domain-Routed Conditional Residual Modulation for Multi-Domain Vision Transformers

  • Ufaq Khan
  • , Umair Nawaz
  • , Massimo Caputo
  • , Muhammad Bilal
  • , Junaid Qadir
  • , Muhammad Haris Khan

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Medical imaging remains challenging due to acoustic shadows, motion blur, and indistinct boundaries, while adapting vision models to such domains often requires heavy task-specific fine-tuning and can degrade general-image capability. We propose DCRM-ViT, a domain-conditioned residual modulation framework for Vision Transformers that preserves general-vision knowledge while adapting to diverse medical and natural domains. DCRM-ViT keeps the backbone frozen and augments each block with lightweight Residual Modulation Blocks (RMBs), whose parameters are synthesized per sample by a Domain Router (DR) and a Parameter Synthesizer Network (PSN). The DR. predicts soft domain weights from input features, and the PSN maps them to low-rank residuals that modulate selected projections and optionally add a domain-aware attention bias. We train the model with a bi-level optimization scheme, where an inner loop adapts RMBs to task supervision and an outer loop updates DR, PSN, and RMB initialization to improve crossdomain generalization. Across fine-grained classification (Food101, SUN397, Stanford Cars) and medical segmentation (ultrasound, CT, MRI), DCRM-ViT consistently outperforms strong baselines with modest trainable compute. Ablation studies confirm the contribution of the proposed components. Overall, DCRM-ViT achieves strong crossdomain performance with low overhead, requiring only 4.7 GFLOPs and 0.3 min/epoch. Our code can be accessed on DCRM-ViT.
Original languageEnglish
Title of host publicationProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Publication statusPublished (VoR) - 7 Jun 2026

Fingerprint

Dive into the research topics of 'Keep It Frozen: Domain-Routed Conditional Residual Modulation for Multi-Domain Vision Transformers'. Together they form a unique fingerprint.

Cite this