MobileNet-v4 feature-extracting backbone (no classification head).
Implements the searched architectures of Qin et al., "MobileNetV4: Universal Models for the Mobile Ecosystem", ECCV 2024 (arXiv:2404.10518). The trunk is a stride-2 stem followed by five stages built from the Universal Inverted Bottleneck (UIB), whose two optional depthwise convolutions let the architecture search pick an inverted bottleneck, a ConvNeXt-like block, an FFN or the new ExtraDW block at every position; a fused inverted bottleneck opens the medium and large models, and the hybrid variants interleave Mobile MQA attention blocks into the two deepest stages. The last stage is a convolution widening the stride-32 map to 960 channels.
Parameters
configMobileNetV4Configmobilenet_v4_conv_small, mobilenet_v4_conv_medium,
mobilenet_v4_conv_large, mobilenet_v4_hybrid_medium,
mobilenet_v4_hybrid_large) for the paper's variants.Attributes
configMobileNetV4Configconv_stemnn.Conv2dactnn.Moduleblocksnn.SequentialSequential containers; stage ends at stride
for , and stage 4 is the 960-channel
widening layer.feature_infolist[FeatureInfo]forward_features returns.Notes
The UIB computes
with the pointwise expansion (norm + activation), the linear pointwise projection, each an optional depthwise convolution and a per-channel layer scale present only in the hybrid variants. The residual is used whenever the block keeps both resolution and width.
Examples
>>> import lucid
>>> from lucid.models.vision.mobilenet_v4 import mobilenet_v4_conv_small
>>> backbone = mobilenet_v4_conv_small().eval()
>>> out = backbone(lucid.randn(1, 3, 224, 224))
>>> out.last_hidden_state.shape
(1, 960, 7, 7)
>>> [f.reduction for f in backbone.feature_info]
[4, 8, 16, 32]