MobileNet-v4
18 memberslucid.models.vision.mobilenet_v4MobileNet v4 family — Qin et al., 2024.
Qin, Danfeng, et al. "MobileNetV4: Universal Models for the Mobile Ecosystem." Computer Vision – ECCV 2024, Lecture Notes in Computer Science, vol. 15098, Springer, 2024.
MobileNet-v4 is built around a single searchable block, the Universal Inverted Bottleneck (UIB). It extends the inverted bottleneck of MobileNet-v2 with two optional depthwise convolutions: one before the expansion and one between the expansion and the projection,
where is the pointwise expansion, the linear pointwise projection, and each is either a depthwise convolution or the identity. The four on/off combinations recover four familiar blocks: the classic inverted bottleneck (IB, middle depthwise only), a ConvNeXt-like block (start depthwise only, a cheap large-kernel spatial mix before the expansion), a transformer FFN (neither — two pointwise layers), and the new ExtraDW block (both), which deepens the network and widens its receptive field at almost no cost. Because the expansion and projection are shared by all four instantiations, a NAS super-network over UIB shares more than 95% of its parameters, and the search simply decides which depthwise layers to keep at each position. The stems use a fused inverted bottleneck — a dense convolution into a projection — which is faster than a depthwise pair at high resolution.
The hybrid variants interleave UIB blocks with Mobile MQA, a multi-query attention block tuned for accelerators: every query head shares a single key and value head, which raises operational intensity when the token count is small relative to the channel width, and the keys and values can be spatially reduced by a stride-2 depthwise convolution,
Sharing keys and values cuts attention latency by more than 39% on mobile accelerators relative to multi-head attention at a negligible accuracy cost. Residual branches of the hybrids are scaled by a learnable per-channel layer scale initialised at .
Five models come out of the search: Conv-Small, Conv-Medium and Conv-Large, built only from UIB and fused-IB blocks, and Hybrid-Medium and Hybrid-Large, which add Mobile MQA to the last two stages. Every model ends with the MobileNet-v3 style head that moves the widest layer after global pooling. They are mostly Pareto-optimal across mobile CPUs, DSPs, GPUs, the Apple Neural Engine and the Pixel EdgeTPU at once, ranging from 73.8% ImageNet-1k top-1 at 0.2 GMACs (Conv-Small) to 83.4% (Hybrid-Large).
Classes
Functions
mobilenet_v4_conv_large→ MobileNetV430.1MMobileNet-v4-Conv-Large feature-extracting backbone.
mobilenet_v4_conv_large_cls→ MobileNetV4ForImageClassification32.6MMobileNet-v4-Conv-Large image classifier.
mobilenet_v4_conv_medium→ MobileNetV47.2MMobileNet-v4-Conv-Medium feature-extracting backbone.
mobilenet_v4_conv_medium_cls→ MobileNetV4ForImageClassification9.7MMobileNet-v4-Conv-Medium image classifier.
mobilenet_v4_conv_small→ MobileNetV41.3MMobileNet-v4-Conv-Small feature-extracting backbone.
mobilenet_v4_conv_small_cls→ MobileNetV4ForImageClassification3.8MMobileNet-v4-Conv-Small image classifier.
mobilenet_v4_hybrid_large→ MobileNetV435.3MMobileNet-v4-Hybrid-Large feature-extracting backbone.
mobilenet_v4_hybrid_large_cls→ MobileNetV4ForImageClassification37.8MMobileNet-v4-Hybrid-Large image classifier.
mobilenet_v4_hybrid_medium→ MobileNetV48.6MMobileNet-v4-Hybrid-Medium feature-extracting backbone.
mobilenet_v4_hybrid_medium_cls→ MobileNetV4ForImageClassification11.1MMobileNet-v4-Hybrid-Medium image classifier.
Weights
MobileNetV4ConvLargeWeightsPretrained weights for lucid.models.mobilenet_v4_conv_large_cls.
MobileNetV4ConvMediumWeightsPretrained weights for lucid.models.mobilenet_v4_conv_medium_cls.
MobileNetV4ConvSmallWeightsPretrained weights for lucid.models.mobilenet_v4_conv_small_cls.
MobileNetV4HybridLargeWeightsPretrained weights for lucid.models.mobilenet_v4_hybrid_large_cls.
MobileNetV4HybridMediumWeightsPretrained weights for lucid.models.mobilenet_v4_hybrid_medium_cls.