mobilenet_v4_hybrid_medium(pretrained: bool = False, overrides: object = {})MobileNet-v4-Hybrid-Medium feature-extracting backbone.
Builds a MobileNetV4 with the Hybrid-Medium architecture of
Qin et al., 2024 (Appendix D, Table 13): a layout close to Conv-Medium's with four Mobile MQA blocks (4 heads of width 64) interleaved into each of the two deepest stages, keys and values reduced by a stride-2 depthwise conv in the first of them, closed by a
widening to 960 channels. 8.56 M
parameters without the classification head.
Model Size
Parameters
pretrainedbool= FalseFalse — no checkpoint is published for the headless
backbone, and True raises rather than returning random
weights. The ImageNet-1k checkpoint belongs to the matching
*_cls classifier, whose error message this names.**overridesobject= {}MobileNetV4Config
(e.g. drop_path_rate=0.1).Returns
MobileNetV4Backbone returning the stride-32, 960-channel feature map.
Raises
NotImplementedErrorpretrained=True.Notes
Qin et al., "MobileNetV4: Universal Models for the Mobile Ecosystem", ECCV 2024 (arXiv:2404.10518). The paper trains and evaluates this variant at 256x256; the network is fully convolutional up to the pool, so any input whose sides are multiples of 32 works.
Examples
>>> import lucid
>>> from lucid.models.vision.mobilenet_v4 import mobilenet_v4_hybrid_medium
>>> model = mobilenet_v4_hybrid_medium().eval()
>>> model(lucid.randn(1, 3, 224, 224)).last_hidden_state.shape
(1, 960, 7, 7)