mobilenet_v4_conv_medium_cls(pretrained: bool | str = False, weights: MobileNetV4ConvMediumWeights | None = None, overrides: object = {})MobileNet-v4-Conv-Medium image classifier.
Builds a MobileNetV4ForImageClassification with the Conv-Medium
architecture (Qin et al., 2024, Appendix D, Table 12) and the
post-pool head (960 → 1280 → classes). 9.72 M parameters;
Table 6 reports 79.9% ImageNet-1k top-1 at 9.2 M parameters and
1.0 G MACs. Classifier dropout defaults to the paper's 0.2
(Table 10).
Model Size
Parameters
pretrainedbool or str= FalseFalse → random init; True → the
DEFAULT tag
(MobileNetV4ConvMediumWeights.E500_R256_IN1K); a tag string
(e.g. "E500_R256_IN1K") → that specific checkpoint. Mutually
exclusive with weights (which wins if both are given).MobileNetV4ConvMediumWeights.E500_R256_IN1K. Takes precedence
over pretrained.**overridesobject= {}MobileNetV4Config
(typically num_classes to retarget the head). Overriding
num_classes away from the checkpoint's 1000 makes pretrained
loading fail the strict key/shape check — load with the
matching head, then call reset_classifier.Returns
MobileNetV4ForImageClassificationClassifier with the Conv-Medium configuration (plus overrides),
optionally initialised from pretrained weights.
Notes
Qin et al., "MobileNetV4: Universal Models for the Mobile Ecosystem", ECCV 2024 (arXiv:2404.10518). Parameter names match timm's implementation, so its ImageNet-1k checkpoints load with an identity key map. The paper evaluates this variant at 256x256.
Pretrained weights are converted from timm's
mobilenetv4_conv_medium.e500_r256_in1k — timm's own ImageNet-1k
training of this architecture, trained at 256x256; the authors released
none — and hosted on the Hugging Face Hub under
lucid-dl/mobilenet-v4-conv-medium. timm reports 79.916% top-1 /
95.188% top-5 for it at 256x256 with the preset that
MobileNetV4ConvMediumWeights.E500_R256_IN1K.transforms
reproduces (256 crop, 269 resize, bicubic, ImageNet mean/std).
Examples
>>> import lucid
>>> from lucid.models.vision.mobilenet_v4 import mobilenet_v4_conv_medium_cls
>>> model = mobilenet_v4_conv_medium_cls(num_classes=10).eval()
>>> model(lucid.randn(2, 3, 224, 224)).logits.shape
(2, 10)
Load ImageNet-pretrained weights:
>>> model = mobilenet_v4_conv_medium_cls(pretrained=True)
>>> from lucid.models.weights import MobileNetV4ConvMediumWeights
>>> model = mobilenet_v4_conv_medium_cls(
... weights=MobileNetV4ConvMediumWeights.E500_R256_IN1K
... )