vjepa_huge_16(pretrained: bool = False, overrides: object = {})V-JEPA with a ViT-H/16 encoder at 224 pixels.
Model Size
Parameters
pretrainedbool= FalseThe released checkpoints are not redistributed here;
True
raises.**overridesobject= {}Optional
VJEPAConfig field overrides.Returns
VJEPAModelContext encoder, target encoder and predictor, untrained.
Notes
Reference: Bardes et al., arXiv:2404.08471, Table 6 — 82.0% on Kinetics-400 and 71.4% on Something-Something-v2, the paper's best at 224 pixels. 630M parameters in the encoder.
Examples
>>> from lucid.models import AutoConfig
>>> config = AutoConfig.from_pretrained("vjepa_huge_16")
>>> config.dim, config.depth
(1280, 32)
>>> config.predictor_dim, config.predictor_depth
(384, 12)