vjepa_huge_16_384(pretrained: bool = False, overrides: object = {})V-JEPA with a ViT-H/16 encoder at 384 pixels.
Model Size
Parameters
pretrainedbool= FalseThe released checkpoints are not redistributed here;
True
raises.**overridesobject= {}Optional
VJEPAConfig field overrides.Returns
VJEPAModelContext encoder, target encoder and predictor, untrained.
Notes
Reference: Bardes et al., arXiv:2404.08471, Table 6 — 72.2% on Something-Something-v2 and 77.4% on ImageNet-1k, the paper's best on both. Three times the tokens of the 224-pixel model: 4608 against 1568.
Examples
>>> from lucid.models import AutoConfig
>>> config = AutoConfig.from_pretrained("vjepa_huge_16_384")
>>> config.image_size, config.num_tokens
(384, 4608)