vjepa_large_16(pretrained: bool = False, overrides: object = {})V-JEPA with a ViT-L/16 encoder at 224 pixels.
Model Size
Parameters
pretrainedbool= FalseThe released checkpoints are not redistributed here;
True
raises.**overridesobject= {}Optional
VJEPAConfig field overrides.Returns
VJEPAModelContext encoder, target encoder and predictor, untrained.
Notes
Reference: Bardes, Adrien, et al., "Revisiting Feature Prediction for Learning Visual Representations from Video", arXiv:2404.08471, 2024, Table 6 — frozen with an attentive probe: 80.8% on Kinetics-400, 69.5% on Something-Something-v2, 74.8% on ImageNet-1k. The paper reports 200M parameters for this encoder.
Examples
>>> from lucid.models import AutoConfig
>>> config = AutoConfig.from_pretrained("vjepa_large_16")
>>> config.dim, config.depth, config.num_heads
(1024, 24, 16)
>>> config.token_grid, config.num_tokens
((8, 14, 14), 1568)