vjepa2_vit_giant_384(pretrained: bool | str = False, weights: VJEPA2ViTGiant384Weights | None = None, overrides: object = {})The released V-JEPA 2 ViT-g/16 backbone at 384 pixels.
Model Size
Parameters
pretrainedbool or str= FalseTrue loads the Lucid FPC64-384 tag; a string selects a declared
weight tag explicitly.Explicit weight enum member; takes precedence over
pretrained.**overridesobject= {}Optional
VJEPA2Config field overrides.Returns
VJEPA2ModelContext encoder, EMA target encoder and predictor.
Notes
Reference: Assran et al., arXiv:2506.09985, 2025. The same network as
vjepa2_vit_giant evaluated at a higher resolution: the rotary
geometry is computed from token positions rather than read from a
learned table, so the wider grid needs no interpolation and no new
parameters — only more tokens, 18432 against 8192.
Examples
>>> from lucid.models import AutoConfig
>>> config = AutoConfig.from_pretrained("vjepa2_vit_giant_384")
>>> config.image_size, config.token_grid
(384, (32, 24, 24))
>>> config.num_tokens
18432