vjepa2_vit_giant(pretrained: bool | str = False, weights: VJEPA2ViTGiantWeights | None = None, overrides: object = {})The released V-JEPA 2 ViT-g/16 backbone at 256 pixels.
Model Size
Parameters
pretrainedbool or str= FalseTrue loads the Lucid FPC64-256 tag; a string selects a declared
weight tag explicitly.Explicit weight enum member; takes precedence over
pretrained.**overridesobject= {}Optional
VJEPA2Config field overrides.Returns
VJEPA2ModelContext encoder, EMA target encoder and predictor.
Notes
Reference: Assran et al., arXiv:2506.09985, 2025. The release lists
ViT-g/16 at 1B parameters and resolution 256. This is the variant the
action-conditioned model is post-trained from, and the only one whose
encoder widens its feed-forward to 48 / 11 while keeping 22
attention heads. The FPC64-256 tag reproduces its source to 1.0e-4
relative.
Examples
>>> from lucid.models import AutoConfig
>>> config = AutoConfig.from_pretrained("vjepa2_vit_giant")
>>> config.dim, config.depth, config.num_heads
(1408, 40, 22)
>>> round(config.mlp_ratio, 4), config.predictor_mlp_ratio
(4.3636, 4.0)