vjepa2_vit_large(pretrained: bool | str = False, weights: VJEPA2ViTLargeWeights | None = None, overrides: object = {})Construct the released V-JEPA 2 ViT-L/16 backbone at 256 pixels.
Model Size
Parameters
pretrainedbool or str= FalseTrue loads the Lucid FPC64-256 tag; a string selects a declared
weight tag explicitly.Explicit weight enum member; takes precedence over
pretrained.**overridesobject= {}Optional
VJEPA2Config field overrides.Notes
Reference: Assran et al., arXiv:2506.09985, 2025. The release lists
this variant as ViT-L/16, 300M parameters, resolution 256 — that figure
counts one encoder, while this factory builds two of them and the
predictor. The FPC64-256 tag reproduces its source to 1.7e-5
relative, the closest of the four.
Examples
>>> from lucid.models import AutoConfig
>>> config = AutoConfig.from_pretrained("vjepa2_vit_large")
>>> config.dim, config.depth, config.num_heads
(1024, 24, 16)
>>> config.token_grid, config.num_tokens
((32, 16, 16), 8192)