vjepa2_vit_huge(pretrained: bool | str = False, weights: VJEPA2ViTHugeWeights | None = None, overrides: object = {})The released V-JEPA 2 ViT-H/16 backbone at 256 pixels.
Model Size
Parameters
pretrainedbool or str= FalseTrue loads the Lucid FPC64-256 tag; a string selects a declared
weight tag explicitly.Explicit weight enum member; takes precedence over
pretrained.**overridesobject= {}Optional
VJEPA2Config field overrides.Returns
VJEPA2ModelContext encoder, EMA target encoder and predictor.
Notes
Reference: Assran et al., arXiv:2506.09985, 2025. The release lists
ViT-H/16 at 600M parameters and resolution 256 — that figure counts one
encoder, while this factory builds two of them and the predictor. The
FPC64-256 SafeTensors tag is hosted by Lucid and reproduces its source
to 5.3e-5 relative.
Examples
>>> from lucid.models import AutoConfig
>>> config = AutoConfig.from_pretrained("vjepa2_vit_huge")
>>> config.dim, config.depth, config.num_heads
(1280, 32, 16)
>>> config.token_grid, config.num_tokens
((32, 16, 16), 8192)