V-JEPA 2
16 memberslucid.models.vision.vjepa2V-JEPA 2 — self-supervised video feature prediction.
Assran, Mahmoud, et al. "V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning." [arXiv:2506.09985](https://arxiv.org/abs/2506.09985), 2025.
V-JEPA 2 learns a video representation by predicting the target encoder's latent features at hidden tubelet positions. A 3-D patch embedding turns a clip into a sequence of tubelets. The context encoder sees only the unmasked indices, while the predictor inserts learned mask tokens and reconstructs the target encoder's representation at the held-out indices. The target encoder is an exponential moving average of the context encoder and is not differentiated through.
The released backbone uses a pre-normalised transformer with three-axis rotary position encoding. Each attention head gives one third of its width to time, row and column — rounded down to an even number, so a head of 64 channels rotates 20 per axis and leaves the last four untouched. Because position enters through that rotation rather than a learned table, a clip of a different length or resolution needs no interpolation and no new parameters: the 384-pixel variant is the 256-pixel network read over a wider grid.
Two details of the released rotation are reproduced rather than corrected. The frequency half is repeated as a block before adjacent feature pairs are rotated, which pairs each rotation with a frequency that a straightforward implementation would not choose; the upstream source marks it as a bug and keeps it, because the published weights were trained through it. Fixing it yields a different model, not a better-implemented one.
The predictor is narrower than the video encoder — 384 channels and 12 heads in every released checkpoint, whatever the encoder's width — and owns ten mask-token slots, one per block the mask generator samples. It sees the context tokens and the mask tokens together, sorted into token order so that a position's identity reaches attention through the same rotary coordinates the encoder used, then answers only at the held-out positions. Targets come from the momentum encoder and are layer-normalised, so the objective compares directions in representation space rather than magnitudes.
Classes
VJEPA2Config4 methodsFrozen architecture configuration for V-JEPA 2.
VJEPA2ForVideoClassification2 methodsV-JEPA 2 read through the probe the paper evaluates with.
VJEPA2Model9 methodsV-JEPA 2 context encoder, EMA target encoder and predictor.
VJEPA2Output1 methodsOutput of a representation pass or masked prediction step.
Functions
vjepa2_vit_giant→ VJEPA2ModelThe released V-JEPA 2 ViT-g/16 backbone at 256 pixels.
vjepa2_vit_giant_384→ VJEPA2ModelThe released V-JEPA 2 ViT-g/16 backbone at 384 pixels.
vjepa2_vit_giant_384_cls→ VJEPA2ForVideoClassificationV-JEPA 2 ViT-g/16 at 384 pixels under the paper's attentive probe.
vjepa2_vit_giant_cls→ VJEPA2ForVideoClassificationV-JEPA 2 ViT-g/16 at 256 pixels under the paper's attentive probe.
vjepa2_vit_huge→ VJEPA2ModelThe released V-JEPA 2 ViT-H/16 backbone at 256 pixels.
vjepa2_vit_huge_cls→ VJEPA2ForVideoClassificationV-JEPA 2 ViT-H/16 under the paper's attentive probe.
vjepa2_vit_large→ VJEPA2ModelConstruct the released V-JEPA 2 ViT-L/16 backbone at 256 pixels.
vjepa2_vit_large_cls→ VJEPA2ForVideoClassificationV-JEPA 2 ViT-L/16 under the paper's attentive probe.
Weights
VJEPA2ViTGiant384WeightsFPC64 ViT-g/16 representation weights at 384 pixels.
VJEPA2ViTGiantWeightsFPC64 ViT-g/16 representation weights at 256 pixels.
VJEPA2ViTHugeWeightsFPC64 ViT-H/16 representation weights at 256 pixels.
VJEPA2ViTLargeWeightsFPC64 ViT-L/16 representation weights at 256 pixels.