V-JEPA
10 memberslucid.models.vision.vjepaV-JEPA — feature prediction for video.
Bardes et al., 2024 (arXiv:2404.08471). A clip is cut into tubelets, two
collections of blocks hide most of them through the whole depth of the
clip, and a predictor says what an averaged copy of the encoder would
report about what is hidden. The image version is
lucid.models.vision.ijepa; the transformer they share is in
vision/_common.
Bardes, Adrien, et al. "Revisiting Feature Prediction for Learning Visual Representations from Video." [arXiv:2404.08471](https://arxiv.org/abs/2404.08471), 2024.
V-JEPA learns from video by predicting features of what it cannot see. A clip of frames becomes tubelet tokens; a mask hides most of them; the encoder reads what is left, and a predictor must produce the output of an exponentially averaged encoder at the hidden positions:
Two mask collections, not one. Eight blocks covering 15% of the frame each, and two covering 70%, both running the full depth of the clip in time. Their union hides about 90% of the tokens. A mask that leaves a tube visible through time makes the task trivial --- the answer is in the next frame --- so both collections extend through it, and the model is asked the two questions at once, with a learned mask token of its own for each.
The target encoder sees everything; the context encoder does not. Masking here removes tokens rather than attending around them, so the context encoder's sequence is genuinely shorter, while the averaged encoder reads the whole clip and its output is normalised before any block is taken from it. No gradient flows into it at all.
Prediction in feature space is what makes the video tractable. A generative video objective must settle every detail that differs between plausible continuations; this one may drop whatever the averaged encoder has already learned to drop. Frozen, with an attentive probe on top, the result reaches 82.0% on Kinetics-400 and 71.4% on Something-Something-v2 --- the latter being the benchmark that punishes models which only recognise appearance.
Classes
VJEPAConfig4 methodsFrozen configuration for the V-JEPA family.
VJEPAForVideoClassification2 methodsV-JEPA under the paper's attentive probe.
VJEPAModel9 methodsV-JEPA: a context encoder, an averaged target encoder, a predictor.
VJEPAOutput1 methodsWhat one pretraining step produced, for both mask collections.
Functions
vjepa_huge_16→ VJEPAModelV-JEPA with a ViT-H/16 encoder at 224 pixels.
vjepa_huge_16_384→ VJEPAModelV-JEPA with a ViT-H/16 encoder at 384 pixels.
vjepa_huge_16_384_cls→ VJEPAForVideoClassificationV-JEPA ViT-H/16 at 384 pixels under the paper's attentive probe.
vjepa_huge_16_cls→ VJEPAForVideoClassificationV-JEPA ViT-H/16 under the paper's attentive probe.
vjepa_large_16→ VJEPAModelV-JEPA with a ViT-L/16 encoder at 224 pixels.
vjepa_large_16_cls→ VJEPAForVideoClassificationV-JEPA ViT-L/16 under the paper's attentive probe.