class
VJEPAForVideoClassification
extends
ImageClassificationModelClassificationHeadMixinVJEPAForVideoClassification(config: VJEPAConfig)V-JEPA under the paper's attentive probe.
Parameters
configVJEPAConfigFrozen configuration;
num_classes sizes the classifier.Attributes
vjepaVJEPAModelThe pretrained networks. Its target encoder is frozen.
poolerModuleThe attentive pooler: one learned query over the feature map.
headnn.LinearThe classifier on the pooled vector.
Notes
Reference: Bardes et al., arXiv:2404.08471, Section 4.3. The paper evaluates a frozen backbone with this probe rather than with a linear map on averaged tokens, and reports the gap at 17 points on Kinetics-400 — so a linear probe compared against its tables would be measuring something else.
Registered under image-classification because that is the task
this zoo has; the input is a clip rather than an image, which is what
the class name says.
Examples
>>> import lucid
>>> from lucid.models.vision.vjepa import (
... VJEPAConfig, VJEPAForVideoClassification)
>>> config = VJEPAConfig(
... image_size=32, patch_size=8, tubelet_size=2, num_frames=4,
... dim=24, depth=1, num_heads=2, predictor_dim=12,
... predictor_depth=1, num_classes=10)
>>> model = VJEPAForVideoClassification(config).eval()
>>> model(lucid.rand(2, 4, 3, 32, 32)).logits.shape
(2, 10)Used by 2
Constructors
1Instance methods
1forward(x: Tensor, labels: Tensor | None = None)Classify clips from the frozen feature map.
Parameters
Returns
ImageClassificationOutputLogits (B, num_classes), and the loss when labels came.