class
VJEPA2ForVideoClassification
extends
ImageClassificationModelClassificationHeadMixinVJEPA2ForVideoClassification(config: VJEPA2Config)V-JEPA 2 read through the probe the paper evaluates with.
Parameters
configVJEPA2ConfigFrozen configuration;
num_classes sizes the classifier and
num_pooler_layers the probe's tower.Attributes
vjepa2VJEPA2ModelThe pretraining networks. Its target encoder is frozen, and that
is the tower this reads.
poolerModuleThe attentive probe: self-attention blocks, then one learned query
cross-attending the feature map.
headnn.LinearThe classifier on the pooled vector.
Notes
Registered under image-classification because that is the task
this zoo has; the input is a clip rather than an image, which is what
the class name says. Averaging tokens and fitting a linear map would
measure something else — the paper prices the difference in points.
Used by 2
Constructors
1Instance methods
1forward(x: Tensor, labels: Tensor | None = None)Classify a clip and optionally return cross-entropy loss.
The backbone is the frozen target encoder, which is what the
paper's probe reads. Freezing is a property of those parameters,
not of this call: wrapping the pass in no_grad would also
stop a caller who unfroze them deliberately.