vjepa2_vit_huge_cls(pretrained: bool = False, overrides: object = {})V-JEPA 2 ViT-H/16 under the paper's attentive probe.
Model Size
Parameters
pretrainedbool= FalseThe probe's own weights are not published;
True raises. The
backbone tags are reachable through vjepa2_vit_huge.**overridesobject= {}Optional
VJEPA2Config field overrides.Returns
VJEPA2ForVideoClassificationThe pretraining networks, the attentive pooler and a classifier.
Notes
Reference: Assran et al., arXiv:2506.09985, 2025. The released
classifiers read a frozen backbone through three self-attention blocks
and one learned query whose cross-attention carries no output
projection — the layout num_pooler_layers defaults to, checked
against the published ssv2 classifier to 4.3e-6.
num_classes defaults to 400, Kinetics-400's count.
Examples
>>> from lucid.models import AutoConfig
>>> config = AutoConfig.from_pretrained("vjepa2_vit_huge_cls")
>>> config.dim, config.depth, config.num_heads
(1280, 32, 16)
>>> config.num_pooler_layers, config.num_classes
(3, 400)