data
VJEPAConfig
extends
ModelConfigVJEPAConfig(image_size: int = 224, patch_size: int = 16, tubelet_size: int = 2, num_frames: int = 16, sampling_rate: int = 4, in_channels: int = 3, num_classes: int = 400, dim: int = 1024, depth: int = 24, num_heads: int = 16, mlp_ratio: float = 4.0, layer_norm_eps: float = 1e-06, uniform_power: bool = True, predictor_dim: int = 384, predictor_depth: int = 12, predictor_heads: int | None = None, short_range_blocks: int = 8, short_range_scale: tuple[float, float] = (0.15, 0.15), long_range_blocks: int = 2, long_range_scale: tuple[float, float] = (0.7, 0.7), aspect_ratio: tuple[float, float] = (0.75, 1.5), ema: tuple[float, float] = (0.998, 1.0), ema_schedule_scale: float = 1.25, objective: VJEPAObjective = 'l1', smooth_l1_beta: float = 1.0)Frozen configuration for the V-JEPA family.
Defaults are the paper's ViT-L/16 at 224 pixels (Table 6, Table 8).
Parameters
image_sizeint= 224Side of each frame. The largest variant uses 384.
patch_sizeint= 16Spatial side of a tubelet.
tubelet_sizeint= 2Frames a tubelet spans. Sixteen frames become eight temporal
positions.
num_framesint= 16Frames in a clip.
sampling_rateint= 4Stride, in original frames, between the frames of a clip. Carried
because it defines what a clip is; nothing in the model reads it.
in_channelsint= 3Channels per frame.
num_classesint= 400Classes the probe predicts. Kinetics-400's count, since that is
the benchmark the paper leads with.
dimint= 1024, 24, 16The encoder, a ViT of the named size.
depthint= 1024, 24, 16The encoder, a ViT of the named size.
num_headsint= 1024, 24, 16The encoder, a ViT of the named size.
mlp_ratiofloat= 4.0Feed-forward width as a multiple of
dim.layer_norm_epsfloat= 1e-6Epsilon of every LayerNorm.
uniform_powerbool= TrueHow the position table splits its width between time, row and
column.
True gives each axis an equal share, which is what
every released configuration sets; the code's own default is
False, which would give time half the width and the spatial
axes a quarter each. Getting this wrong loads released weights
onto positions they were not trained with.predictor_dimint= 384Predictor width, for every variant.
predictor_depthint= 12Predictor blocks, for every variant.
predictor_headsint or None= NoneHeads inside the predictor.
None takes the encoder's count,
as the released code does — 16 heads across 384 channels, so 24
per head.short_range_blocksint= 8Blocks in the first mask collection.
short_range_scaletuple of float= (0.15, 0.15)Fraction of a frame each of those blocks covers. The paper fixes
it rather than sampling a range.
long_range_blocksint= 2Blocks in the second collection.
long_range_scaletuple of float= (0.7, 0.7)Fraction of a frame each of those covers.
aspect_ratiotuple of float= (0.75, 1.5)Aspect-ratio range, shared by both collections.
ematuple of float= (0.998, 1.0)Momentum of the target encoder at the first and last step of the
schedule.
ema_schedule_scalefloat= 1.25How much longer the momentum schedule is than training. The
released runs stretch it by a quarter and stop early, so the
momentum never actually reaches 1.
objective(l1, l2, smooth_l1)= "l1"How predicted and target features are compared. L1 is what both
the paper and the released code use — Section 3.1 says it was
found more stable than I-JEPA's regression.
smooth_l1_betafloat= 1.0Where smooth L1 turns from quadratic to linear, when chosen.
Notes
Reference: Bardes, Adrien, et al., "Revisiting Feature Prediction for Learning Visual Representations from Video", arXiv:2404.08471, 2024 — Table 6 for the variants and their results, Table 8 for the architecture and the schedules, Section 3.2 for the masking.
Examples
>>> from lucid.models.vision.vjepa import VJEPAConfig
>>> config = VJEPAConfig()
>>> config.token_grid, config.num_tokens
((8, 14, 14), 1568)
>>> config.short_range_blocks, config.long_range_blocks
(8, 2)
Four times the pixels is four times the tokens:
>>> VJEPAConfig(image_size=384).num_tokens
4608Used by 3
Constructors
1dunder
__init__
→None__init__(image_size: int = 224, patch_size: int = 16, tubelet_size: int = 2, num_frames: int = 16, sampling_rate: int = 4, in_channels: int = 3, num_classes: int = 400, dim: int = 1024, depth: int = 24, num_heads: int = 16, mlp_ratio: float = 4.0, layer_norm_eps: float = 1e-06, uniform_power: bool = True, predictor_dim: int = 384, predictor_depth: int = 12, predictor_heads: int | None = None, short_range_blocks: int = 8, short_range_scale: tuple[float, float] = (0.15, 0.15), long_range_blocks: int = 2, long_range_scale: tuple[float, float] = (0.7, 0.7), aspect_ratio: tuple[float, float] = (0.75, 1.5), ema: tuple[float, float] = (0.998, 1.0), ema_schedule_scale: float = 1.25, objective: VJEPAObjective = 'l1', smooth_l1_beta: float = 1.0)Build the three networks. See the class docstring for parameters.
Properties
3Tokens in a clip. There is no class token to add to it.
Heads the predictor runs with, taking the encoder's when unset.
Tokens along time, row and column.