data
VJEPA2Config
extends
ModelConfigVJEPA2Config(image_size: int = 256, patch_size: int = 16, tubelet_size: int = 2, num_frames: int = 64, in_channels: int = 3, num_classes: int = 400, dim: int = 1024, depth: int = 24, num_heads: int = 16, mlp_ratio: float = 4.0, layer_norm_eps: float = 1e-06, qkv_bias: bool = True, init_std: float = 0.02, use_silu: bool = False, wide_silu: bool = True, use_rope: bool = True, rope_base: float = 10000.0, uniform_power: bool = False, drop_rate: float = 0.0, attn_drop_rate: float = 0.0, drop_path_rate: float = 0.0, predictor_dim: int = 384, predictor_out_dim: int | None = None, predictor_depth: int = 12, predictor_heads: int = 12, predictor_mlp_ratio: float = 4.0, predictor_num_mask_tokens: int = 10, zero_init_mask_tokens: bool = True, predictor_return_all_tokens: bool = False, num_pooler_layers: int = 3, normalize_targets: bool = True, objective: str = 'l1', ema: tuple[float, float] = (0.99925, 0.99925), ema_schedule_scale: float = 1.25)Frozen architecture configuration for V-JEPA 2.
The defaults are the released ViT-L/16 checkpoint geometry. The public
factories select the other paper variants by overriding dim, depth
and num_heads; callers may still use a small explicit override for
tests or experiments.
Parameters
image_sizeint= 256Height and width of each video frame.
patch_sizeint= 16Spatial side of one patch.
tubelet_sizeint= 2Number of adjacent frames in one token.
num_framesint= 64Clip length used by the released pretraining configuration.
dimint= 1024, 24, 16Width, block count and head count of the video encoder.
depthint= 1024, 24, 16Width, block count and head count of the video encoder.
num_headsint= 1024, 24, 16Width, block count and head count of the video encoder.
use_ropebool= TrueUse the released three-axis rotary attention geometry.
predictor_dimint= 384Width of the latent predictor.
predictor_depthint= 12Number of predictor blocks.
predictor_headsint= 12Predictor attention heads.
predictor_mlp_ratiofloat= 4.0Predictor feed-forward expansion. The ViT-g encoder ratio is wider
than its predictor ratio in the released checkpoint.
predictor_num_mask_tokensint= 10Learned mask-token slots used by the official mask generator.
num_pooler_layersint= 3Self-attention blocks the attentive probe runs before its query
attends. Every released classifier checkpoint carries three.
ematuple of float= (0.99925, 0.99925)Target-encoder momentum at the start and end of the schedule.
The released pretraining configuration holds it constant, which
is where V-JEPA 2 departs from V-JEPA 1's ramp.
ema_schedule_scalefloat= 1.25Stretches the momentum horizon past the run's own length, as the
released
ipe_scale does.Used by 3
Constructors
1dunder
__init__
→None__init__(image_size: int = 256, patch_size: int = 16, tubelet_size: int = 2, num_frames: int = 64, in_channels: int = 3, num_classes: int = 400, dim: int = 1024, depth: int = 24, num_heads: int = 16, mlp_ratio: float = 4.0, layer_norm_eps: float = 1e-06, qkv_bias: bool = True, init_std: float = 0.02, use_silu: bool = False, wide_silu: bool = True, use_rope: bool = True, rope_base: float = 10000.0, uniform_power: bool = False, drop_rate: float = 0.0, attn_drop_rate: float = 0.0, drop_path_rate: float = 0.0, predictor_dim: int = 384, predictor_out_dim: int | None = None, predictor_depth: int = 12, predictor_heads: int = 12, predictor_mlp_ratio: float = 4.0, predictor_num_mask_tokens: int = 10, zero_init_mask_tokens: bool = True, predictor_return_all_tokens: bool = False, num_pooler_layers: int = 3, normalize_targets: bool = True, objective: str = 'l1', ema: tuple[float, float] = (0.99925, 0.99925), ema_schedule_scale: float = 1.25)Properties
3Number of tokens in a full configured clip.
Target width emitted by the predictor.
Number of tubelet positions along time, height and width.