data
IJEPAConfig
extends
ModelConfigIJEPAConfig(image_size: int = 224, patch_size: int = 16, in_channels: int = 3, num_classes: int = 1000, dim: int = 768, depth: int = 12, num_heads: int = 12, mlp_ratio: float = 4.0, layer_norm_eps: float = 1e-06, predictor_dim: int = 384, predictor_depth: int = 6, predictor_heads: int | None = None, num_target_blocks: int = 4, target_scale: tuple[float, float] = (0.15, 0.2), target_aspect: tuple[float, float] = (0.75, 1.5), context_scale: tuple[float, float] = (0.85, 1.0), min_keep: int = 10, allow_overlap: bool = False, ema: tuple[float, float] = (0.996, 1.0), objective: IJEPAObjective = 'smooth_l1', smooth_l1_beta: float = 1.0)Frozen configuration for the I-JEPA family.
Defaults are the paper's ViT-B/16 (Table 1, Appendix A.1).
Parameters
image_sizeint= 224Input resolution. The 448-pixel variant is the one exception.
patch_sizeint= 16Patch side;
image_size must be a multiple of it.in_channelsint= 3Image channels.
num_classesint= 1000Classes the linear probe predicts. Pretraining ignores it.
dimint= 768, 12, 12The encoder, which is a ViT of the named size.
depthint= 768, 12, 12The encoder, which is a ViT of the named size.
num_headsint= 768, 12, 12The encoder, which is a ViT of the named size.
mlp_ratiofloat= 4.0Feed-forward width as a multiple of
dim. ViT-g/16 is the one
variant where the paper uses a different ratio.layer_norm_epsfloat= 1e-6Epsilon of every LayerNorm.
predictor_dimint= 384Predictor width, for every variant (Appendix A.1; Table 14 measures
384 against 1024 and prefers it). The predictor is deliberately
narrower than the encoder.
predictor_depthint= 6Predictor blocks: 6 for ViT-B/16, 12 for the larger encoders.
predictor_headsint or None= NoneHeads inside the predictor.
None takes the encoder's count,
which is what the released code does — so a 16-head predictor 384
wide has 24-dimensional heads.num_target_blocksint= 4Target blocks predicted per image (Section 3). They may overlap
each other.
target_scaletuple of float= (0.15, 0.2)Fraction of the image each target block covers.
target_aspecttuple of float= (0.75, 1.5)Aspect-ratio range of a target block.
context_scaletuple of float= (0.85, 1.0)Fraction of the image the context block covers. Its aspect ratio
is 1 — not stated in the paper, fixed in the released code.
min_keepint= 10Patches a sampled block must exceed, or it is resampled. Not
stated; the released configs set it.
allow_overlapbool= FalseWhether the context may keep patches that fall inside a target
block. The paper removes them; the released code makes it a flag
and sets it False.
ematuple of float= (0.996, 1.0)Momentum of the target encoder at the first and last training step,
moving linearly between them.
objective(smooth_l1, l2)= "smooth_l1"How predicted and target representations are compared. ⚠️ The
paper states an loss; the released code — which
produced the released checkpoints — uses smooth L1. The default
follows the code, and
"l2" is the paper's.smooth_l1_betafloat= 1.0Where smooth L1 turns from quadratic to linear.
Notes
Reference: Assran, Mahmoud, et al., "Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture", CVPR 2023 (arXiv:2301.08243), Table 1 for the variants and Appendix A.1 for the masking and the momentum schedule.
Examples
>>> from lucid.models.vision.ijepa import IJEPAConfig
>>> config = IJEPAConfig()
>>> config.num_patches, config.grid_size
(196, 14)
>>> config.predictor_dim, config.num_target_blocks
(384, 4)
The predictor takes the encoder's head count when it is not given:
>>> IJEPAConfig(num_heads=16, predictor_heads=None).resolved_predictor_heads
16Used by 3
Constructors
1dunder
__init__
→None__init__(image_size: int = 224, patch_size: int = 16, in_channels: int = 3, num_classes: int = 1000, dim: int = 768, depth: int = 12, num_heads: int = 12, mlp_ratio: float = 4.0, layer_norm_eps: float = 1e-06, predictor_dim: int = 384, predictor_depth: int = 6, predictor_heads: int | None = None, num_target_blocks: int = 4, target_scale: tuple[float, float] = (0.15, 0.2), target_aspect: tuple[float, float] = (0.75, 1.5), context_scale: tuple[float, float] = (0.85, 1.0), min_keep: int = 10, allow_overlap: bool = False, ema: tuple[float, float] = (0.996, 1.0), objective: IJEPAObjective = 'smooth_l1', smooth_l1_beta: float = 1.0)Build the three networks. See the class docstring for parameters.
Properties
3Patches along one side of the image.
Patches in an image. There is no class token to add to it.
Heads the predictor runs with, taking the encoder's when unset.