ijepa_base_16(pretrained: bool = False, overrides: object = {})I-JEPA with a ViT-B/16 encoder.
Model Size
Parameters
pretrainedbool= FalseNo weights are published for this size;
True raises.**overridesobject= {}Optional
IJEPAConfig field overrides.Returns
IJEPAModelContext encoder, target encoder and predictor, untrained.
Notes
Reference: Assran, Mahmoud, et al., "Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture", CVPR 2023 (arXiv:2301.08243), Table 1 — 72.9% ImageNet-1k linear probe after 600 epochs. Its predictor is 6 blocks deep, the shallowest of the four.
Examples
>>> from lucid.models import AutoConfig
>>> config = AutoConfig.from_pretrained("ijepa_base_16")
>>> config.dim, config.depth, config.num_heads, config.predictor_depth
(768, 12, 12, 6)
>>> config.num_patches
196