ijepa_huge_14(pretrained: bool = False, overrides: object = {})I-JEPA with a ViT-H/14 encoder.
Model Size
Parameters
pretrainedbool= FalseThe released
.pth.tar checkpoints are not redistributed here;
True raises.**overridesobject= {}Optional
IJEPAConfig field overrides.Returns
IJEPAModelContext encoder, target encoder and predictor, untrained.
Notes
Reference: Assran et al., arXiv:2301.08243, Table 1 — 79.3% ImageNet-1k linear probe after 300 epochs, 73.3% on 1% of ImageNet (Table 2), and the transfer results of Tables 3 and 4 (CIFAR-100 87.5, Places205 58.4, Clevr/Count 86.7).
A 14-pixel patch over 224 pixels is a 16×16 grid, the same token count as the /16 models at that resolution.
Examples
>>> from lucid.models import AutoConfig
>>> config = AutoConfig.from_pretrained("ijepa_huge_14")
>>> config.patch_size, config.grid_size, config.num_patches
(14, 16, 256)