ijepa_huge_16_448(pretrained: bool = False, overrides: object = {})I-JEPA with a ViT-H/16 encoder at 448 pixels.
Model Size
Parameters
pretrainedbool= FalseThe released
.pth.tar checkpoint is not redistributed here;
True raises.**overridesobject= {}Optional
IJEPAConfig field overrides.Returns
IJEPAModelContext encoder, target encoder and predictor, untrained.
Notes
Reference: Assran et al., arXiv:2301.08243, Table 1 — 81.1% ImageNet-1k linear probe, the paper's best, and 77.3% on 1% of ImageNet (Table 2); 87.1% when fine-tuned on all of it (Table 15).
Four times the pixels is four times the tokens: 784 patches against 196, which is where the accuracy comes from and what it costs.
Examples
>>> from lucid.models import AutoConfig
>>> config = AutoConfig.from_pretrained("ijepa_huge_16_448")
>>> config.image_size, config.num_patches
(448, 784)