I-JEPA
12 memberslucid.models.vision.ijepaI-JEPA — self-supervised pretraining that predicts representations.
Assran et al., CVPR 2023 (arXiv:2301.08243). A context block is encoded, and a narrow predictor must say what an exponential moving average of that encoder would report about four blocks it was never shown. Nothing is reconstructed in pixel space, which is the point.
Assran, Mahmoud, et al. "Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
I-JEPA asks a question about an image and grades the answer in representation space. One context block is encoded by ; four target blocks are encoded by a second network , and a predictor must produce the target encoder's output for each target block from the context alone, told only where each block is:
The target is a representation, not a pixel. A generative objective must account for every detail that distinguishes one plausible completion from another --- the exact texture of grass, the grain of a wall --- and spends capacity doing so. Predicting 's output lets the model drop whatever that encoder has already learned to drop. The paper's ablation is unusually clean: the same architecture trained against pixels reaches 40.7% on 1% of ImageNet, and against representations 66.9%.
The target encoder is the model's own past. is not trained; its weights are an exponential moving average of , with the momentum moving linearly from 0.996 to 1 over training. There is no gradient path into it at all, which is what stops the pair from agreeing on a constant --- the target moves only as fast as the average lets it.
The masking carries the difficulty. Four target blocks covering 15--20% of the image each, at aspect ratios from 3:4 to 3:2, with a context block of 85--100% from which every target region is then removed. Predicting a large connected region from a large connected region is what forces semantics; the paper's ablation against rasterised or small random masks (54.2% versus 15.5% and 17.6%) is a statement about the task, not the architecture.
No token. Evaluation average-pools the target encoder's patch outputs.
Classes
IJEPAConfig4 methodsFrozen configuration for the I-JEPA family.
IJEPAForImageClassification2 methodsA linear probe on I-JEPA's representation.
IJEPAModel8 methodsI-JEPA: a context encoder, an averaged target encoder, a predictor.
IJEPAOutput1 methodsWhat one pretraining step produced.
Functions
ijepa_base_16→ IJEPAModelI-JEPA with a ViT-B/16 encoder.
ijepa_base_16_cls→ IJEPAForImageClassificationA linear probe on I-JEPA ViT-B/16.
ijepa_huge_14→ IJEPAModelI-JEPA with a ViT-H/14 encoder.
ijepa_huge_14_cls→ IJEPAForImageClassificationA linear probe on I-JEPA ViT-H/14.
ijepa_huge_16_448→ IJEPAModelI-JEPA with a ViT-H/16 encoder at 448 pixels.
ijepa_huge_16_448_cls→ IJEPAForImageClassificationA linear probe on I-JEPA ViT-H/16 at 448 pixels.
ijepa_large_16→ IJEPAModelI-JEPA with a ViT-L/16 encoder.
ijepa_large_16_cls→ IJEPAForImageClassificationA linear probe on I-JEPA ViT-L/16.