FPC64 ViT-L/16 representation weights at 256 pixels.
Converted from facebook/vjepa2-vitl-fpc64-256: separate query, key
and value projections fused, and the context encoder mirrored into the
EMA target. Reproduces its source to 1.7e-5 relative over the
encoder and the masked predictor — the closest of the four.