Add MetaCLIP 2 to transformers: new model package with text/vision/joint towers, image-classification head…
Add MetaCLIP 2 to transformers: new model package with text/vision/joint towers, image-classification head…: a task in HF ML Bench v0 (Harbor dataset). Add a new vision-language model, MetaCLIP 2, to the transformers library. MetaCLIP 2 is a CLIP-family multilingual worldwide model: same overall…
The task
Add a new vision-language model, MetaCLIP 2, to the transformers library. MetaCLIP 2 is a CLIP-family multilingual worldwide model: same overall bi-encoder shape as CLIP (independent text tower + vision tower, contrastive pretraining objective, projection heads yielding image and text embeddings), differing mainly in…
Part of AdithyaSK/HF_ML_Bench_v0.