HF RL Explorer

Add MetaCLIP 2 to transformers: new model package with text/vision/joint towers, image-classification head…

Add MetaCLIP 2 to transformers: new model package with text/vision/joint towers, image-classification head…: a task in HF ML Bench v0 (Harbor dataset). Add a new vision-language model, MetaCLIP 2, to the transformers library. MetaCLIP 2 is a CLIP-family multilingual worldwide model: same overall…

The task

Add a new vision-language model, MetaCLIP 2, to the transformers library. MetaCLIP 2 is a CLIP-family multilingual worldwide model: same overall bi-encoder shape as CLIP (independent text tower + vision tower, contrastive pretraining objective, projection heads yielding image and text embeddings), differing mainly in…

Part of AdithyaSK/HF_ML_Bench_v0.