HF RL Explorer

Implement DINOv2 with Registers as a new first-class model in the transformers library, including config…

Implement DINOv2 with Registers as a new first-class model in the transformers library, including config…: a task in HF ML Bench v0 (Harbor dataset). Vision Transformers (ViTs) develop artifacts in their attention maps: certain image patches are repurposed as internal "registers" for computations…

The task

Vision Transformers (ViTs) develop artifacts in their attention maps: certain image patches are repurposed as internal "registers" for computations that are unrelated to their spatial content. This produces high-norm tokens in low-information background areas and corrupts downstream feature maps.

Part of AdithyaSK/HF_ML_Bench_v0.