AdithyaSK/HF_ML_Bench_v0
HF ML Bench v0: Harbor dataset on Hugging Face with 16 tasks. Generated by Repo2RLEnv — turning real GitHub repositories into verifiable RL environments.
Tasks
- Introduce a ParallelismConfig dataclass that composes DDP / FSDP / HSDP / TP / CP into a single named torch DeviceMesh with the expected…
- Support dynamic (variable-length) batch sizes in BatchSamplerShard, including the even batches=True padding round; keep split batches…
- Add Flux.2 Klein text-to-image / image-to-image pipeline, transformer backbone, and VAE to diffusers.
- Add discrete-diffusion text generation to diffusers: a block-refinement scheduler, a LLaDA2 pipeline, and a confidence-aware training loss.
- Add a first-class conversion path in PEFT that distills non-LoRA adapters (LoKr, LoHa, and similar) into equivalent LoRA adapters over…
- Integrate the PVeRA parameter-efficient adapter (probabilistic vector-based random matrix adaptation) into PEFT alongside the existing…
- Bridge PEFT ↔ transformers weight-conversion contract so LoRA adapters trained against one transformers model-key layout load cleanly…
- Guard against targeting nn.Parameters on the root module in LoRA by detecting the degenerate empty-path case and raising a descriptive…
- Allow multiple LoRA adapters to use target parameters on the same model, provided all such adapters target the same set of parameters —…
- Add a Moonshine (RoPE + SwiGLU) automatic-speech-recognition model with encoder-decoder architecture and Auto registration.
- Implement DINOv2 with Registers as a new first-class model in the transformers library, including config, full modeling stack (Model…
- Add MetaCLIP 2 to transformers: new model package with text/vision/joint towers, image-classification head, Auto registrations, and a…
- Add a DeepSeek Sparse Attention (DSA) causal-LM family with MLA-shaped KV cache and a sibling GLM-MoE-DSA architecture to the library.
- Add chunked LM head for memory-efficient log-prob computation for AsyncGRPOTrainer
- Add a memory-efficient chunked cross-entropy loss mode to TRL's SFT trainer that avoids materializing full batch x seq x vocab logits.
- Add a picklable, length-shaped cosine-scaled correctness reward for RLVR trainers, reusing math-verification for correctness.