HF RL Explorer

Add a DeepSeek Sparse Attention (DSA) causal-LM family with MLA-shaped KV cache and a sibling GLM-MoE-DSA…

Add a DeepSeek Sparse Attention (DSA) causal-LM family with MLA-shaped KV cache and a sibling GLM-MoE-DSA…: a task in HF ML Bench v0 (Harbor dataset). Extend this transformers codebase to natively support a new decoder-only causal-LM architecture whose distinguishing feature is DeepSeek Sparse…

The task

Extend this transformers codebase to natively support a new decoder-only causal-LM architecture whose distinguishing feature is **DeepSeek Sparse Attention (DSA)**: a lightweight learned *indexer* module produces per-query scores over the key-value cache, a hard top-k selection picks the most relevant tokens, and the…

Part of AdithyaSK/HF_ML_Bench_v0.