Add a DeepSeek Sparse Attention (DSA) causal-LM family with MLA-shaped KV cache and a sibling GLM-MoE-DSA…
Add a DeepSeek Sparse Attention (DSA) causal-LM family with MLA-shaped KV cache and a sibling GLM-MoE-DSA…: a task in HF ML Bench v0 (Harbor dataset). Extend this transformers codebase to natively support a new decoder-only causal-LM architecture whose distinguishing feature is DeepSeek Sparse…
The task
Extend this transformers codebase to natively support a new decoder-only causal-LM architecture whose distinguishing feature is **DeepSeek Sparse Attention (DSA)**: a lightweight learned *indexer* module produces per-query scores over the key-value cache, a hard top-k selection picks the most relevant tokens, and the…
Part of AdithyaSK/HF_ML_Bench_v0.