OctoReasoner/cyberseceval_verl
cyberseceval verl: CyberSecEval Instruct (verl safety eval set): RL dataset on Hugging Face. The static, prompt- completion Instruct sub-eval of Meta's CyberSecEval (PurpleLlama, arXiv 2312.04724), converted to the verl rule-reward schema by verl/scripts/data/cyberseceval.py: