HF RL Explorer

jucamohedano/oxford-pets-grpo-think

oxford-pets-grpo-think: Oxford-IIIT Pet — GRPO training data (structured reasoning): RL dataset on Hugging Face. GRPO (verl) training data for Oxford-IIIT Pet breed classification with a structured-reasoning prompt: the model emits a scratchpad tagging visible properties (HasProperty), parts…

jucamohedano/oxford-pets-grpo-think on the Hugging Face Hub