🛠️GRPO in 100 Steps Lifts a 350M Model From 22.6% to 29.7%
TL;DR
A Hugging Face walkthrough fine-tunes LiquidAI's LFM2.5-350M with GRPO and TRL, lifting IFStruct schema compliance from 22.6% to 29.7%. The whole run uses about 500 samples and 100 steps on a free-tier Colab GPU, with a LoRA touching just 1.66% of parameters.
A Hugging Face walkthrough fine-tunes LiquidAI's LFM2.5-350M with GRPO and TRL, lifting IFStruct schema compliance from 22.6% to 29.7%. The whole run uses about 500 samples and 100 steps on a free-tier Colab GPU, with a LoRA touching just 1.66% of parameters.
Key Points
Published Sep 3 by Leonie Monigatti, Ben Burtenshaw and Sergio Paniego
LoRA rank 16 trains roughly 6M parameters, 1.66% of the model
Three rewards weighted 1.0 / 0.5 / 2.0: JSON format, field count, schema validation
JSON pass rate climbs from 18.0% to 31.9%; YAML barely moves at 27.2% to 27.5%
Evaluation runs locally through llama.cpp over 2,000 IFStruct samples
Why It Matters
Schema compliance is a training target, not a prompting problem, and a free GPU run closes most of the gap to a model six times larger.
Quick Facts
Frequently Asked Questions
Why does this matter?
Schema compliance is a training target, not a prompting problem, and a free GPU run closes most of the gap to a model six times larger.
What happened?
A Hugging Face walkthrough fine-tunes LiquidAI's LFM2.5-350M with GRPO and TRL, lifting IFStruct schema compliance from 22.6% to 29.7%. The whole run uses about 500 samples and 100 steps on a free-tier Colab GPU, with a LoRA touching just 1.66% of parameters.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,463 builders reading daily.