Skip to content
daily-hour-news·

🛠️GRPO in 100 Steps Lifts a 350M Model From 22.6% to 29.7%

TL;DR

A Hugging Face walkthrough fine-tunes LiquidAI's LFM2.5-350M with GRPO and TRL, lifting IFStruct schema compliance from 22.6% to 29.7%. The whole run uses about 500 samples and 100 steps on a free-tier Colab GPU, with a LoRA touching just 1.66% of parameters.

A Hugging Face walkthrough fine-tunes LiquidAI's LFM2.5-350M with GRPO and TRL, lifting IFStruct schema compliance from 22.6% to 29.7%. The whole run uses about 500 samples and 100 steps on a free-tier Colab GPU, with a LoRA touching just 1.66% of parameters.

GRPO in 100 Steps Lifts a 350M Model From 22.6% to 29.7% — daily-hour-news

Key Points

1

Published Sep 3 by Leonie Monigatti, Ben Burtenshaw and Sergio Paniego

2

LoRA rank 16 trains roughly 6M parameters, 1.66% of the model

3

Three rewards weighted 1.0 / 0.5 / 2.0: JSON format, field count, schema validation

4

JSON pass rate climbs from 18.0% to 31.9%; YAML barely moves at 27.2% to 27.5%

5

Evaluation runs locally through llama.cpp over 2,000 IFStruct samples

Why It Matters

Schema compliance is a training target, not a prompting problem, and a free GPU run closes most of the gap to a model six times larger.

Quick Facts

GRPOTRLfine-tuningLoRAstructured outputsLFM2.5Hugging Face

Frequently Asked Questions

Why does this matter?

Schema compliance is a training target, not a prompting problem, and a free GPU run closes most of the gap to a model six times larger.

What happened?

A Hugging Face walkthrough fine-tunes LiquidAI's LFM2.5-350M with GRPO and TRL, lifting IFStruct schema compliance from 22.6% to 29.7%. The whole run uses about 500 samples and 100 steps on a free-tier Colab GPU, with a LoRA touching just 1.66% of parameters.

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,463 builders reading daily.

Also get