Skip to main content

Post-training a VLM for reasoning with GRPO using TRL

Authored by: Sergio Paniego 🚨 WARNING: This section requires substantial computational resources. A Colab run uses an A100 GPU. This section shows how to post-train a Vision Language Model (VLM) with GRPO. The implementation uses the Hugging Face Transformer Reinforcement Learning library (trl). We fine-tune Qwen2.5-VL-3B-Instruct on a subset of the lmms-lab/multimodal-open-r1-8k-verified dataset. Each example contains an image, a problem description, a solution, and a reasoning trace. We use this data format and the GRPO reward functions to train the model to reason toward the solution. fine_tuning_vlm_grpo_trl_diagram.png

1. Install dependencies

Install the libraries required for fine-tuning. Install trl from source because, at the time of writing, the official release does not include the VLM GRPO trainer.
Authenticate with your Hugging Face 🤗 account so you can save and share the trained model.

2. Load the dataset 📁

This section uses lmms-lab/multimodal-open-r1-8k-verified. The dataset contains 8,000 multimodal reinforcement learning examples for mathematical reasoning. GPT-4o generated the data. Each sample includes image, problem, solution, original question, and original answer. The data comes from this project. We use image and problem as inputs and solution as the output so the model learns to reason from images. To reduce training time, we use 5% of the dataset and split it into training and test sets. A full training run would use the complete dataset. Load and split the dataset.
Inspect the dataset structure.
Inspect one sample:
In addition to the problem and image columns, we include a custom system prompt that specifies the output format. The system prompt is adapted from DeepSeek R1. See this previous recipe for details. We convert each dataset sample to the conversation format expected by the GRPO trainer. Each conversation contains the system prompt, one image, and one problem description. We set padding_side="left" so each generated completion follows its prompt directly during training. GRPO requires this arrangement to compare token-level probabilities between preferred and rejected responses.
Inspect a converted example:
Remove the columns that are not needed for training.
Verify that the columns were removed.

3. Post-training the VLM with GRPO

The diagram below shows the main difference between PPO (Proximal Policy Optimization) and GRPO (Group Relative Policy Optimization): GRPO does not use a value model. See this further explanation for more detail. The training pipeline uses trl, the Hugging Face reinforcement learning library. We use the GRPOConfig and GRPOTrainer classes. Custom reward functions specify the behavior required by the training objectives. First, load Qwen/Qwen2.5-VL-3B-Instruct, a VLM developed by Qwen. A model with more parameters may produce better results. Other VLM projects with reasoning capabilities include: ppo_grpo.jpeg

3.1 Load the baseline model

Load the Qwen/Qwen2.5-VL-3B-Instruct baseline model.

3.2 Configure LoRA

Configure LoRA for training.

3.3 Load the reward functions

The reward can come from a pretrained reward model or from functions defined in code. The DeepSeek-R1 authors used an accuracy-based reward to evaluate whether a response is correct. They also used a format-based reward to require reasoning inside <think> </think> tags. See the Open R1 implementation for details. Here, we implement the rewards as Python functions. We use the following reward functions from the Open R1 implementation:
  1. Format enforcement: Requires the generation to use <think> </think> <answer> </answer> tags for reasoning.
  1. Solution accuracy: Compares the generated solution with the solution column in the dataset.

3.4 Configure the GRPO training parameters

Configure the GRPO training parameters. You can vary max_completion_length, num_generations, and max_prompt_length to compare training configurations. These parameter values fit the hardware limits of a Google Colab session. Larger models, more generations, and a larger, diverse dataset would allow more complete training. These changes may improve the rewards, especially for the second objective, and may improve reasoning performance.

3.5 Train the model 🏃

Configure the trainer and start training. Pass the model, training arguments, dataset, and the two reward functions defined above to the trainer. The diagram below shows the training procedure from the Open-R1 project. image.png
Train the model.
[307/307 1:48:37, Epoch 1/1]
StepTraining Loss
100.000000
200.000000
300.040757
400.000000
500.067859
600.000000
700.000000
800.000000
900.000000
1000.000000
1100.000000
1200.000000
1300.000000
1400.000000
1500.000000
1600.015370
1700.000000
1800.000000
1900.000000
2000.000000
2100.000000
2200.052498
2300.000000
2400.000000
2500.000000
2600.013925
2700.000000
2800.000000
2900.013467
3000.000000
The training metrics are available in TensorBoard on the [model page]((https://huggingface.co/sergiopaniego/Qwen2.5-VL-3B-Instruct-Thinking/tensorboard). The loss curve may appear irregular, but the reward increases during training. Save the results to your account 💾

4. Evaluate the model

Evaluate the trained model qualitatively.
Restart your session to release the resources used for training.
Use the test subset for evaluation. First, load the trained model and its processor.
Define a helper function that accepts a problem and an image and returns the model response. The response should contain a reasoning trace and a final answer.
Run the model on a test example.
The answer follows the format imposed by the reward functions: <think>reasoning</think><answer>solution</answer>. Compare it with the reference solution to determine whether it is correct.
The model now produces a reasoning trace with its answer. Check the inference time and generated token count to further assess its performance.

5. Further reading 🧑‍🎓

The following resources cover GRPO, reasoning, and VLMs: