← All posts

Geek Out Time: PPO vs GRPO on Google Colab: Two Ways to Align Language Models

May 2025·~926 words in full

Introduction

When it comes to aligning large language models (LLMs) with human values, reward modeling is key. Two notable methods are:

This geek-out time, we explore both, simulating how they score outputs without training, to make their differences clear and concrete.

Background

This is an excerpt — the full article continues on Medium.

Read the full article on Medium →

© 2026 Nedved Yang

Vibe-coded with AI + Next.js + Tailwind CSS

Singapore