"Standard PPO in RLHF often suffers from training instability and high sensitivity to hyperparameters. I built Tpo-torch to provide a clean, modular PyTorch implementation of Target Policy Optimization (TPO) for more stable alignment. Would love to hear thoughts from anyone working on LLM post-training!"