It’s great to see someone focusing pretraining data and not just post training. It's probably a harder problem.