Training a model this small on them is distillation
When models of this size were last studied seriously such corpora did not exist
X discussion: https://x.com/GregoryDiamos/status/2096873745420075020?s=20
I added some of the main points to the thread so they are easier to read.