Presumably they only have empirical results and if so, wouldn't you need to train multiple successive models to even test such a hypothesis? Even then, it's not like an inductive proof. An improvement in one generation does not imply an improvement in the next (unless you define improvement to mean exactly that).