> Realized reasoning gains, hardware efficiency, and RL scaling remain to be established.
Yeah so the first entry in that list is - if I may be so bold - typically what you'd start with at a small scale _before_ writing up and publishing your "genius" idea. This is the usual crank with delusions of grandeur presenting something straightforward that he hasn't tested as though it were a working breakthrough.
Ironically the cost of testing such theories has fallen to an all time low given the capabilities of coding models. I wouldn't be surprised if a frontier model could one shot a test of this.
Edit: Thanks to your link I've now learned about the zigzag transformer which is yet another design that somehow feels like cheating reality. https://en.wikipedia.org/wiki/Zigzag_transformer