But in this programming-only example study, the refactor changes the functionnality very seriously, and it has to go deep in the core of the system, find all the traps, plan everything. And the agent succeeds without human code review. Code review was impossible because of the complexity and size of the task.
Code review was not necessary, nor useful, nor essential, and even practically impossible.
Don't get me wrong, this is impressive and capabilities do still seem to be on a generally upward trend with nothing "hitting a wall" despite all the parroting of that phrase, but there's a huge gap between the first time a machine manages something impressive enough to document, and that machine becoming so good at that task that humans need not apply.
"Is this the end of human code review?" != "the end of human code review"
Now since the problem in this study is more complex than 90 to 99.5% of what a normal ticket is, the question makes sense. The agent does not have knowledge of the project, at the begining of each session, and manages to handle something more complicated than virtually anything a developper has to do. The question holds IMO.