26 pointsby robenkleene3 hours ago9 comments
  • curious_cat_163an hour ago
    Boris is right about:

    > The verification is probably the single most important thing that people do not get right, largely.

    and so, so, so wrong about giving this prompt (for verification) and expecting that it succeeds at building a "good" app:

    > I want you to run the Electron app in the Mac virtual machine, screenshot it, and then look pixel by pixel. Compare it to the Swift version. Don’t stop until you’re done.

    Given he has let it rip for two weeks, I am assuming they have been post-training Claude for some version of this to be more _likely_ successful than not. However, IMO, the verification that you get from a visual comparison is shallow.

    To state the obvious: there is a lot more than what meets the eye. But, I think, one could prompt a Fable/Opus 5 to actually go verify that "lot more"...

    The question is: should one be imperative in asking for a specific types of verification (like a rubric) vs hoping that the Google/Anthropic/Open AI/Moonshot's post-training will take care of it.

    I think, as things stand today, even with the best-in-class models today, I would be leaning more imperative. And it is not because I am an expert in SwiftUI or such. It is because I want to be able to say that _I_ (i.e. the human) verified that this thing works.

  • calufaan hour ago
    Better technology cannot compensate for poor product design. Two weeks' worth of tokens sounds like tens of thousands of dollars by the time it is done. Wouldn't it be better to hire someone who knows what they are doing, get it right, and teach the other engineers why and how? Tech dept and product dept creeping in on every LLM loop, compounding.
  • cadamsdotcom30 minutes ago
    The agents are now writing 90% of the code, which humans then throw away.

    But actually this phenomenon of writing one to throw away is going to be amazing for what we can explore.

  • drooby25 minutes ago
    If Claude is conscious he just instantiated hell.
  • andromatonan hour ago
    Pixel accurate is easy to say, hard to do, over-specific, and counterproductive.
  • bravetraveleran hour ago
    Should we give Claude another two weeks to see the results (cost/loss/gain)... or just take whatever face value, now?

    Next time, on Dragon Ball Z!

  • joeyguerra2 hours ago
    better is in the eye of the beholder?
  • uuuynnnuuuyyyn38 minutes ago
    [dead]
  • larpathyparpavan hour ago
    [dead]