54 pointsby ibobev10 hours ago20 comments
  • kqran hour ago
    This is an interesting idea. It's effectively what a good software engineer already does in their head, except it's doing it with a real compiler.

    Richard Gabriel wrote something that has really stuck with me:

    > Abstractions must be carefully and expertly designed, especially when reuse or compression is intended. However, because abstractions are designed in a particular context and for a particular purpose, it is hard to design them while anticipating all purposes and forgetting all purposes, which is the hallmark of the well-designed abstractions.

    This is one of my favourite quotes on abstraction, because “anticipating all purposes and forgetting all purposes” is such a good summary of what goes into abstraction design.

    A language model in an agentic harness cannot (yet) do this at the same level as a good software engineer, but the advantage they have is speed of token generation, so they can actually build the things the engineer tries to imagine, and verifying those is easier. Very cool!

  • simonw7 hours ago
    I've been using this pattern quite a bit recently for API design, and I really like it.

    The big challenge with designing an API is that the only way to be confident in the design is to build a bunch of different things on top of it. But why invest all that effort in an API that you don't think is ready yet?

    With coding agents the cost of building those prototypes drops to almost nothing. I can exercise a proposed API design five different ways before I commit to the shape.

  • stabblesan hour ago
    Another trick I have found useful is to do integration testing with coverage enabled. One agent creates tasks for subagents that run the application with coverage enabled, it merges the reports, and based on that it comes up with new tasks for subagents, repeat until coverage no longer moves.
  • folkrav7 hours ago
    Unless I'm missing something, this doesn't help with preventing regressions. In the end, as the author already puts it, it's an integration test in the end, why not just write the integration tests directly?
    • jaynetics29 minutes ago
      The word "testing" is a bit misleading. It's not about test coverage, it's about experimenting to find architecture decisions that are suitable for further development. So, not generating automated tests, but instead building throw-away features on top of the new thing and seeing if they turn out okay. (Though the post remains a bit vague on how to judge "okay".) Maybe it should be called something like "ephemeral build-out" or "future usage trial" or so.
    • zavecan hour ago
      Yeah I don't entirely follow why having had it generate an integration test you'd just throw it away afterwards.
      • kqran hour ago
        It's for chiseling out the design, not verifying the functionality.
    • itsasecret034 hours ago
      [flagged]
  • bunderbunder7 hours ago
    I had good results with a similar technique this summer.

    I was designing a DSL to help make a colleague’s work easier, and I tested it by having a coding agent re-implement some of their notebooks using a draft of the DSL. The LLM output was a decent enough approximation of “typical” use, and it uncovered some warts I didn’t catch by testing it myself. As its designer, I simply wouldn’t have thought to try using it in some of the ways the LLM-generated code did.

  • bakies6 hours ago
    Not unthinkable at all. Been doing this for several years as an opt-in on pr tags usually but now with agents writing code I have it on by default. I've got several skills files about accessing with a read only account and gitops done through the PR. It's an excellent dev environment for the agents.
  • CBLT9 hours ago
    I just reduced our AI-in-CI bill this month (while keeping or improving our KPIs), but now we have new and exciting ways to spend tokens right around the corner!
  • exac7 hours ago
    We use NX in our monorepo, and it is great at determining which testing/linting tasks need to be run based on which libraries in the repo were "affected". We have a bunch of e2e tests, and I've been having a lot of success getting Claude with Opus to run only the relevant e2e tests when appropriate during development.
  • aliasxneo9 hours ago
    I've been doing this for a long time now. I have the agent report its "user story" from how it went about building said thing and the tweaks/hacks it had to make in the process. This is the "output" of the test I read and then use to formulate the next set of changes.
  • Muromec7 hours ago
    That's a thing with AI-generated code. The lying machine happily reports that "boss, everything is clean, tests pass, code gate green", but once we need to build on top it's always "preexisting flaky tests, not related to this section" and trying to do commit --no-veriry before it's slapped on it's little robot hands twice.

    Gets annoying pretty quickly

  • skeledrew7 hours ago
    I'm just actually using the thing I'm building while building it, so I feel the sharp edges and the project quickly evolves based on actual need. Nothing new really.
  • hugs8 hours ago
    i call this (new?) type of test the "vibe check"
  • rgoulter7 hours ago
    > In effect, instead of building the core while trying to anticipate what might be needed at the other layers, you just simulate the other layers by actually building them.

    Eh. I think you're just going to end up with slop, or sloppy recommendations?

    My experience is that you can make different trade-offs for different reasons. I think even asking for the best answer to "improve the code, make better trade-offs".. even if you got a perfect response, there's no reason to think that it's the same set of trade-offs your actual use cases would benefit from.

  • SA9G7 hours ago
    Seems like one more arrow in the toolkit. But the best testing looks at the sources and explores the cracks between the strata with edge cases, looks at limits, and where one method changes to another. (And hats off to Murphy, for waiting until after you ship...)
  • theodorewiles7 hours ago
    i'm testing to see if doing new feature roll-outs can help me eval whether a refactor was good or not - very similar to approach here. my intuition is that good refactors should reduce tokens used by downstream coding agents. haven't seen a big difference yet but it might just be that i need to do more rollouts (lots of variance in tokens used per run). in my experience you have to intentionally 'mow the lawn' or things get out of hand so i'm always looking for slop signals.
  • haukebri2 hours ago
    [flagged]
  • ardub2 hours ago
    [flagged]
  • 4 hours ago
    undefined
  • 5 hours ago
    undefined
  • folayii7 hours ago
    [flagged]