74 pointsby Xeophon3 hours ago7 comments
  • embedding-shapean hour ago
    LLM-generated code that seemingly went without much review or design is always such an interesting dive into just how bloated you can make code. Multiple files are close to 10K LOC, one file contains a switch statement that has so many case statements it spans more than 1000 lines.

    I guess it depends on the model you're trying to use, but seems most of them prefer smaller codebases, they work a lot better with less code, which kind of makes sense. With that in mind, I'd probably aim for something way smaller to bootstrap a self-improving agent. Then I'd use this "Prime Agent" as an example to my self-improving agent for what it should not evolve to.

  • riddlemethatan hour ago
    I built one of these RLM harnesses and a local MCP server along with logging, memories, and project rules based on directories. It worked great for a while but the foundational models have largely caught up to the point where they don't need this harness anymore. At least for my use cases. I can basically just store context in .md in the directories we work out of together and accomplish what I need.
  • supermdguyan hour ago
    It'll be really interesting when they run RL training on the harness self-improvement loop. I've tried using LLMs for harness engineering, but it often creates too much bloat that weighs things down in the end. Guessing it's just not something the models are tuned to do by default.

    Curious if anyone's tried using RL for harness engineering? I think we're still pretty far away from the optimal harness, especially when it comes to long-context memory management.

  • axusan hour ago
    • 43 minutes ago
      undefined
    • EarlKing24 minutes ago
      For everyone downvoting: It's literally a story about the creation of mankind's first artificial general intelligence, Prime Intellect, and the consequences of that discovery.
  • staredan hour ago
    It is impressive that it (almost) saturates ARC-AGI-3, https://x.com/PrimeIntellect/status/2085087000764568010.

    I am curious - how does it fare for other benchmarks, or everyday programming?

    • tintoran hour ago
      PrimeIntelect is not on official ARC-AGI-3 leaderboard: https://arcprize.org/leaderboard
      • staredan hour ago
        Good to know!

        Is it that it wasn't accepted yet, or are there issues with how it was run?

        • noahbp37 minutes ago
          It’s a self-improving harness, and ARC-AGI-3 is explicitly a few-shot benchmark. It’s likely that it gave itself more than the maximum number of tries to learn the games, or even hardcoded the answers.

          There’s a lot of improvement to be had from the benchmark harnesses, but sometimes, like with ARC-AGI-3, the limitations are intentional.

  • woah2 hours ago
    Might actually try this
  • 2 hours ago
    undefined