42 pointsby cyfyifanchen6 hours ago14 comments
  • riskable4 minutes ago
    I always figured that the free/open weights models like qwen3.8:27b would perform just as well if not better than Claude's latest if you just fed it back into itself enough times. This project seems to prove that this is indeed the case.

    What I'd like to see now is how good it can get when you feed the micro models like qwen3.5:0.8b into itself to solve problems. Will it be like toddlers discussing neighborhood politics at a pretend tea party or will it actually get some decent results?

    Another game-changer (if this style works out): Just get a model like qwen3.8:27b onto one of those model-on-a-chip cards that makes it 1000x faster and see how fast it can go using the same method.

  • calebhwin44 minutes ago
    Dumb but honest question - do repos like this buy stars? How do they have thousands of stars with very little presence across HN/Reddit/X?
    • bitpusha few seconds ago
      An accusation wrapped as a genuine question. Bravo.

      Let me turn this around to you. Do you think your timelines on Reddit/X are indicative of the tech scene of {London,China,India,Indonesia}? What makes you think you know of every popular project out there?

  • julesrmsan hour ago
    It's all very glitzy, but I'm failing to understand how "One prompt in. One result out." is of any importance.

    This and other recent AI hype-fests all seem to be obsessed with making agents do more work unattended.

    But surely in the real world, anyone who's got a real product to make is going to want to steer what's happening. It's ridiculous to think that anyone with a deadline would write a prompt so perfect that they walk away for 4 days and come back to find the finished product ready to ship.

    If you really can write a prompt so complete and perfect that it needs nothing further, then any regular harness could probably also do the job. But if like normal people you need to try something, think about it, iterate, and repeat.. then you also just need a regular harness.

    • J03daSchm0an hour ago
      The Venn Diagram of people building the AI hype-fests and people building a real product is just two circles
    • hughwan hour ago
      Not saying you're wrong. But right now I'm trying to craft some skill prose to instruct an agent how to optimize a certain process based on my own heuristics, and it's failing. Maybe I can let a super agent divine the right skill prose, iterating to see what works. It's worth a try.
      • julesrmsan hour ago
        OK, but by the time you've worked out the correct prompt to tell your meta-agent how to iterate on optimising the right prompt for your actual agent, perhaps you could have just tried a few things and got it going!
        • conartist636 minutes ago
          Nice try, but there's a meta-meta-agent that just tries a few things and gets it going while the other agents optimize the prompt.

          Or put another way: "You can't fool me, it's turtles all the way down!"

  • ssddanbrown3 hours ago
    RSI here is for "recursive self-improvement", instead of a harness being built to help users with repetitive strain injury like I first thought when reading the post title
    • broodbucket2 hours ago
      We gotta do something about these clashing TLAs.
      • poloticsan hour ago
        Well for starters the use of the term "Recursive" is very dubious.

        This is as far as the eye can see all very iterative, there is no tail-call or anything fancy, it is loops. Also "Recursive" I feel kind of tries to imply the LLM's weights are being pushed around in some feedback, because to recurse you have to invoke the thing you're recursing into at its very start right? That would be reinforcement probably, and there is none of that in any such RSI so far, at least not public. Please someone contradict me with examples.

        So please: "ISI" for iterative self improvement is fine. Also "ISI" does not fit the (outdated!) Vernon Vinge "singularity" trope, and that is a good thing!

        • IanCalan hour ago
          I'm unsure.

          The example in the docs of improving nanochat is iterative. It's a looped process in one thing altering a second thing.

          What would be recursive is raven updating raven to make it better at doing things. For what I picture as RSI the important part would be that it's able to make itself better at doing things and better at improving itself.

          Now there's an "evolver" part that improves the harness over time but I don't know how far that goes or what scope it has to update things.

          I guess it's that if you have a function called "optimise" that takes functions and makes them better, calling optimise(my_process) is iterative regardless of how many times you do it. Calling optimise(optimise) is inherently different.

      • gpm2 hours ago
        Perhaps modelling acronyms in TLA+ will help?
        • daveguyan hour ago
          Well, we could add a + after any acronyms after they have a formal definition in TLA+. Except for TLA itself. That would need to be TLA++.
      • brookst2 hours ago
        A register is the obvious choice, but of course then you get a gold rush and people squatting on XPZ and stuff without even having invented an opaque term behind it.
      • peddling-brinkan hour ago
        I agree, we can’t have all these three letter agencies running amuck.
      • flexdan hour ago
        Time for FLAs?
        • daveguy42 minutes ago
          IIRC there are already FLAs. FLAs too. ROTFL.
    • wafflemaker2 hours ago
      A very popular yoga kata called sun salutations (available in various versions depending on your fitness/advancement level) helped me get rid of wrist and thumb RSI. It stretches and gently stresses these load bearing ;) joints, thus strengthening them.

      Hope this helps the last few folk before searching for RSI will be impossible due to the term being taken over.

      For those afraid of Satanism in yoga, there's also non satanic variants where you don't greet each other saying Namaste or say Shanti anywhere during the practice.

      • brookst2 hours ago
        As long as we’re burying things for archaeologists to find, let it be known that AGI once stood for “adjusted gross income”, a measure used for income taxes.
  • arminluschin3 hours ago
    This looks similar to https://paseo.sh/, if I understand correctly. I’ve recently tried it and liked it a lot. Would be nice to see a comparison. When’s the harness of harnesses of harnesses coming?
    • jfaat2 hours ago
      I'm using paseo heavily too. The main differences here seem to be around how opinionated Raven's orchestration is. They're providing agents, workflows, memory, skills, etc. Paseo gives you some orchestration tools but it's mostly letting the 'native' harnesses do the work. For my workflow, I'm interested in some of what they're doing here but I'm not in buying into their whole system (markdown for memory and calling it RSI, as someone else called out, is not doing it for me). The follow up actions and meta-harness tuning look pretty cool.

      They also don't ship an app which is one of the best parts of paseo. Then again Paseo's perf leaves a lot to be desired.

  • jeffnash2 hours ago
    Reminds me a lot of omnigent (which I am a huge fan of) with a persistent memory layer. Unlike omnigent's subagent threads, the DAG it uses to coordinate other harnesses doesn't look to be durable; I am curious as to whether this is by design or is a forthcoming feature, as this essentially makes or breaks my use case of long-running project-sized implementation sessions.

    In any event, it's great to see competition in this meta-harness space, which is likely one that none of the frontier labs will touch since it, by definition, would utilize their competitors' products.

    • marginalxan hour ago
      I'm curious if you have a few mins for feedback, what are the top 2 things here that omnigent does that is significantly better for you than latest cc/codex which can launch subagents, auto save memory of a project.

      I'm wondering that as these core tools continue to enhance and add these capabilities, how much of a benefit these meta harnesses actually provide.

    • lin7can hour ago
      [flagged]
  • fraywing44 minutes ago
    It feels like the only thing left people are building are harnesses? Harnesses of harnesses?
    • yuck3935 minutes ago
      Harnesses are only useful when frontier models are incapable of designing their own efficient interfaces with systems which they are improving at rapidly. This will be seen as a transitional artifact of a specific time in the development of general intellegent systems
  • mawadev2 hours ago
    Why do none of the projects it built work on my machine? Even the fps game looks cool but can't be downloaded? :0
  • Kuyawa2 hours ago
    That's one of the most beautiful readmes I've seen in my whole life. Threshold is stunning. I am so impressed even RSI got overshadowed
  • revexos3 hours ago
    Recursion has started. Here we go!
  • mpalmer3 hours ago
    Agents editing their prompts and writing memory to markdown files is not and will never be RSI.
  • aatd863 hours ago
    Interesting. Only thing, from someone who has built something similar, is that it moght tend to duplicate certain capabilities that those harnesses handle on their own. Some overlap is bound to happen.
    • brookst2 hours ago
      This, and also they have to chase changes in model tuning and capabilities. I have a nice little meta-harness focused on product development (requiremwnts, acceptance criteria) for Claude code, and when major new models come it it’s weeks before I can find and fix constraints that are no longer necessary + constraints that have become necessary.
  • PcChip3 hours ago
    I wish they had also compared it to omp and dsh, I’m curious how it stacks up
    • monkmartinezan hour ago
      I think the only way to do that is test yourself... otherwise, you are asking for bias. DSH's "cordis" is very interesting, I haven't tried it yet. I have spent lots of time with pi.dev, omp and hermes. Aspects of all harnesses are great, but there is always something that bugs me.

      If you have the chops to evaluate different harnesses, you have the chops to build one that is perfect for you.

  • cyfyifanchen5 hours ago
    [dead]