This core problem remains unsolved. The solution presented in the article with Human In The Loop and some skill-magic such as "Write principles, not rules etc." is unsatisfactory because it offers no guarantees whatsoever. I find it difficult to harness agents into deterministic workflows which need to produce reliable outcomes.
You want workflows where the human gates are properly placed, not a "software factory" that you never place eyes on.
Where an LLM with “run lint every time a subagent completes their task” might do it 99% of the time, a hook tied to the worker ending will run 100% of the time.
I founded a company called Aide where our goal was to help support teams reliably deploy customer-facing agents without worrying about poor interactions. The first problem we needed to solve was making them deterministic and eliminate the variance that comes naturally with base models.
Getting them to always adhere to brand policy, eliminate hallucination, and stay grounded in data was a fun challenge. Proud to say that we’ve devised a solution that runs well and it’s worked out quite nicely in compliance-heavy and regulated environments.
This is not at all related to the problematic circular financing stuff that I suppose you're trying to allude to.
It’s an open tent. IMO it’s good to know these types of thinking exist, so you can think ahead about how to respond to them quickly if they ever come up in the real world, without being stunlocked by the sheer volume of obvious nonstarters.
I don’t see OPs post as a “complaint about capitalism” but rather a concern about centralisation and monopoly. Such discussions make HN a better place.
The I in AGI stands for "IPO".
wired: this company could have been a prompt
How is the invalidation handled in Warp? Is it actually self-improving, or just better retrieval?
Btw it’s a fair challenge, I will probably hit it one day so something like a “compact” skill for decision log would be useful.
They do get to it later on - casually mentioning it opens a PR for changes that a human accepts in the middle of a paragraph somewhere. Even there though, the focus isn't on risk mitigation (eg: against embedded prompt injection) but rather just "check if this is a good idea or not".
It seems to me that, in engineering terms, identifying self-modifying agent loops and managing the risk of them is going to be one of the key aspects that will emerge in best practises for how these systems are eventually designed.
> In our series, we highlight how startups are transforming their industries with AI.
I'm sorry, but I don't see Warp transforming shit.
Most impressive.
Ah, now that's a name I haven't heard in many moons. Looks like they found their niche... editing markdown files?
(now I understand why cloudflare, had to?, change the name of their warp)
aka "make no mistakes" in various random Markdown files that "self-improving agents" are free to ignore at any moment
Enterprise.
Look at their case studies, this is why organisations pay for Warp.
https://www.warp.dev/enterprise
I think you would be pretty much sued if you faked your testimonials especially in Enterprise.
We do so in a chaotic and unreadable way since unlike Hitler we could not afford editors for "Our Struggle".
Do you mean they found Claude's output, full of smoking-guns and honest caveats which are all load-bearing and genuinely bite -- they found it "low-quality" by default? Wow. Color me surprised. /s