~~[EDIT: you need to put "only for macOS" in way more prominent places, all over -- that offends my soul greatly and may Linus frown upon you all]~~ [EDIT2: I was mistaken!]
This all looks really solid. That said, two remarks:
1. The integration with OS LSPs is quite fun and commendable. Is it possible that the diagrams might get their own LSP, someday? Or is better just staying as direct TS callsites?
2. The choice of the word "IDE" seems like it might get you in trouble, given the small "cannot edit files" detail. Any comments on the decision there, as opposed to, say... "brainstorming tool"? Or hell, "[architectonic] harness"?
3. The psuedocode "semantic diff" thing is an incredible idea, wow. Props there.
4. This language kinda concerns me: "how the requirements that you set were implemented". In my highly-arbitrary development flow, it ideally goes `idea -> spec/reqs -> plan -> test -> impl -> eval -> land -> review`, and this kind of tool seems explicitly targeted towards just the second and third with some partial coverage of their neighbors on either side. More concretely: by adding implementation, don't you lose a powerful specificity selling point and now have to compete with all full harnesses?
5. Suggesting "GPT-6 Luna and Claude Opus 5.5" is presumably a typo? Cause the equivelant of Opus 5.5 is Astra, and even then not really.
1. Yeah, we've thought about this too - at the minimum, we're going to implement a vibe-codeable extension system so that you can add your own diagram types without rebuilding the app. definitely hopeful for some sort of common schema or lsp-shaped thing in the future
2. this is a good point and is something we've considered and struggled with. we ultimately settled on the term 'IDE' because we've found that most people use vscode to review code nowadays, almost exclusively (so it makes it a bit easier to draw the comp in your mind). we're also strongly considering adding an editing feature, but not sure what the exact shape of it is, so decided to go in favor of not shipping it yet - editing tends towards a conductor / superset shape of product. maybe we could try "canvas" instead of "ide" or something? will mull it over more.
4. Yeah, this is definitely tricky. This is why we didn't end up shipping edits as part of this release. Part of the solution here, we think, is tighter integration with the "spec" part of the lifecycle - where whiteboard makes it easier to understand if a spec (like a formal proof) is extensible, generalizable, etc. will mull on this more.
5. yes that's a typo! we're fixing right now - we meant "GPT 6 Sol" (basically, fast TPS, don't need the limits of intelligence really).
Edit: we do have a Fedora Linux build out if that's what you use! releasing stuff on every distro requires some care, so please let us know what you'd want to see it on (re: appimage lol) Edit 2: (5) is fixed! thank you
I'll be honest that my initial belief it was for OSX exclusively left me feeling indignant, so knowing I was just mistaken (based on the link up top, tbf) turns me around completly. I do in fact run Fedora, so I'll be trying this ASAP!
I feel Fedora+Debian+Ubuntu+Arch covers all but the long tail of devs, based on vibes alone? You might get bullied if you don't support Nix, but you'll probably be bullied by them anyway lol so no advice on navigating those waters.
Along similar lines, I also don't know what we're calling tools like T3 or Superset. They're basically harnesses for harnesses.
Whiteboard is targeted at almost the opposite problem of the ADE (which is targeted to context switching) - having a dedicated tool to help you understand & participate in the development process in places where humans are high leverage.
maybe also useful: https://x.com/ThePrimeagen/status/2101869827266596973?s=20
Looking at the code the "no" seems to relate to the context expiring, so you wouldn't want to wait more if the context already expired, you'd want to stop. Is there a reason that label exists?
I'm pretty wary of LLM development tools hallucinating and wasting my time, is that whats happening in the lease broker example?
Real diagrams are all linked to code, so hallucinations don’t really happen in practice (hallucinated architecture really bothers me too!)
Will update shortly with an actual gif of the app. Sorry about that!
The cool thing is even though this tool addresses a few, use case for reducing cognitive load, they happen so frequently that they add up.
I think they key with cognitive load is that the agents often produce a lot of trace/docs etc, in the end only a small % of the traces really are important to the final changes, because concise changes are typically very small and self contained.
One thing i see this becoming important with is maintaining internal tooling built using this. I maintain an internal docs system that helps me do designing before i build code, that docs system is completely vibe coded and i can add features to it very fast, but now its all grown up. A challenge then is can i understand just enough about a new proposed change to approve it? That is key to the doc system not becoming a burden in of itself, while maintaining a tight core feature set.
yes, totally. i think this is mostly 1 piece of the broader puzzle. we found that, esp. for complex changes where i need to spend my brainpower anyways:
> A challenge then is can i understand just enough about a new proposed change to approve it?
whiteboard has been a powerful tool for us! i think much more of this needs to be instrumented as part of a larger system, as you said (we are thinking the same way btw: https://dev.fast/about/)
From the website demos i definitely think this is a clean interface, although I don't know how much better this is compared to some simple custom Mermaid format, which the agent can write as artifact files and present to users. Zooming out, this app seems like 1 feature (a MCP with a GUI attached to it) rather than an entire product.
Also, I don't know if asking the agent to write specific code changes into the plan is a good idea. I think maybe that a "plan -> approve -> write code" would let the agent write higher quality code than "plan which contains code -> approve". But maybe you can make it work when combined with some specific prompting marking the code as clearly work-in-progress and subject to change, and that the agent should surface any parts implemented differently relative to the plan to the user, etc.
1. "is this a feature" - it could be! in fact, we will expose this as an MCP UI next so that you can view the info directly in Codex Desktop or Superset/Conductor/Emdash for example. that aside, we found that the big things that matter for us are: (1) good code navigation (diagram/spec -> code), (2) beautiful diff viewing, and (3) visualizing agent traces as they connect to code. we found that these problems were hard enough, and enough folks that were using platforms that didn't easily map to these requirements - e.g. TUIs like claude code - that a dedicated product that was just focused on these problems exclusively makes sense.
2. "using whiteboard for plan mode": hmm, i think our wires are crossed a bit here. how people mostly use whiteboard today is:
plan -> approve -> agent codes -> use whiteboard to explain the code.
(or just omit the plan phase as a formal artifact -> just emit a plan + code together, like a golang design draft [A]).
we are exploring an explicit "put the plan in whiteboard first" mode (there's a scratchpad feature that's experimental right now), but it's definitely not ready for prime time yet.
[A] we were heavily influenced by golang's practice of "design drafts" as a way of scaling engineering velocity, e.g.: https://go.googlesource.com/proposal/+/master/design/draft-i... (thanks to Russ Cox, the legend)
I guess it's interesting and useful for now, but I don't think people are going to work at the code level much longer.
In my opinion current coding agents + automatic review systems are already at superhuman reliability during the implementation phase (as in they will not fail something in the plan during implementation and not tell you about it, so there's no need to look at the actual code beyond maybe a cursory glance). I literally just use plan mode + CC's /code-review in each task so it's not like I'm doing anything special. So I think the main human interaction surfaces to target in the future will be in the planning process.
yes, agreed. we're working on more stuff in that direction (a plan / scratchpad mode), but what i personally like the most is eliminating / shrinking the plan/review gap.
i think reviewing a plan without an implementation doesn't feel that useful anymore, at least to me, because key tradeoffs often only surface during implementation that effect the top-level spec.
in some sense, the code writing process is just a cheap effort which makes the spec better and more thorough?
An easy upgrade (ime) is to be intentional about a process, move the planning artifact to a file, use multiple research/propose/review sessions to dial it in. Still tuning my vibes for when to add in some actual exploratory implementation elements, because there's always something you didn't foresee when getting to the actual implementation, while also not having them implement the solution as a "plan" in markdown
Yeah, I think we need something like that as well. I am actually working on an virtual artifact filesystem in my orchestrator to enable this. So agents can create a persistent, versioned plan artifact separate from the codebase (maybe a HTML) and iterate it alongside the user, much like what ChatGPT/claude.ai can already do but for a coding agent. Then you'd need to define a process and get the agent to follow it, but that's much easier and mostly a mix of prompt and orchestration primitives.
> exploratory implementation elements
This is a good point, I've ran into a lot of instances as well where my agents in plan mode would like to explore something but can't because of permissions. I wonder if there should be some kind of system like a "experiment subagent" to handle it.
Commit it to git, it's not far off from an llm-wiki
I have no orchestration primitives, just a skill tied to a .design/*.md
Unless we consider opencode sometimes using a subagent as a primitive? Maybe the problem is leaving the clankers to their own devices for too long/much
I intentionally block almost every tool for the design/review agents, letting them "do" things is a distraction. I will use the build agent and tell it what to do if I need that experiment. I don't want to have to read through the wasted tokens a bunch of dumb bots burned through to create walls of markdown. They go on way too many side quests
Maybe I’m missing something.
Edit: just updated the GitHub + marketing site to reflect this!
YAT (Yet Another Tool) is something I use (sometimes pejoratively, sometimes as a reality), arising from the general trend and burnout in the developer tools space
One of the nice things about Ai is that it can deal with all that and I don't have to go through YAT docs and code to figure out how to use it
start brainstorming with their agent -> agent writes code -> agent writes Whiteboard session (connected to the underlying code) for them to iterate on.
(eliminates the middle plan mode phase)
(2) i don't think every piece of work fits in that way, so we're working on a "scratchpad" mode your agent can use for planning via versioned artifacts for architecture + behavioral specifications. Then it can update that plan when the implementation is complete! when this feature rolls out I think an "architect" would be using that feature more in collaboration with an engineer.
(the boundary b/w architect and engineer does seem a bit fuzzy to me, esp. these days, so hopefully we share a similar mental model)
AI note taking is a scourge on society and needs to go.
wrote a blog about this if you're interested! https://dev.fast/blog/youre-still-going-to-have-a-job-in-5-y...
I wrote some thoughts on this a while ago. They're not cleanly organized (sorry!) but it's my raw thinking on AI: https://nonograph.com/some-disorganized-thoughts-about-artif...
Also wrote this on the state of VC if interested: https://nonograph.com/write-some-software-give-it-away-for-f...
If you can create an aur it'd be awesome for the arch crowd :-)
Speaking from personal experience, I still find myself reaching for Whiteboard. It’s helpful when it’s critical for me as a developer to understand the implementation, which is certainly not every change!
In the future, we want to deliver a hosted product that addresses the N+1 concern you raised. The problem with the existing tools is that plans don’t stay up to date with what the agent decided to implement, and the back-and-forth that happens after the initial prompt isn’t captured. We believe a single whiteboard canvas can be used to capture not only a plan at design time, but what happens after.
You building something else and this scratches an itch?
That aside. I actually love this. Anything that helps with the "wtf did you just do?".
If I can still learn I will. If an agent can't reason about the changes made, then they are not good changes.
Sid of GitLab raised for Kilo, and GitLab is also of YC
https://capwolf.com/former-gitlab-ceo-launches-kilo-in-ai-co...
The headline says "IDE", but what you wrote here does not sound like an IDE, why would I want my agent calling an MCP / API to do the things your feature list suggests? What I'm seeing here would/could be better/replicated as a VS Code extension
I'm only interested in an "Integrated Developer Experience", winner takes all kind of thing, tool sprawl is out of hand
/facepalm
as a comparison, vanilla cursor / vscode is ~1GB and zed is ~400Mb.
in the future we will definitely rewrite this app as fully native and get it way, way down. in the meantime, we're working to get the size down in other ways (e.g. our diff viewer can definitely be optimized - it's 138Mb, yikes)
we've put in a lot of time and attention into the app especially and hope it shows in the details - e.g. the diff viewer, the rendering animations - we want development to feel human again while still enjoying the speed boost of agents
So many launches don't show what they are. This on the other hand - got it straight away from the animation.
Great job