(OP, nice project! Sorry for my rant.)
Whenever I see these vibecoded apps, my move is to just do what you've described: point an agent at it and tell it to reproduce it (usually with my own little customizations). It's a bit absurd, but I don't trust that they vibecoded the app as well as "I" could lol.
I think the next generation of exploit will be putting up code all over the internet that does x, y and z and then when asked to one shot whatever that code does the agents bake in the exploit that was indirectly included in their training corpus.
Everything your building is now streaming through a third party who could steal it, the same way Amazon steals open source software and repackages it.
All you need to do is look at palentir's forward deployed engineers. Once you let a third party observe your processed, you're cloneable and replaceable _as an entity_.
I think we can look at our own human psychology as it developed into social groups. If our brains were obviously deterministic, it would be easy for people to prey upon us via structured inputs. Some level of randomness in our thought process is almost evolutionarily required to avoid becoming part of the larger super structure, similar to an ant colony.
So these AI companies are definitely more dangerous than just the "AI is super intelligence". They're basically a intelligence platform everyone is wiring into.
That's today. Tomorrow, they'll be run by an MBA which is the enshittification.
If you can make a facebook clone it doesn't mean you can steal facebook's lunch. there are exceptions of course.
It makes sense to not trust AI models with the ability to read and alter your screen for many many many reasons, but the developer of the tool knows that as well. A tool that lets your ai do something doesn't also let it do anything.
So, as always, it's a question of if you trust the software producer. You're giving them the ability to draw on your screen too. Would you have the same fear if there were no AI involved?
For stuff like this I imagine AI misbehavior as a novel failure trigger not a novel failure type. It can only do the harm that the software was able to do on its own already.
Brilliant. This is exactly the kind of content I seek when I visit Hacker News.
Truly art.
It's just the epitome of living in an era where code is so cheap to generate that we don't invest any effort into even fully fleshing out the idea. It's just "Claude, make me a program that draws arrows on screen". And HN, I guess, is dominated by two groups: one still stuck in the mentality of "code = effort" and trying to recognize that; and another stuck in the mentality of "AI = cool".
A big arrow to an interface element which ... has a label that explains what it does?
So...a label for a label?
Helping people deal with bad UX.
Everybody who has worked in a traditional corporate environment has encountered poorly tutorialized, poorly labeled, counterintuitive user interfaces where the only way to know how to get something done is to have somebody sit over your shoulder and tell you which buttons to click.
Oh, you want to figure out which OIM entitlements your coworker has that you don't to debug an access issue? You just need to know that you have to click the "compliance" button to get there. Why is it called "compliance" when what it actually does is let you export comparisons from the user tables? Because fuck you that's why. Memorize it.
An AI assistant that can read the procedures for you, or sometimes even read the source code, or the website code, can show you how to accomplish something that the UI on its own fails to do.
For the first experiment, Claude stepped me through how to edit its settings.json to remove warnings from its configuration: a little clunky as I had to respond to Claude's CLI prompts to get it to auto run.
In the second experiment, Claude walked me through how to use Apple's Compressor.app to speed up a video. It walked me through dragging the file into compressor and clicking on the appropriate UI elements to do this.
So, one use case for the bigarrow skill enables Claude to step users through a relatively complex UI to accomplish a task.
There are likely more efficient ways to accomplish the tasks at hand (e.g. ffmpeg for the second experiment), but this use case is interesting if using a GUI tool is required.
That would be a godsent for those of us with aging parents.
Right yeah? I mean I have freagin remote desktop installed on all my parent's devices. I am the agent for them every time they call frantically for help.
It's Time to redesign the apple genius bar for the modern aging boomer using this.
- rainbow dripping arrows
- angrily pointing arrows
- flame-surrounded text boxes with particle effects
- the ability for the agent to play airhorn.wav at max volume, overriding existing volume or mute settings
EDIT: ayy lmao https://github.com/franzenzenhofer/big-arrow-on-the-screen/p...
https://hyperties.org/cabinet/symelec/
PIXIE and FORTH on PDP-7 Cabinet Emulator with Type 340 Vector Graphics Display:
It's already got the power to be obnoxious at you
Maybe isoprophlex will add that to his PR!
Whenever I use web tools that don't have "deep linking" [0], I love to throw out this quote:
"Have we thought about using hyperlinks? They are these SUPER useful things that were invented in Switzerland back in the 1990s."
For web apps, WebMCP is a way where you can make a frontend way more discoverable to an LLM rather than that it is going to read all your code. I sometimes create this in my apps at home and sometimes in my apps at work and it makes LLM assistants way more usable. It also helps if they can highlight certain things of a visual.
We humans can do this too on paper. We can point at something, we can highlight something with a marker. Why not give an LLM these capabilities as well?
We built something to help with that [1].
Why does this need to be a skill? Literally just ask Claude to “call out actions in these screenshots with numbered stylized arrows” and you get the same thing without burning tens of thousands of tokens in context.
> burning tens of thousands of tokens in context.
~1400?
> Literally just ask Claude to “call out actions in these screenshots with numbered stylized arrows”
Is there a macos native thing it uses? Or are you talking about screenshots? This draws on your screen instead.
> "It's this window, not that one." You have 14 Chrome windows. The agent knows which one it means: --window "Google Chrome:Pull request". It even picks the right tab: --app "Google Chrome:Pull request".
> Remote help. "No, the other gear icon." Point at it instead of describing it.
What terminal / agentic workflows spawn the demonstrated dialogue boxes that require the User to click/reject the action? Aren't most such flows actually inline UIs? And finally, if the core issue is that these confirmation boxes are tied to the terminal that triggered them, which does not autofocus, what's the point of said "big arrow" that is also lurking behind without focus?
If the said "big arrow" automatically gains focus, isn't the real fix here to just make the dialogue boxes themselves gain focus automatically? Both are similarly disruptive anyway.
"Teach me to do this" kinda stuff. "Guiding agent" rather than "doing agent".
Honestly, @franze, if you're reading this, maybe update your screenshots to show an agent in tutorial mode on some complicated app?
To me this feels like it takes away from what the human is supposed to do (read, understand the consequences of the action, then.. consent or abort)
There is a reason your AI Agent won't automate these clicks for you
https://laughingsquid.com/yahoo-neon-billboard-in-san-franci...
Grim. If you're just there to click sudo buttons for the bot, you might as well give it root access and be done.
I absolutely loathe this phrasing now. I don’t even know what part of it is good or notable.
But I feel dumber just by looking at it. If this is how this will look like in 10 years, why to make desktop at all. Just connect mic and speaker to your PC or talk to the phone: 'I need new pair of socks. Order 10 for me. In your favourige color it will be 21.37. Should I charge your credit card?'.
It is not like most of the people enjoy computers. I am pretty sure they do not. They just need them to operate systems they need: government websites, banks, maps, restaurant menus etc. If some agent will do that for them, why bother looking at screen at all? Rich people have it with their own personal assistants.
% @(#)handy.ps
%
% Handy Pointer
% Copyright (C) 1989.
% By Don Hopkins. (don@brillig.umd.edu)
% All rights reserved.
https://donhopkins.com/home/archive/psiber/cyber/pointer.psPSIBER Space Deck and Pseudo Scientific Visualizer Demo:
https://youtu.be/_fqCeuue5Ac?t=213
The Shape of PSIBER Space: PostScript Interactive Bug Eradication Routines — October 1989:
https://medium.com/@donhopkins/the-shape-of-psiber-space-oct...
I found this awesome vibe-coded Pixel News Network project recently. How fucking cool is this?? Would never have been created otherwise. https://pnn.watch/
As a human stochastic parrot, don't you find it embarrassing and humbling to be so easily outdone by an LLM?