244 pointsby franze5 hours ago45 comments
  • dr_kiszonka4 minutes ago
    I take this opportunity to shame GitHub for completely ignoring the mobile experience in their own Android app. The project's README is pages upon pages of largely blank space. In general, the app does not render mermaid diagrams and does not allow for zooming in, so smaller pictures are unusable. Even if you access pictures directly in a repo's source, you still can't zoom in. iPython Notebooks, which are extremely common in data science, are an "unsupported file type" and are not rendered.

    (OP, nice project! Sorry for my rant.)

  • hn87264 hours ago
    I tried to read the "Does it need Screen Recording or Accessibility?" part, but it's slopped to the point I have no clue what it's trying to say. But if it can draw on top of permission prompts, what's stopping it from drawing box that hides the "decline" button and changing the "approve" button copy?
    • causal3 hours ago
      I wonder if someday we will get to the point where Github repos are just markdown files describing the project and then you just let your own agent implement it because why the hell would I trust your agent's implementation?
      • theropost38 minutes ago
        Isn't it already kind of like that? Except the agent reads the code as if it's markdown. There's really no difference anymore, is there?
      • nater5000an hour ago
        Seriously. I imagine this will be the case soon enough in some shape or form.

        Whenever I see these vibecoded apps, my move is to just do what you've described: point an agent at it and tell it to reproduce it (usually with my own little customizations). It's a bit absurd, but I don't trust that they vibecoded the app as well as "I" could lol.

        • Abimelex38 minutes ago
          A 100%! Just imagine you could describe every solution to a problem just using language! Of cause you would to make sure to avoid ANY missunderstanding, but therfore you could invent a language specified for avoiding ambiguity. I could imagine just to reduce the words to a very small corpus so everybody can remember it and have a very strict grammar so a program can effortlessly check correctness of its sentences.
      • fennecfoxy30 minutes ago
        Eh I think it'll be more like Gibson's defensive ICE in that ICE is "my swarm of agents scans your code for nasties, then compiles it from source".
        • pydry12 minutes ago
          First you'd have to stop the agents from flagging hundreds of irrelevant nasties.

          I think the next generation of exploit will be putting up code all over the internet that does x, y and z and then when asked to one shot whatever that code does the agents bake in the exploit that was indirectly included in their training corpus.

      • cyanydeez3 hours ago
        if all your models are in the cloud, why would you trust anything your agent builds
        • jerf2 hours ago
          That sounds cynical today, but that's just because the AI models are currently outrunning enshittification. I'm already pondering personal plans about what to do when that turns around. We haven't seen enshittification yet that is going to be like the enshittification of AI. It may even deserve a new term of its very own, it's going to be such a big problem. The AI companies are leaving a lot of value on the table to entice us on to their systems but at some point that's going to turn around.
          • cyanydeez2 hours ago
            The existential danger exists today: you're providing your entire business toolchain to models in the cloud, owned by businesses who are for profit entity. Even if the model itself doesn't care about your data, the business model does.

            Everything your building is now streaming through a third party who could steal it, the same way Amazon steals open source software and repackages it.

            All you need to do is look at palentir's forward deployed engineers. Once you let a third party observe your processed, you're cloneable and replaceable _as an entity_.

            I think we can look at our own human psychology as it developed into social groups. If our brains were obviously deterministic, it would be easy for people to prey upon us via structured inputs. Some level of randomness in our thought process is almost evolutionarily required to avoid becoming part of the larger super structure, similar to an ant colony.

            So these AI companies are definitely more dangerous than just the "AI is super intelligence". They're basically a intelligence platform everyone is wiring into.

            That's today. Tomorrow, they'll be run by an MBA which is the enshittification.

            • jack_ppan hour ago
              software businesses are not like watches (recently saw a youtube video about fakes being made virtually identical at 10% of the price).

              If you can make a facebook clone it doesn't mean you can steal facebook's lunch. there are exceptions of course.

    • hannasanarionan hour ago
      The fact that the utility doesn't give it the ability to do that?

      It makes sense to not trust AI models with the ability to read and alter your screen for many many many reasons, but the developer of the tool knows that as well. A tool that lets your ai do something doesn't also let it do anything.

      So, as always, it's a question of if you trust the software producer. You're giving them the ability to draw on your screen too. Would you have the same fear if there were no AI involved?

      For stuff like this I imagine AI misbehavior as a novel failure trigger not a novel failure type. It can only do the harm that the software was able to do on its own already.

    • SwtCyber2 hours ago
      Nothing stops it, except the fact that the agent is already executing arditrary code in your shell. If it's malicious, it'll just steal your shh keys directly instead of bothering with button masking
    • smugglerFlynn37 minutes ago
      Don't tell me you are reading readmes with your eyes in 2026. Blasphemy!
    • tkdb3 hours ago
      ...and there we have it. Slop assumes a verb form.
  • tangotayloran hour ago
    "It is an arrow, so we spent an unreasonable amount of time on how it looks."

    Brilliant. This is exactly the kind of content I seek when I visit Hacker News.

    Truly art.

    • socializer42 minutes ago
      Have you actually looked at this? In the first screenshot, two of the five arrows are obviously and badly misaligned. The LLM-generated text may be saying one thing, but there's clearly zero effort spent on... anything. There's no punchline, there's no aesthetic angle, there's no conceivable purpose.

      It's just the epitome of living in an era where code is so cheap to generate that we don't invest any effort into even fully fleshing out the idea. It's just "Claude, make me a program that draws arrows on screen". And HN, I guess, is dominated by two groups: one still stuck in the mentality of "code = effort" and trying to recognize that; and another stuck in the mentality of "AI = cool".

  • usrbinbash3 hours ago
    SO the point of this is ... what exactly?

    A big arrow to an interface element which ... has a label that explains what it does?

    So...a label for a label?

    • hannasanarionan hour ago
      It seems silly but I think there's a good use case for this:

      Helping people deal with bad UX.

      Everybody who has worked in a traditional corporate environment has encountered poorly tutorialized, poorly labeled, counterintuitive user interfaces where the only way to know how to get something done is to have somebody sit over your shoulder and tell you which buttons to click.

      Oh, you want to figure out which OIM entitlements your coworker has that you don't to debug an access issue? You just need to know that you have to click the "compliance" button to get there. Why is it called "compliance" when what it actually does is let you export comparisons from the user tables? Because fuck you that's why. Memorize it.

      An AI assistant that can read the procedures for you, or sometimes even read the source code, or the website code, can show you how to accomplish something that the UI on its own fails to do.

    • mistersquidan hour ago
      I ran two experiments which require the `claude` CLI tool be installed.

      For the first experiment, Claude stepped me through how to edit its settings.json to remove warnings from its configuration: a little clunky as I had to respond to Claude's CLI prompts to get it to auto run.

      In the second experiment, Claude walked me through how to use Apple's Compressor.app to speed up a video. It walked me through dragging the file into compressor and clicking on the appropriate UI elements to do this.

      So, one use case for the bigarrow skill enables Claude to step users through a relatively complex UI to accomplish a task.

      There are likely more efficient ways to accomplish the tasks at hand (e.g. ffmpeg for the second experiment), but this use case is interesting if using a GUI tool is required.

    • voidUpdate3 hours ago
      It's so your agent can make a big arrow on the screen saying "click this, human" so that it can keep doing things
    • ghm21802 hours ago
      Glad you asked. Pointing my aging aunt to the right place on the screen to click without having to take control of her laptop, of course.
  • arshxyz4 hours ago
    The README is geared towards technical people (complete with the HN screenshot) but when I see a tool like this all I can think of is how helpful this would be for my mom when I'm trying to tell her how to download and print a document over the phone
    • conception4 hours ago
      Honestly the giant arrow annotation Zoom has makes it worth any amount of money compared to the competition.
  • lbreakjai4 hours ago
    I would pay good money for something like this on iPad. It wouldn't even need to be agent-driven, just a big "I want to make a bank transfer" button, that would launch the correct app and guide through the interface.

    That would be a godsent for those of us with aging parents.

    • ghm21803 hours ago
      > That would be a godsent for those of us with aging parents.

      Right yeah? I mean I have freagin remote desktop installed on all my parent's devices. I am the agent for them every time they call frantically for help.

      It's Time to redesign the apple genius bar for the modern aging boomer using this.

  • antonyragleap25 minutes ago
    Clever approach for agent debugging. Visual cues > logs when multiple agents run. How do you handle overlapping annotations?
  • isoprophlex4 hours ago
    Literally unusable as it is. Some minimal extra features this would need:

    - rainbow dripping arrows

    - angrily pointing arrows

    - flame-surrounded text boxes with particle effects

    - the ability for the agent to play airhorn.wav at max volume, overriding existing volume or mute settings

    EDIT: ayy lmao https://github.com/franzenzenhofer/big-arrow-on-the-screen/p...

    • htrp4 hours ago
      This is where the agent says did not understand sarcasm coded and shipped features
      • DonHopkins3 hours ago
        Last time I accidentally said something sarcasticly over-ambitious to an LLM, it shipped this popup callout tooltip feature on a PDP-7 Type 340 vector graphics display emulator that shows you the meaning of the drawing you're pointing at, as well as the address of the instruction that drew it.

        https://hyperties.org/cabinet/symelec/

        PIXIE and FORTH on PDP-7 Cabinet Emulator with Type 340 Vector Graphics Display:

        https://www.youtube.com/watch?v=lo8kdY-5i6c

    • dihinbuttan hour ago
      W developer
    • thih94 hours ago
      At the risk of stating the obvious - let's not do that, the goal of this repo is to be useful and not to give agents the power of the `<blink>` tag.
      • voidUpdate3 hours ago
        > ""I need you, and you're making coffee." --say reads the sign aloud. Your Mac will literally call you back to your desk."

        It's already got the power to be obnoxious at you

      • koalacola4 hours ago
        Oh dear, they were making a joke.
        • yen2234 hours ago
          if only there was a way to make a subtle thing obvious
          • isoprophlex4 hours ago
            such as... angry flaming rainbow textboxes and arrows?
            • ale423 hours ago
              I thought that the dripping rainbow ones were enough. Maybe you have to ask for rainbow unicorns flying on the screen.
      • jaapz2 hours ago
        the repo is one big joke, of course they should add this
      • DonHopkins3 hours ago
        For the humor impared, it would also be useful to have a colorful animated "WHOOSH" overlay with sound effects for every time a deadpan joke goes over your head. ;)

        Maybe isoprophlex will add that to his PR!

      • ipsod4 hours ago
        under_construction.gif
        • isoprophlex3 hours ago
          just submitted the airhorn PR; ~second rainbow arrow slop grenade incoming~ BOOM slop cannon fired
  • alexpotato2 hours ago
    > Arrows have existed since roughly the Paleolithic.

    Whenever I use web tools that don't have "deep linking" [0], I love to throw out this quote:

    "Have we thought about using hyperlinks? They are these SUPER useful things that were invented in Switzerland back in the 1990s."

    0 - https://en.wikipedia.org/wiki/Deep_linking

  • cyberjunkie4 hours ago
    I'm just as impressed by this as any other LLM-generated project.
  • LoneRanger102442 minutes ago
    I think this makes a lot of sense for tutorials, but permission prompts need different treatment. Besides pointing to a button, the assistant should explain what clicking it will do, so the user can make their own decision.
  • vessenes5 hours ago
    Interesting. When I read the headline I imagined this would be a sort of thinking trace booster -- letting the agent focus its own attention on different parts of the screen. But this is cool in a different way. I bet agentic harnesses would find it useful for communicating with other agents / themselves as well.
  • alansaber4 hours ago
    This might be goofy, but it underscores that there's potential for more visual agent UIUX than reading off a sidebar/opening modals.
  • melvinroest4 hours ago
    My message to the world is that LLMs should be able to point anything they see in the application they're in or even the whole computer (if you give it that kind of access).

    For web apps, WebMCP is a way where you can make a frontend way more discoverable to an LLM rather than that it is going to read all your code. I sometimes create this in my apps at home and sometimes in my apps at work and it makes LLM assistants way more usable. It also helps if they can highlight certain things of a visual.

    We humans can do this too on paper. We can point at something, we can highlight something with a marker. Why not give an LLM these capabilities as well?

  • FinnLobsien4 hours ago
    This could be great for documentation. Screenshots in docs are frequently useless because they show me a screen and say "click X" where I still have to search X visually. And I could just to dhat in the other tab I have open.
    • peaxkl3 hours ago
      It doesn’t make sense that you have to read a whole article and then still search for the buttons in the UI afterwards. And with longer articles, you always have to keep the article open next to your product to follow the whole flow.

      We built something to help with that [1].

      [1] https://www.happysupport.ai/en/in-app-messaging

  • m-s-y3 hours ago
    Genuine question…

    Why does this need to be a skill? Literally just ask Claude to “call out actions in these screenshots with numbered stylized arrows” and you get the same thing without burning tens of thousands of tokens in context.

    • IanCal3 hours ago
      There is a skill, but you don't have to use it, though you'd have to explain how to use the app.

      > burning tens of thousands of tokens in context.

      ~1400?

      > Literally just ask Claude to “call out actions in these screenshots with numbered stylized arrows”

      Is there a macos native thing it uses? Or are you talking about screenshots? This draws on your screen instead.

    • ballofrubber13 hours ago
      With a skill claude will know when to use it without you specifically prompting it.
  • ghm21803 hours ago
    Man, Ive lost count of How many times have I had to repeat this over the phone to my parents; The repo has all the punch lines in the README

    > "It's this window, not that one." You have 14 Chrome windows. The agent knows which one it means: --window "Google Chrome:Pull request". It even picks the right tab: --app "Google Chrome:Pull request".

    > Remote help. "No, the other gear icon." Point at it instead of describing it.

  • satyanash5 hours ago
    Am I missing something here?

    What terminal / agentic workflows spawn the demonstrated dialogue boxes that require the User to click/reject the action? Aren't most such flows actually inline UIs? And finally, if the core issue is that these confirmation boxes are tied to the terminal that triggered them, which does not autofocus, what's the point of said "big arrow" that is also lurking behind without focus?

    If the said "big arrow" automatically gains focus, isn't the real fix here to just make the dialogue boxes themselves gain focus automatically? Both are similarly disruptive anyway.

    • inanutshellus5 hours ago
      The first example (of HN) is the one that feels like it has the most potential to me.

      "Teach me to do this" kinda stuff. "Guiding agent" rather than "doing agent".

      Honestly, @franze, if you're reading this, maybe update your screenshots to show an agent in tutorial mode on some complicated app?

    • ravila43 hours ago
      I think this would’ve been very handy to me a couple years ago when I was learning to use Blender and asking LLMs for help performing certain actions like “how do I display the normals of all the vertices in my mesh?” I spent a lot of time trying to figure out which button the model was talking about.
  • pimlottc3 hours ago
    Even the example arrows in the first screenshot are wonky and unnaturally weird.
  • flr034 hours ago
    That would have been handy 20 years ago to point to that one valid 'Download' button.
    • Maken3 hours ago
      How could you tell apart the legit arrow pointing to the download button from all the fake ones?
  • tilemarch2 hours ago
    An arrow on the screen solves “where do I click”; it doesn’t solve “should this happen”

    To me this feels like it takes away from what the human is supposed to do (read, understand the consequences of the action, then.. consent or abort)

    There is a reason your AI Agent won't automate these clicks for you

  • erickhill21 minutes ago
    I don't know why exactly but something about the arrows feel repulsive and oddly gross.
  • harrouet3 hours ago
    There is no limit to burning tokens :)
  • TekMol5 hours ago
    Swift, Shell, Python and Objective-C

    Does one need 4 programming languages to draw something on a mac?

    • sitzkrieg5 hours ago
      welcome to zombocom. err i mean modern HN :-(
  • xyzsparetimexyz4 hours ago
    Seems like a pretyu useful way to help infants use desktop computers
    • ipsod4 hours ago
      Have you met users?
  • ForHackernews5 hours ago
    "and they keep hitting the same wall, the part that only a human may do"

    Grim. If you're just there to click sudo buttons for the bot, you might as well give it root access and be done.

    • franze5 hours ago
      Claude refuses to do certain actions (enter passwords, change security settings, create new accounts on external services) even in Yolo mode running as sudo. (I tested it all on its own mac machine)
      • gwerbin5 hours ago
        You can add custom auto-mode classifier rules and even disable the built-in ones, if you want to live on the edge like this.
        • jaapz2 hours ago
          isn't `--dangerously-skip-permissions` enough?
          • sitkack36 minutes ago

                --norms-taboos-off
    • 5 hours ago
      undefined
  • ex-aws-dude4 hours ago
    If you can’t even take the time to understand what you’re clicking why even go through the formality of “approving”
  • amelius3 hours ago
    Because AI can paint pelicans on bicycles quite well, and not arrows?
  • lapestenoire5 hours ago
    I love it.
  • okasakian hour ago
    Look at what Windows and Mac users need to mimic a fraction of our power
  • Retr0id4 hours ago
    Reminds me of something from Idiocracy (2006)
  • dihinbuttan hour ago
    I love this
  • intended2 hours ago
    > big-arrow-on-the-screen (bigarrow): one small macOS CLI and an agent skill. Click-through, never steals the focus, gone by itself. MIT.

    I absolutely loathe this phrasing now. I don’t even know what part of it is good or notable.

  • npodbielski2 hours ago
    seems useful for most of the folks that just want to do stuff without going through (over)complicated UIs...

    But I feel dumber just by looking at it. If this is how this will look like in 10 years, why to make desktop at all. Just connect mic and speaker to your PC or talk to the phone: 'I need new pair of socks. Order 10 for me. In your favourige color it will be 21.37. Should I charge your credit card?'.

    It is not like most of the people enjoy computers. I am pretty sure they do not. They just need them to operate systems they need: government websites, banks, maps, restaurant menus etc. If some agent will do that for them, why bother looking at screen at all? Rich people have it with their own personal assistants.

  • DonHopkins4 hours ago
    Ha ha, I love it! I wrote a pointing hand annotation overlay in PostScript in 1989 for NeWS and the PSIBER Space Deck's Pseudo Scientific Visializer:

      %  @(#)handy.ps
      %
      %  Handy Pointer
      %  Copyright (C) 1989.
      %  By Don Hopkins. (don@brillig.umd.edu)
      %  All rights reserved.
    
    https://donhopkins.com/home/archive/psiber/cyber/pointer.ps

    PSIBER Space Deck and Pseudo Scientific Visualizer Demo:

    https://youtu.be/_fqCeuue5Ac?t=213

    The Shape of PSIBER Space: PostScript Interactive Bug Eradication Routines — October 1989:

    https://medium.com/@donhopkins/the-shape-of-psiber-space-oct...

  • sehw3 hours ago
    Fuck you, I won't do what you tell me.
  • mococa3 hours ago
    Another slop project on front page.
    • Gareth3213 hours ago
      While it's vibe-coded, I'm immensely enjoying the creativity which AI has enabled. These apps would never have been built before, and I find it so fun to see how AI is being used to provide (semi) useful things for people who need it.

      I found this awesome vibe-coded Pixel News Network project recently. How fucking cool is this?? Would never have been created otherwise. https://pnn.watch/

      • an hour ago
        undefined
  • quikroofficial3 hours ago
    best
  • coldbootHqan hour ago
    [flagged]
  • SwtCyber2 hours ago
    [flagged]
  • mouldloft2 hours ago
    [flagged]
  • colinmarc4 hours ago
    [flagged]
  • einpoklum5 hours ago
    More LLM-authored items about LLM slop.
    • inanutshellus5 hours ago
      As long as it has value... I'll allow it.

      ~guywithnopowertodisallowit

    • DonHopkins4 hours ago
      Yet your post is even lower effort and quality, with much less useful information, and a tragic waste of everyone's time.

      As a human stochastic parrot, don't you find it embarrassing and humbling to be so easily outdone by an LLM?

  • nixosbestos5 hours ago
    What a time to be a radical centrist - the AI haters seem out of touch, the AI thought leaders can't stop huffing their farts and being condescending, and somehow this is on the top of HN. What a silly time.
  • VCFundedGenYer4 hours ago
    Please don't submit low quality content like this to HN.