I found myself asking my AI assistant to help me with some non-trivial tasks (describing visual changes to an app, filling in forms, setting up a server instance in a new provider, changing some dns settings in another one, mostly things I know how to do but that would take me extra time to figure out on an unknown UI). I was solving this by taking screenshots all the time and pasting them, when I got tired of that I just copy&pasted the whole text.
With DesktopVisionMCP, I just ask my AI to "look" and it takes a screenshot on-demand of all my monitors, passes it directly to the AI (no file), so it gets the context instead of me figuring out the best way to describe it and my intent.
I'd like to hear your thoughts and feedback.
It is live on Product Hunt today.
PS, this whole thing was written by me, by hand but proofread using DesktopVisionMCP ;)