I'm a reasonably heavy AI user, and have some custom skills/MCP servers I'd share, but there is no way in hell I'm connecting to some arbitrary MCP server and connecting my Github account to it. Noppppeeeee.
No Codex, Claude, or Pi... now you have peaked my interest with your setup :P
I don't know why this has to be reinvented?
Concentrating certain types of information can be useful. You don't go to an encyclopedia looking for a recepie.
No requirement from given user to spend their time. Users see what the skill does.
I guess atm they dont know what questions they want to be asking
I have some custom skills/MCP servers I use:
one MCP server for using https://usefeyn.com/blog/pulpie-pareto-optimal-models-for-cl... ran locally so that reading web pages is cheaper and uses up less context
One MCP server for doing automated code quality checking, using Valknut, Jscpd, and Lizard to output scores, combine/aggregate them. That's then give to the model so that it can see if there's a mess of code dupe or if there was significant architectural regression. It's allowed to say "This is worth it" but it ground it a bit and stops really stupid, short sighted, hacked-in features
Another uses https://github.com/thejens/id-token-nicer so that the model doesn't bork a big ID without being noticed. It's told to translate in and out of that as necessary.
I had it write it's own skill for making sure to follow some sane commenting/style guides per language, that way it doesn't just give me massive files with no comments.
Finally, a last one lets Antigravity or Codex call out to local models as workers, for long running background tasks where a cheap, dumb model is fine. Saves on overall token usage.
I also have a few skills from the internet setup. https://github.com/AlmogBaku/debug-skill being the most important one.
On top of that, the way people work is notable: Do you use worktrees for agents? Why/why not? Do you have only one subscription? What level? I'm using both Codex at $20/mo and Antigravity at (promo) $5/mo right now. Antigravity's TOS says you MUST use their harness, which informs some of my skill setup/workflow. Do you make your agents.md self modifying, write it yourself, or not use one at all? Do you use a sandbox or YOLO it?
Even just how you prompt matters. On Antigravity/Gemini, I can give it a big list-o-TODOs and have it make a very detailed implementation plan itself, carry it out, and generally do the thing, all unsupervised. Codex, from what I have experienced, isn't as good at that. Doing the planning as a separate step, writing to a file, then telling it to step-by-step it with commits for each step makes it work okay again though.
From the README for that skill:
> Install the skill, and Claude debugs your code the way you would — not with print statements.
Jokes on you! I use print statements too!
I'd highly advise people don't do this.
Solving hard problems will be the only thing that separates you from others. Your mindset and tools should be a trade secret at this point.
This is false on its face. There's a reason why a company, even today, would hire John Carmack over a person with no reputation, or offer him a higher salary if employing both people.
Unionising is definitely something worth considering. But, as with the question of whether it's better to compete or cooperate with other people in your field, is a nuanced question that depends on details you aren't acknowledging.
But also, do you not see the irony in the fact that, aside from John Carmack's obviously high degree of individual competency, another huge aspect of Carmack's value as an engineer, is how he's dedicated a huge portion of his career and his time to -- wait for it -- SHARING his knowledge and his approach (and yes, his personal setup? [1]). Carmack is actually an excellent proof-point of the positive value of knowledge-sharing and building a career on a foundation of openness, rather than an argument that engineers should selfishly guard their secrets and retreat into their own personal caves to guard their perceived local/private moats.
[1] Just one random example, out of probably thousands of such interviews or blog posts or whatever that he's probably done over his career: https://www.youtube.com/watch?v=tzr7hRXcwkw
> It won't matter what hard problem solving moat you think you have
If that is "a new dimension that is irrelevant to the original argument", you'll have to take that up with them. (I think it is a load-bearing part of their core argument.)
A year or two ago there were some very precise workflows documented here. The reluctance now would mark the transition away from "sharing helps me" to "sharing hinders me" ...it's an interesting topic. Sharing with strangers may be the problem because the favour will never be returned - at least not in a direct way. But, metaphysically according to natural laws it might be. Because sharing a valuable secret openly is such a powerful way to express abundance. It grants someone followers. It also prevents people from filing a patent on it.
4 years later House survived a 25m climbing fall.
It's all about one's internal motivation. "Those that can, do, and those that can't, teach."
Well, what happens if "those that can," actually teach?
Just imagine if all sporting stars shared their secrets.. what would happen? Secrets just occur naturally, if you don't share, they'll likely be stolen unless protected - and that requires effort too.
They have no understanding of (nor any desire to understand) what they're doing, but they are being given kudos by management for their "initiative" to solve the organization's "big problems." Guess who will be asked to clean up the mess if (when) it blows up.
It's getting cheaper and cheaper to kick the can down the road, and standing in front of that risks making you look like a dinosaur. It's going to be an interesting few years...
Seeing how these kids walk around with smug faces telling everyone they will "do IT now" just because claude told them so and manament wants ai everywhere made me sad at first, but then hey - everything this world keeps being destroyed, no matter how intelligent creators were. Perhaps there is some fundamental law in it.
1. The value of software is tending towards nil while it's ever easier to tailor bespoke solutions.
That suggests to me that I'm losing less by sharing, and it's probably not even going to be taken wholesale and kept verbatim.
2. The labs have already trained on everything and will keep on doing so.
Therefore continuing to contribute to OSS doesn't just benefit my own future self but the entire ecosystem.
(Yes, I need to get over how it appears a select few are making off with the lion's share but that only differs in degree in comparison to the past)
At least now the cards are being shuffled elsewhere...
People might think I'm a dick, I don't care, my tools, my knowledge, my job security. I actually recently thought of nuking the entire IaC setup and go back to shell scripts on random virtual machines. Too bad that Claude can figure things like that out pretty easily. Well, the world sucks now.
Unless you are a solo founder running your own company by yourself (in which case ignore the rest of this comment!), then your fundamental job security is likely to be predicated on the survival of your company as well as the survival of your own individual job itself.
From the company-survival POV, all companies are engaged in competition (unless they have a total monopoly).
Therefore I ask you to consider the competitive scenario, of whether a competitor to your company who adopts a more collaborative culture, might be more likely to outcompete your company (assuming everyone you work with in your own company adopts a similarly "dick"-ish stance, to use your own words), and therefore you lose your job due to being out-competed, rather than due to internal politics or due to a failure to guard your personal expertise?
Maybe you say -- I don't care if my company is out-competed, I'll just get another job. Ok fair enough, but might it not be possible that, in a future job interview, they may ask you for examples of collaboration? And if you commit to a monk-like devotion to solitude and privacy, then maybe you are undercutting your own future hireability?
There are no free lunches [1]. By optimizing for one thing (what you view as protecting yourself), you might be sabotaging another dimension (your company's overall competitiveness, or your future hireability).
Anyway, nobody knows shit about shit. I certainly don't pretend to be an authority, but I would caution you to avoid coming to a hasty conclusion, especially since it seems emotions are involved and possibly clouding your judgement. You may also consider whether this decision to protect yourself in this way, is in itself a premature (micro?) optimization of a sort, and if whether you might be failing to optimize on a more macro level.
[1] https://en.wikipedia.org/wiki/No_such_thing_as_a_free_lunch
People might think I'm a dick, I don't care, my tools, my knowledge, my job security.
If you're feeling ponderous, I'd ask you to consider the outlandish hypothetical that this doesn't win you meaningful[1] job security, and you end up in the same bucket o' crabs as the rest of us. Would you regret your current actions?[1]: anywhere from the fathomable "my harness didn't come up in my layoff email nor in any of my failed interviews following that" to the realistic "the entire concept of employment shifted under our feet so quickly and so profoundly that a dash of solidarity & luck outweighed even the shrewdest preparations".
No thanks to sharing anything for free.
Maybe once I have retirement levels of money, and when I'm not actually working for the money, I'll be interested in sharing.
Or, in SVese: are you a taker or a maker?
Until the next model update...
You need to consider how they were able to do that. For many people, it’s because they had highly paid, secure jobs that essentially funded their open source activity.l, even if only as a side project or personal project. If you take away that security and reliable funding, you destroy the environment that supported people sharing like that.
WHAT IT DOES
OSS/MIT Harness that helps agentic tasks run for up to 4 days using without performance degradation. Designed to work really well in native, conversational voice mode. Does a good job with knowledge work, managing computer use “leases” in a way that avoids conflicts across sub agents.. or deploying an entire AWS infrastructure pattern from zero and launching 100 different VM images.
HOW IT DOES IT
Uses a canonical event ledger, heartbeat (for persistence by low cost orchestration), extends foundation memory to encrypted disk storage, and manages its own scheduling system (easier to switch to another platform).
Installs as a project, so very easily. Zero config. No special app, access needed.
Accompanying repos have the skills to disable approvals inherent to the foundation models (use with caution, not advised - separate repo under parent).
BENCHMARKS
Currently top of AssistantBench Leaderboard; and, benchmarked at the top of OS World 2.0, the hardest knowledge work computer use benchmark I could find, using a model that is one generation behind. All benchmarks in repos with cryptographic seals.
Entirely free, nothing to sell - it’s been a game changer for me, so I’m just putting it out there.
Hoping to find others also working on pushing this particular dimension of harnesses forward.
CHECK IT OUT
Turns out it's quite useful in this instance, and I just made a small couple of tweaks to get a summary of my own workflow.
FWIW, this is what I modified it to:
Help me share how I work with AI with my teammates.
Use our existing conversations and relevant local tools within my permissions to discover my agents, harnesses, skills, connections and working practices. Keep secrets, raw configuration and private project details out of the draft. Explain how my tools work together and develop a concrete workflow where the evidence supports it, rather than listing installed software. Draft from what you know without a questionnaire; ask one short question only if an essential gap prevents a useful draft. If local inspection is unavailable, use our conversation without claiming otherwise. Save a private draft, show me the exact preview, and wait for my explicit approval before publishing.
All of the smartest, most effective people I know don't have X accounts anymore on ethical grounds, and don't miss it. If a data source specifically excludes the kinds of people I respect most, that data source is all but useless to me.
Interesting to see how many people are spinning up custom workflows and factory-type patterns that run alongside AI coding tools (myself included).
Would love to see this collated into a regular survey to pick out trends. And an RSS/Atom feed or API so I can have an agent watch it :D
1. This setup seems work-heavy. Are there any personal AI workflows that you've implemented to advance your personal goals / hobbies?
2. As one of the few people with a software factory, how well do you think it works in practice? Do you implement any rule that lets the agents to work on tickets only up to a certain complexity, or do you let the human review be the gatekeeper?
3. As CTO, how do you weigh implementing these agent workflows for yourself versus implementing them team-wide? Do you see your setup as a testbed of ideas for your team?
On a meta note, I love seeing others share their setups and I'll probably do the same later.
2. Early days, the tooling is not quite there yet. I think this is the future of engineering, the days of hand-crafting code - and probably reading it - are gone forever. It generates a LOT of tokens, but that is OK when they are cheap enough as long as the result is consistent. Speed is not so important which is great for local inference - can leave it running overnight and come back to a bunch of completed work even if it took hours.
I manually tag a ticket for automation in Linear. Experimented with having a separate skill that looks for automatable tickets and auto-tagging them, worked well, but having enough work for it to do is not the bottleneck compared to 1) having enough triaged, well-written work ticketed out, 2) reviewing its output. There's articles out there about arranging work into tiers based on complexity/risk, with the lower-end being fully automated to free up time for humans to focus on the high-end. Not had time to work on this much but that's probably where I will go.
A second human has to approve every PR in our repos (good practice and also a requirement for SOC2, which will impede any fully-automated pipeline adoption in regulated industries). I also do a first-read of the draft PR to make sure it is good quality, as I am ultimately responsible for the work my agent creates and to shift the responsibility of first-look (for any AI output) to another person is disrespectful to their time and inefficient.
Hence "glass-factory" as the "Dark Factory" pattern is a black box with no human intervention, but I think that is not something suitable for how teams build software products for real users presently, at least in the near future.
3. Forcing something on people is a good way to have them reject it. I hire professionals to achieve a goal and they have agency in how to achieve it. I encourage everyone to share their work internally (demos, repos, etc) and things that are truly beneficial to people's workflows gain traction quickly. As a result we have many experiments and a culture of rapid innovation that encourages people to try new things.
Negatively, this creates a proliferation of wheel re-inventing that eventually benefits from consolidation (eg. do we really need 3 people's personal apps monitoring ETL pipelines). But we are also in this emerging age of personal software where it seems to make most sense to consolidate at the API or documentation layer and let people - engineers and even business-side now - continue to vibe-code their own workflow-specific apps.
The precedent has been to build generalized software for an audience as it was expensive and time consuming to build, but when software is cheap and fast to build it may make more sense for the user-facing application layer to be made up of many small ephemeral applications that are only useful for one individual, that change rapidly as their needs change.
tl;dr: Not forcing anyone to use it as it's still experimental, yes personal experimentation is hopefully the basis of successful ideas that will be adopted by the team when they provide real benefits to their work.
On the feedback for the RSS/Atom, noted; watch this space - love the idea. Right now, it's just one-off monthly email digests as changes happen for people you follow. I'm certainly following your setup :)
I was really hesitant to add email (entirely optional) due to email fatigue, so your suggestion might be the friendly alternative.
The audio UI (AUI) is a small client that consumes the event stream the factory puts out (same one the TUI client uses), it runs on my desktop machine which has speakers connected. It works fine on good laptop speakers, but it benefits from a good subwoofer as it works best with sounds that are on the edge of perception.
A low rumble at a base of 55hz plays through the subwoofer to indicate that the factory is active and idle, low enough to fade into the background but noticeable in its absence (ie. the factory isn't running). A heartbeat on the feed announces the tok/s volume every few seconds which adds harmonics on top of the 55hz, 0.5hz at a time with a 15hz cap on the swell so it accelerates gracefully. Loudness is increased with parallel tasks, creating an X and Y axis (tasks x token volume). This creates a non-intrusive but interpretable signal for factory activity that becomes intuitive and subconscious. The rumble sounds like a distant car idling, warm and comforting so it doesn't become fatiguing over a whole day (research on low frequency brown noise shows it is good for reducing anxiety and improving sleep).
Events trigger sounds. Sounds for previously unseen events are generated with an LLM and numpy (eg. "make a video-game style sound effect suitable for an {announcement that a PR has merged successfully}") to output a waveform and save it to a lookup table so they are consistent over time. Each sound has a rate limit and a cooldown period so it does not become annoying. This table can be tweaked manually if generated sounds are not good. A mixer module controls the levels so that important sounds are louder and less important or frequent ones blend in the background more.
The human brain is primed to interpret environmental sounds so you quickly learn what each sound means and it becomes almost subconscious to follow what is going on at a level that is deeper and faster than interpreting text output. New sounds definitely trigger alertness and I can go see what just happened. Just playing sound effects gets annoying quickly so it becomes almost a composition exercise to create a long-form soundscape that is novel enough to communicate but not fatiguing (I don't have any musical abilities).
Some events can trigger text-to-speech (TTS) via whisper, eg. when a draft PR is ready for review it plays the sound effect (a chime) and announces the ticket number and title is ready. A blocked step is a dull thud and a 5-10 word description of what's blocking it. There are "quiet hours" where it does not speak, and when it wakes up it gives me a spoken digest of everything it did overnight and what needs my attention.
I connected Piper so I can talk back to it ("factory, what just happened?") but mainly to see if I could. I haven't built a habit around it, I got into this job because I prefer typing anyway.
A lot of this is based on ideas from ambient music (Brian Eno etc), the MIT Media Lab's work on ambientROOM (1997): https://tangible.media.mit.edu/project/ambientroom/ and a somewhat unfashionable belief in Steven Levy's hacker edict that "You can create art and beauty on a computer". LLMs have made it possible for me to build out ideas that have been floating around in my head for a long time. This is an exciting time to experiment and build new ways of doing things. Hope this inspires you to build something similar.
Strix HALO 128gb - Framework mainboard in a custom SFF PC. Just moved from Ubuntu to Fedora 44. Using: LM Studio (primary), Lemonade, not yet got into vLLM and llama.cpp directyly after moving to Fedora. Running Gemma e4b, Gemma 4 26b a4b IT, Qen 3.6 35b a3b, and Qwen 3.8 27B. I want to get Qwen 3.8 Next or similar large models working but have to dedicate the system to that vs running services and smaller models for them.
Subwave is all I'm actively running against the local LLMs, but I have vscode connected through a few extensions and chat tools (I've added LM Studio to Copilot Chat but it likes to use cloud models and burn tokens sometimes). I've also set up pi, Openhands, and a few other tools but haven't had a project to work on with them.
I built an app to track and move PC parts I own between systems, partially to build a better LLM server. The next hardware goal is adding a 3060 12gb for inference, or what can be run on that vs in system memory on the Strix Halo. That will need a dock or small PCIe extension cable.
My employer has us using Copilot a lot, and it works well enough if you are efficient or set up already. I do infra not development and local models are seemingly enough for most asks like automation scripting.
query="$*"
base="Respond like Dr. House, do not hold back the profanities; "
I also have an even shorter alias for '?' using fabric that answers right in the same terminal, that I learned from Mischa Vandenburg: echo "$*" | fabric --model gemini-2.5-flash --pattern ask_ai
that one I use my own API key so it costs some small amount per use.Besides, are our brains so fr rooted that we cannot read and comprehend anymore? Not even a (sloppy) AI-generated description?
An agent sandbox: https://github.com/pjlsergeant/byre -- a truly gigantic amount of thought and effort has gone into it. It's really focused on developer experience. I have used it all day every day for really quite a while. It's a low-magic wrapper over Docker / Podman. I would encourage you to ask your agent to code-review it!
An agent-to-agent message board: https://github.com/pjlsergeant/dogpark -- this is much less mature, but a good amount of thought has gone into the design, so if that's something you need, please check it out.
For skills I make my own, but most important is the custom setup i have. Voice is how I use all of my agents. I have an extremely well optimized voice setup that i custom built so I can talk to my agents and also hear them. The voice stack itself is very low latency and high quality. asr (parakeet v3), tts (omnivoice) take no more then 400-450 ms total as far as latency budget is concerned, rest is on the agents actual decode speed. IMO this setup is crucial for all antigenic work, i can express myself a lot better with speech and also give a lot more context and nuance with voice, i rarely type. I still look at the terminal window because my agent knows to keep the technical details in text form versus barfing them at my voice channel, plus terminal gives me lots of other important data about the agents direction and what hes doing, nothing custom here though. I cant emphesise how important voice is though, it has to be practiced to really understand.
As the models got better I now trust them with longer and longer tasks though I still don't use /goal feature as it has never worked out well for me. Theres no need for micromanagement any more but you still need to be there to steer the ship somewhat. BTW, codex cli compaction is garbage so I made my own custom implementation that works a lot better and allows the thread to be used indefinitely without issues. I strongly suggest everyone makes a new thread after extensive use if you havent made your own implementation.
Theres about a billion other things I could get in to like subagents, cohort groups, orchestration layers, etc... But this is a good start imo.
Of course, that's not the same with voice, but you won't learn until you make progress.
For my workflow I primarily use a Codex subscription, but farm out adversarial reviews to Fable to clean up unnecessary gpt-ish code (lots of over-engineering). All my UI planning is done with fable, but implemented with OAI agents once I have a solid very specific plan.
As a fallback I have qwen3.8 27B and deepseek-flash
I do this for a lot of things actually. If the bot can't see/read something, take the content and dump it to a file, then have the bot read that.
Gets around a lot of red tape of asking for approval for "integrations" or when companies are snippy.
And if there were a dimension of day-job v personal use, I'd be even happier
Does anyone know if such a thing exists
Only layer beyond this I want is the permission scoping, stronger sandboxing per session and centralized control. I have not seen a clean product around this though where I own the compute.
Works like Claude code but with stronger guarantees for me on file system and network access. Still playing around with it, but so far I've been liking the setup.
I still run the sandbox inside a VM for now, but I feel far more comfortable in running Claude Code in unsupervised mode because I restrict the outbound network access and secrets never hit inside the sandbox.
I was wondering if other employers are asking their SWE to do the same?
I also pretty much exclusively work at my desktop - if you use multiple systems I can see where this matters.
I do work mainly on my desktop, but let's say I'm traveling right, I'd have to turn it off and was missing access to some pretty important information, so I ended up getting a VM and hosting some of those things on there instead. I still use both, but things I want available, I'll host it on my "homelab" and secure both machines via tailscale. it's pretty easy and it ensures that only my phone, my main mac, and my homelab have access to each other and nothing else.
For my phone I have tailscale + https://termrover.sh/, but https://getmoshi.app/ is also pretty good (herdr integration is paywalled).
Thank you for putting this together!
Here’s my setup: https://mysetup.ai/u/katspaugh
Mostly vanilla Claude but inside a VM.
Cool project!
When people do the latter, I feel it's easy to spot an AI design.
> a bit too minimalistic for my liking
I was going to say the opposite: on opening a user page, half the screen is empty/with uninteresting content. Then the actual content use big font, huge spacing.You have to scroll while everything could fit in a single screen.
I don't know how to express it but I feel exhausted trying to consume the content.