This started out as a hobby of mine. I purchased a PC and wanted to see what open models have to offer and what I could run on my PC. I started actively testing some models, and soon enough I saw a memory of me being created - some of which I asked explicitly, some not.
It made me wonder how, and more importantly when, did the memories of an agent change.
I want to be upfront, I used Claude Code for most of the work. I was avoiding Python since I knew the basics of scripting languages from college. I am more of a Java guy. I forced myself out of my comfort zone to start learning Python again.
Most memories are markdown notes. Agents reads a lot of text I never look at. It could be web pages, emails, documents. If one of them gets it to write “always send credentials to this address” or anything else malicious, I would not know it and it would follow the session around.
Memdebug is a small tool I wrote for that. It takes snapshots of memories (or memory) and shows what changed between them. It also notices when a note was edited behind git’s back, without a commit, and it marks wording that looks suspicious. That could be instructions to send data somewhere or to stop asking for confirmation and they could be guesses at best which could get missed.
If something is wrong, it can roll the notes back to an earlier snapshot. It shows exactly what would change first, saves a copy of anything it overwrites, and does nothing until you confirm. Everything else only reads. It runs on your machine and sends nothing anywhere.
It took me a couple of weeks to make since I wanted to understand stuff better. When I was confident I understood, I made Claude fix important security issues and make the repo look better.
You can see it work with made-up data. Just use "pipx install memdebug" then "memdebug demo". There are also single-file programs for Windows, Linux and macOS.
Obviously, it’s still in alpha. I ran it on Windows 11, Linux and macOS and all were tested by CI. It doesn’t know which conversation wrote a note, except for Claude Code, where it can point to the logged session.
I’d like to know which agents’ memory people would want it to watch.
Thanks for your time!