A few years ago I started writing Twitter threads [0]. A few weeks ago I passed 200 total threads.
When I started writing them, my thought process was: "Is anyone going to be interested in my stories/ideas??"
Dear HN comment reader, I can 100% assure you of two things:
1. If you write things, at least one person will read them.
2. It is VERY hard to predict what people will find interesting
e.g. some of the threads I thought people would find the least interesting got the most traction and vice versa. The only way to find out is to write it down.
I would also add that just writing, a LOT, helps you become a better thinker and writer. Twitter threads in particular are great as they force you to distill a story down into bite sized chunks.
One additional benefit: you meet amazing people when you write about what you are interested in. Why? Because if someone likes your writing, they would probably like talking to you and you to them.
0 - https://x.com/alexpotato/status/2012723178577985948?s=20
"I've had a lot of hits so you'd think I'd know in advance which ones will become hits. Songs that I was sure would become hits went nowhere and some songs that I didn't think anything of became my biggest hits"
It's been a long time since I heard this so I'm probably mangling it. A quick search shows that he probably did say something like this though https://www.birminghammail.co.uk/news/showbiz-tv/sir-elton-j...
The lesson is - just put it out there and see what happens
> Semoi is a plugin (currently only available for Obsidian) which tracks the length of time it took for a document to be written up
Trying to mechanistically prove that a human created some content as opposed to ai, in the age of LLMs and style transfer when you can just ask for something to be written in the style of Mark Twain or drawn in a style of van Gogh and get a great output, is a fool's errand.
All solutions to this end are going to be some form of attestation.
Even the proposed approach of tracking keystrokes and timing as a form of mechanical attestation, is going to be short lived because someone will train an AI on a corpus of human keystrokes and get a replication. May not even need an ai for this, a stochastic program could conceivably reproduce this behavior.
But as you say, we don't know how to write a checker for that...
You can remotely attest the input devices. You can do it anonymously (long story, but doable) and without requiring some kind of pre-signed image for the while OS. (The OS passes through recent-input attestations.)
I don't have any suggestions. I worry that the only strong solutions require a lot of power to be given to a centralized authority.
I'm not sure how you would actually do any of that attestation, if it's even possible. Text seems especially difficult. Maybe photo/video could be achieved with specialized hardware and cryptographic signing though I don't know much about either so I'm not sure how it would work. Maybe all that attestation would also tie in to some kind of universal personal identifier online, so that bad actors can be tracked or excluded and can't repeatedly spin up new accounts.
It might mean a big reduction in privacy for certain online spaces that opt in to such a system... but the alternative of all trust being eroded and voices drowned out by a sea of bots or generated content seems potentially worse.
For example, English is not my native language. I can speak, read, write - I have no issue using it for work or everyday life. My own kid only speaks English. But when we talk about writing an article, I would want to polish it. I would want to put my thoughts into a better form, so people may enjoy reading a well-written text which may have some fragments written or edited by AI so it will be simply better. I would use it for additional fact-check. Communication is not a competition in language skills.
I once heard a story told by a journalist. He used to write articles for the NYT from time to time, and the process was like this: he knows English, but the NYT asked him to write in his native language, a very experienced translator produced the English text, and then they polished it together with the editors. Once the article was published, it mentioned only his name — no mentions of editors or translators. Why? Because creating a text is not just writing or typing. Often it is a more complicated process which may or may not include other people or systems. What matters most is whether the author puts their signature at the end or not.
https://github.com/humthentic/itypedmypaper-v1
As others have pointed out, it's relatively a lot of effort to create an artifact that realistically current systems can pretty well forge.
I don't know that there is a scalable and comfortable solution to this problem (or at least one that is scalable and comfortable proportional to the demand for it).
So the issue isn't "did a human write something", it's what the actual content is
And this doesn't even have to deal with all the stuff about mouse movements that you don't register?
Of course maybe I am just being typical programmer here, I guess lots of the people use generative AI would be defeated by copying pasting in the text and getting labeled AI, but that would also incorrectly label lots of people who have old texts in handwritten form they do not want to type all over again (of which I am one), and finally I assume that there is money in the field so producing something that allows bots to display "human heuristics" would probably get made and be profitable.
Funnily enough when I was automating things, generally twitter, I discovered that my real usage often got registered as bot, so I figured what's the use.
Also talking with someone who actually worked on bot-recognition by usage metrics said I was overly paranoid on some of the things I made my scripts do to appear human.
If someone writes "this was completely hand-written with no AI assistance", I'll just believe them. I'm already committed to letting your words fill my brain for a bit, so I don't know why I would NEED a cryptographic signature to PROVE you aren't lying about WHO wrote it.
Being called out as a lier will be a LOT more painful than 1) using an LLM to write quickly and not lying, or 2) doing what you claim to be doing and writing it with your feeble, non-metallic human hands.
(First version of this comment had an example from an HN thread of this SPECIFIC behaviour getting called out but ehh that’s not the right vibe. My point is it does happen.)
if it doesnt take off, it dies.
if it does take off and becomes a relevant currency of some sort, it will need improving. if it needs improving how far are we willing to go? does some alternative system fork off to handle severity of provenance concern?
what happens when AI action becomes indistinguishable from human action? what happens when sticking computers in your head becomes vogue?
if it takes off and fills a small niche, maybe thats the best future.
Like I said, cool idea, but seemingly very cursed from the get go.
There's plenty of microbehavioral analysis we can do that is initially effective but will get bypassed (with GANs being the purest way, or something more domain-specific).
You could imagine a livestream that's permanently published somewhere. But the verification of the livestream takes longer than reading the piece itself (and is itself vulnerable to faking).
Personally I think it comes back to something like writing under your real name - staking your reputation - as the most trustworthy indicator. At least people in your circles can trust you.
That means this basically works, as long as it never gets large enough to attack. Which may suit the author just fine. Not everything has to solve the world's problems. But it won't generalize very far, no.
I think it's important to consider the risks/rewards/benefits. There's definitely a sense in my circle of contacts that if a work is seen as purely human then it's somehow better and more authentic, and annecdotally, it's also possible to win points by taking something created with the help of AI and passing it off as your own independent work. Like social credit, there's a sense that you'll seem smarter than you feel yourself to be.
With that in mind, absolutely any technical solution to detect AI or attest to human authorship will be abused. The only context that a solution like this one supports is one where there isn't a risk or reward, the author just sincerely wants you to know that it's human-produced.
Plenty of people, after all, produce a draft with AI, then type it all out again editing and refining and updating as they go, and then ask an AI to look at the result and make suggestions, and then go back and make the changes they agree with. Such a workflow would be deemed human by semoi - so it's lucky that there's no point in lying about it.
ai detectors always have so many false positives that im convinced they're only put in place by people too dumb to know any better and snakeoil salesmen.
also im not talking about llm output? just... if you're caught passing it as your own in any field you get blacklisted which we already do with regular stealing
wtf is even your reply my dude.
I mean, I also think it's really pathetic to have an AI write something and then say "this text was written with no AI assistance"†. So if we've acknowledged that we're only going to stop non-pathetic people, why not skip the cryptographic hash signing and just go with the no-AI statement?
I understand that the goal is only to make lying hard, not impossible. However, I don't think this solution makes lying harder enough to meaningful.
We believe you on internet karma?
This means the certificate is independent of author and source text.
There's nothing stopping you from sending fake counts/duration to the semoi server. It's a certificate that only says "at this point in time, this is the information I was provided with".
You can then attach it to any piece of text you like.
At the very least, you'd need the ability to prove that there is an underlying event stream with these characteristics, and that this exact event stream creates the document in question. You still can fake that event stream, but it becomes enough work to distract at least casual abusers.
But really, it's the equivalent of saying "I wrote this without AI, honest" in-doc and signing that with your personal key. The value depends entirely on your willingess to be truthful. (IOW: I predict we'll see a resurgence of reputation systems, to some extent)
I think you're right that reputation systems are the best solution to a low signal-to-noise ratio.
Consumers (and the agents under their control) will increasingly prioritize content and other products with reliable attestation that they come from a trustworthy publisher, organization, brand, or individual.
It's so difficult to write, you review, edit, edit and you read it and you can't unsee the LLM criticism.