Anthropic certainly isn't planning to take responsibility for Claude's mistakes - that's the user's responsibility. Describing everything as 'collaborating' is I think part of their efforts to emphasise the user's role in the process.
I trust the output of colleagues I am collaborating with and have an assumption of some shared responsibility. But if I use a tool like numpy/matplotlib then I am accountable for the conclusions I come up with. I can't make an excuse that "matplotlib" created a plot so my decisions were incorrect and not due to my own negligence.
Things like:
> Not allowed: Prompt: "Write my answers to the application questions for an AI safety researcher position at Anthropic." Result: Generic content with experiences you haven't actually had
I haven't seen other policy documents from them that are as relevant to the question anywhere else.
name= a variable name or fields
load bearing= coupled dependencies that can break other things if you change it
Random example from a session I have open:
> *One new failure mode the doubling opens, and how it is closed.* A ring never shrinks. With two 8-byte index words per record instead of one, the ring's doubling comes inside the budget's reach in the blank regime: a doubling taken [...]
It takes extra deciphering cycles to see "the doubling", "budget's reach", "the blank regime", etc. and figure out what it's referring to. I had to read that opening sentence multiple times. At first I parsed it like "One cat the table yawns".
Sometimes it's so encumbered I can't tell what it's saying at the directional level: good or bad, fast or slow? "Your blank regime negated the pre-armed run's dilemma but clawed back the overall metrics."
It's not how I'd phrase things if I were trying to be easily understood, though Claudese is probably great for LLMs due to ad hoc jargon usage.
Minor inconvenience != load bearing, yet Claude consistently uses it while missing actual loading bearing things.
Likewise, it tends to jump to terminology that’s technically correct but practically meaningless.
tree is green = tests built and ran without error
landed = surprisingly, not "made a commit" but rather "finished the code." It might be confused because we're using Perforce and not git.
Claude will often comment on it using the phrase wrong when it gets context.
Every damn Claude article does the “tell ‘em what you’re going to tell ‘em, tell ‘em, tell ‘em what you told ‘em” routine.
The real way to actually get a good output has always been in the prompt, not these parameters
No, I do not. Good for me I guess.
> The real way to actually get a good output has always been in the prompt, not these parameters
What an absurd claim.
It's an interesting model/tool. Powerful but still chock full of trade-offs. I hope the next model's language is more like e.g. OpenAI models in terms of language use. (oh, and the code comments, yikes).
"excursion" (means: a spike/jump — a value that rises or deviates from baseline, just say spike or outlier)
"legitimate majority-normal baseline" (means: a real majority of normal pixels)
"matched pool of pure-noise ('normal') pixels" (means: the same number of normal pixels, don't bring pools into this)
"ablation" (means: comparing before vs. after — turning a thing on/off to see what changes)
I actually used Claude Code to try and find examples like these but not entirely unexpectedly it had a very hard time detecting these, even though I encounter them like every other sentence. I can imagine this style is great when writing prose, not so much when writing documentation.
Similarly, I tried using Claude Design to make some slides, and the individual slides were fine, but just about every other slide it used different words to denote the same concept, which I would say is the fastest way to lose your audience.
Did you try specifically telling it that? I haven't used Fable but I do write lots of docs for a non native audience and I get good results from stating that.
And they probably know basic LLM tricks like "write it in the style of X".
When you read a blog post with Claude voice, you're seeing the result of someone who couldn't even be bothered to do that which is why they deserve extra lashings.
You can probably get it to do a surface-level impression of Mark Twain that way, but from what I've heard, it's not that easy to prompt Claude out of the "voice" described in the article for technical writing, and all existing methods, like asking it to ELI5 or tropes.fyi, only have partial success.
There are all sorts of ways to do it though, like pointing it to an example doc.
Good enough for essentially blogposts? Easier to finetune?
> They aren't paying for the tokens.
They are paying for electricity, though.
Are there any creative-writing LORAs published for open-weight models?
Oddly enough standalone LoRA adapters are very popular in image generation world and utterly unpopular in LLM world - there it is customary just to distribute fully merged models on HF.
Having said that I personally use a finetune of Qwen 3.6, Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic, which was was not designed for creative writing per se but is far better than any other 24B-32B model I've tried for writing, let alone awful vanilla Qwen 3.6.
Not sure I agree about e.g. GPT not sounding this formulaic though, imo it definitely does.
I don't actually have a problem with the common terms Claude emits, I think they're appropriate for a coding agent. But I do find them overused.
/s?