----------------
It's going to be awkward if you share a youtube link with somebody and what they see is significantly different from what you saw, perhaps even to the point of them replying, "Why on Earth did you send this to me? Are you on crack?"
More importantly, this will almost inevitably lead to content creators being given the ability to not just randomly A/B test versions of a video, but produce different versions of the same video that are shown to users based on their data.
e.g. Shania Twain used to produce different versions of her albums with different instrumentation based on which section of the music store they'd be sold in. There was a Country version for the Country section and a Rock version for the Rock section. She's still bootin' around today, so she could produce different versions of music videos targeting users based on whether Google thinks they like Rock or Country more. This would be relatively harmless, although one might be surprised by the version that appears on a friend's phone.
Musical taste isn't what really divides people these days. What might content creators do if they could show different videos to people based on their political views? This might be good for their business, but it undermines objective reality. People would be shown different "facts" based on their beliefs. This is precisely the opposite of what needs to happen to reduce political polarization and bring people closer together. A common reality is necessary for society to function.
Give Google, Meta and the rest an award showing they beat human psychology and now we can encorage people to put the phone down and go outside again.
Data-driven optimization in general, and not just when it comes to things like "online content".
Like, it certainly benefits us to a certain point, but after a while it starts to create fragile systems and we're getting deep into the fragile systems phase.
See: the complete breakdown of the supply chains of just about everything, algorithmic price discovery contributing to out of control inflation, etc. This is beginning to impact nearly every aspect of our lives and IMO rarely in a good way.
Yep, and A/B testing as experienced by uninformed, unaware end-users is a dark pattern.
It undermines the perception of (and trust in) continuity which is necessary to make effective use of a tool. The best way I can describe it to the skeptical is: imagine the dials on your car's dashboard rearrange themselves occasionally overnight, and on some commutes to work you suddenly can't work the radio or the AC while moving at ≥35mph. Of course, since the widespread use of touchscreens, that example became very literal.
So the car manufacturer has figured out the "optimal" arrangement of dials and buttons on their dashboard for their preferred levels of user engagement. Great. How many of those users now associate their car's brand with inconsistency? "I can't trust the damn buttons to be in the same place the next time I drive."
Is it unethical to do phased rollouts (where a small percentage get the new version) as a way to do safe deploys? If the issue is that two users making requests at the same time might see different things, then this would also be unethical? Yet, these sorts of phases rollouts is the best way to release something safely. When I worked at a large CDN with 50,000 servers around the world, we ALWAYS did phased releases, to make sure we didn't take down everything all at once, and to make sure we caught any performance regressions right away.
Is your issue that the user might be getting a version that won't stick around? That seems always the case, whether you do A/B or not. You might rollback if there is an issue, and you will certainly roll forward at some point, meaning users will get a new version at some point.
Would it be an issue if the A/B test was temporal? Like all users got one version today, and a different version tomorrow?
I guess I am just confused by this statement:
> If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all.
This seems contrary to so many other complaints we see all the time, that companies push changes out without taking into account what users actually want. So, do we want companies that push out changes with no user feedback because they are confident that they know what users want, or do we want companies that get feedback from users on whether new changes are helping or hurting.
> or do we want companies that get feedback from users on whether new changes are helping or hurting.
You don't get that feedback. The feedback you get is whether some telemetry KPI goes up or down. That's not the same as actual utility for the user.
-(not) Henry Ford
When I have two free hours to turn on the XBOX for the first time in half a year, I want it to turn on right away and play my game. I don't want the box I paid for and have been looking at to figure that it needs an OS update, and a game update, and that the game should now be slower and glitchier than it was the last time I played it.
When I play music on my phone in my car, I don't want to find out that the "Start Mix" button moved, or that showing the upcoming playlist now takes one more swipe, or that the UI won't load because YouTube Music doesn't cache it's UI anymore and when you have a cell network reporting 1 bar but it's actually zero bars, you get a spinner for ten minutes.
The six CD changer in my dashboard has worked exactly the same way since 2006, the discs play when the key turns on, and nothing moves. There's no engagement to be had other than "my music plays when I turn on the car in my driveway which also has spotty cell service". There isn't a KPI to be measured, a PM to be promoted, or anything. It's just a car radio.
Most consumer goods are solved problems. Nothing's changed since 2015. Even tech from 2015 is just a convergence of 2005 tech, like MP3 players, digital cameras, and Blackberries. People don't have radically different problems to solve in their day to day lives. People take pictures, share them, do email and group chat, voice calls, read the news, watch TV, pay for parking, do some banking. Watch a 90s TV show and all of those activites required different physical places and tactile goods. (Hell, that's why screens in Android are called *Activities*.).
When I pay for a product, I expect to be paying for a finished product, not some psychological experiment that's someone else's promo packet.
I constantly experiment on users in my work. It’s all around extracting the most money you possibly can. Meanwhile we have mountains of UX interview material where people tell us exactly what’s wrong with our site, and we don’t implement any of it lol.
Profits are up though! In a big way! And our users continue to hate us more and more.
Are you not aware how unpopular practically all recent changes on YT are among its users? Almost none of the changes done on YT in the last 5+ years would have happened if they took into account what users & creators want. So how exactly does telemetry and A/B testing help when it either tells them the opposite of reality or they simply interpret the data however they like anyway?
It could be that YT is just making everything worse for everyone, but I also know they have data that you and I don't have on how people actually use their product. I don't think we can assume they are just bad at making a product just because all the people we talk to agree with us that it is bad.
> Users don't want their shit changing all the time.
Eliminating A/B testing won't solve this problem. Even without A/B testing, they make updates, etc.
You might as well just say "Updating an online service without asking the user first is unethical."
Oh, and how much are you paying for that service...?
I paid YT Premium once. It unlocked a playback queue in the app. That queue had 5 separate bugs I found within an hour. It's literally a simple playlist and yet not a single feature (adding, removing, reordering etc) worked reliably. The playback queue also randomly emptied itself sometimes. Meanwhile I get a superior version of this in the browser by simply opening a video in another tab, for free.
Why should I pay for that while they only have AI support designed to never solve any issues, do nothing against bots, and then warn you that you may get banned when you report too many bots?
I actually discussed it with a friend a few years back while debating the declining quality of Google's engineering.
The craziest thing to me - it appears to be using some kind of eventual consistency, so additions/reorders/deletions have to go through some complex process server-side that takes several seconds to update in the UI (and often the order is wrong, or silently fails to add videos). And yet, the whole queue disappears without a trace if the YouTube app gets unloaded by iOS, and is unavailable on other devices, so it could have just been stored locally all along.
I thought that perhaps they were storing it server-side because eventually it would allow the queue to be restored or transferred, but it has been about 3 years now and I don't think that day is ever coming. I just lost a whole queue of several vids this morning (on the plus side it was a good incentive to get off YT, so I'll give it that).
My guess is it's using eventual consistency or some complex multi-microservice chain of RPCs because that's just what you have to do at Google. I'm sure there are engineers who want to fix the feature or go back and complete the rushed launch but likely can't convince the decision makers that it's necessary.
So we get a subpar experience from one of the largest companies in the world with thousands of the best engineers, while Google keeps getting their $16/month because there are no alternatives
Updates are one thing, changes made specifically targeted towards manipulating users into spending more time on the site in ways they can’t opt out of (youtube shorts is a good example) are different. No one is bothered if YouTube updates to allow 8k streams. I am extremely bothered that YouTube does not allow me to disable shorts in the app. I don’t want shorts, they are a distraction and one more thing i have to guard against getting sucked into. Let me use the app how I want to use it, not how you want me to use it.
(I have this weird feeling that you're not upset with blue/green deployments...)
Look I get it feels weird but in practice most A/B tests are stuff like “does this copy change if ppl use this feature”.
The reasons they don’t is the same reason RCTs for new drugs don’t tell patients either. You end up with selection bias.
> If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all
This is a bit hyperbolic, empirics is something that should be used more by decision makers not just for their own sake but for people who don’t understand why they are making them, especially in government (although It’s harder because finding cases where it’s appropriate is hard).
An A/B test isn’t just about what’s better, it’s about understanding all other things being equal how does one change to X affect Y. Which is information that can be used to inform the design of yet to be build features.
A lot of ppl have bad takes on what makes a product better, and they would otherwise have a greater say in the product design. Some product managers are just really stupid and are there due to nepotism so it’s an external equaliser and allowing the thoughtful ones to have more of a say.
> Users don't want their shit changing all the time.
Yep that’s why you don’t ask them.
I get if you have a specific flow that your use to. It would annoying for me too if that changed (as I’m pretty stubborn don’t like ppl making changes on my behalf), but that doesn’t mean it’s an objective better experience for all users or users who have yet to be familiar with the apps process.
When these products operate in competitive markets and not some winner takes all market these are often about improving users experience.
If this was something more high stakes like a medial trial I’d get it, but for stuff like filling out a document or watching a piece of media. The stakes for most SASS app are really low.
In my experience working in software, empiricism is the only thing valued anymore. Intuition and thoughtfulness is out the window because it's not scientific enough. I would say most software now reflects that - it's almost all bland and statistically optimized to maximize engagement or revenue.
> If this was something more high stakes like a medial trial I’d get it, but for stuff like filling out a document or watching a piece of media. The stakes for most SASS app are really low.
There are plenty of subtle patterns used in SaaS form filling things too. For example notice that the "primary button" is always chosen as the one that will make the company the most money or collect the most data.
Likewise, popups are annoying, but they result in more conversions. Forced logins are the same (how many form filling apps now force you to sign up with an account that you'll never use again, so that you can become a potential lead in future).
Google recently started doing all of these on anonymous searches with a modal overlay and a big bright blue "Continue" primary button that takes you to a login screen, while a "don't sign in" button appears as far less noticeable text above it.
It's at the point now where I'm surprised when any software gives you an option without blatantly telling you which one they want you to pick for their own benefit.
A lot of it seems innocuous but I feel we're at the stage of death-by-a-thousand-cuts at this point.
> In my experience working in software, empiricism is the only thing valued anymore. Intuition and thoughtfulness is out the window because it's not scientific enough.
With things like A/B testing, its not entirely an objective as you need to make assumptions which can be difficult to measure (although typically randomisation solves a lot of them), but you can only measure what you've decide to measure (which isn't random), so you don't know when you're in a local max. So IMO intuition and thoughtfulness is necessary. Sometimes product managers don't listen to data scientists when they say you can't measure Y with X, or the research design violates the required assumptions to make a causal claims (like reverse causality or controlling on a post treatment effect, e.g. employment as control when measuring income after hospitalisation (the treatment)). I think the worse offences I've seen have been from marketing teams.
But proper research design does require intuition and thoughtfulness, because statistical models require thought, like other forms of supervised learning.
I've seen both
- PMs use questionable experiments to justify shipping something.
- PMs dismiss experiments when it was a null result and shipped anyways.
In either case I don't think the methodology is the cause of problems here, although I think shipping with a null result is justifiable if it's a larger unit of work (provided its not a regression).
Sometimes things that have heterogenous effects get measured as a homogenous effect, Like say:
- Your primary user base is X1 and X2 is a larger consumer base but makes up a small portion of your user base.
- Your experiment does poorly with X1, but say there was an increase in user base X2.
- However because X1 dominates the user base and your signups (because say you target ads to X1 over X2), no one drills into the effects on these different user bases, the result gets discarded as it seems to be a bad outcome.
There's valuable information in the experiment outcome but without thought and attention you can miss it.
> Google recently started doing all of these on anonymous searches with a modal overlay and a big bright blue "Continue" primary button that takes you to a login screen, while a "don't sign in" button appears as far less noticeable text above it.
I mean that sucks, but IMO with their market share, the way Google chrome is inclined to develop their product is very different to firms in more competitive spaces.
Perhaps A/B Tests, allows google to optimise the things they are incentivised to pursue, but in the hands of smaller firms with different incentives are willing to tweak things to be more appealing to users when they have far less market power, which I think is probably more the issue in the case of Google.
I just don't think this is a universal problem with the methodology
That’s software though—I see your point for something like content. I’m already used to seeing the title or thumbnail of a YouTube video change as the result of an experiment “winning” but the content itself…that would very jarring.
There are small A/B tests which absolutely make sense. You often see marketing sites making small tweaks to banners and copy. It's not that one has worse UX or even that one is objectively worse, just that different users have different preferences and it's often difficult to know exactly what will work best.
Similarly you can be very confident of something, but A/B testing it still reduces risk. Any significant change should probably always be rolled out to a small fraction of the user base first in case you accidentally change something for the worse.
I agree if you're talking about some BS experiment where a company uses A/B testing as an alternative to putting the hours into product design and user research.
> The funny thing about scoring systems is they are kind of little dictators. They tell you what you’re supposed to want and value. And that’s the weird thing. Scoring systems are little definitions of success and failure. I think one of the biggest differences is that, in games, those definitions are temporary and playful and under your control. And if you don’t like it, you can throw it away and you never have to play again. And in institutions, they’re authoritarian. […] After a period of time, [metrics] seem to drain what’s genuinely valuable from the system because they point people at something that’s very easily and mechanically checkable and measurable.
[0] https://99percentinvisible.org/episode/673-the-score/transcr...
I don't think I am the audience for YouTube any more.
YouTube is converging towards where all the social media sites are:
* shorts
* photo/text posts
* longer videos too - typically 12 minutes approx
* AI videos - some of which are fine but I want to be able to filter them out
* their algorithm/feed is very bad at letting me explore my interests, when I choose to - instead it feeds me stuff that leaves me unsatisified
But I came here years ago to watch TV made by people NOT bound to 12 minutes. I watch pretty much nothing else at night on the couch except YouTube but I am coming to realise it's no longer what I want.
That "original YouTube" seems to be gone. Nothing has replaced it.
Youtube certainly took a lot of unpopular decisions lately, but being able to stream any hi-res video ever posted in an instant and for "free" is a miracle. I will be happy with youtube for as long as content creators I care about are happy and I can get their fresh content via browser, yt-dlp, or other means.
Alternatively, it's there in my RSS reader app, where I can see new stuff showing up without even visiting or interacting with YT.
But it's not that simple. We definitely need to "pave the bare spots" (desire paths); It's just that we need to start off, at what we sincerely believe to be an optimal place, knowing that it isn't, in fact, optimal.
Absolutely, and unfortunately in my experience many people implementing A/B testing actually believe that it is finding the "best outcome for users".
I'm sure if we blind tested people to see whether they consumed more when unknowingly given cocaine vs protein powder, we would see cocaine win the A/B test.
intro (2x 15 second clips pulled from random places in all the other sections)
section 1 (a/b/c)
section 2 (a/b/c/omit)
section 3 (short/long)
section 4 (a/b/c)
It would be like a choose your own adventure video, without the choosing, or the adventure.
this seems wildly naive
Truth is: “number go up” is the only valid strategy for growing because that’s what favours the platforms the most.
Unfortunate. Depressing.
But the vast majority are engaged in attention baiting; outrage, grievance, conspiracy, FOMO.. Take your pick.
Even some oldies like GN are churning out grievance and drama.
If almost everything is different, it’s hard to learn for next time what exactly what led to a change in which ever dependent variable your observating.
It's critical that specialists who know whatever they do, work with AI to see the path ahead for their area, while the mega models try to be a single mega model for everything.
And even in the software world where there's a place to test say accessibility and whatnot this attitude is so prevalent that everything looks the same. Nobody's making a website like Larry Wall any more[1]. Can people please start making things out of their own volition again instead of following this brain dead attention economy
It seems easy enough to say "keep the youtube we love" but how do you think this variant came about in the first place? I can assure you there have been numerous A/B tests that have led to the current feature set. And even if its a local rather than global maximum, at least there are measurable qualities by which it is preferable. Also - do you believe that everyone who says this is harkening back to the same historical reality? This is quickly approaching "make youtube great again" territory - when exactly was it great again? and why? This statement is easy to agree with and hard to prove.
It also seems easy to handle links to different variants; just supply a URL parameter. This is a non-issue.
If someone out there has a non-tea leaf divination style alternative to measuring things as a way to determine success, I'm all ears, but I don't believe in fortune tellers, and people who make software shouldn't either.
Just sounds rubbish both for viewers (look how terrible titles and thumbnails are after a/b tests) and uploaders alike