Manaus is in the middle of the Amazon.
Needless to say, a bit scary to hear that, but we landed without issue.
They told us we had two choices: the nice hotel with a shared room, or the lesser nice hotel with no roommate. I chose the latter. When we go there, they said, "oops, sorry, short on rooms!" So I had a roommate.
Wandered around Manaus, took a skiff out on the Rio Negro. Saw pink river dolphins. A little boat approached us and a kid handed me a sloth, and then demanded I return it with a twenty dollar bill.
The airline got us another plane 24 hours later. Made it back to the US safely.
A few weeks later, the airline reached out and said "Here is $100 for your trouble."
I declined to take that offer. I had missed several business meetings that cost me actual money. I couldn't donate blood for years because I had been to the Amazon and was tagged a malaria risk.
During the many arguments with the airline I threatened to take them to small claims court.
I got a really strange response over email which I clearly wasn't supposed to see. A representative from that airline was asking internally if they could put me on the no-fly list. That was really chilling.
But, this is the kind of information I'm worried about when a vendor sells my data. If Google wanted to sell a product to the airlines that offered to keep annoying people like me from purchasing flights, they could do that with that email chain. I'm skeptical it'll be wiped correctly. Isn't my poor writing style basically my signature? How do you wipe that?
What if they only pretended to forward it to you by mistake and you /were/ supposed to see it?
As in "it would be a shame if you could never fly again".
https://www.nytimes.com/2006/08/09/technology/a-face-is-expo...
1. Google could do it.
2. <This space is intentionally left blank>
3. Therefore, Google is doing it!
(Step 2 needs to be filled in a bit for it to be a good argument. Generally, analogies don't quite make the cut.)
I don't know exactly what gets sent to google, but it's certainly enough to identify and track (retrospectively) a huge part of the world's population.
Now I know you get tracked by the celltowers anyway, but still. Navigation works fine with the accuracy offered by just using GPS and it doesn't need all the wifi scanning, it's pure data harvesting.
However, you can increase gps accuracy using wifi. GPS is not that precise (part of which is government regulations) and WiFi does help immensely with accuracy.
You can argue that its a bad deal, shouldn't be allowed period, etc. But that is a different argument than saying that there is no scrutiny.
Sure 50 M or even 1 B might be peanuts for faang but still there is real progress.
Support Noyb at all costs
When you join a company, you typically sign an agreement that talks about how the company owns all your output. Thumbs upping a Teams message is work output and they own it.
Every email sent and received. Every keystroke. Etc etc etc.
If you don’t want your employer to log and sell it, start your own company. Or use a personal device. I do the latter.
In the US, whoever owns the computer owns the data on it. Courts have routinely ruled that you have no say in what other people collect about you. The goal of bankruptcy courts is to minimize the losses of the creditors. And bankruptcy courts routinely rewrite contracts except where statute prevents it (like mortgages).
In the EU, you own the data about yourself. A lot of people utterly hate GDPR, but that's reason that you own the data about yourself.
Pretty easily. Do you look sort of like their enemy of the week?
That's all the excuse they need to grab you.
They don't actually give a shit about being accurate in that identification. Accuracy in repression isn't the goal. The goal is a visible, public show of force.
I think this is a far more important question than do you believe in right or left economic policy. idc, I want you to know, do you think a human’s rights, especially many humans together, outweigh that of non living entities like large tech companies.
They don’t.
That's just untrue.
They may have rights that you do not like that they have, but they do not have MORE rights than humans.
Similarly, companies face all sorts of penalties and restrictions that constrain them in particular ways, including very severe penalties. There might be some actions that you think ought to have more severe penalties than they get, but that doesn't make your statement that companies are immune to real consequences false.
1) +95% of the population live on the coast very far away from the Amazon. Most of the population has not been there. Most of the coast has a very different jungle biome called Mata Atlantica and the countryside close to the coast is not that different from temperate forest of Europe. That is what most all Brazilians are used to. There is a significant population in the arid northeast though and the cold south as well (which is even more similar to europe).
2) Manaus is the biggest city in the Amazon and it is huge developed place (and has been for decades). You are not in the middle of the jungle if you land in the airport. The countryside around the city is jungle though.
3) Brazilian people do not necessarily like or are used to tacos and spicy food. Mexico is _really_ far away from Brazil.
I had a similar blood donation issue for visiting a particular island in the Philippines, and could not donate for 4 months.
https://www.cdc.gov/tick-borne-encephalitis/prevention/tick-...
Your points are valid. And there probably are plenty of Americans who needed the correction. But still.
Do you think Canadians know all these facts about Brazil? Do Indians? Or Swedes? It's only Americans that are smugly called ignorant for not knowing about the entire world.
(Which is to be expected, proximity and all)
My grandparents came to Brazil right after WW1 way before the Nazis came to power. High ranking Nazis fled to south america because there were a lot of germans living there already. Nearly all german people who moved to south america did it way before WW2.
I just run into this stuff a lot living in Europe.
The OP article is quite poor in terms of information provided, but the buyer (Google) had to explicitly agree not to attempt to re-identify users. https://www.axios.com/2026/08/17/google-spirit-airlines-bank...
> Facebook has been fined €110m (£94m) by the EU for providing misleading information about its 2014 takeover of WhatsApp. (...) When Facebook took over the WhatsApp messaging service in 2014, it told the commission it would not be able to match user accounts on both platforms, but went on to do exactly that.
https://www.theguardian.com/business/2017/may/18/facebook-fi...
ahhh, there seems to be different no-fly lists? The one Im aware of is the one for terrorists and moneylaunderers, and usually they will not tell you who put you on that list :-D
They may look cute and docile, but also have a panic reflex to grab onto the nearest tree-like thing they can feel with their claws, including their would-be attacker. Or if they just feel like they’re about to fall off their “tree” since they have horrible eyesight and get confused easily.
Which then leads to a cycle of pain and violence as they just dig-in harder and harder while trying to free them and/or fling them around wildly due to the human's “get this thing off me” response to sharp claws digging into their flesh.
There are some painful-to-watch videos on YouTube of this phenomenon from unsuspecting passersby tying to “help” them off the road/beach/etc and ultimately making things worse for everyone involved.
If someone else shares information about me, without my consent, and someone uses that to nose in on me, that feels creepy and problematic.
If they want to do that they already can, thanks to the public’s overwhelming appetite for “FREE” overriding every single other possible concern.
Which is funny to point out on a post about Spirit, since that was an airline built to serve the customers for whom cheapness was the overwhelming single concern.
In the early pioneering days of the commercialized Internet, email addresses were inextricably linked to your ISP. You paid for an ISP connection and you got an email box, with MTA and MUA service to match. You were reluctant to switch or leave your ISP, because that also meant leaving behind your email address. Of course, a minority of nerds got around this with their own domains, etc.
However, it seems that Google, AOL, Yahoo!, Hotmail, and other players got into providing free email services and eventually grew into giants that supplanted every other MTA service. This was not an accident and it was not merely our appetite for “FREE” but it was a very calculated plan by the industry. Those early ISPs did not have a business model that admitted monetizing our private data; they seemed to have a more respectful attitude for keeping it private. Perhaps that was a result of being telecommunications-based companies, rather than advertising or entertainment.
If a service like email provides such endless treasure troves of personal data, including a social graph and glimpses into our private daily lives, why not provide it for free and monetize opportunistically on the data itself? The free email services killed the paid services, not by being better or cheaper, but by being bigger, centralized, and more persistent. The main reason I signed up for Yahoo! was because it would be an utterly stable presence. I saw my parents and others so hopelessly attached to an ISP-based email, but I couldn't end up like that.
After streaks of losing my home and non-payment of bills and moving around over decades, the most stable point of contact for me has been a "free" email address.
"If you’ve flown Spirit and worry that Google will soon know about a testy conversation you had with the airline’s call center, you’re being told not to worry. The court filing says the data was deidentified before being put on sale and Google has promised to scrub any PII it finds in the trove."
Accidentally making a machine learning system that happens to (potentially) do it is a different matter.
This part is worrying.
That email you accidentally received really bothers me. I don't understand why a CS rep would get this invested to the point of wanting to cause you real harm. They're not the airline. The psychology is fascinating. There are people out there who feel like a mild short-term inconvenience to them where they have no stakes somehow justifies life-changing harm is kinda frightening, honestly.
I'm reminded of the Yahoo search data fiasco that was allegedly anonymized. Turns out, it wasn't so anonymous [1]. For one thing, people tend ed to search their home address. Whoops.
You mention writing style. We already have LLMs quite capable of copying a writing style. It's a natural extension to say we can fingerprint writing style too.
But here's another aspect. Imagine you're in a relationship with someone and you somehow fingerprint their personal data with a company. For example, you use their Netflix to like 5 very obscure movies, to the point where it's likely unique. Now imagine that Netflix's data gets released in an "anonymized" form and you can now find it based on those obscure likes. I can imagine many scenarios like this. And there's no text involved here at all.
[1]: https://www.vice.com/en/article/yahoos-gigantic-anonymized-u...
This is the kind of power tripping that easily corrupts people, especially those who don't have much power outside of work. The US national security apparatus is vast and powerful. Many people get giddy at the thought of inflicting punishment on those who "deserve" it. I can easily imagine the kind of CS rep who would delight in putting a customer who was merely rude on the US federal no fly list. Police are the most notorious for this kind of petty "you will respect my authority" But it's the same kind of person who, as. fast food worker, thinks they're the hero by spitting in the food of a customer who doesn't deserve it.
HN / hacker culture celebrates this in the 'Bastard Operator from Hell' [0], the sysop who will ruin your work and life with their IT wizardry, no matter if you're an intern or CEO, if they do not feel respected.
Two possibilities come to mind... 1 - The CS rep has been instructed to do this. Scary, but corporate leaders can be assholes and wield lots of power within their orgs, so doesn't seem completely unlikely to me.
2 - The CS was just a dick.
Frankly, given the behavior of various SuperMegaCorps over the past few decades, I'm going with #1.
Surely if this is the scenario you're worrying about, and you accept the lack of regulatory protections, Google buying Spirit's data is a good thing, right? Much better them than the airlines who you already know to be corrupt?
Now think about all the Gmail data Google has.
"Deidentification" seems really murky and imprecise at best.
But I think this is a place where we should apply bounded distrust: there are lots of places where we should distrust Google, but reidentifying people in an explicitly deidentified dataset isn't one of them.
E.g. the parent wrote that he fears, he could be identified by his writing style, which is totally plausible. How would you "deidentify" this?
Of course, takes a lot more effort than not doing proper deindetification in the first place but if they wanted to appear like caring about data privacy they still have enough data points to correlate the sets later on (and/or over time).
Don't worry. Spirit probably lost all of the emails from the customers (or they were devnulled) and 90% of the data is probably autoresponder messages promising the company would respond.
The other 10% was probably the meme collection of the executive management team.
> There’s also operational data in the trove, describing over 763,000 flights, five million crew pairings, more than 1.2 million fuel slips, and records describing purchases of 787,452 parts.
> Google has reportedly said it bought the data to improve its AI services.
Gives "this call is being recorded for training purposes" new meaning.
I got a call that started with the usual automated message, "This call is being recorded." After the person joined, I pushed the record button on my iPhone, "This call is being recorded."
They were surprised and asked why I'm recording.
I said, "You're recording, so I'm recording too."
The rep insisted that the company doesn't like this but that they will continue with the call anyway.
Anyway, I wish people took this stuff much more seriously. It always seems to boil down to, "I don't have anything to hide" type of conversations and I've never managed to convince anyone that privacy as a concept isn't about having something to hide.
I kid, but...
It's probably the precursor to insurance denials and job screening.
I got banned from r/technology a few weeks back for decrying tracking in AI content. The community was piling on saying it was okay because it removed AI content or made it easy to spot. I made the counter argument that watermarks would find their ways into everything and eventually be bound to attestation. The mods didn't like that. (Yet another structural problem with the lack of p2p self-service town squares.)
The socials are training the next generations for broad acceptance.
Yup.
Elsewhere in another front page thread today: "oh but apps blocking screenshots because of 'sensitive content' can be bypassed by taking a photo of your screen with another phone".
Any tech-savvy person with two brain cells reading this and that: "gee, I wonder if the same magic imperceptible watermark that survives multiple rounds of cropping and printing and scanning, that's used to tag AI-generated content, could also be used to tag sensitive data, or ads, or which app is rendering it on screen, and then the camera app could refuse photographing it...".
I don't know why people don't see that AI watermarks are DRM, and DRM is universal, and there are many clients...
I really doubt all this stuff was “de-identified”
De-identified but far from useless.
as an example, they can remove the names off these sales data, so you can't identify who purchased what items. However, the purchaser would be identified by some sort of number, and you would be able to extract information about purchasing habits, and aggregate these habits into usable information for advertising purposes (like targeting and profiling).
And that's before AI training for LLM purposes.
90% of e-mails and Teams communications are inane. Polite banter, "thanks for taking care of that, I appreciate it" "please route the forms to Janet this week because Bill is on vacation" "unit will be un available until the parts come in" . I can't see the intrinsic fact value of this kind of communication without screening it. And after screening, the gold nuggets would be minimal.
(Something something we will add your distinctiveness to our own, you will be assimilated, ...)
(Hell, the fact that it's all from one org would make it a great dataset to have in the open for sociological studies. I bet that today, aided by LLMs to sift through it, you could use it to map how information flows through a large org - how incident on the floor travels through time and layers of management until it reaches the C-suite, what of it survives, how it gets reacted to, how the reactions flow down...)
I have a hunch what this is for. AI companies want to make bigger inroads into nontechnical work settings. But LLM progress outside of fields where verifiable rewards for RL post-training can be synthetically generated (coding, math) has been pretty flat. Buying years of operational data from a company like an airline could be used to reconstruct long-horizon task trajectories in areas like customer service or marketing.
This. Training data companies like Mercor are even creating simulated companies to create similar data, so they can better automate even more white collar work:
https://www.nytimes.com/2026/07/10/business/ai-white-collar-...:
> The data-training start-ups see a lucrative opportunity in recreating workplaces in miniature: controlled environments in which their gig workers can evaluate and reproduce emails, memos and slide presentations in context. The information emerging from such a setup, the companies boast, will help shrink the gap between what A.I. models can accomplish and what office workers actually do from one minute to the next, as ideas and instructions flow between meetings, documents and applications.
> Scale, for example, has said that “our environments replicate real-world workflows,” and that its contractors “curate artifacts that capture the complexity, ambiguity and edge cases of real professional work.”
> Executives see the models’ shortcomings as a sign there’s more for them to do.
> “I often use Claude Cowork, right?” said Edwin Chen, the founder of Surge. “And even though Claude Cowork is incredibly smart, oftentimes it doesn’t quite understand the nuance of Slack. It doesn’t quite understand this ambiguous question I have. It doesn’t quite understand where to go and find this Google document.”
> According to the Bloomberg Billionaires Index, Mr. Chen’s stake in Surge — and his vision for what it could become — makes him the 258th-richest person in the world.
> “I often think about us as essentially, like, the school for A.G.I.,” Mr. Chen said, referring to a prophesied level of A.I. that surpasses human intelligence. “A.I. comes to us, and A.I. learns to run the world.” He and his customers at the big A.I. labs, he said, are designing the curriculum.
If it comes with the context, yes. More data the better. Someone considered it served some purpose at some point. Thus it contains, no matter how tiny, a sliver of information.
The perfect data to train an agent swarm on how to run a company.
Maybe it's innefficient and inane, but it's how you start.
The first LLMs, GPT-1, 2, were trained on complete garbage, the average document from the common crawl is random non-sense, yet they worked, and now we can use LLMs to filter the data for the next training run.
3 years ago, the Wall Street journal covered a company trying to use AI to generate documents required for the approval of new nuclear reactor reactor designs, which sounds dangerous, until you get to the point where they'd cite 2 million pages as necessary for a typical application. [1] I don't need to explain why no individual or institution could read that, much less examine it in detail. I remember a rather funny question I found in a comment to that story - "How many pages of those 2 million could contain pornographic images before anyone notices?"
It's obvious why companies are so desperate for training data - a text generator of sufficient quality is more conductive to the growth of the nuclear industry than any scientific breakthrough in nuclear physics. (And of course, if you want to prevent the development of a nuclear reactor by your competitors, being able to produce millions of pages of high-quality objections will do the trick.)
And it's not just nuclear power. When it comes to stuff like building rail lines, apartments or power plants (both conventional and renewable), you'd find that the main bottleneck is the necessity to produce documents. And of course, many documents can be subject to judicial review - a process that consumes even more text.
[1] https://www.wsj.com/tech/ai/microsoft-targets-nuclear-to-pow...
There also is the liability angle, where no one necessarily reads the document until the document is relevant to some situation in the future. Then you better have that document on hand.
I guess the appropriation "any sufficiently advanced stock market is indistinguishable from entertainment" doesn't quite stop at equating the trade floor with a casino. At some point, entertainment also becomes the modus operandi of corporations.
(The PDF mentions "the standard for deidentification set forth under the California Consumer Privacy Act", which suggests this is all pretty well legislatively understood.)
> For example, one initial bid requested certain customer list information; however, by the first round of the Auction, the most competitive bidders had agreed to bid on an asset schedule that expressly excluded PII.
Huff, what a relief!
> 100 million emails
How does one deidentify 100 million emails?
Absolutely scary
On real life data on operations of a real large company.
Internally, most big companies are probably just as big of a mess, if not worse. But you can't get that data easily.
Yeah this (and lawsuits/investigations) are why, the employer owns the data, in some contexts (like this one) it can become an asset (or a liability) but in either case it's not yours.
Of course that only gets you part of the way there anyway see Twitch recently opting in all users by default to mined for AI and only adding an opt out after backlash with a quote that was so on the nose it made me stop "If we'd have asked them to opt in, they wouldn't have opted in" (paraphrasing but it was that blunt).
> If you’ve flown Spirit and worry that Google will soon know about a testy conversation you had with the airline’s call center, you’re being told not to worry. The court filing says the data was deidentified before being put on sale and Google has promised to scrub any PII it finds in the trove.
… and even if someone can prove that they didn't, the only consequences will be a teeny-tiny slap on the wrists.
This kind of stuff needs to come with promises to pay the P in the PII big bucks if the I is indeed I.
It should really make us appreciate living in a country where freedom is the default.
No I don't want to just use your enshittified service for free. Fucking pay me and watch me all you want :)
Martha Wells hit it nicely in The Murderbot Diaries.
And how long before Google and Microsoft add to their T&Cs that all your anonymized email will be used to train their A.I.?