It shocked me how there is absolutely no "right" answer.
If you are teaching English for travel, then you're prioritizing a lot of stuff around bathrooms, transportation, menu items, etc.
If it's for understanding TV, it's a lot of words like "murder", etc. Depending on which TV shows you want to understand.
If it's for reading the newspaper, you don't ever need to know "bathroom", but you sure do need to know words like "congressman".
While if you are living somewhere, it's really important to know a lot of basic supermarket items that you wouldn't prioritize for other usages.
Also, while it's easy to calculate word frequencies for stuff like newspaper articles, there aren't any good statistics (last I checked) around just normal everyday conversation. Because that stuff isn't getting recorded and transcribed. And the substitutes -- transcribed speech from TV, radio, podcasts, etc. -- is not the same context as the random stuff you say at home and during an average day.
Linguists have resorted to spying on people in parks and shops and restaurants, to get "natural" language unperturbed by observer effects, audience tailoring, the filter of writing, etc. (In most cases that would not be considered ethical today!)
The gold standard is something like: get acculturated to the language community you wish to study, make some acquaintances and friends, somehow finagle an invite to their homes or workplaces or social venues, and then with their permission, record their conversations as unobtrusively as you can.
There is so little data on how language is really used it is quite bedevilling. As soon as people know you are listening they change how they speak.
Hence the common phrase, found in most language phrasebooks: “is that a linguist hiding behind that bush?”
Ever wonder who decides what's self-explanatory and what resonates?
For reasons unknown to me, instead of using the word with the meaning they probably are trying to convey, they rather use vague analogies or metaphors.
It always made me question if Americans in contrast to the British are a bit afraid to be precise.
Of course, American corporate-speak is unique from normal conversations, and people intentionally use vagueness to make sure there's general agreement before saying something concretely that might be contentious or cause problems.
This is more corporate/middle class.
There's also a working class register which is very direct and sweary, and extreme insults are actually meant affectionately.
Even more confusingly, the two can overlap, and the same people will switch between them - sometimes in the same setting - with relative equals.
Americans don't seem to do understatement. "My business died" becomes "My incredible journey." Relationships are cross-bred with hustle, so you can be friends with someone and they'll ask you to invest in their MLM scheme.
Instead of using understatement US corporate speaks uses noise to act out hierarchy, to gaslight and to misdirect.
I was talking to someone today who was asked to produce a report. They sent it on time, and then the team lead called a meeting to ask why it hadn't arrived.
The report was literally sitting in their inbox unread. Everyone could see it, but they all went through the motions of pretending it wasn't there and developing an action plan for future delivery.
It's all very David Graeber. So far as I can tell the US seems to specialise in situations where corporate insanity is happening but the sharp edges have been filed off, so it looks like productive activity higher up the pyramid.
This is not specific to "Americans". Learn any language and you will have the same impression, and experience the same gulf between textbook usage and everyday usage.
I live and work in a foreign country, and have a modest but functional grasp of the language here. Which is to say, I know how to say 'sustainability', 'union-negotiated collective agreement', and 'offensive conduct in the workplace'.... but if you asked me the words for 'cedar', 'robin', 'pond', or 'linen' I would be struck dumb. And yet presumably every ten year old I walk past in the street would know those words as comfortably as I knew them in English as a ten year old in England.
Just earlier I was talking to them, trying to say "when life gives you lemons" and being unable to find an equivalent phrase. I hope they understood my lemons reference regardless.
Also, both because of the familiarity and how English form words, English assign very specific non-compound words to them (rather than "xxx tree/bird"), making it more difficult for people from a completely different language family to learn them.
But of course this is not only English. In Chinese a child is taught to act 孝 (filial piety, respect and obedience to parents) and know how your uncles 舅 (mother's brother) 伯 (father's elder brother) 叔 (father's younger brother) relate to your parents. I'm sure it's OK for a foreign learner not to master them.
They likely show up in Western works a lot because they’re mentioned in the bible a good bit.
Copying lyrics from my favorite songs of the 80s was helpful. And then later on TV series and movies in english.
I've watched Friends, HIMYM and Californication* more than once (it's probably the reason I swear so much) and those really helped.
Now I raised my kid in english: from the beginning she was only ever allowed to watch TV shows in english. And she's always been to international schools, in english too.
The other day she taught me a simply word which, surprisingly enough, I didn't know: a "coaster" (a drink coaster).
My absolute favorite is when she uses english grammar and applies it to french: the rules about where adjectives do go vary between french and english and she regularly messes it up in french (supposedly her "native" language but at this point I wonder whether english ain't a first language, and french second).
And one thing she has that I'll never have is a beautiful, proper, english accent. Mine shall forever be a strong french one.
1. TV vocabulary is 2,000 words
2. High school vocabulary is 10,000 words
3. College vocabulary is 30,000 words
4. English language has 1,000,000 words
This is good news - you can be proficient in a foreign language by learning only 2,000 words!
Depends what you mean by "proficient". I currently know ~6k Spanish words and would hesitate to self-apply that word. Others might, but damn languages are hard.
I was pretty confident in my English but I couldn't understand half of the sentences since I didn't know any of the vocabulary related to cooking. From tools (spatula) to ingredients (parsley) or even actions (frying). My delusions of proficiency shattered as I found out my everyday vocabulary is smaller than an elementary schooler.
Why? Having spent a good amount of time living in Shanghai, I found it important to be able to understand menus. But there's no pressure to know the words for supermarket items; you can just go to the supermarket and look for the item.
Otherwise your point is correct; all semantic words are equally difficult and which ones you know depends on the things you like to talk about. Grammatical words are more difficult, and more important, but this is so widely understood that language-learning material already treats them as an entirely separate class of things to learn.
> Also, while it's easy to calculate word frequencies for stuff like newspaper articles, there aren't any good statistics (last I checked) around just normal everyday conversation. Because that stuff isn't getting recorded and transcribed.
(1) You seem to want COCA, which includes a bunch of transcribed telephone calls.
(2) Word frequencies are still the wrong concept. If you want to understand a particular document, you need to understand almost all of the words that appear in that document. (You'll be able to learn some of them from their use in the document.) If you decide to learn a list of "frequent" words, you're unlikely to be able to understand more than a couple of isolated sentences in any given document.
I didn't learn English by reading vocabulary lists or dictionaries. Pretty much all was learned by context.
I would blame inequality on this one. In a more unequal world tribalization is a survival strategy and language follows.
When you see everybody else as your equals then focusing on describing that individual person, instead of their group, makes more sense.
Economic inequality affects deeply how we think about others.
What I see is shifting from a focus on individuals and close personal relationships to a focus on the functioning of society and, well, identity.
It seems like a swing of the pendulum to me. Encouraging people to create their own identities, creating vocabulary to describe one's identity, and advocating for a restructuring around society that is relatively collectivized economically, but culturally celebrates individual diversity. I'm not a professional historian, but I think that's somewhat of a new phenomenon historically.
And it's clearly a reaction to the perception that the ideals of equality and being uniformly kind to each other we're not enough to bring about a just and equitable society, and if anything were abused in order to suppress diversity. We still have to be kind to each other at a basic level, but it's not enough anymore.
So maybe now there's too much focus on that stuff and not enough focus on the part where you should actually just be nice to your neighbors and be humble and all that. Or, maybe people are disproportionately writing about it, because it's a new thing, and a lot of people know a lot to say about it, and a lot of people want to talk about it.
I don't think it makes a shred of sense to conclude that somehow the world is now more equal than it was in the 50s. It's disappointingly far from equal, though, and I think the change in lexicon is a change in people's focus specifically because the old system clearly wasn't working, and people are eager to try something different.
In 1953 people were not exposed as much to different groups of people far away.
I used the examples of Latin to Spanish and English, or Old English to new English.
I say that to say, languages changing over time is to some people not actually something they believe happens. The facts are right there in front of you, but some people cannot have their minds changed no matter how much sense you make.
> (Formerly in Israel, when a man went to inquire of God, he said, “Come, let us go to the seer”; for he who is now called a prophet was formerly called a seer.)
I love that one because it translates so well into English because we also have an older and a newer word for close enough to the same thing. Though when saying older and newer I should note that the words “prophet” and “seer” have been used consistently since the very first English translation, Wycliffe’s in ~1382. Maybe it’s more the feel of the words. (In Bible Gateway’s catalogue of English translations, only the CEV doesn’t use both of those words: <https://www.biblegateway.com/verse/en/1sam9.9>; I disqualify OJB as not English.)
That’s the only one I can think of off the top of my head that’s specifically about language, but there are plenty of others about changed customs over time, in Old and New Testaments.
There are some databases but e.g. they are biased towards Wikipedia and web which makes some very obscure words at the top of popularity (like some technical words which are present on each wiki page like Datenschutz or Impressum).
Having said that, the categories that shrank all did so by a big enough percentage to also shrink in absolute number of words, so at least that isn't a problem.
I'm not going to finger-scroll or down-arrow the whole thing.
Don't break scrolling. Please.
That seems not to be available at either the Internet Archive nor Archive Today:
<https://web.archive.org/web/20200000000000*/https://ello.co/...>
<https://archive.is/https%3A%2F%2Fello.co%2Fdredmorbius%2Fpos...>
Sic transit gloria data.
Update: At least some of the Wayback captures have the actual page content if you view source. See for example: <https://web.archive.org/web/20210524202824/https://ello.co/d...>.
(Another reason I'm pretty salty about breaking basic HTML page functionality for slick trix.)
Must just be the combination of that increase as well as other words decreasing in usage I suppose. E.g. perhaps we're a bit less keen, but also much less passionate, so keen ends up making the cut.
Saying "I'm keen to do x" or "he's keen to do y" is a fairly common normal usage but having said that, wouldn't normally be my first choice of language as a millennial, maybe does sound a bit old fashioned
Oughtn't really affect emotive language like 'keen', though.
This is something I've been thinking a lot about. We have trended from subjective language to objective language. Why?
Computing. Software is written with objective language. Everything is clearly unambiguously defined. Blue is no longer a category, it's #0000FF. Logic must always reduce to a binary truth value. Most of what we have to talk about is somehow relative to software. Software even structures most of what we write! We don't just talk to each other, we tweet, email, message, post, search, etc. These structures each imply a specific set of phrase structures that can make sense.
Lately, it's hard to go even a day without reading some complaint that such and such was written by "AI" (an LLM). Why is this so obvious? Well, the core advantage that LLMs provide is that they don't compute. Inside an LLM, there is no arithmetic, no logical branches, no truth values. Phrases aren't generated to define or to resolve. They are generated to continue. Sure, we can direct the story to follow the steps of logical deduction, but that isn't anything like calculation. An LLM simply isn't invested in logic, precision, correctness, etc. the way we expect modern writers to be. It's not the em-dashes or the word choice that illustrates this, it's the fundamental perspective of the system.
We are sorely missing subjectivity. Natural language never was, and never will be, computable. You can't reduce a natural story to binary truth values without choosing an arbitrary perspective that resolves its ambiguity. The more precisely abstract our language gets, the more detached from reality our stories become. The more objective our assertions about reality are, the less relevant they can be.
My answer to this is to make the arbitrary choice of perspective a first-class feature. If we can explicitly decide what meaning is relevant, we should be able to weakly solve natural language processing. It seems like a pretty simple and obvious idea, but so far is easier said than done.
Language has to do with our Monkeysphere, that humans over long periods of time in the past were limited to a very small subset of people they interacted with socially and closely. A few institutions likely had a large effect on your language like the church depending on where you were. After that it was the people you interacted with to stay alive. Because almost everything was in person or person to person transfer of information a lot of socially encoded clues were involved which lessened the need for well defined words.
Books were the first stage of homogenizing language as they could be shared over long distances and to many people, but more sequentially than latter forms of communication. After that radio and TV had a huge effect, for example the 'General American' used in broadcasts that was based heavily on a midwestern accent.
As we encroached on the 70s and 80s the previous technological advances and things like high speed interstates and trucking shrank America to something you could drive across in less than a week, and you could reach anywhere by voice nearly instantly. Suddenly people in California, Texas and New York all could be in the same meetings and local colloquialisms would need explained, so people would trend to a shared vocabulary.
It's also odd to me to say an LLM isn't subjective. Each LLM has it's own behavior, it's that there are like 20 or 30 big LLMs in all, and people are using them millions to billions of time so we're getting that one LLMs language everywhere. And that's why I disagree and will say natural language is computable, but it's also lossy and probabilistic. And for the most part it's single prompt and being ran by the user for the cheapest price possible.
https://news.ycombinator.com/pool
But not in this case. It's a slow Sunday afternoon and getting a few upvotes quickly is enough
Would you mind explaining to us how it works?