English can be traced back to Proto-Indo-European (PIE), a reconstructed ancestral language shared by nearly all European languages (including Germanic, Romance, Celtic, Slavic, Baltic, Albanian, and Armenian) and many Asian languages (such as Hindi, Pashtu, Kurdish, Farsi, and Bengali). This linguistic family tree was established through comparative linguistics, where scholars like Sir William Jones in 1785 noticed striking similarities between Sanskrit, Greek, and Latin, leading to the hypothesis that these languages descended from a common source. Linguists reconstruct PIE by identifying sound correspondences across related languages—for example, the word for 'hundred' shows variations (hundred, centum, hekaton, sto, shatam) that can all be traced back to the PIE root *kmtom-. While PIE represents our best guess at what this ancient language might have sounded like, some linguists propose even older ancestors like Nostratic or Proto-World, though these theories remain controversial due to the speculative nature of reconstructing languages without written evidence.
Tracing English to Its Proto-Indo-European Roots
Added:How far can we trace English back?
Really far.
Not just through Middle English and Old English, but y back thousands of years.
Allow me to introduce you to English’s oldest ancestor.
Welcome to another RobWords.
This is the English language family tree, and as we trace it back, we essentially go further back in time.
English sprung out of the West Germanic languages - the same as modern Dutch and German. But the West Germanic languages share a common ancestor with all the other Germanic languages, the Scandinavian languages and Gothic, which is extinct no.
An ancestor that gets called the “proto-Germanic” language: a language that was theoretically spoken by a single group of people who would eventually go on to become the Swedes, the Germans, the Dutch, the English and more.
But can we go even further back than that?
Yes we can because linguists believe that an even older language is not only the shared ancestral language of English speakers and all of the Germanic peoples, but also speakers of the Romance languages, the Celtic languages, the Slavic languages, the Baltic languages, of Albanian, of Armenian, and of Greek. An ancestor shared, not just by almost all of the European languages, but also of Asian languages like Hindi and Pashtu and Kurdish and Farsi and Bengali.
Just look at this map.
Languages that developed as far West as in Iceland and as far East as in India are thought to share a single, common ancestor.
And that ancient ancestor is known as Proto-Indo-European.
Humanity appears to have caught the scent that would lead us to Proto-Indo-European as much as two millennia ago.
In the first century BC a Greek scholar called Dionysius of Halicarnassus suggested that Latin might actually be a mixture of Greek and other languages.
He was actually trying to prove that the Romans were in fact Greeks, which was nonsense, but he was right that the two languages were more closely related than had been previously thought.
However, it took us many, many centuries to eventually start joining up the rest of the dots and to see that the links weren’t limited to southern Europe either.
One of the leading figures in this was a fella called Sir William Jones, a Brit living in India, who had studied Latin and Greek and now found himself taking an interest in Sanskrit a language that holds a similar historic and prestige position in South Asia to Latin in Europe. He not only adored Sanskrit but also felt he’d picked up on something that should have been obvious.
In 1785 he wrote: “The Sanscrit language, whatever be its antiquity, is of a wonderful structure; more perfect than the Greek, more copious than the Latin, and more exquisitely refined than either, yet bearing to both of them a stronger affinity, both in the roots of verbs and in the forms of grammar, than could possibly have been produced by accident;” Intriguing.
Go on… “so strong indeed, that no philologer could examine them all three, without believing them to have sprung from some common source, which, perhaps, no longer exists.” And there it is. He’s doing it. He’s describing Proto-Indo-European.
And he went a little further too, saying: “there is a similar reason, though not quite so forcible, for supposing that both the Celtic, though blended with a very different idiom, had the same origin with the Sanskrit, and the Old Persian might be added to the same family.” Weird use of “the” there.
Anyway, Sir William wasn’t the only linguist to make an observation along these lines and he certainly wasn’t the first to link Sanskrit with Old Persian.
But from then on, this idea of a common ancestor linking languages across two continents is established with the Germanic languages also included and therefore, English!
This great-great-great-great-grandmother tongue was given a handful of different names.
Some scholars called it Japhetic a Biblical reference to one of Noah’s three sons, who are supposed to have fathered the three divisions of humanity.
Others called it Aryan. But that term fell out of favour some time around the 1940s.
It’s hard to imagine why.
It does however still get used to describe a smaller family of languages, safely far away from Germany.
The new, linguistic superfamily was also termed Indo-Germanic in early descriptions a name that describes the wide geographical spread from as far south and east as India and as far north and west as Northern Europe.
And indeed, it’s still called Indo-germanic in German. It’s indogermanisch.
But linguists - writing in English, at least - eventually settled on describing it as Indo-European and the original language as Proto-Indo-European.
Proto is a Latin prefix meaning earliest or original.
So how similar are all the many languages that come under that massive Indo-European umbrella?
Are they similar enough that say, an English speaker way over in western Europe could understand Sanskrit from over in India?
Probably not, right?
Well… you would be surprised.
Let’s take those examples, and throw in Latin for the sake of further comparison.
Now these are the numbers one to ten in those three languages and though the similarities aren’t always obvious, they are very much there.
Just look at two: two, duo and dve.
They are so similar - particularly when you bear in mind that “t” and “d” are produced in a very similar way in the mouth, you just wiggle your vocal chords to turn a t into a d.
Try it: “t”, “d”.
“t” “d” See?
Check out three as well: “three”, “tres” and “treeni” - there is clearly a pattern there.
I honestly think an English-speaker could hear the words ekam, dve, treeni and recognise that the other person is counting to three.
And they’re doing so in Sanskrit.
Look at sexy six too. And nine, novem, nava.
It’s amazing, right? And these similarities simply don’t exist between Indo-European languages and non-Indo-European languages, even when they are on each other’s doorsteps.
Finnish is not an Indo-European language, despite being surrounded by them, and its numbers 1-10 bear no resemblance whatsoever to those that are.
And it’s in universal concepts like numbers that these similarities between the languages really shine through.
Perhaps the shiniest of them is the word “sun”.
In German it’s Sonne, in Spanish it’s “sol” in Czech it’s “slunce” and in Sanskrit it can be “Surya”.
All ss words.
The moon is also a constant in skies across Europe and Asia.
And the words for it are consistent across the Indo-European languages although, they fall into two groups the M-sounding ones, like our moon, Swedish “måne”, and Lithuanian “ménuo” are some examples.
And the L-sounding ones like French “lune”, Welsh “lloer” and Ancient Greek “selene” although you have to rummage a bit to find it in there.
I don’t know if you’ve ever tried to learn one of these European or Asian languages, but this idea that they’re all related to English, I think makes that a whole lot less daunting.
Because you know that there are patterns of similarity to be found.
I highly recommend trying to learn another language and I recommend starting your language-learning journey with Babbel, who’ve kindly sponsored this video.
Babbel is one of the top language learning apps in the world. Its intuitive lessons help you learn a language through real-life conversations.
[In Swedish] “I’m learning Swedish.” I’ve been using it to learn Swedish not just because all the Swedes I know are awesome and I want to impress them, but because it’s a fun language with lots of surprising similarities to English.
And now: Jag pratar lite Svenska!
With Babbel you can set yourself a goal and it’ll give you all the tools to meet it: live teaching sessions, flashcards, podcasts, games - which I love - and of course: lessons you can do whenever and wherever you want.
[Babbel] “Har du barn?” [Rob] “Har du barn?” [Babbel] “ett barnbarn” [Rob] Did you know that in Swedish, grandchild is just “childchild”?
I’ve found I can make fast progress with Babbel because you’re very quickly learning sentences that you can actually use.
it teaches real-world conversations.
So click the link in the description or scan the QR to get 60% OFF on your subscription! There’s a 20-day money-back guarantee by the way So just give it a go.
Tack!
Now, the sun and moon comparisons we just did show a couple of the easier-to-spot relationships between words in different Indo-European languages.
But over the last couple of centuries, linguists have found many, many more of them that are less obvious, because pronunciations change.
Take all these different words for “hundred” across the Indo-European languages: hundred, “centum” - yes, it’s kentum, not sentum - “hekaton”, “sto” and “shatam”.
These don’t appear to have very much in common with one another but linguists believe they are all related and that, for example, the /h/, /k/, /s/ and /sh/ sounds at the start of them all started off as the same sound.
They believe that all of these words can therefore be traced back to a single, source: the Proto-Indo-European *kmtom-.
That star at the start, by the way, is important to mention.
It’s used by linguists to flag up that this word has never been seen written down. It is not “attested” and that is the case for all Proto-Indo-European terms because it is a reconstructed language.
It is our best guess at what a common ancestral language could have been like.
And linguists remind you of that fact every time you see a P-I-E term, by sticking that asterisk at the start.
Anyway, it’s quite easy to reverse engineer that *kmtom into all of those words for hundred.
If you swap that nasal “m” for an equally nasal “n” *kmtom becomes “kntom” which is extremely close to Latin “centum”.
Particularly if you take for granted that vowel sounds shift pretty easily over time.
Likewise, if you accept that at some point, one group of Indo-Europeans started pronouncing the /k/ sound as /sh/ and that they might also have stopped bothering with that little nasal flourish in the middle, you get from kmtom to Sanskrit’s “shatam”.
What about our word “hundred”, then?
Well, suppose that another group of Indo-Europeans started to pronounce that /k/ as a /h/ Not a stretch: those two sounds often switch around between languages.
And we’ve also discussed how the “mm” could turn into a “nn”, plus we’ve also talked about the linguistic similarity between /t/ and /d/ Drop the “-om” because humans tend towards the lazy and you get “hnd”.
And do you know what the Old English word from hundred was?
Hund.
Tell me that isn’t satisfying.
And what we’ve just done there, is the perfect inverse of what linguists have done to try to reconstruct Proto-Indo-European.
They’ve looked for patterns of sound change between languages among words with similar meanings then recreated what they think would have been the common, original sound.
So a lot of the sound variations we saw with the words for hundred are also present in the different words for “heart”.
Look at how between all of these we again see the same pattern of /h/ /k/ and /s/ sounds at the start, as well as that switching around of /t/s and /d/s in the middle or end.
So historical linguists looked at all of those words together and concluded that the common denominator, the common ancestor in Proto-Indo-European must have been something along the lines of *kord- or *kerd-.
Through this same process they hypothesised the Proto-Indo-European for “father”, from which Latin “pater” is also descended. As is Irish “athair”, and Old Persian pita, and many others.
They worked out the P-I-E for “mother” too and for “brother”. And for “head” and for “foot”, and for “mouse” and for “goose”. And many, many more.
But as well as pronunciations changing over time, meanings of words change over time too.
So you find that words that no longer mean the same thing across the Indo-European languages share a single Proto-Indo-European root.
Take this term from P-I-E: *weyd.
It is the theorised origin of the English word “wit” meaning “intelligence”.
But it is also the origin of Latin “vidēre” meaning “to see”, a Sanskrit word meaning “found” and the Lithuanian for “face”.
So how can that be?
Well it isn’t actually that much of a stretch to link the concepts of seeing and knowing - in English, to see something can mean to understand it, right?
You see?
And in a lot of languages the word for face essentially means “thing you see” like the French “visage” and German “gesicht”.
In fact, the word “face” means the thing that is “facing” whoever is looking at it.
So you see how meaning changes over a few millennia make these relationships harder to spot, but they are still there.
Here’s another example I like. In Proto-Indo-European, this means “sharp”.
But it’s thought to be the origin of the English word “edge”, the Latin for “sour” and the Albanian for “blade”.
However, you can see how the concept of “sharpness” links all of these words.
By the way, that same P-I-E root also gives us the words “acute”, “eager” - a sharp keenness - and vinegar, which is from the French for sour wine.
Actually, this is where the fun with Proto-Indo-European really kicks in for a dork like me: finding the words in English that are unexpectedly related to one another, because they share a P-I-E origin.
For example, did you know that the words “hound” and “cynic” are related?
They can both be traced back to the Proto-Indo-European *ḱwon- meaning dog.
*ḱwon- passed through a Germanic filter and arrived in English as hound, with its meaning almost unchanged.
But *ḱwon- also turned into the Ancient Greek for dog, kyôn.
The Greeks then used that word to - rather unkindly - describe followers of a specific branch of philosophy. They called them kynikos, meaning “dog-like”.
And then that tem passed through Latin and then French, and into English as “cynic”.
Cool, right?
That’s a pairing where the two meanings are very different, but there are a shed load where you can kind of see how the concepts are linked.
For example, “head” has the same proto-indo-european root as the word “chief”, the P-I-E *káput.
Along the Germanic branch it became hēafod in Old English, which became “head” in Modern English.
And along the Italic branch it became caput in Latin, chef in French and chief in English.
The same root also gives us “chapter”, “capital”, “captain”, “cape” and “cap” too.
Another to get our teeth into is the pairing of tooth and dental.
We know that the words obviously have related meanings, but I don’t think it’s obvious that they actually share the same root.
Excuse the pun.
They both originate in this P-I-E term.
Pedal and foot share a Proto-Indo-European ancestor as well.
As do the words “tongue” and “language”.
So speaking of both, how would the Indo-Europeans - those early folk probably living somewhere around the Black Sea at the nexus of Europe and Asia - how would they have got their tongue around this extraordinary language?
What did Proto-Indo-European sound like?
No one alive has ever heard Proto-Indo-European being spoken by a native speaker.
Indeed, even based on the most conservative estimates, no one born within the last four and a half thousand years could have heard it either.
However, plenty of people have attempted to represent what it could have sounded like. And you’ll find lots of examples of people doing it here on YouTube.
Check out this from the Quellant YouTube channel: [Voice in video] “ueuked leukos deiuos Uerunos “Nu hyreks potnih suhxnum gegonhe.” An impressive vocal performance if nothing else.
It sounds like a cross between Spanish and Swedish or something.
It sounds surprisingly modern, doesn’t it?
And why shouldn’t it, right?
The human brain hasn’t changed and our mouths work the same.
We would have still had the capacity to speak a complex language back then.
So that is Proto-Indo-European - our best attempt at reconstructing our original language.
But hold on.
Is it possible to go back any further?
Well some linguists say “yes” There is a proposed even older ancestor - a hypothetical language that was somehow shared, not just by all of the Indo-European languages, but by the Afro-Asiatic languages, the Dravidian languages, the Altaic, The Uralic and the Kartvelian.
This notional super-family has been named the Nostratic family a name that comes from the Latin for “compatriots” or “fellow countrymen”, or just “us”.
What the linguists who believe in this theory have done is compare words from all the different hypothesised prehistoric languages the “proto” languages like Proto-Indo-European and looked for patterns.
And they found them: for example, they found correspondences between certain sounds in P-I-E and in Proto-Kartvelian.
And they found similarities between specific terms in Proto-Indo-European and Proto-Uralic.
Look at how similar the P-I-E for water is to the Proto-Uralic equivalent and ditto for the two language’s words for “name”.
They looked for these pairings within universal concepts because water is everywhere and, you know, we all have names, no matter where we live.
But not only did they find patterns within pairs of proto languages, they also looked for things that were common to all - or at least many - of these ancient languages.
And they came across some.
Another apparent pattern is the common use of “m” words to refer to oneself.
In English we have “me”, right?
And lots of other Indo-European languages have first-person pronouns beginning with mm sounds.
But so also do the Kartvelian, Afro-Asiatic, Uralic and Altaic proto-languages.
Nostraticists also point to the similarities between the Proto-Indo-European, Uralic and Altaic words for hand.
The problem is that not all of the hypothesised members of the Nostratic superfamily fit that pattern.
For the theory to work, you have to pick and choose the ones that do, which is something that makes the whole notion controversial.
Another major reason why the Nostratic scenario is seen as a bit wobbly is that it involves reconstructing a languages based on other reconstructions, some of which are already based on reconstructions.
Nostratic, of which there’s no written proof, is partially based on Proto-Indo-European, of which there’s no written proof, which is partially based on Proto-Germanic, of which there’s no written proof.
However, there are some very dedicated linguists doing their best to identify the pitfalls in that theory and fix them.
So watch this linguistic space, I guess.
But if the Nostratic theory doesn’t cover every language in the world, could there be more of these super-languages?
And if so, could they have a notional common ancestor themselves?
Well, believe it or not, there are people who have hypothesised this very thing and called it “Proto-Human” or “Proto-World”.
But it requires a rather large leap of the imagination.
And the fact is we’ll never know if is existed because if it did, it was being spoken tens of thousands of years before humans wrote anything down.
However I have done my best in this video to explain the very earliest origins of English.
If you’ve enjoyed it, I think you’ll like this video too.
And I also think you’ll enjoy my wordy nerdy podcast - Words Unravelled - which you can watch here or just download wherever you get your podcasts.
I’ll catch you in whatever you choose to watch or listen to next.
Cheerio.
Up Next

How We Know Proto-Indo-European Languages Existed
@simonroper9218
105.7K views•2023-09-04

Conversation Analysis: Key Concepts & Research Domains in Linguistics
@pointstoponder5186
9K views•2020-12-30

The Complete Origin of Every Letter in the Alphabet
@RobWords
3.9M views•2023-02-11

Accent Expert Explains U.S. Regional Dialects | Part 1
@WIRED
9.3M views•2021-01-21
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Linguistics






































