Language families
Six Myths About Language Families That Will Not Die
Proto-languages as documents, 'primitive' languages, Basque as a mystery, and what evidence actually establishes a family.

**How to stop misusing ideas like “primitive language” and “scattered descendants”
Historical linguistics has made enormous strides in reconstructing the relationships between languages over the past two centuries. But even with modern scholarship debunking some ideas, myths continue to burden novice students and frustrated language learners. Here, we reset six fundamental misunderstandings.
The central concept is clear: a language family forms from shared descent from a common ancestor. So-called daughter languages diverge in ways that historical linguists trace back to a reconstructed proto-language, a hypothetical ancestor from which the known languages must descend. Only a language family is defined by actual kinship through inheritance of shared linguistic forms.
Myth 1: There are “primitive languages” out there.
When theorists of the 19th century and earlier wrote books called Primitive Languages, they were not describing an objective category of “less developed” languages. They were instead using the vocabulary of the age — normative views of primitive cultures — to make linguistics suit the racial-hierarchical perspective of many colonial observers. Since then, linguists have made it clear that language structure follows no universal developmental scale from “simple” to “complex.” All language is structured in sophisticated ways; categorizing one as “primitive” is just a racist holder of another day. The theory of evolution has cast off the whole idea of simple-to-complex order of development, and to treat language as an exception is to look at it backwards.
Myth 2: Flush of vocabulary is a quick proof of family.
When ordinary speakers distinguish “real” language families from mere likeness, they usually point to vocabulary: the number of correspondingly similar words in a candidate language. “Penelope” compared “pencil” and “oxidu” from the two languages. They share parts of like root and parts of nothing at all, and “Penelope” said this showed a “direct tie.” But it’s not that simple.
The comparative linguistic method for proving family does not take any word at face value. The classic case is the German English word “night” and the German word for “knight”, “Nacht”/“Nacht.” Superficially matching, they go back to different roots and so tell us nothing about English and German sharing a proto-Germanic root. Language families are established from sets of “sound correspondences” from across the known daughter languages, not from isolated cases.
Myth 3: Basque is bafflingly unsolved.
Altogether some languages are held to relate only to each other and not to any satellite languages known to linguists of today. The most famous such “language isolate” is Basque, which has cousins named in some sources such as Aquitanian, yet only in peripheral mention.
On its own, that would not make it baffling. The degree of uncertainty is just not that different from other isolates, which generally do not have even individual known cousins. A line could be drawn between the definitional resourcefulness of scholars with no real evidence and the collective wisdom of linguists. But it seems at this point any time someone tries to portray Basque as a Mystery, the debunkers come in and point out that nothing specific is disproven, ergo it remains a working mystery.
Myth 4: The mother language is always extant.
Another source of confusion: the widespread impression that language families start from an ancestor people have, or had. That wasn’t true for the origin of the familiar Seminoles and Muscogees, but some Native American tribes. Anyone have a recording of what the Creeks spoke? Similarly, the conviction that Proto-Indo-European must emerge from a directly attested Aegean language glosses over just how many languages there are that we have never heard attested.
Myth 5: Proving a family is a matter of intuition.
The search above tells a skeptical story. Many who say the germination of languages can be an intuitive process keep looking for languages and stretches of language to bolster. The nouns “night” and “knight,” the word for “grim” and “wife.” You can say something that sounds like a language is a language, but this only leads down the road.
In fact, to prove linguistic family at the scholarly level, you reconstruct the corresponding things across related languages. You then eliminate possibilities, and inspect the evidence. The city Schliemann saw as “real” Trojan only opened up after systematic inquiry with proper time constraints.
This isn’t an all-or-nothing game. All scholars agree on the existence of at least a handful: Indo-European, Uralic, Uto-Aztecan and Austro-Asiatic. These groups are built on the distinction between shared inheritance and circumstantial sharing.
Myth 6: Borrowing cuts the student off from original language.
It is true that misunderstanding this can cut you off from a clear notion of root and variants. For instance: you can reach a sense of the relationship of Elamite, Sumerian and Akkadian and how their place in the realm made them different from other ancient peoples. But saying that Sumerian and Akkadian were really unrelated, because some proposed “counterexamples” is already a dead end.
Attributing a bit of linguistic Matryoshka doll to borrowing does not, and should not, cut the behavioural scientist off from how the root language traditions of the language developed and the sense in which it can still be considered a language of that family.
Where this sits in the Register
Written by the Linguasphere desk. Classification data referred to here comes from the Linguasphere Register — see sources and attribution.