Other

"Lektorium": Who are our relatives? Looking for answers through language

In the first lecture of the "Lektorium" project, launched jointly by Gazeta and "Yangi O‘zbekiston" University, linguist Eldor Asanov analyzed the most interesting scientific and pseudoscientific hypotheses regarding the origin of Turkic languages. Additionally, the scholar explained how the field of comparative-historical linguistics works.

Eldor Asanov, a linguist, anthropologist, and head of department at the Institute of Anthropology of the Academy of Sciences, participated in the first open lecture of the "Lektorium" project, launched in cooperation between Gazeta and "Yangi O‘zbekiston" University. The scholar analyzed the most interesting scientific and pseudoscientific hypotheses regarding the history and origin of Turkic languages. He also explained in detail how his narrow specialty—comparative-historical linguistics, or comparativism—works.

This is a collaborative project between Gazeta and "Yangi O‘zbekiston" University, which aims to organize a series of free online and offline lectures for everyone.

Hello, dear friends. We will begin the lecture. In this project started by Gazeta.uz, I was asked to give a lecture on my narrower specialty. The time allocated to us is one hour.

For this lecture, I chose the field of comparative-historical linguistics—comparativism. We will also talk about some other related fields, for example, glottochronology. Glottochronology refers to the lifespan of a language. It is a science similar to genetics; in fact, it is considered an exact science rather than a humanity. We will talk about it now as well.

If you pay attention, there are many similarities among languages. For example, you can find similar words in Uzbek and Russian, and you can also find similar words in English and Uzbek. In such situations, people start thinking about genetic relationships.

I will give an example from my own life: a neighbor in my village once asked me, "I am from the Ming tribe, and there was a Ming dynasty that ruled in China. Does that mean that dynasty is related to us?" The similarity of the word created a parallel for my neighbor. He is someone very far from linguistics, yet he draws such a conclusion based on a single similar word.

There are also people among us who are very passionate about linguistics; they find a million parallels and draw various conclusions. For example, you have probably heard this conclusion: the oldest written language in the world is Sumerian, which has a 6,000-year history. Many people consider Sumerian to be related to Turkic languages. The reason for this is the abundance of similar words in these languages.

For example, people have found hundreds of parallels between Sumerian and Turkic languages, such as "ota" (father), "ona" (mother), and "tangri" (god).

But do 5, 100, or 300 words prove a genetic relationship or not? What do you think? If you dig deeper, you can find similarities not only with Sumerian but with other languages as well. For example, I recently realized something I had never noticed before. In English, there is a word very similar to one in our language: the word for "ho‘kiz" (bull/ox) is "ox" in English.

If you look at the older form of the word "ho‘kiz", it is "öküz"; after the initial "h" is dropped, it becomes even closer to "ox". What conclusion can be drawn from this? Can we conclude that Uzbek and English are related? Some people make such claims. What does science say about this? Comparativism, that is, comparative-historical linguistics, answers this question.

Comparing languages with one another has existed among people since ancient times. For example, Greek and Chinese sources contain information stating that the language of a certain people is similar to the language of another, but differs slightly, and that they can or cannot understand each other. However, the systematic comparative study of languages began in the 19th century.

What happened in the 19th century? European empires conquered and colonized the entire world, and in parallel, the study of the world began. Anthropologists and Orientalists began studying the peoples of Africa, America, and Asia, and their languages. At that time, European linguists noticed something: Sanskrit, the ancient language of India, was very similar to Greek or Latin. This amazed them.

I am explaining this in simple terms, but the similarity is obvious. Several words are presented here as examples.

Thus, in the process of studying Sanskrit, they compared it to Greek and Latin and discovered their similarity. After that, they concluded that languages so distant from one another could still be related, and that this relationship could be preserved over thousands of years. This led to the concept of the Indo-European language family. Before that, they were called European languages. Since it started with Sanskrit, the word "Indo" was added to it.

After that, they studied other language groups, such as Iranian languages, then Anatolian languages like Hittite, Luwian, and so on. Written sources of the Tocharian language were found in Central Asia and China, and it also turned out to be an Indo-European language. In short, concrete similarities were found, on the basis of which the Indo-European language family was identified, and comparative-historical linguistics was born.

Comparing languages and finding similarities became fashionable. But as I mentioned earlier, systematic, scientific comparison began, rather than just comparing a word like "ming" to the name of the Ming dynasty. What does systematic, scientific comparison look like? What is the problem with the gentleman I gave as an example? The thing is, he is taking a word from modern Uzbek and comparing it to a word from, say, Chinese of a thousand years ago.

The main principle of comparative linguistics is that the oldest variant of each language must be used. The oldest known variant must be compared, because language changes and shifts with every generation. Therefore, modern Hindi is not compared with modern Greek; Sanskrit is compared with ancient Greek because they are closer to each other.

If we take Vedic Sanskrit (3500 years ago) and Ancient Greek (3000 years ago), there is a 500-year difference between them. They are relatively close to each other. If we take Greek and modern Hindi, what is the difference between them? There is a 3500–4000 year difference. They may have diverged completely. As a result of studying based on this principle, truly related languages were identified, and the fashion of "language families" emerged in linguistics. Grouping languages into branches and branches into families became a trend.

Language families are taught in school, right? At the very least, in native language classes, we were taught that Uzbek belongs to the Turkic language group, and this group belongs to the Altaic language family. In fact, there is a problem even with this, which we will address later. We started with Indo-European, as they are the best-studied languages. The reason they are the best-studied is clear: linguistics as a science began in Europe, and Europeans primarily studied the languages they were interested in.

Which languages do you think came next? Semitic languages came next. They are also called Shemitic or Semitic languages. The Semitic languages include several famous languages we know, such as Arabic, Hebrew (called by various names), as well as Akkadian, Syriac, and many other languages. They were studied second.

By "second," I do not mean they were studied later, but that they received attention next. The main focus was on Indo-European languages, followed by Semitic languages. The reason was, firstly, that holy books are in Semitic languages. The Quran is in Arabic, and if we take the Bible, the Old Testament is in Hebrew. In addition, there are many religious books in Syriac and Aramaic.

Therefore, naturally, Semitic languages were of interest to Muslims, Christians, and Jews. In addition, another important aspect is that the oldest written languages in the world are also Semitic languages. We said the oldest is Sumerian. While Sumerian has a history of about 6,000 years, starting from approximately the third millennium BC, the Akkadian people, who spoke a Semitic language, settled in Mesopotamia, adopted Sumerian cuneiform, and began writing in Akkadian.

Akkadian is also a Semitic language, but it differs slightly from other languages because it is very ancient. Thus, Semitic languages were studied and grouped. Later, relatives of Semitic languages were also found. The closest relative of Semitic languages is the language of ancient Egypt. The Egyptian language, written in Egyptian hieroglyphs, is a close relative of Semitic languages.

In addition, there are Berber languages and others. They were combined into one family called Afroasiatic languages. Thus, two families emerged. After that, they began studying other languages, and several other language families were identified, but they were still poorly studied, their specific features were not well known, and their vocabulary and etymology had not been worked out.

At that time, the German linguist Max Müller proposed a classification, and this classification is familiar to you—the Turanian languages. A huge language family called the Turanian languages was identified. This concept appeared in the 19th century. The name is self-explanatory: it is taken from Central Asian, Iranian, and Turkish mythology, from the Shahnameh. Because by "Iran" they understood mostly Indo-European, they separated Iran's enemies as Turan—other languages that were not Indo-European.

Look—Ural-Altaic, and within Altaic there are also Turkic languages, and look at the languages grouped as related to it: Dravidian, Tibeto-Burman, Tai-Kadai (languages in Thailand), Malayan (languages in Malaysia and Indonesia). Such a massive family was identified. What was the reason for identifying this family?

As I said earlier, they were not well studied, but there was one feature common to these languages—agglutination. Agglutination means expressing grammatical meaning with the help of suffixes. For example, in Russian or English, grammatical meaning is often expressed with prepositions, right? Like "v", "na", "pod". It is similar in English.

In Uzbek, prepositions do not exist as a category. We usually express tense, person, and location with suffixes. "Maktabimiz" (our school)—in one word, with the help of a suffix, we express that the school belongs to us. But we cannot say it like that in Russian. We have to say "nasha shkola". That makes two words. This is agglutination, and after observing Ural-Altaic, Dravidian, and other languages, they assumed that these must all belong to one common family, with a common origin, and they classified them as Turanian languages.

Why is this interesting to us? This name "Turan" later influenced our intellectuals in a different way, and they found inspiration in this name for their various theories. However, in Europe, by "Turan" they understood something slightly different, which we will also get to.

Another thing: Ural-Altaic languages are called the northern branch of the Turanian languages, and the rest are called the southern branch. Because they are distributed further south. The Dravidian language was the main language of India before the arrival of Indo-European languages. It is the language of the Indus Valley Civilization. It is still preserved in southern India today. Indian films made in Dravidian languages are not called Bollywood. They are called Kollywood and Tollywood. There is a very large, separate cinema industry in the country in languages like Tamil and Telugu.

So, those languages are also very widely spoken in India to this day. The rest are spoken in Indo-China and Oceania, which is why they are called southern. I said their specific feature is agglutination.

Expressing grammatical meaning with the help of suffixes is called agglutination, and this was initially considered a characteristic feature of Turanian languages. As a result, in 19th-century linguistics, another language was considered one of the Turanian languages. This language is also familiar to you—the Sumerian language.

Our own intellectuals have written books citing similarities between Sumerian and Uzbek or Turkic languages, claiming that it is related to us, and they did not pull this theory out of thin air. In the 19th century, such a theory existed in science, and Rawlinson himself, the first scholar to read Sumerian-Akkadian cuneiform, said this.

After deciphering, reading, and becoming familiar with Sumerian cuneiform, he said that Sumerian, in its structure, resembles Mongolian-Manchu languages. Later, when he became familiar with Turkish (Ottoman), he said it was also similar to Turkish. That is, this theory was widespread in the 19th century. Even in the mid-20th century, this theory was taken seriously.

Thus, the 19th century can be called the wild era of comparative-historical linguistics. People and scholars studied various languages, collected material, and, inspired by the regularities they found, put forward theories about many language families and macrofamilies. As a result, until recently, there was a view in linguistics that languages consist of families.

Mainly, about 30 families are identified, excluding the Americas. The Native American languages themselves are divided into about another 30 families. Languages in Eurasia, Africa, and other adjacent regions were divided into about 30 families. This was a kind of trend. What is the attitude toward these families today?

Let us continue. The next stage. Before moving to this stage, I must mention another debate in linguistics. There are two views in linguistics: monogenesis and polygenesis. The meaning of this is: monogenesis claims that all languages in the world originated from a single language. Polygenesis claims that languages never had a single common ancestor. According to my observations, monogenesis was more popular in the 19th century. There are several reasons for this.

The first is the achievements of comparativism. That is, since we discovered the relationship of the Indo-European family, identified the relationship of the Semites, and learned of the relationship of the Afroasiatic languages, the view prevailed that if we dig even deeper, everything will trace back to a single root. This view is called evolutionism. Its most famous representative is Charles Darwin.

Charles Darwin's evolutionary theory held a dominant position not only in biology but in all fields, including linguistics. Linguists thought in a similar way to biologists. That is, if all living creatures, all life, originated from a single root, then languages must also trace back to a single root.

In modern science, this is not 100% accepted. For example, yes, let us say all living things have a common cell; the cell is common, so they have a single ancestor. But before that, there were thousands of forms of life, and one of them won. Or if we take economic and cultural history, they are now saying that humans did not simply transition from a primitive communal system to slavery, from slavery to feudalism, and from feudalism to capitalism. Humans were making political experiments even 300,000 years ago; democracy was born somewhere, capitalism was born somewhere—quite free views are being expressed.

A similar situation exists in linguistics, which we will also discuss. For now, we will talk about linguistics at the end of the 19th century, the mid-20th century, and the period after. So, monogenesis is coming into fashion, and the theory I mentioned earlier is being put forward: almost all languages of the world trace back to a single root. And that root is the Borean language.

Borean means northern in Greek. It is said that some language in the north spread throughout the world and took various forms. This theory has two ideologues: Starostin, a representative of the Moscow school, and the English scholar Harold Fleming. The two developed two variants of this theory and had their own classifications.

They tried to answer the question of which language, in what form, belongs to the Borean macrofamily—not even macro, but hyperfamily. We cannot dwell on each hypothesis in detail now. I will try to introduce them briefly. There is a variant, a theory, called the Borean language. The next one is Turit.

The Turit languages cover an even wider scope, including even the smallest languages on the planet. If we look at the origin of the word Turit, it is derived from the Latin word "turris". This means tower. Which tower? Of course, the Tower of Babel. Earlier, I spoke about the influence of Darwin's theory. The second influential factor was the Christian religion.

In science, research was mainly done by European scholars, and their religion was mostly Christianity. Especially in the 19th century, they sometimes introduced their religious views into science without fully verifying them. Therefore, the myth of the Tower of Babel was very famous in linguistics.

For those unfamiliar with the myth of the Tower of Babel, I will briefly summarize its content. A very long time ago, all of humanity spoke one language, and people developed and lived in abundance. At some point, humanity became so proud that they decided to build a tower reaching to heaven to see God and reach God. This tower was built in Babylon.

The reason it was built in Babylon is that Babylon is a frequently used image in the Bible, in the Old Testament. It was used as an image of pride, immorality, and so on—the image of a city that became very rich and very corrupt. In later periods, the image of Rome was used instead of Babylon.

So, a tower is being built in Babylon, and the goal is to reach heaven. God becomes angry at this and confuses the people's languages, as a result of which people stop understanding one another. If they need to pass a brick to each other, they cannot say it; if it is something else, they cannot communicate, they cannot ask for anything from each other, and the construction stops. Later, humanity scatters across the world. This is how today's languages appeared, it is said. This is why each country has its own unique language.

This myth greatly inspired European and Russian scholars, and thinking that perhaps there is truth behind this myth, and that indeed all languages trace back to a single base, they put forward theories like this: the Borean language, Turit languages, and so on.

Even the name Turit is derived from the word for tower, but the author of the name, Militarev, noted that he came up with it as a joke. Look, it is divided into four large groups: Amerind, Indo-Pacific, Australian, and Eurasiatic/Afroasiatic. It covers almost every language. But at that time, was there really a basis for such coverage? After all, not all languages had been sufficiently studied. Despite this, such theories were put forward at that time.

Borean and Turit were hypotheses that were very difficult to prove, but there is also the Nostratic theory, which was almost a proven theory. This term was introduced to science by the Danish scholar Holger Pedersen. Representatives of the Moscow school of linguistics, Vladislav Illich-Svitych and Aaron Dolgopolsky, along with Starostin, continued and developed the Nostratic theory.

What is the essence of this theory? They concluded that Indo-European, Afroasiatic, Semito-Hamitic, Ural-Altaic, and Kartvelian languages are related to one another, and unlike the theories above, they conducted massive research. They actually demonstrated thousands of parallels among hundreds of languages.

In the classification given by Illich-Svitych, the first variant included the following families: Indo-European, Afroasiatic, Uralic, Altaic, Kartvelian, and Dravidian. Later, some linguists also added Eskimo-Aleut and Chukotko-Kamchatkan languages to this. This has not been fully proven. The Nostratic theory itself has also not been fully proven.

In short, they created and developed such a large macrofamily. If you actually look, parallels among them are everywhere. As I gave an example earlier, "father" and "mother" are pronounced similarly in Sumerian and Turkic languages. If you study different languages, "father" and "mother" are pronounced almost the same way in nearly every language.

I gave an example from Sumerian, but take even Gothic, one of the Germanic languages, or Hittite, or Elamite, Dravidian—whichever language you take, it is approximately the same. Ata, ama, forms like this. It is the same in Japanese. They call mother "ama". There is also a goddess named Amaterasu. This is why Starostin includes Japanese in the Altaic family. He has a famous book dedicated to this issue. Indeed, these words are similar in almost all languages. What is the reason for this? Are we really related?

As a result of searching for the answer to this question, at the first stage, the answer was that we are related. Today, there are slightly different answers. Nostraticists reached such a level that, having fully studied the material of the languages belonging to these families and compared every word, they reconstructed what the oldest form would look like when compared, and they worked out the Proto-Nostratic language.

And they wrote poetry in the Proto-Nostratic language. What does Proto-Nostratic language mean? For example, if there are conditionally 10 language families, with 10 languages in each, they compared the material of each and reconstructed how each word was pronounced 7000, 8000, or 10000 years ago. And they provided such a reconstruction. The most interesting thing is that this reconstruction is very difficult to read.

In general, the main problem with language reconstructions is that if you fix phonetic changes according to precise laws and trace them to the end, such complex forms emerge that it is hard to believe an ordinary human could pronounce these words.

For example, how do you pronounce the first word? There is a dot under the K, so it is not an ordinary k. There is a capital h next to the L, so this is not an ordinary h sound; this is aspiration, meaning you have to pronounce the l with a slight breath. It should not be L, but lh. There are two dots over the A, so it must be read softly. Try reading it. Then there is a reversed question mark, what is that?

The reversed question mark is called a glottal stop in English; in our language, it can be called a "tutuq" or glottal stop. It is the sound ʼa, which is very common in Semitic languages. The word starts with that. How do you read it? Then look below, the first word of the second line. It is not an ordinary l; it gives the Greek letter lambda. This means it should be read not as an ordinary l, but differently. In short, transcription is very difficult. They probably cannot read it precisely themselves, but for some reason, reconstruction yields only such results. I think people probably did not torture themselves to this extent.

But even if these reconstructions do not reflect 100% of the reality, they are significant for the development of linguistics. So far, we have talked about comparative-historical linguistics, comparativism—the study of languages by comparing them with one another.

Later, another field developed—glottochronology. What is the essence of glottochronology? It is a science more like genetics. In it, the antiquity of a language is calculated based on how many generations it takes for changes to occur in the language. How does this work?

Two types of segments are distinguished in language. The first is the active vocabulary, and the second is the passive vocabulary. The two change at different speeds. Because frequently used words change more slowly. Passive vocabulary changes faster because it is used less. Based on the speed of these changes, they predict when languages originated.

According to this prediction, for example, when did Turkic languages originate? There is Starostin's prediction. He says the point of divergence of Turkic languages from one another was 2300 years ago. That is, 2300 years ago, a people speaking Turkic languages lived in one place, then scattered from that point, and their languages began to change. Is this calculation accurate?

In the case of Turkic languages specifically, it is quite accurate, because around the fourth-third centuries BC, there is information about the Hun state in Chinese sources. They are building a state, fighting with China, migrating, and spreading as far as Europe. Glottochronology data also corresponds to that period. So, it can be calculated approximately by comparing it with historical data.

Various opinions have been expressed about other languages as well. Only in the Altaic family have very big problems arisen. The Turkic languages themselves spread only 2300 years ago, but their relationship with Mongolian languages is set at 6000 years ago. That is, Turkic peoples separated from Mongolians 6000 years ago, and lived in one place for another 3500 years without dividing. Then they scattered. It has strange aspects like this.

In comparativism, besides Nostratic, there are many other similar theories. Because there are still languages that have not been well studied and whose relationships have not been determined. Scholars study them and try to establish their relationship to one another. For example, there is such a great model, Sino-Caucasian, that is, Chinese-Caucasian languages. In an unexpected place, there are similarities between Chinese and Caucasian languages. Even their numbers are similar. Based on this, the idea is put forward that Chinese and Caucasian languages are related to each other.

There is such a model, and the languages included in it are: North Caucasian languages, Basque (a language spoken in Spain with no relatives), Burushaski (a language spoken in the valleys and mountains of Quetta in Pakistan), Hurro-Urartian languages (languages spoken in the Near East), Yeniseian languages, Sino-Tibetan languages, Na-Dene languages, languages in America, Native American languages. You have probably heard of a tribe called the Navajo. The Navajo speak Na-Dene languages. They are unique, and we will return to this.

Sometimes people find similarities even in Native American languages, but in reality, Native American languages are also not related to one another. There is reliable information about a relationship with Eurasian languages only for the Na-Dene language family. The reason is that speakers of Na-Dene languages are considered the very last wave of migration to America. There are no clear parallels connecting other Native American languages with Eurasian languages.

Scholars searching for relationships among languages were inspired to create a project. You can find it on the internet. I did not leave a link because the link changes. The project is called the Tower of Babel. This is a creation of the Moscow school. It was created under the inspiration of the Nostratic theory and the relationship of language families to one another. Within this project, studies comparing distant languages with one another are published, but the most valuable part is this.

If you enter the site, databases of almost all language families are located there. Then you enter them, and each has its own etymological dictionary. As an example, I have provided screenshots from the etymological dictionary of Turkic languages and Altaic languages. This is truly a treasure. Having studied all the languages in so many language families, they created such etymological dictionaries consisting of hundreds of thousands of words.

Even if it does not prove the relationship of languages, this is a treasure. From this, you can study the vocabulary of any language in the world. You enter, you are interested in the origin of a certain word, you find it in the dictionary, and you find parallels from other languages and information on its etymology. True, the Nostratic theory is not 100% proven, and it needs to be checked against other data, but it is hard to imagine how much work went into creating the database itself. I highly recommend this project to everyone.

We talked about how comparativism developed. Now, if we talk about its impact on our linguistics, I personally have heard theories like this. Some people believe in these theories completely, while others reject them sharply, but if you look at the history of comparativism, these were not invented by Uzbek or Kazakh scholars.

Similar ideas were expressed by Rawlinson, Max Müller, and Hommel. These ideas were expressed 150 years ago. At that time, science was not as developed as it is now, and today's theories and materials did not exist. Our own intellectuals had not yet reached this issue at that time, because they had other problems. In the 1960s, control over the colonial republics in the Soviet state loosened slightly, and as a result, nationalist movements appeared in the republics, including Uzbekistan.

The understanding of Uzbekness, of one's origin, and the search for roots began. Intellectuals searching for their roots turned to various sources, studied their own written sources, studied European scholars, and in the 1970s and 80s, they came across books by European linguists written 150 years prior. What was written in those books? Various views were reflected, such as that the Sumerian language was related to Mongolian, related to Turkish, and that Saka and Scythian languages were Turkic.

Naturally, after seeing their own name, they became interested in these views, studied them, and indeed found some of the parallels I mentioned earlier, and a book emerged: "Az i ya" by the Kazakh intellectual Olzhas Suleimenov. This is a book in two parts. The first part is dedicated to the analysis of a specimen of Russian literature called "The Tale of Igor's Campaign".

The second part is called "Shumernoma" (Sumerian Book). The author takes some words from the Sumerian language and compares them to the Kazakh language. The sun was called "utu", and in Kazakh, fire is called "o‘t". In Sumerian, grandchild was called "nimara", and in our language, it is "nevara", and so on.

Olzhas Suleimenov's book became famous in Uzbekistan, Azerbaijan, and Tatarstan. In Turkey, in parallel to him, Osman Nedim Tuna worked in this direction. As a linguist, he also compared the Sumerian language to Turkic languages and found parallels. Thus, the Sumerian-Turkic theory emerged, and it is still popular today. In addition, there was a state called Elam in the south of today's Iran, and in the Elam state, they wrote in cuneiform in a unique language.

The Elamite language is also an agglutinative language, and it also has words similar to Turkic languages. In it, they also call father "atta". They also compare it to Turkic languages. Then what other unaffiliated languages are there? In Italy, before Rome, the Etruscans ruled, right? The Etruscan language also has no relatives; its existence has not been proven to this day. In our country, they compared it too and found parallels.

I must emphasize that these languages were not compared only with Turkic languages. For example, I myself have seen the Etruscan language compared with the Albanian language. They also compared the Etruscan language with the Hungarian language and found parallels. They compared the Sumerian language with the Georgian language and found a great many parallels. They also compared it with Indo-European languages.

Therefore, these are not views unique to us. Intellectuals all over the world search for their roots like this and find similar words somewhere. And if you look, similar words can be found everywhere. In addition, the Elamite language was studied not just by intellectuals, but also by Starostin, whom we mentioned earlier, who linked it to Dravidian languages.

They also find parallels in Native American languages. What is the problem with those parallels? Native American languages themselves are divided into 30 families, and the relationship of Native American languages themselves has not yet been 100% proven. Intellectuals of Turkic peoples take a similar word from any Native American language. For example, they take one word from the Mayan language, one word from the Navajo language, and say, look, they are related, whereas there may be a 5000-year difference or an 8000-year difference between Navajo and Mayan. These are highly unreliable things. The above were romantic theories spread among intellectuals.

These are scientific theories. The first, put forward back in the 19th century, assumes a single large family called Ural-Altaic. Which ones are included in the Altaic languages? Turkic, Mongolian, and Tungusic-Manchu languages are included. Uralic languages include Finno-Ugric languages, Hungarian, and other languages.

Later, they separated Uralic and Altaic, saying they are not related to each other but are similar, and only Altaic languages remained. At the initial stage, Turkic languages, Mongolian languages, and Tungusic-Manchu languages were included in the Altaic languages. Later, Korean and Japanese languages were also included. So, it consists of five branches. This Altaic theory became very popular and widespread at the time.

Later, doubts about the relationship of Altaic languages increased, and now many scholars do not agree with this idea. In recent years, a group of scholars in Germany led by Martine Robbeets conducted research on this issue and, slightly updating the Altaic languages theory, put forward the Trans-Eurasian languages theory.

Trans-Eurasian languages consist of the same languages: Turkic, Mongolian, Tungusic-Manchu, Korean, and Japanese, with a changed name. Theories were put forward that these languages might either be related to each other or have lived together in one region. If we talk about scientific theories about the origin and relatives of the Uzbek language, we can mention these three.

Why are these not fully accepted? What is the problem? I will give one example: numbers. Numbers are considered one of the parts of speech that appear earliest in languages. The reason for this is that even the most ancient hunter-gatherers had a need to count. Therefore, numbers are usually the same in languages belonging to one group or one family.

There are indeed many similar numbers in Indo-European languages. Here they are listed from one to ten. In Sanskrit, Ancient Greek, Latin, Gothic, and Old Church Slavonic. In this respect, the interconnectedness of Indo-European languages is excellently proven.

Now, if we look at the numbers in Altaic languages, they do not resemble each other at all. Let us look at them in Old Turkic, Classical Mongolian, Manchu, Korean, and Japanese. The reason they are called native words is that in Korean and Japanese, there are two types of numbers. One is native, and the second is borrowed from Chinese. I have also listed them from one to ten.

Numbers in Altaic languages do not resemble each other at all. Two kinds of conclusions can be drawn from this. First conclusion: Altaic languages scattered before they learned to count. When did they scatter? When did humanity learn to count? They must have scattered tens of thousands of years ago.

Second conclusion: they are not related. Then why do we still find similarities? We will dwell on this in more detail now. There is a second method that demonstrates relationship. This usually does not prove relationship, but shows the degree of closeness of the relationship. There is a list by a scholar named Swadesh. The first list consists of 100 words. The extended list consists of 200 words. What kind of list is this? It is a vocabulary consisting of 100 basic words. According to Swadesh, these words change most slowly in a language.

The more these words match each other in languages, the closer they are considered to be. Therefore, in linguistics, until recently, when measuring relationship, just any similar word was not taken. This basic list consisting of 100 words was taken. If these are similar, then we can talk about a relationship. A word that happens to be similar by chance might be a loanword, right? Such a view prevailed. Now, even this method is outdated.

There is an article by a British linguist named Clauson on the extent to which Altaic languages are related to each other. Having studied the parallels of Turkic, Mongolian, and Tungusic-Manchu languages according to the Swadesh list, he says that Altaic languages are not related to each other. This is not a conclusion, but the opinion of an ancient scholar expressed at that time.

That is, even though we say there are so many parallels, there is similarity, we are related, our numbers do not match. Our basic vocabulary also does not match. How then do we explain this? Now, we will dwell on the new stage of comparativism, today's views. How was relationship studied previously? First of all, vocabulary was compared. When comparing vocabulary, the oldest forms, the oldest reconstructed forms, were taken, and basic vocabulary, numbers, and other things were compared. On this basis, relationship was proven or not proven.

In addition, grammatical morphemes in words were compared. There are also many similarities in grammatical morphemes. For example, morphemes are also similar in Indo-European languages or Semitic languages. At the next stage, they compared grammar. They said grammar changes quickly. Because some Indo-European languages became inflected languages, some became analytical languages, and in some languages, suffixes were radically reduced.

This can be observed even among clearly related languages. Take Russian, how many suffixes it has. Bulgarian has become an analytical language, meaning suffixes have been significantly reduced. Therefore, grammar was considered less reliable. Now in linguistics, they say even the Swadesh list cannot be trusted. Because there is no such thing as a basic vocabulary. It is said that any word can change, even numbers.

For example, the number seven itself is considered a loanword in Indo-European languages. It was taken from Semitic languages. In Semitic languages, "sab'" means seven, and forms derived from that are used in Indo-European languages: septem, sapta, hapta. In Tajik, "haft". Right? They say the number five in our Turkic languages might also have been taken from Indo-European languages. They say even our seven might trace back to that. In short, even numbers can change. Basic vocabulary can also change. There is even an opinion that the word for heart, which is the same in all Indo-European languages, was borrowed from Proto-Kartvelian, that is, the ancestor of the Georgian language. Card. A familiar word, right? The Russian "serdse" also traces back to the same root. Heart, card, serdse—all trace back to one root.

They also link the word for star to Semitic. The word "star" and the Babylonian goddess Ishtar, the name of the star Venus, they compare it to that. This is not 100% proven. It turns out even basic vocabulary can change. What then proves relationship?

Grammar is variable, vocabulary is also variable. But they say language preserves its features. What is there in language that does not change? It turns out there are word-formation models in language that do not change; this is also one of the existing hypotheses.

Word-formation models are of two types: there is the Altaic and the American model. The American model has its own variants, and the Altaic model has its own variants. So, there are attempts to determine relationship by comparing according to that, and as a result of this comparison, do you know what was identified? Almost no language turns out to be related to another.

Interestingly, 150 years ago, when linguistics began to systematize and develop as a science, people were so inspired by this that they saw relationships everywhere. Now that linguistics has reached today's level, the opinion has become dominant among people and linguistic specialists that "relationship cannot be proven. Languages are very diverse. Every language originated from different roots. There is no single ancestor language."

True, I cannot say that all scholars hold the same opinion now. But the dominant opinion is like this. If we compare the views of 150 years ago with today, there is such a difference. Today, relatively complex but reliable variants are being put forward. First of all, it is said that language itself does not have to trace back to a single root. It is not a cell. That is, if cells are similar, then they clearly have one root. In the past, we were overly interested in biology, overly interested in evolutionism. As a result, we transferred those biology and evolutionary principles to language as well.

Language does not have to develop like that. They say languages can appear separately in each tribe. Let us imagine that 30,000 years ago, two tribes are living in Australia and Africa. Both need to count. So, numbers appear independently of each other. In both tribes, people need to warn each other of danger. Some interjections appear. There are father-mother, kinship relations. Kinship terminology appears. Humans who are very far from each other still belong to the same species. Everyone is Homo sapiens, the organism is the same, the brain is the same, the needs are the same. That same pattern is repeated. Wherever anyone lives, they invent a language for themselves.

Earlier, we talked about the words for father and mother. Why are the words for father and mother similar in all languages? They explain it like this: the words for father and mother appear first. These are the first words a baby says when born. We ourselves teach babies to say "dada" (daddy), "aya" (mommy), and these words always consist of sounds that the baby can pronounce, and consist of repetitions of those sounds. Like dada, mama. Repeating one easiest syllable twice, it makes a word.

There was even a theory that the oldest languages consisted of repetitions of syllables, and they call these "banana languages". There is another view, which is about interjections and onomatopoeic words. For example, the word for "qarg‘a" (crow) is also the same in all languages. In Sumerian, in Japanese, and in Indo-European languages. Why? Because it sounds like "qarr". They explain the similarity with patterns like this. Fine, we explained the similarity of father-mother, some onomatopoeic words.

Why do complex words look similar? For example, the word for star. Or the word for heart, or numbers, why do they resemble each other? Today's linguists give an excellent answer to this question. So, they say we made a mistake by starting the history of languages from the appearance of the first families—10, 12, 14 thousand years ago. That is, 14 thousand years ago, these languages did not live in isolation. So they conclude that languages have been in contact for tens of thousands of years and have been passing words to each other for 300,000 years.

But we cannot take language archaeology back 300,000 years. There are no such reconstruction tools. Now they are saying that language families did not consist of one pure language at the beginning, but mixed together to create a language family. So, what we call the ancestor language was itself full of loanwords. Saying that Turkic languages split 2300 years ago does not mean that 2300 years ago the Turkic language was in a pure state. Not 2300, but even 3000 years ago, that Proto-Turkic language probably contacted the Chinese language, the Tocharian language, and Iranian languages.

Perhaps it had already been taking Chinese and Tocharian words for 700 years, and then split. Such a view is being put forward, and it explains many similarities. If a word in Kartvelian resembles a word in Basque, they do not have to be related. There is a view that in that period we do not know of, 70,000 years ago, through some cultural contacts, a word migrated from here to there.

I will not list all the theories and all the names now. To explain briefly, what is the picture? They say that in ancient times, before states, people lived in the form of certain small and medium groups. Someone lived in a cave, someone began to build small villages, and often each tribe had its own unique language. They do not necessarily have to originate from one another. Everyone, even tribes living side by side, might have spoken such diverse languages that they did not understand each other.

Later, this is called the pre-family period. That is, each tribe has a separate language. These languages did not originate from one root, but they have contact with each other. For example, if one tribe invents wheat, it delivers wheat to the next tribe, exchanges it, and along with the wheat, the word also passes into their language. The word for wheat also spread widely like this. It is similar in Semitic languages and Indo-European languages, and so on.

Therefore, they say language in ancient times was an innovation distribution network. Innovations also spread through language. Therefore, for example, 15 years ago, along with technology, words might also have come from the Near East to China. Perhaps this event occurred in an even more ancient period.

So, the pre-family period was like this, language diversity and contacts were more intensive than we thought, they were in very close cultural and economic relations with each other. Then what changed? The Neolithic demographic explosion occurred. There is a view that in the Neolithic period, in societies that began to engage in agriculture and animal husbandry, food accumulated in excess. As a result, they found the opportunity to increase sharply. For example, if resources are limited, a tribe plans its children. Let us say, if there is one child in the family, we can feed them, and they will help us produce more additional products and food. But if we have two children in the family, we will starve, resources will not be enough. Coming to the conditions of animal husbandry and agriculture, excess products are accumulating, and people are saying we can easily multiply.

Sharply, the number of farming and pastoral communities increases, and the Neolithic expansion begins. In the Neolithic period, farmers and pastoralists spread sharply throughout the world. Because their numbers are growing and they need new places to graze cattle, to plant wheat, to plant barley.

They are spreading rapidly throughout the world or spreading their innovation to other peoples, and they say that during this expansion, language diversity disappeared in many regions. That is, in fact, not now, but before the Neolithic, there were more languages. Now we are seeing the remnants of pre-Neolithic language diversity. A great many languages have disappeared. Perhaps 80, perhaps 70, perhaps 90 percent disappeared. Only its relicts remained.

For example, because there was the oldest written culture in the Near East, we know the languages that have survived to this day: Semitic languages, Indo-European languages. But Urartian, Hurrian, Akkadian, and Elamite languages left no descendants. Or let us say, in India and Pakistan, mainly Indo-European languages are spoken, but small languages like Burushaski and Werchik have remained in the mountains. They are not related to anyone. In Europe, the Etruscans are not related to anyone. In Spain, the Basques are not related to anyone. In the north of Russia, there are small peoples, not related to anyone. They say these are relicts, remnants of language diversity. This diversity is relatively well preserved in the Caucasus. Therefore, in the small Caucasus, there are very diverse, many small languages, and none of them are related to each other either.

Kartvelian languages are not related to North Caucasian languages. Other Caucasian languages are divided. There is an opinion that Semites, Indo-Europeans, Dravidians, Altaic, and Uralic peoples spread to all other places starting from the Neolithic period. If this assumption is correct, then our languages are invader languages. Having an advantage in the Neolithic period, they destroyed the diversity in the world.

Just as our cell destroyed other forms of life, and a few remained. Our languages are also like that. True, now in terms of numbers there may be many languages, but in terms of origin, they all belong to certain families. Little diversity remains. This now explains a lot of things.

Firstly, it explains how families originated. Secondly, it explains where similar words came from. So, the point is not in relationship, the point is in those contacts that have existed since ancient times. I tried to explain very simply. Within this, there are also various debates. A certain linguist will definitely watch our video and say, no, you spoke incorrectly in this place, you generalized in that place.

So, within this view, there are also several schools, hypotheses, and debates. The general picture is like this. Now I want to give you an excellent example. It looks like a formula, I will try to explain now. At the beginning of the lecture, I gave the similarity of our word "ho‘kiz" with the word "ox" as an example.

Look, if I were a linguist living 150 years ago, perhaps I would have become interested in the relationship of Indo-European languages to Turkic languages. At that time, there were comparison methods that were fashionable. Today, because etymological databases have been created and the ancient forms of all words have been reconstructed, what do I do? As I said earlier, I do not compare today's word "ho‘kiz" with today's word "ox". I take the form in the oldest Turkic language and I take the Proto-Indo-European form. Then no similarity remains.

Look, pay attention to how the word "ho‘kiz" changed. This excellently explains how Turkic languages developed. At the end of the word "ho‘kiz", there is the sound z, and in the ancient language, in that Proto-Turkic language of 2300 years ago that Starostin spoke of, there was the sound r in place of that sound z. That is, the sound r later transitioned to z. This is called the phenomenon of zetacism.

In modern Uzbek, wherever the sound z stands at the end or in the middle of a word, you can easily transition it to r. One of the bases proving this is the materials of the Bulgar languages, in particular, in the Chuvash language, which is alive today, it is like this. They still use r instead of z to this day. Where we say "to‘qqiz" (nine), they say "toqur". This serves as the basis for ancient reconstruction.

So, we reconstruct it as *hökür. The similarity to "ox" is gone. But there is another deeper parallel. In Turkic languages, the sound p does not occur at the beginning of a word. There are one or two examples, for example, "pichoq" (knife). The ancient form of "pichoq" was bičäk. It was B. In Bukhara, they still say "bichoq". So, the sound p does not occur at the beginning of a word in Turkic languages.

The reason for this, if we are to believe a number of Turkish, Finnish, and Hungarian Turkologists and Altaicists, is that a very long time ago, at the Altaic stage, the initial *p was reduced and transitioned to *h. The sound h itself is rarely found in Turkic languages. But there are 20 or 30 words where h stands at the beginning of the word, aspiration. It is pronounced softly, as h. It is not even an independent sound, because it does not occur in other positions.

It is assumed that the reason the sound *p disappeared at the beginning of the word is that *p dropped out and turned into *h. The basis for such assumptions is that in the Mongolian language and other languages belonging to the Altaic languages, if there is p at the beginning of a word, it corresponds to h in those Turkic languages. For example, forms like pökür or püker are indeed preserved in Mongolian languages, and in our languages, they come in the form of *hökür or höküz.

So, we reconstructed the oldest form of the word. This is a form of at least 3000 years old, a form before the division of Turkic languages. With *P. It becomes *püker. If we trace the English "ox" back to Proto-Indo-European languages, it becomes *uksen. They do not resemble each other at all. That is, the further back you trace them, the more they do not resemble each other. So, they might have become similar to each other only by today. This is why linguists always call for comparing the oldest form of a word.

But when we reconstructed the oldest Proto-Turkic form, unexpectedly, a similarity to another root in Indo-European languages emerged. In Latin "pecus", in Sanskrit "pasu", in Proto-Indo-European "*peku". This means cattle or property/wealth. It has two meanings. First, livestock in general, second, in the sense of wealth. The word *peku is compared to the word *püker. Look, we are not finding relationship in a place similar in form, but if we reconstruct the oldest form, a parallel emerges from somewhere.

How do we explain this? In linguistics, this is called a wanderwort. As I said earlier, languages served as innovation distribution networks. Once a technology was invented, that technology reached other peoples, and the name of that reached technology also entered along with it. For example, if conditionally the Indo-European peoples taught animal husbandry to the Altaic peoples, they would also have brought this word along with them.

This does not indicate a relationship, but shows how many thousands of years ago our contacts were. There were contacts even in a period deeper than we thought. In exactly the same way, we can explain the Altaic family. Earlier, we said its relationship is not being proven, not being confirmed. But parallels like this are everywhere in Altaic languages. True, their numbers do not match, there are problems in the basic vocabulary, but there are hundreds of similar words. Grammar, word-formation structures, vowel harmony, agglutination match completely. How do we explain this?

They explain this excellently. They say that in the north of China, there are several ancient Late Neolithic and Bronze Age cultures. We are coming to approximately that Neolithic period and the period slightly after it, there are several cultural complexes. Ancient Chinese culture grew more out of the complexes in the south. The question stands, who lived in the north?

A frequently cited example is the Xinglongwa culture. It has quite an ancient history, dating to 6000 or older. Today, it corresponds to the territories of Manchuria in China. Chinese culture is considered the culture where the image of the dragon was first created. It is assumed that tribes speaking languages not related to each other gathered in the Xinglongwa culture and lived together for thousands of years. Later, over thousands of years, their languages became similar to each other. In linguistics, this is called convergent evolution.

For example, the Uzbek language and the Tajik language are not related, but they are very similar, right? The pronunciation, phonetics are similar, the vocabulary is similar. Perhaps a dozen languages gathered there, and their vocabulary, vowel harmony, agglutination, and other features became similar because they lived together, it is said. There are such examples in the Uzbek-Tajik language, and in many parts of the world.

For example, it is called the Balkan sprachbund, not a political alliance. In the Balkans, languages not related to each other have become very close to each other. Bulgarian, Greek, and Albanian have become very close to each other. There are extremely many parallels. Being like that, perhaps Altaic languages are not related, did not spread from one language, but are friends that underwent convergent evolution.

That is, they are not relatives, but friends. This is called a polyphyletic family, that is, a family that does not have a common root but has become close. Perhaps they are such a family. Initially, there were dozens of languages, they lived together at that time and became very close, and there is a view that some tribes gained dominance, assimilated other languages, destroyed them, and spread widely.

That is, Altaic languages consist of five language groups that are not related to each other. But these five language groups are the languages that survived by destroying and wiping out others. In fact, there might have been dozens of languages not related to each other there. That is, as I said earlier, our languages are invader languages. Having gained an economic advantage, we converted it into a political advantage, and destroyed other small languages. Perhaps now the Altaic family could have consisted of 15 groups.

Today, the descendants of five have remained, and we are speaking in one of them now. Because in Chinese sources, dozens of tribes are mentioned. If we do not have precise material about the language of ancient peoples, then they might also be languages that left no descendants. Today, we search only from the variants we know. It is called "survivorship bias", right? We are the survivors. In fact, in ancient times, languages were extremely numerous and diverse.

When we reconstruct our history, we do not take them into account now. In fact, modern linguistics is coming to the correct point. We are not related. Our deep root does not trace back to one language, but we have communicated as friends for 100,000 years, and at a certain point, we wiped each other out, and we are merely the survivors. They say we destroyed 90 percent of other languages. And they say we must always take those destroyed languages into account in our reconstructions.

Cookies on xabarchi

We use cookies to remember your language and theme, and to count how many people are reading right now — that count is anonymous, lasts only while your browser is open, and cannot be tied to you or to another visit. With your permission we also measure how the site is read: Microsoft Clarity, which records page views and on-page interactions, and our own count of returning readers. Nothing that recognises you across visits is measured until you accept.