• A
  • A
  • A
  • ABC
  • ABC
  • ABC
  • А
  • А
  • А
  • А
  • А
Regular version of the site

'We Would Like Our Corpora to Be Used More Widely'

Expedition to Dagestan

Expedition to Dagestan
Photo: Chiara Naccarato

The Linguistic Convergence Laboratory and the School of Linguistics at HSE University have created corpora of the Abkhaz-Adyghe languages spoken in the Western Caucasus. The corpora serve as valuable resources for studying these languages with their unique features and demonstrate the potential for their modern use. The corpora were developed through a series of field expeditions to the Caucasus conducted by HSE University researchers and students, combined with modern linguistic processing methods and collaboration with colleagues from regional universities. In this interview with the HSE News Service, Yury Lander, Leading Research Fellow at the Linguistic Convergence Laboratory and Associate Professor at the School of Linguistics, discusses the work of linguists.

— Please explain what a language corpus is.

— A language corpus is an electronic collection of texts that can be searched using a wide range of parameters, such as individual words, word combinations, or abstract grammatical features. It is essential for a corpus to contain a large and diverse body of texts. This allows researchers to conduct statistical analyses and, in some cases, identify patterns in language development. The main purpose of a language corpus is to show how a language is used in real life.

Yury Lander

Yury Lander

People learn certain rules about language at school. We are taught that words should follow a particular order, that pronunciation should conform to established norms, and that some words are considered more appropriate than others. In reality, however, language develops naturally and does not always follow prescriptive rules. A corpus containing millions—or even tens of millions—of words reflects how a language is actually used.

Corpora have been developed for many languages and language varieties. In Russia, the best-known example is the Russian National Corpus (RNC).

— What corpora are being developed by your laboratory and by you personally?

— The Linguistic Convergence Laboratory at HSE University is actively involved in creating corpora for a wide range of languages, but my own research primarily focuses on languages of the Western Caucasus.

Corpora exist for many languages, ranging from major world languages spoken by hundreds of millions to the smallest minority languages, both living and extinct. Many languages of the Caucasus are represented by corpora. For example, corpora of the Ossetian and Armenian languages have existed for quite a while. Our laboratory has developed corpora for several languages spoken in Dagestan, as well as for other languages, including Abaza and Adyghe. Work is currently underway on creating a corpus of the Abkhaz language.

Among our most recent achievements, I would highlight the large corpus of the Kabardian-Circassian (East Circassian) language. Although it does not contain grammatical annotation, the corpus includes tens of millions of words and was developed in collaboration with Kabardian language activists from Djarez, a language club in Nalchik, who collected a substantial body of texts. The corpus has not yet been officially launched, as we are still making some final corrections. However, a version called the Circassian Corpus is already available.

— What is the significance of the language corpora created by you and your colleagues, and how much does their availability facilitate the work of researchers?

— A corpus is one of the main tools used by researchers. It is valuable not only for linguists, but also for literary scholars, historians, ethnographers, and language teachers. A corpus provides direct material for understanding not only a language itself, but also its styles, individual features, and the history of ideas. It is also a fundamental element of language documentation, as it captures the current state and ongoing development of a language. After all, it often turns out that actual language use differs not only from prescriptive rules but also from existing grammatical descriptions.

— How did you begin studying the Abkhaz-Adyghe languages, and what attracted you to them?

— It often happens that our interests develop by chance, and this can be a good thing because otherwise we might simply follow well-worn paths. I graduated from the Russian State University for the Humanities (RSUH), where my main field of study was Indonesian. I joined the Indonesian language study group by chance as well, but 25 years later, this experience allowed me to lead an expedition of HSE University to Sumatra (Indonesia’s largest island and the sixth largest island in the world) to conduct research on the Toba Batak language. Back in 2003, however, the outstanding linguists Nina Sumbatova, Yakov Testelets, and Svetlana Toldova invited me to join a project on the Adyghe language. The project later expanded to include other languages of the Western Caucasus. We quickly realised that these languages are both remarkable and little studied. They challenge many established linguistic rules and generalisations, including ones that are highly specific yet often taken for granted. What other language allows you to express something like 'the person who saw whose friend'? In these languages, however, we are not even certain that ‘friend’ is a noun and ‘saw’ is a verb. Phenomena like these are of great significance for linguistic research.

The Western Caucasus project was initially based at RSUH before gradually moving to HSE University. Over the years, the team conducted numerous field expeditions to study a wide range of languages and dialects—as linguists themselves put it, 'expeditions into a language.'

— Which languages are part of the Western Caucasian language family?

— The Western Caucasian, or Abkhaz-Adyghe, language family includes Abkhaz, Abaza, Adyghe, and Kabardian-Circassian. The Adyghe and Kabardian peoples often regard themselves as a single people speaking the same language—Adyghe, also known as Circassian. Until recently, the family also included Ubykh, whose last native speaker died in Turkey in 1992.

The peoples who speak these languages live not only in Russia and Abkhazia but also in diaspora communities that emerged following the mass displacement after the Caucasian War, primarily in Turkey, Syria, and Jordan. There are even two villages in Israel where our laboratory once carried out a small field expedition. More recently, during an expedition to Adygea this April, I visited a village inhabited by Adyghe people who had relocated there from Kosovo.

— How different is the language situation between those living in Russia and those who left their homeland after the end of the Caucasian War in the 19th century?

— The development of the Adyghe languages has followed different paths. In the Soviet Union, these languages continued to evolve, being used in everyday communication, literature and newspapers, and television and radio broadcasts. However, younger generations began to switch to Russian in their daily communication.

In Turkey, first under the Ottoman Empire and later under Atatürk and his successors, everyday use and, especially, the teaching of these languages were discouraged—to put it mildly—as they were viewed as expressions of Circassian nationalism. As a result, the Adyghe people in Turkey gradually lost their language. However, in the 21st century, Turkey has allowed Adyghe language classes in schools and the teaching of Adyghe languages at universities. The current situation with the language in Syria remains unclear, but those who have left the country in recent years continue to speak their native language. In Israel, the situation of the local Circassian community is favourable, with the language being taught in schools. In the Adyghe village of Kfar Kama, as we observed during our 2017 expedition, street names are displayed in three languages: Hebrew, Arabic, and Adyghe in Cyrillic script.

Expedition to Adygea
Photo: Yury Lander

— What are some of the distinctive features of the Abkhaz-Adyghe languages that set them apart from other languages spoken in the Caucasus?

— The Abkhaz-Adyghe languages differ significantly from those around them because they are polysynthetic, meaning that they can incorporate an enormous amount of information into a single word—although linguists may debate to what extent these 'words' correspond to what is usually understood as a word in European languages. It is therefore not surprising that the Abkhaz-Adyghe languages contain some extremely long word forms. For example, the longest word in the Adyghe language corpus is къызэрэригъэпшIыкIутIукIыщтыгъэхэр (some character combinations in this word represent single sounds, though). It can be translated as 'the way he bashed out [tunes]' and literally means 'the way he turned them into twelve.'

Native speakers can produce these words (or word-like units) spontaneously in speech, and they can do so in different ways. As a result, remarkable linguistic phenomena emerge, such as pauses occurring within words. These languages therefore function differently from the so-called Standard Average European languages and feature a different approach to processing information. Consequently, developing their corpora, which required the analysis of previously undocumented forms, presented a significant challenge for us.

— Which one of the Western Caucasian corpora do you consider the most important?

— Our main achievement is the large Adyghe Corpus, which now contains about 12 million words—a substantial amount for a language corpus. It is essentially a corpus of the so-called literary language (although the concept of a literary language can, of course, be interpreted in different ways), but it has been significantly enriched with folklore materials thanks to our colleagues and friends, the outstanding folklorists from Adyghe State University. The Adyghe Corpus was officially launched in 2018, although it had actually been created two years earlier. It is probably important to note that this is the first corpus of a polysynthetic language to include detailed grammatical information. The excellent computational linguist Timofey Arkhangelskiy and, of course, Irina Bagirokova, a linguist and native speaker of Adyghe, played a crucial role in creating the Adyghe Corpus. Without their contributions, this project would not have become a reality.

In addition to the large corpus, there is also the Oral Corpus of the Adyghe Language, now a joint project between HSE University and Adyghe State University in Maykop. The oral corpus is directly linked to field expeditions, during which speech samples are recorded in villages and then transcribed and subjected to grammatical analysis at the Linguistic Convergence Laboratory.

— How did researchers in the Caucasus republics and local residents respond to your interest in their native languages?

— The development of the Adyghe Corpus was supported not only by HSE University but also by Adyghe State University. As for local residents, they were initially surprised and often asked whether I had relatives in the region. However, we can always explain why our projects are important both for them and for us.

— Is creating a corpus mainly a matter of fieldwork or computer-based work?

— Collecting material for a corpus requires long and painstaking fieldwork.

We travel extensively, and linguistic expeditions play an important role in the work of HSE University’s fundamental linguists. In particular, we conduct numerous expeditions to the Caucasus, where we have recently launched several projects. In the Western Caucasus, these include projects on Adyghe, Abaza, and Kabardian. We have also carried out expeditions to study Ossetian and many Dagestani languages.

For example, expeditions focused on documenting one of the Avar dialects and the relatively rare Andi language have recently returned.

— How do the administration of the Linguistic Convergence Laboratory and the HSE Faculty of Humanities respond to your initiatives?

— This is one of the laboratory’s main areas of activity, and the laboratory administration fully understands the importance of fieldwork. The Faculty of Humanities also supports our initiatives, providing both financial and organisational assistance.

— How actively do students participate in work on the corpora?

— Students make a major contribution to the collection and processing of texts, especially oral materials. This is an essential part of field expeditions, which involves recording texts, processing them, translating them, and analysing the results. Moreover, students are actively involved not only in creating corpora but also in their further development and expansion.

— How is your research used in the educational process at HSE University?

— It is essential for the School of Linguistics to incorporate linguistic diversity into its educational programmes, because modern linguistics cannot be based solely on the study of the best-known European languages. We can understand how language works only by examining data from a wide range of languages—from small to large, including those that differ fundamentally from the ones we are most familiar with.

Our students use corpora to write term papers and theses. This year, for example, two bachelor’s students have based their theses on materials from the Adyghe Corpus.

In addition to term papers and theses, bachelor’s and master’s programmes in fundamental linguistics include courses that introduce students to minority languages, and many students choose to take them. Our goal is not to teach students to speak these languages but rather to help them understand the languages’ grammatical structures and other distinctive features. I also occasionally teach an Adyghe language course myself.

— What are the practical outcomes of your work?

— In 2019, Irina Bagirokova and I gave a talk to principals of Adyghe schools on how the corpus could be used in teaching the native language. It would be good to develop a CPD programme for teachers, but more generally, we would like our corpora to be used more widely in research and education. We want people to recognise the remarkable opportunities these resources offer.

George Moroz

George Moroz

'The creation of morphologically annotated corpora of languages spoken in Russia, including dialects and non-standard varieties of Russian, is one of the most important areas of our laboratory’s work. Our team is among the leading groups in the country in this field. Russia is home to around 155 languages, many of which are endangered. Developing corpora for these languages is an important tool for their preservation. Corpora make it possible to document languages, capture their distinctive features and linguistic data, and create resources for further research and support of these languages, which are a vital part of the country’s multi-ethnic cultural heritage. All of our laboratory’s resources, including not only corpora but also dictionaries and linguistic databases, are available on our website. Examples of spoken corpora's practical application include both internal and external uses. Internal applications are related to the needs of the language itself, since corpora serve as excellent resources for educational purposes, including use in schools. For example, the Russian National Corpus has created a dedicated section that provides a frequency dictionary, a practice example generator, and a 'word at a glance' feature, where users can access both grammatical and diachronic information about a particular linguistic unit. External applications of corpora, on the other hand, involve research conducted using data from different corpora and languages. For example, our colleague Irina Politova defended her master’s thesis in which she modelled different word orders in sentences containing several units. Thanks to the morphological annotation of our corpora, we can conduct similar studies on entirely different languages, such as Adyghe, with relatively little additional effort. Since we have access to corpora of very different languages, we can even attempt to measure the extent to which speakers of other languages reflect trends observed in Russian and model the degree of Russian influence on other languages spoken in Russia. Of course, our contacts with teachers and language activists from various ethnic communities in Russia are still limited. However, Yury Lander’s work provides a positive example of how collaboration with members of a language community can help us discover, develop, and create valuable resources that benefit both native speakers and academic research.

See also:

‘Speech, Facial Expressions, and Gestures Cannot Lie’

Would you like to know whether a speaker’s trembling voice or an accidental gesture can give them away? At HSE University in Nizhny Novgorod, researchers are developing an algorithm that analyses speech, facial expressions, and gestures, and determines whether information is truthful with 92% accuracy. The project has applications ranging from forensic examination and bank recruitment to fundamental research. Anna Khomenko, head of the research group and Senior Research Fellow at the Centre for Language and Brain at the HSE Faculty of Humanities in Nizhny Novgorod, explains how students and researchers are working together to create a corpus of video recordings, train a classifier, and prepare to introduce computer vision technology.

Biologists Discover 'Molecular Fingerprint' of Preeclampsia

Researchers at HSE University employed a new method to model hypoxia in placental cells during pregnancies complicated by preeclampsia and identified molecular markers of tissue hypoxia. Since hypoxia is one of the key mechanisms underlying preeclampsia, these findings are important for a more accurate and timely diagnosis of the disease and for the development of effective treatment methods. The paper has been published in Placenta.

Social Integration: At the Crossroads of Knowledge and Values

The International Laboratory for Social Integration Research (ILSIR) at HSE University studies the challenges faced by vulnerable groups and explores ways to help them participate fully in everyday life. To develop effective solutions, the laboratory’s researchers combine cutting-edge methods with practical fieldwork. In this interview with the HSE News Service, Laboratory Head Elena Iarskaia-Smirnova discusses the laboratory’s work.

Physicists Discover What Happens Inside a Stable Vortex

Large vortices with characteristic spiral arms are often observed in the atmosphere and the ocean. Physicists from HSE University have explained how these structures form and why they retain their shape. The researchers found that velocities at points located along the same vortex arc remain correlated even over long distances. At the same time, this correlation weakens rapidly with increasing distance from the vortex centre. These differences help explain the formation of spiral arms and may improve models of atmospheric and oceanic currents. The findings have been published in Physical Review Fluids.

‘Science Is Universal—It Knows No Borders’

Fuad Aleskerov, Tenured Professor and Director of the International Centre of Decision Choice and Analysis at HSE University, together with his colleagues, has developed methods of network analysis in bibliometrics that have made it possible to identify patterns in the appearance and citation of publications in academic journals, as well as their influence on each other. When one or a number of studies are frequently cited by a wide range of journals, this is an indicator that the research is of high quality. By contrast, extensive cross-citation within a limited group of journals increases the likelihood of identifying a network of predatory publications.

Scientists Propose Method for More Efficient Resource Use in Machine Learning

An international group of researchers, including mathematicians from the AI and Digital Science Institute at the HSE Faculty of Computer Science, has provided a theoretical justification for a simple and computationally efficient method of estimating uncertainty in Stochastic Gradient Descent (SGD). The paper has been published on the scientific preprint server arXiv.org and presented at AISTATS 2026.

Biologists Discover Unique Properties of MiR-93-5p MicroRNA in Prostate Cancer

Researchers at the International Laboratory of Microphysiological Systems of the HSE Faculty of Biology and Biotechnology investigated how different isoforms of the same microRNA influence gene function in prostate adenocarcinoma. The study found that in some cases, microRNAs can reinforce each other’s effects by targeting and suppressing the same genes. This finding offers a fresh perspective on the molecular mechanisms underlying tumour development and on the search for disease biomarkers. The results have been published in PeerJ.

HSE Develops App for Assessing Phonological Processing in Children

Researchers at the HSE Centre for Language and Brain have developed a new digital tool for assessing children's phonological processing skills—the ZARYA (Sound Analysis of the Russian Language) test battery. It is the first standardised application in Russia designed to provide a fast and reliable assessment of children's ability to distinguish speech sounds, retain them in working memory, and perform phonemic analysis. The app runs on Android tablets and smartphones and is available for download from RuStore. Details of the test validation have been published in the Journal of Speech, Language, and Hearing Research.

‘In Science, You Are Your Own Boss’

Polina Nasledskova is interested in identifying gaps in linguistics and topics that have been overlooked by other researchers. In an interview for the  Young Scientists of HSE University project, she spoke about rare ordinal numerals in Nakh-Daghestanian languages, the benefits of knitting for concentration, and the beauty of the Patriarshy Bridge.

Scientists Explain How Emotions Shape Attitudes Toward Digital Governance

Today, interactions between citizens and government increasingly take place through digital governance platforms, including digital public services, AI-powered systems, and algorithmic decision-making tools. Until now, however, these technologies have largely been viewed as technical instruments, with their effectiveness assessed primarily in terms of efficiency and user-friendliness. The authors of a new study propose a broader perspective, arguing that digital governance should also be understood as an emotional experience that directly shapes citizens' trust in public institutions.