Morphology studies how morphemes—irreducible units of meaning—combine to form words, distinguishing between free morphemes (standalone words like 'dog') and bound morphemes (attached elements like '-ing'). Languages vary in their morphological complexity along a spectrum from analytic (few affixes, relying on syntax—English) to synthetic (many affixes, marking multiple grammatical categories—Turkish, Hindi) to polysynthetic (entire sentences in single words—Greenlandic). Key concepts include roots/stems, derivation (changing meaning), inflection (marking grammatical categories), and alignment systems (nominative, ergative, split) that determine how sentence roles are marked. Glossing provides standardized notation for analyzing language structure.
Conlang Morphology Tutorial: Morphemes, Affixes & Alignment
Added:When you learn a language, your learning can usually be divided up into three main parts - pronunciation, grammar and words. In order to become proficient in a language, you need to master all three parts. It follows that in order to make a language, you need the same. So far in this series, we've looked at pronunciation, or phonology. Now it's time to begin on grammar.
"Grammar" is a term used by both linguists and non-linguists to describe all the bits of language that stick words together. But as linguists, we more often divide it up into separate fields, namely "morphology", which studies how smaller elements come together to form words, and "syntax", which studies how whole words combine to form larger units, like phrases and sentences.
Even though we do this, it's important to understand that they aren't strictly separate categories - as you'll see in the next few episodes, they do often interact with one another and with phonology to form combined fields, such as "morphosyntax" and "morphophonology". However, they are still useful to distinguish, and this is the approach taken by most linguistics courses, so I will follow suit in this series, with the next few episodes focusing on various parts of morphology and related topics and a subsequent episode dedicated to syntax. It's just worth knowing that a lot of the ideas I'll introduce in this video are also applicable to syntax.
Just as phonology focuses on phonemes, morphology focuses on morphemes. Where a phoneme is an irreducible unit of sound, a morpheme is an irreducible unit of meaning.
Morphemes can be classified as either "free" or "bound". A free morpheme is one which can stand on its own, forming a full word. For examples, think of any of the basic, unmodified words of English, like "dog", "sing", "smooth" or "with". Bound morphemes, however, cannot stand alone, relying on free morphemes to be used. Think the "-ing" in "eating", the "-es" in "wishes" and the "-er" in "younger". These elements can only be used when attached to host words - "-ing" could never be used on its own.
Just as phonemes can have allophones, morphemes can have allomorphs - different phonological realisations of the same underlying morpheme. Both free and bound morphemes can have allomorphs.
We've discussed the allomorphs of the English plural morpheme in a previous episode - it has different forms after voiced sounds, voiceless sounds and sibilants. Since these realisations are separate phonemes in English and don't vary like this in other words with the same sequences of sounds, they're not allophones, but they do show allomorphic variation.
A free morpheme with multiple allomorphs can be seen in the indefinite article "an".
Before vowels, this takes the form "an", while before consonants, it loses the N to become "a".
A few times already in this video and in previous episodes, I've used words like "noun", "verb", "article", "adjective" and "pronoun". But what do these mean exactly? Well, they're what we call "parts of speech" or "word classes". Basically, not all words are created equal - different ones do different things within a sentence.
We can categorise them roughly based on "form" and "function", or "how they look and are internally composed" and "what they do in the sentence and how they interact with other words". This is, very roughly, the morphology/syntax distinction we've discussed already.
In the work of the Ancient Greek grammarians, there were considered to be eight parts of speech - nouns, verbs, participles, pronouns, adverbs, articles, prepositions and conjunctions. This system has been expanded and modified over the last two millennia to give the general classification we use today. But crucially, this system is based on English, Latin and Greek, and influenced by various other languages of Europe, but it does not work for all languages around the world. Some languages, for instance, lack adjectives, adverbs or determiners. Others make only marginal use of certain classes, while others still have parts of speech that European languages don't. The Wagiman language of northern Australia, for instance, only has a few true verbs, but has an additional word class, the "coverbs", which must combine with regular verbs to express their meaning.
But that aside, let's take a look at some of the more common parts of speech and the formal and functional behaviours that define them.
A noun is, generally, a term for a person, place, thing or idea. It can be placed in relation to a verb to show the entities involved in an action or state. It can also have properties attributed to it via adjectives and be related to other nouns with adpositions. In English, nouns can be either singular or plural and can take, or include, morphemes such as diminutives, denoting a smaller version of something, and agentives, showing the doer of an action. In other languages, they might also take cases, which show their role in the sentence, class (suggesting particularities of their behaviour) and definiteness (showing whether the specific instance of a noun is already known).
Examples in English are "donkey", "woman", "book", "darkness", "river", "fun" and "extinction".
A verb is a word expressing an action or state. They define what the sentence is about and, along with nouns, are one of the only two word classes found in all known languages. Functionally, they usually take nouns or pronouns to define what's happening. They can also be delimited with adverbs to specify their meaning. In English, they change their form based on person and number to show who's involved, tense and mood to show if and when it's happening and can combine with other verbs to show further tense, aspect, mood and voice to describe the nature of the event. In other languages, you may also encounter features such as evidentiality, to show knowledge of the event and negation, to show an event as having not happened. English examples are "think", "eat", "sleep", "want", "describe", "have" and "be".
Adjectives describe nouns. They can mark comparison, to compare one noun to another, and in other languages may "agree" with nouns, mimicking their plurality, case, class and definiteness. Examples in English are "big", "happy", "yellow", "wooden" and "Canadian".
Similarly, adverbs describe verbs and adjectives. They often end in "-ly", especially when derived from adjectives, though many also don't. Examples are "slowly", "happily", "now", "well" and "here".
Pronouns are, in some languages, a subclass of nouns. They stand in for nouns, often to reference something without repeating or naming it in full. In English, they are similar in function to nouns, but slightly more restricted. For instance, they usually can't take adjectives, although this is because technically, they replace whole noun phrases, including a noun and its associated adjectives and determiners. However, formally, they're a little more elaborate. Unlike nouns, they vary based on case and their forms aren't as clearly derived as in nouns. English examples are "you", "we", "her", "what", "no-one" and "those".
In English, we have a part of speech known as "prepositions".
These define the relationship of nouns to other parts of the sentence. They're mostly simple, not containing any smaller morphemes or able to have morphemes added to them. However, there are also "complex prepositions", made up of phrases involving a noun and preposition, which together fill the same role. We call them "prepositions" because they come before the noun, but in other languages they may come after, so we term the whole word class "adpositions". Examples are "with", "in", "to", "regarding" and "in spite of".
Other common parts of speech are conjunctions, which join clauses or sentences together, determiners, which specify a noun in some way, and particles, which don't include or take special morphemes and add grammatical information to a sentence, rather than physical meaning.
We can also describe parts of speech as either "open class" or "close class". This describes how easily new words of that type can be created. Open classes allow new words to be created easily, while close classes do not. Open classes are usually lexical; that is, their words have clear, easily definable meanings. Close classes, however, are usually grammatical; their words are harder to define and perform functions to bind sentences together. Examples of open classes are nouns, verbs, adjectives and adverbs. Close classes include pronouns, determiners, adpositions and conjunctions.
We'll go into more detail on each of the parts of speech in the coming episodes, but as a beginner, you're probably best just sticking to basic parts of speech in English or shown in this list. Try to understand the basic essence of what each is so you can assign appropriate behaviours to it when we reach that point in the series.
In case you haven't noticed, parts of speech represent the categories of free morphemes. They're the types of independent word which can occur in a language. But what about bound morphemes?
Well, one basic distinction we can make is in their function. When a bound morpheme attaches onto a free one, we refer to this free morpheme as either a "root" or a "stem". A root is the most irreducible form of a word, composed of a single free morpheme. A stem can actually be composed of more than one morpheme, but it's just the lexical part of a word, without any grammatical elements added on. For example, in the word "brotherhoods", we can find three morphemes - "brother", "-hood" and the plural "-s". The root of this word is "brother" - the smallest part of the word with an independent non-grammatical meaning. Onto this, the morpheme "-hood" is added. This doesn't add grammatical information, but it does modify the meaning of the word. Adding a morpheme to alter meaning like this is called "derivation". In this case, no more derivational morphemes are added, so "brotherhood" is the word's stem. However, we can also add the "-s" on, to form the word "brotherhoods". This morpheme represents the grammatical category of "plurality", indicating that we're talking about more than one of the thing in question.
It's adding to the grammatical status of the word, but isn't adding actual lexical meaning, as the essence of "brotherhoods" doesn't differ from that of "brotherhood". This is therefore not a case of derivation, but rather one of "inflection" - that is, adding grammatical morphemes to a word.
In this example, we saw both morphemes added onto the end of the root. But this isn't the only option. Bound morphemes coming after the root are called "suffixes". If they come before, however, they're "prefixes". Examples in English are the negative prefixes "un-" and "in-" or "co-", meaning "together". Most English "affixes", as they're neutrally called, are suffixes, including all grammatical morphemes. But in the Zulu language of South Africa, this is reversed, with most affixes being prefixes before the root and only a few suffixes.
Other languages may be intermediate to this, using neither more often than the other, such as is the case in the Choctaw language of the southern United States.
While prefixes and suffixes are most common, affixes can also occur in places other than before or after the root. In Tagalog, a language of the Philippines, certain aspects of verbal morphology are marked through "infixes", which occur inside a word stem, with the stem being broken up in some predictable way.
In German, past participles (the forms of verbs equivalent to English "eaten", "been", "sung" and "gone") are formed with circumfixes, where morphemes are added both before and after the stem, but where one or both have no meaning on their own.
Many languages also show grammatical meaning by "reduplication", where all or part of a stem is doubled, or sometimes even tripled. The element which is added in these cases in known as a "duplifix". This can be seen in the Igbo language of Nigeria, where gerunds (noun-like verbs) are formed by reduplicating part of the verb.
Most of the time, languages use some sort of affix to derive and inflect words. But some use other means. Affixing involves adding a consistent, unitary morpheme to a root or stem.
Given that they're just stacked together like this, we call this "concatenative" morphology.
However, some languages change the internal structure or elements of a root. This we call "non-concatenative" morphology. The most famous example of this is in the Semitic languages, like Arabic, Hebrew and Aramaic, where roots are not blocks of consonants and vowels like in many other languages, but instead consist of just consonants. The vowels of the word are then added in at specific places to form derived and inflected forms. For instance, Arabic has the root "S-L-M" which denotes concepts relating to safety, peace and surrender. The root never appears in this vowelless form, but all derived and inflected stems can be traced back to this base. By inserting a short and a long "A" into the middle, we can form the word "سَلَام" ("salām"), which is a common greeting in the Arab world. But by inserting other vowels, we get "إِسْلَام" ("ʔislām"), the source of the English word "Islam", the religion of the Arabs.
Another derivation makes use of yet another vowel template, as well as a prefix, to give "مُسْلِم" ("muslim"), the source of English "Muslim" - a follower of Islam.
But non-concatenative morphology isn't restricted to this consonant-root template system of the Semitic languages - we actually also use it in English! Think about nouns like "mouse", "man" or "foot". To form their plurals, we don't add "-s" like most other nouns; the vowel in the root gets changed - "mouse" to "mice", "man" to "men" and "foot" to "feet". Our verbs do funky things too, with "sing" having past tense "sang" and participle "sung" and "have" using "had" for both past tense and participle.
Importantly, not just any change to the stem is non-concatenation, only those cases where certain phones change in predictable or semi-predictable ways. The past tense form of English "go", for example, is "went", which is not a case of non-concatenation, but rather one of "suppletion". Suppletion is where a certain form of a word is not related to the reference form, but rather started out as a separate word, then got co-opted as, say, a plural, past tense or diminutive. Suppletion is incredibly rare as a d regular strategy of derivation or inflection, but crops up in many, or most, languages as an irregularity. In English, we also see it in the plural "people" compared to singular "person" and the comparative and superlative "better" and "best" compared to the usual "good".
So, now we've looked at what morphemes are and where they go, but there's still the question of how many you need. As it turns out, this is another place we see variation. In some languages, words can have many inflectional morphemes added to them, while others don't allow any.
This forms a spectrum. Languages with a low morpheme-to-word ratio are termed "analytic".
These languages don't use many affixes per word, if any, and rely instead on syntax to convey this information. English is an analytic language. We don't have much in the way of inflection, with our nouns only showing plurality and possession and our verbs showing only tense, person, number and mood, with other forms being shown through long strings of separate verbs.
Note that this is only relevant for inflection and not derivation. English has a rich range of derivational morphemes available, but only a few inflections, so it remains analytic.
Languages which are very analytic, to the point of having no inflectional morphology, are called "isolating" languages. The Yoruba language of Nigeria and surrounding parts of West Africa is isolating, as any grammatical functions are filled by separate words and placed in certain positions relative to the main headword.
On the other hand, some languages mark far more categories on their words. These are called "synthetic languages". Synthetic languages have two subtypes - fusional and agglutinative.
In agglutinative languages, each grammatical category has its own morpheme and these are stacked together separately. Turkish belongs to this class. Most morphemes denote just one thing and are stacked together at the end of the word. In fusional languages however, many grammatical categories are condensed into just a few morphemes. In Hindi, there's a single suffix conveying the first-person singular masculine future indicative form of a verb, giving that morpheme five categories combined into one.
The line between fusional and agglutinative isn't a hard one. Even Turkish uses a single morpheme for both person and number in its verbs and the fusional language Persian forms its imperfect tenses by adding a separate prefix to the perfective form. And just as the extreme of the analytic side is an isolating language, a highly synthetic language can be called a "polysynthetic" language. These may be fusional or agglutinative, but are generally more agglutinative. An example language is Greenlandic, in which entire sentences can be expressed by a single inflected verb.
Next, let's talk about alignment. This is a topic that's important for several parts of speech, both in morphology and syntax. It'll underline several of the things we talk about in the next few episodes, so it's important to get a basic understanding of now.
We've seen that nouns and verbs can work together to make a sentence. But nouns aren't all the same in how they do this. We can define two main types of sentence - those where the verb requires two noun phrases and those where it requires only one. We call these "transitive" and "intransitive" respectively.
Transitive verbs generally denote an "agent" ("A") doing something to a "patient" ("P") (think "the girl eats the cake", "David sees John" or "my cat killed God"). Intransitive verbs generally show something happening to an "experiencer" ("S"), or the experiencer doing something without affecting anything else. Think "the nurse runs", "the snake dies" or "the Moon rises". It is actually, as ever, a little more complicated than this - if you want to challenge yourself, look up "semantic roles" (I'll leave a link in the description), but just stick with this for now.
All languages need a way of distinguishing the agent from the patient, so speakers can tell who is doing the action to whom. "The girl eats the cake" is very different to "the cake eats the girl". Languages have a variety of different ways of doing this, which we'll explore more in the coming episodes. As you can see, in English we use the order of the constituents. Swapping the order swaps the meaning. In other languages, word order is less important. In German, you can arrange the sentence either way around and still know which noun is filling each role. That's because German relies partly on noun morphology to distinguish A from P. Other languages instead use verb morphology. In Greenlandic, there are different forms of the verb according to both the agent and the patient. If these are swapped, a different suffix must be used.
But how does this relate to the intransitive experiencer? Well, if a language is going to mark the role of the nouns in transitive sentences, it generally likes to keep things symmetrical and do the same for intransitive sentences, especially since you can sometimes turn transitive sentences intransitive or vice versa - keeping some element of marking is helpful for this conversion. So we end up with these three roles and find that in different languages, they are grouped in different ways.
In English, we treat the experiencer of intransitives the same way we do the agent of transitives. Since our main marking method is word order, we put both the A and the S roles before the verb, while only the patient comes after. Our verbs also use morphology to agree with these elements. In the present tense, we can add an "-s" suffix to the verb if the agent or experiencer (which together we call the "subject" of the verb) is third-person singular, while if it's a pronoun "I", "you", "we", "they" or a plural noun, it doesn't do this. Languages which behave like English in this way are said to have a "nominative" alignment.
But, as you might expect, there are other options. Basque is what's known as an "ergative" language. It uses the same marking for the experiencer as the patient. This stems from the idea of the experiencer being affected by the action, rather than instigating it. In English, the patient of the verb "kill" and the experiencer of "die" are filling the same role, but we structure our sentences nominatively. In Basque, this relationship is more obvious, as exactly the same marking is used for the patient and the experiencer. The agent is treated independently.
Some languages play on the potential double function of the experiencer. These are called "split" languages. In some instances, a nominative pattern is used, while in others, an ergative one occurs. The reasons for using one pattern or the other vary between languages.
Sometimes it is related to whether the experiencer is more active or passive, but sometimes it's down to grammatical or pragmatic conditions. Hindi is an example of this, where the marking system of the noun depends on the "aspect" of the verb.
Other systems are rarer. One is with "direct" languages, where no distinction is made between agent, patient and experiencer. In word order, this may mean there's no fixed word order, while in terms of morphology, it simply indicates a lack of this.
In "tripartite" systems, A, P and S are all marked separately.
Each may have distinctive morphology, or there may be special word order for each one.
One of the coolest systems is probably "direct-inverse" or "hierarchical" alignment, where the marking depends on where the participants fall on a hierarchy of importance.
If a constituent is of a certain slot relative to another, they're always marked the same, but there may be an additional marked to show that the roles are reversed. This is the system used by Ojibwe.
The rarest system is that of "transitive" languages, where the agent and patient are marked the same, but the experiencer is marked separately. Given that this doesn't actually provide extra information about the roles, since you can already tell if a verb is transitive or not, not many languages use it. We mainly see it when a language which historically used one system is changing over to another, or where there is some other reason to distinguish the cases.
Now, a lot of the time, languages use the same alignment across the board, but some have different systems depending on the type of marking. For instance, Nepali uses ergative noun morphology, but nominative verb morphology. Verbs show the same marking for the agent and experiencer, but nouns show the same marking for the patient and the experiencer.
One thing that's actually very common is for one or more marking strategies to use a direct system, while others use something else. This is because you usually don't need all three systems to show the same thing, so it can be simplified by ignoring this. Mandarin uses a nominative-style word order, with agent and experiencer coming before the verb and the patient after. However, it uses no verb or noun morphology to show this, so could be said to have direct alignments for these.
The final general thing you need to know is "glossing". We won't look too much at this in this video - I'll leave some links in the description and I'll do a full video on the topic at some point in the future, but you do need the basics for the coming videos, so let's take a quick look.
When we're looking at languages and their structure, it's hard to tell what's going on if we don't speak that language, which, usually, we don't. So we note down the meaning and function of every morpheme under the language sample - a kind of "running, in-text translation" that also shows all the important information. This is called a "gloss" and is something you've probably seen quite a lot already in this series.
The basic principle is to translate every lexical word and give a functional description of every grammatical morpheme. We do this using a set of standardised abbreviations which are written in capital letters.
Let's gloss the English sentence "The man eats his chips". We start by translating (or in this case, just rewriting) all the lexical words - "man", "eat" and "chip". Note that we use the unmarked form of "eat" and "chip" - what's known as the "lemma" form. All the other words and morphemes are grammatical, so we leave these for now. Next, let's look at the word "the". This is the definite article. It's filling the function of making the noun definite - referring to a certain individual. This is the only role it's filling and it's use doesn't preclude any other functions, so we gloss it just as "definite", the abbreviation for which is "DEF". Next, let's do the same with "his". This is the third-person singular masculine possessive pronoun, showing the chips as belonging to the man. This morpheme shows four grammatical functions - third person (i.e. not involving me or you"), singular (denoting possession by in individual, not a group), masculine (denoting possession by a male individual) and possessive (showing ownership or holding of the chips). Combinations of person and number occur often enough that we can combine them into one abbreviation. "Third person singular" is denoted by either "3SG" or just "3S". "Masculine" is "M" and "possessive" is "POSS". Since these are all in a single morpheme, we combine them using dots without spaces, so the full gloss for this word is "3SG.M.POSS". Now, onto the bound morphemes. The two most obvious of these are the "-s"s on "eats" and "chips". These are filling different roles. On the verb, it's showing a third-person singular agent and the present tense, while on the noun, it's showing plurality. However, these are separate morphemes to their host word, so we shouldn't join them on with dots. Instead, we use dashes. A dash indicates a morpheme break, a dot indicates the same morpheme. We see both in "eats". In reality, there is actually a bit more going on with this word - it shows more than just the categories we've discussed, but for simplicity, let's stick with this. Ideally, all categories should be marked in a gloss, but you can usually omit some if they're not relevant to a discussion.
So, that's it, right? Well, not quite. "Man" does actually show something, even though we can't see it. When we use this word, it's obvious we mean only one person. Two or more would be "men". We can therefore add an extra gloss here, showing that this morpheme also codes the singular. "Man" is a special case, because the morpheme itself shows the singular. But other nouns are different, as they use a separate morpheme. If we had the word "chip", we'd know it was only one because of the absence of the plural suffix. We can therefore suppose an invisible morpheme with no sound, which exists just by virtue of another's absence. We call these "zero-morphemes" and denote them using an "O" with a line through it. We can then gloss this just the same as any other.
When working with other languages, we usually put the gloss in between the sentence and it's translation. We also often separate out morphemes with dashes, so we can see where the morpheme boundaries are. If you want to get to grips with glossing, the standard system uses the "Leipzig glossing rules", which are linked in the description.
So, that brings us to the end of this episode. I realise this isn't a very practical episode, but hopefully you can start thinking about how you might want your conlang to behave.
If you're following along, feel free to create some temporary words using your phonology and phonotactics and start sketching out some of what we've discussed in this video. The next episodes will go into various topics on specific parts of speech, so until next then, thank you for watching.
Up Next

Case & Morphosyntactic Alignment Explained: Ergativity & More
@ColinGorrie
2.7K views•2023-01-05

Constraint Interaction in Optimality Theory: Epenthesis & Deletion
@Vidyamitra
180 views•2018-08-20

Grammatical Case: Definition, Types & Examples in Linguistics
@LexisLang
3.8K views•2023-03-12

Accent Expert Explains U.S. Regional Dialects | Part 1
@WIRED
9.3M views•2021-01-21
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Linguistics





























![[Шесть лекций о лингвистике] Как можно изучать другие языки, при этом не говоря на них?](https://i.ytimg.com/vi_webp/zROBHR6e4DQ/maxresdefault.webp)




![[Wikipedia] Nonconcatenative morphology](https://i.ytimg.com/vi/Ojg8g_ouoyY/maxresdefault.jpg)







