Unit5 Research Methods in Historical Linguistics
Unit5 Research Methods in Historical Linguistics
HISTORICAL LINGUISTICS
Common Ancestors
A primary reason for linguistic similarity is a shared common ancestor. Historical linguistics focuses
on tracing languages back to proto-languages. For example, Proto-Semitic is the common ancestor
of Hebrew and Arabic, explaining why these languages share similar words like “shaloom” (Hebrew)
and “salaam” (Arabic), both meaning “peace.” This is not coincidental but rather a result of regular
sound changes that diverged as the two languages evolved independently. Hebrew retained the
consonant cluster “sh”, while Arabic simplified it to “s”.
Language Contact
Not all similarities arise from shared ancestry; some are the product of direct or indirect contact
between language-speaking communities. Borrowing occurs when languages influence each other
through trade, conquest, migration, or cultural exchange. For instance, Hebrew and Yiddish share
words like “shalom” due to prolonged cultural and religious contact. Similarly, modern Arabic
includes numerous words borrowed from Persian, Turkish, and European languages, reflecting
historical exchanges.
Areal factors also contribute to shared features among unrelated languages. The Balkans
Sprachbund, for example, comprises Albanian, Greek, Macedonian, and others, which share
grammatical structures like the postposed definite article, despite belonging to different language
families. These similarities arise not from common ancestry but from prolonged contact in the same
geographic area.
Language Universals
Another explanation for linguistic similarities lies in universals—features that emerge independently
in many languages due to shared cognitive, physical, or social constraints. For example, the sound
[t], produced by placing the tongue against the alveolar ridge, is found in both Hebrew and Dutch
despite their unrelated origins. This is because [t] is one of the simplest and most universally favored
consonants, given the human articulatory system.
Similarly, the tendency for languages to develop pronouns, negation, and basic word orders reflects
universal patterns in human cognition and interaction. These features are not the result of shared
ancestry or contact but rather arise from the intrinsic properties of human language.
Typological Similarities
Typology, the classification of languages based on shared structural features, provides yet another
lens for examining similarities. While Hebrew and French are from different language families, both
possess two grammatical genders—masculine and feminine. This typological trait may arise from a
historical tendency in their respective families (Afro-Asiatic and Indo-European) to categorize nouns.
Gender as a feature also highlights cultural influences, as languages often evolve systems that reflect
their speakers’ social and cognitive frameworks.
Typological comparisons reveal broader patterns of language structure. While Hebrew and French
share two genders, Mandarin Chinese lacks grammatical gender altogether, showing that gender
marking is not universal. Conversely, some languages, such as Swahili, feature extensive noun classes,
demonstrating the diversity in typological organization.
Linguistic similarities have long fascinated scholars, prompting various approaches to understanding
their origins. Historical linguistics, language typology, and areal linguistics each offer distinct
frameworks to explain why languages resemble one another. Historical linguistics emphasizes genetic
relationships, language typology seeks universals and classification, and areal linguistics highlights
geographic proximity and cultural interaction. Together, these approaches provide a nuanced
understanding of how and why languages may converge or diverge.
The study of linguistic similarity originated in the 19th century with historical linguistics, which posits
that languages resemble one another due to a shared common ancestor. This approach emphasizes
the reconstruction of proto-languages, tracing modern languages back to their roots. For instance,
the similarities between Sanskrit, Latin, and Ancient Greek led to the reconstruction of Proto-Indo-
European, a theoretical ancestor of many Eurasian languages.
Through systematic methods, historical linguists identify cognates—words that descend from a
common source—by observing regular sound correspondences. For example, the word for "father"
in English (father), German (Vater), and Latin (pater) showcases systematic phonological evolution.
Such findings have solidified genetic relationships as a cornerstone of comparative linguistics.
However, genetic relations alone cannot account for all linguistic similarities. Historical linguistics
focuses primarily on diachronic (historical) changes, often overlooking synchronic (contemporary)
convergences influenced by geography or typological features.
In the 20th century, the field of language typology emerged, offering a different lens to study
linguistic similarity. Rather than tracing languages back to a common ancestor, typologists classify
languages based on shared structural features and universal tendencies. This approach highlights
how languages can exhibit similarities despite having no genetic relation.
Typological studies categorize languages into "types" based on properties such as word order (e.g.,
Subject-Verb-Object or SVO in English versus Subject-Object-Verb or SOV in Japanese), phonemic
inventories, or morphological complexity. For example, many unrelated languages, including English
and Mandarin, share an SVO structure, demonstrating a typological rather than genetic similarity.
Typology also investigates language universals—features that arise across languages due to cognitive
and communicative constraints. Joseph Greenberg’s work identified universal tendencies, such as
the near-universal presence of pronouns or the relationship between word order and the placement
of prepositions. These universals highlight the shared cognitive framework of human speakers,
offering insights into why languages sometimes resemble one another independently of ancestry or
contact.
Areal linguistics, also a 20th-century development, emphasizes the role of geography and cultural
contact in shaping linguistic similarities. Unlike genetic or typological approaches, this framework
examines how languages in close geographic proximity influence one another, regardless of genetic
relationships.
One of the most notable examples of this phenomenon is the “Balkan Sprachbund”, a linguistic area
comprising languages like Albanian, Macedonian, Romanian, and Greek. Despite belonging to
different families (e.g., Indo-European versus Afro-Asiatic), these languages share features such as
postposed definite articles and similar case systems due to prolonged cultural interaction.
Similarly, the Baltic languages (e.g., Latvian and Lithuanian) exhibit structural and lexical similarities
that reflect not only shared ancestry but also centuries of geographic and cultural cohesion. Areal
linguistics underscores how shared environments and histories can lead to convergence, even among
genetically distinct languages.
Integration of Approaches
While historical linguistics, language typology, and areal linguistics offer distinct perspectives, they
are not mutually exclusive. In fact, a comprehensive analysis often requires integrating these
approaches:
• Genetic similarity provides the foundation for tracing long-term linguistic evolution and
shared ancestry.
• Typological similarity offers insights into structural convergence and universal principles.
• Areal similarity explains local interactions and borrowings, especially in regions of intense
cultural exchange.
For example, Slavic languages share features due to their genetic ties within the Indo-European
family, but they also exhibit areal influences from neighboring Turkic and Uralic languages. Likewise,
Hebrew and Arabic are genetically related through Proto-Semitic roots, but contact with Greek,
Persian, and European languages has also shaped their modern forms.
In linguistics, the question "Why?" arises frequently when we encounter a particular phenomenon,
pattern, or observation in language. Providing a satisfactory answer requires drawing from diverse
explanatory frameworks, as linguistic phenomena are shaped by historical, cognitive, functional,
developmental, and even coincidental factors.
Historical Explanations
Many linguistic phenomena can be explained by their historical development. Languages are not
static; they evolve over time due to processes like sound changes, grammaticalization, and lexical
borrowing. A historical explanation traces a feature back to its origins, often reconstructing the steps
that led to its current form.
For instance, the irregular past tense forms in English (“go/went”, “be/was/were”) can be explained
historically: these forms come from different linguistic roots in Old English and earlier Germanic
languages. While they may appear arbitrary today, their irregularity reflects historical shifts rather
than a lack of logic.
Historical explanations also account for why related languages share certain features. For example,
the similarities between Spanish, Italian, and French vocabulary arise from their shared Latin
ancestry. This diachronic perspective highlights the cumulative effects of time on linguistic structure.
Cognitive Explanations
Language is fundamentally a product of the human mind, and cognitive explanations focus on how
linguistic structures and patterns are shaped by mental processes. These explanations address
questions about how language is represented, processed, and understood by the brain.
For example, the preference for subject-verb-object (SVO) word order in many languages can be
linked to cognitive efficiency. SVO mirrors the natural way humans conceptualize actions: the agent
(subject) performs an action (verb) on an object. Cognitive explanations also shed light on
phenomena like phonological simplification, where speakers favor easier-to-produce sounds and
sound combinations due to articulatory constraints.
Neurolinguistic research further reveals how specific areas of the brain, such as Broca’s and
Wernicke’s areas, are involved in producing and comprehending language. Understanding these
processes helps explain why certain structures are universal or why some linguistic features, like
recursion, are uniquely human.
Functional Explanations
Languages are tools for communication, and their structures often reflect their social and
communicative functions. Functional explanations examine how linguistic features serve practical
purposes, such as conveying meaning efficiently or maintaining social relationships.
For example, the use of honorifics in languages like Japanese serves a clear social function: it encodes
respect and social hierarchy. Similarly, redundancy in language (e.g., adding plural markers even
when the context is clear) can enhance clarity and reduce the likelihood of misunderstanding in noisy
environments.
Functional explanations also address the evolution of grammatical features. For instance, the
development of tense markers (e.g., “-ed” for past tense in English) helps speakers express temporal
relationships more explicitly, facilitating clearer communication.
Developmental Explanations
Another way to answer "Why?" in linguistics is by considering how language is acquired by children.
Developmental explanations focus on the learnability of linguistic structures, emphasizing how
features are adapted to the cognitive capacities of young learners.
For example, children acquire phonemes and morphemes in a specific order, starting with simpler
sounds like /m/ and /p/ before mastering more complex ones like /r/ or /θ/. This sequence reflects
the natural progression of motor and cognitive development.
The regularity of patterns, such as the conjugation of regular verbs in English (“walk/walked”), makes
them easier for children to learn. Conversely, irregular forms often take longer to acquire but persist
in languages due to their historical roots.
Coincidence
Sometimes, linguistic similarities or patterns arise without any deeper historical, cognitive, or
functional rationale. Coincidence, while seemingly unsatisfactory, is often the most straightforward
and accurate explanation.
For instance, the word for "mother" in many languages includes similar sounds like “mama” (e.g.,
“mama” in English, “maman” in French, “mamá” in Spanish). While this might seem universal, it is
more likely coincidental, as these sounds are among the first that infants can produce. Parents often
adopt these sounds as terms of endearment, creating a superficial resemblance across unrelated
languages.
Similarly, some structural similarities between unrelated languages might arise purely by chance,
rather than from shared ancestry or contact. Recognizing coincidence as a valid explanation helps
avoid overinterpreting patterns that lack a clear causal link.
Linguistics, the scientific study of language, is a discipline of profound intellectual appeal, rooted in
centuries of inquiry and diverse motivations. Its evolution reflects humanity’s enduring fascination
with language as a tool, a historical artifact, a social construct, and a cognitive phenomenon.
Philosophical Foundations
The roots of linguistic inquiry lie in philosophy, where questions about language were deeply
intertwined with broader questions about human existence and thought. Aristotle and other ancient
philosophers laid the groundwork by asking fundamental questions:
During Late Antiquity and the Middle Ages, linguistic study evolved into “philology”, which
emphasized the interpretation of old, often sacred, texts. This period reflected a practical motivation:
understanding ancient languages like Latin, Greek, Sanskrit, or Hebrew was essential for accessing
religious, legal, and literary traditions.
Philology also revealed the historical layering of languages, offering insights into cultural and
intellectual history. For medieval scholars, deciphering a text’s language was a way to unlock its
meaning, message, and moral authority. This interest persisted into the Renaissance, underpinning
the humanist revival of classical learning.
By the late 18th and 19th centuries, language began to be studied as a historical phenomenon.
Historical linguistics emerged as scholars like Jacob Grimm and Franz Bopp sought to reconstruct the
evolution of languages over time.
• Language as a tool for understanding history: Linguists realized that the history of a language
often mirrored the history of the people who spoke it. The reconstruction of Proto-Indo-
European, for example, illuminated ancient migration patterns and cultural exchanges.
• Regularity in language change: The discovery of systematic sound correspondences (e.g.,
Grimm’s Law) demonstrated that language evolved in predictable ways, enabling linguists to
trace connections across time and space.
This era firmly established linguistics as a scientific discipline, with methods grounded in empirical
observation and analysis.
In the early 20th century, linguistics shifted its focus from historical processes to the internal
structure of language and its role in society. This period was shaped by two groundbreaking
approaches:
A transformative shift occurred in 1957 with Noam Chomsky’s introduction of generative linguistics,
which framed language as a biological phenomenon. Chomsky proposed that humans possess an
innate universal grammar, a set of principles hardwired into the brain that governs language
acquisition.
This "cognitive turn" aligned linguistics with psychology and neuroscience, positioning language as a
key to understanding the human mind. Key developments include:
The biological perspective continues to drive research into language acquisition, neurolinguistics,
and the evolutionary origins of language.
While Chomsky emphasized the universality of language, scholars like William Labov shifted attention
to its variability, emphasizing language as a social phenomenon. Sociolinguistics, emerging in the
1960s, examines how language is shaped by and reflects social factors such as class, gender, ethnicity,
and geography.
• Variation and change: Labov’s studies on dialects in urban settings revealed systematic
patterns in how language varies and evolves within communities.
• Context and identity: Sociolinguistics explores how speakers adapt their language to fit social
contexts, signaling identity and group membership.
This approach underscores the dynamic interplay between language and society, highlighting how
linguistic diversity reflects broader social structures.
Today, linguistics balances these diverse approaches, integrating insights from historical linguistics,
structuralism, cognitive science, and sociolinguistics. Contemporary linguistics addresses complex
questions about:
• Multilingualism and language contact: How do languages influence each other in an
increasingly globalized world?
• Technological advances: How can computational linguistics and artificial intelligence model
and process human language?
• Endangered languages: How can linguists document and preserve the world’s linguistic
diversity?
Modern linguistics recognizes that language is at once historical, social, cognitive, and biological—a
uniquely human phenomenon requiring interdisciplinary study.
The study of language operates on two fundamental axes: synchrony, which examines language at a
single point in time, and diachrony, which analyzes language change over time. These dimensions
allow linguists to explore both the static and dynamic aspects of language, encompassing its
phonology, morphology, syntax, semantics, and lexicon, often drawing evidence from literature and
other sources.
Diachrony
• 1500 BCE: Proto-Indo-European had a complex vowel system, including long and short
vowels.
• 500 BCE: Greek had developed a unique set of vowels, while Sanskrit preserved much of the
Proto-Indo-European inventory.
• 1200 CE: the Great Vowel Shift in English was underway, dramatically altering the
pronunciation of long vowels.
• This diachronic change shaped the phonemic distinctions in Modern English by 1948 and
2012.
Synchrony
At any given moment, the phonological system of a language exhibits an internal structure. For
instance, in 1948, Hebrew was being revitalized with a fixed phonological system based on Biblical
Hebrew but influenced by modern spoken languages, while English already exhibited stress-based
prosody and a simplified vowel system compared to Middle English.
Morphological systems, such as noun plurals and verb tenses, often simplify or innovate over time:
• 1500 BCE: Proto-Indo-European used a complex system of inflections to mark noun plurals
and verb tenses.
• 500 BCE: Greek retained rich morphology, but some Indo-European languages, such as Old
Persian, had begun simplifying.
• 2012: English had reduced its inflectional morphology, relying more on syntax (e.g., auxiliary
verbs) than morphology to convey tense. In contrast, languages like Finnish maintained
complex case systems.
Synchrony
Synchronically, morphological systems show patterns of regularity and irregularity. For example, in
1948, Modern Hebrew featured a blend of Biblical morphological roots and new constructions
influenced by European languages, such as regularized plural forms for neologisms.
Diachrony
• 1500 BCE: Proto-Indo-European likely had a relatively free word order, with SOV (subject-
object-verb) being dominant.
• 500 BCE: classical languages like Latin and Ancient Greek had flexible word orders influenced
by morphology, while word order in Chinese had solidified as SVO.
• 1948: English was firmly SVO, while languages like German retained verb-final (SOV)
structures in subordinate clauses.
• 2012: global languages exhibited increasing uniformity in syntax due to language contact and
globalization, but significant variations remained.
Synchrony
At any point, the syntax of a language reflects a set of grammatical rules. For example, in 200 CE,
Classical Latin exhibited highly inflected morphology that allowed flexible word order, while in 1948,
Modern English’s relatively fixed SVO order reflected its reliance on syntax rather than inflection to
convey meaning.
Diachrony
Lexicons expand and transform due to borrowing, innovation, and semantic change:
• 1500 BCE: Proto-Indo-European languages shared a core vocabulary, but contact with
neighboring languages introduced borrowings.
• 1200 CE: Norman French had profoundly influenced Middle English, introducing thousands of
loanwords.
• 2012: English, influenced by technology and globalization, featured numerous neologisms
(e.g., “selfie”) and borrowings from languages worldwide.
Synchrony
At any moment, a language’s lexicon reflects the cultural and social environment. For instance, in
1948, Modern Hebrew incorporated terms for modern concepts, while in 2012, English embraced
internet slang and technical jargon.
Both synchrony and diachrony rely on written and oral sources to reconstruct linguistic states:
• 1500 BCE: Texts like the Rigveda and Linear B inscriptions provide insights into early Indo-
European languages.
• 500 BCE: Greek philosophical texts and early Buddhist scriptures offer detailed records of
phonology, morphology, and syntax.
• 200 CE: Roman literature and early Chinese texts reflect linguistic norms and sociolinguistic
contexts.
• 1200 CE: Middle English literature, such as “The Canterbury Tales”, illustrates ongoing lexical
and grammatical shifts.
• 1948: Newspapers, radio broadcasts, and written works document modern languages'
structure and vocabulary.
• 2012: Digital communication, including social media and online corpora, captures real-time
linguistic change and usage.
Languages are living systems, constantly evolving and transforming as they are shaped by human use,
cultural shifts, and external influences. This process happens at varying rates, from immediate
changes observable within a generation to gradual transformations that unfold over millennia.
Understanding how languages change over time sheds light on why speakers of related languages,
or even distant dialects of the same language, may eventually become mutually unintelligible.
Over the course of hundreds of years, more substantial changes in syntax, grammar, and mutual
intelligibility emerge. A notable example is the evolution of word order in Classical Chinese, which
shifted from Subject-Object-Verb (SOV) to Subject-Verb-Object (SVO) over several centuries. This
gradual reorganization reflects the natural tendencies of speakers to prioritize clarity and efficiency
in communication.
Mutual intelligibility between dialects can also erode significantly over time. For instance,
Appalachian English and British English, despite sharing the same linguistic roots, have diverged to
the point where they can be challenging for speakers from each region to understand one another.
These changes, while slower than generational shifts, still highlight the fluid nature of languages
within a historical timeframe.
Given enough time, the differences between related languages become so pronounced that they are
no longer mutually intelligible. After a thousand years, dialects that originated from a common
"mother tongue" can become entirely distinct languages. The Germanic language family illustrates
this phenomenon: modern English, Dutch, and German all descended from Proto-Germanic, yet they
now differ in vocabulary, syntax, and pronunciation to the point where comprehension across them
is impossible without prior study.
Even more striking are the relationships between English and more distantly related languages like
Pashto, spoken in Afghanistan. Both belong to the Indo-European language family, but their shared
ancestry, which dates back thousands of years, is no longer obvious without linguistic analysis. This
process illustrates how languages evolve to reflect their own unique histories and contexts, leaving
behind their common roots.
Over ten thousand years, languages change so extensively that their historical relationships may
become indistinguishable from chance similarities. The cumulative effects of sound changes,
grammatical shifts, and the replacement of vocabulary obscure linguistic connections. At this stage,
languages with a common origin may appear to be completely unrelated unless deep linguistic
reconstruction techniques reveal the hidden ties.
Several forces drive these linguistic transformations. Internal pressures, such as the simplification of
grammar or natural variations in pronunciation, play a significant role. External factors, including
contact with other languages through trade, conquest, or migration, introduce new elements that
influence language evolution. Random events, like the isolation of a speech community, can also lead
to significant divergence. Together, these forces shape the trajectory of languages, ensuring their
ongoing transformation.
In historical linguistics, determining whether languages are related involves uncovering systematic
correspondences between vocabulary items in different languages. This method relies on the
principle that the connection between sounds and meanings is largely arbitrary. The word for "dog,"
for instance, appears as “dog” in English, “chien” in French, and “gǒu” in Mandarin, showing no
apparent similarity across these unrelated languages. Such differences arise from the independent
evolution of each language and highlight why accidental similarities between unrelated languages
are highly unlikely.
Systematic Correspondences
When two languages are related, systematic sound correspondences can be observed across their
vocabularies. These correspondences suggest a shared linguistic ancestry. For example:
• Romance languages (descended from Latin), we see systematic correspondences for words
like "father": “pater” (Latin), “padre” (Spanish), “père” (French) and “padre” (Italian)
These similarities are not random; they follow consistent patterns that reflect the phonological and
morphological changes the languages underwent after diverging from a common source.
The relationship between sound and meaning in language is arbitrary, meaning there is no inherent
reason why a particular sound sequence should represent a specific concept. This principle, central
to linguistic theory, makes systematic similarities in related languages more compelling because they
are unlikely to arise by chance. For example:
• Indo-European family, we observe similar forms for the word "two": “dvá” (Sanskrit), “duo”
(Latin) and “two” (English)
The recurring pattern (e.g., "d-" or "t-" sounds) across these languages suggests a shared origin rather
than coincidence.
Occasionally, unrelated languages may have words that sound similar and share meanings due to
pure coincidence or borrowing. Historical linguists account for this by:
Historical linguistics, as a field, seeks to uncover the evolution of languages over time using
incomplete and indirect evidence. It combines theoretical frameworks with practical methods to
reconstruct linguistic changes and trace the connections between languages. The ideas and
references provided highlight the tools and challenges of this process, emphasizing its interpretive
and detective-like nature.
• Arm-chair Method: This traditional approach relies on analyzing written records, historical
texts, and inscriptions. It focuses on interpreting documents that have survived from earlier
stages of a language.
• Tape-recorder Method: A more modern method involves collecting data directly from
speakers, particularly in communities that preserve archaic or endangered linguistic features.
This approach, while more common in sociolinguistics, aids historical linguists by capturing
contemporary remnants of past language states.
This statement underscores the idea that historical linguistics is not an exact science. Linguists
work with incomplete or ambiguous data, much like detectives or archaeologists piecing
together clues to form a plausible picture of the linguistic past.
For example, while we lack audio recordings of Middle English or Classical Latin, we can infer
pronunciation and usage through evidence like spelling variations, poetic rhymes, and scribal
errors.
Labov’s remark highlights the inherent limitations of historical linguistic data. Examples of
“bad data” include:
• Spelling errors which often reflect how speakers actually pronounced words.
• Written texts which may be overly formal or stylized, failing to capture colloquial speech.
Despite these challenges, historical linguists extract valuable insights about phonetic shifts,
grammatical changes, and language contact.
This principle emphasizes that no language is ever “complete” or static. All languages are
constantly evolving, making them protolanguages for what they will become in the future.
Foe example, Modern English can be seen as a protolanguage for future linguistic stages, just
as Old English was for Middle English and later stages of the language.
Aitchison (Chapter 2) discusses examples of how historical linguists identify linguistic changes
using indirect evidence. These clues form the foundation for reconstructing earlier stages of
languages:
In Middle English: “wane” for “when”, indicating the loss of the /ʍ/ sound in some dialects.
In Latin: “consul” vs. “cosul”, suggesting a more relaxed pronunciation in informal contexts.
• Puns, Rhymes, and Poetry: Literature offers clues about pronunciation and usage. For
example, in Shakespeare’s “A Midsummer Night’s Dream”, rhymes like “seen” and “queen”
suggest that these words once rhymed before certain vowel changes occurred.
• Representation of Animal Sounds: Imitations of animal noises (e.g., "woof" or "meow") reflect
how speakers perceived certain sounds in their language and how these perceptions changed
over time.
• Social Climbers: Social mobility influences linguistic change, as individuals adopt prestigious
speech patterns to signal their status. For instance, the standardization of English in the British
court led to the decline of many regional dialects.
• Texts as Letters, Sermons, and Homilies: Personal letters, sermons, dialogues, and homilies
are invaluable because they often reflect everyday speech rather than formal or literary
language. These texts provide a more authentic glimpse into how people actually spoke.
Historical linguistics relies on a blend of traditional and modern methods to interpret fragmentary
evidence and reconstruct the evolution of languages. Clues like spelling mistakes, poetic rhymes, and
informal texts enable linguists to fill in the gaps and make sense of incomplete data. As Haas and
others suggest, all languages are in flux, making them “protolanguages” for future changes.
Ultimately, historical linguistics is both a science and an art, requiring creativity, patience, and careful
interpretation to unravel the mysteries of linguistic history.
Main Sources
The study of historical linguistics relies on a variety of sources to reconstruct the evolution of
languages and uncover their relationships. These sources can be broadly divided into two categories:
external and internal. External sources include historical records, archaeological findings, and insights
from present-day linguistic theory, while internal sources focus on written texts and the
reconstruction of spoken language. Together, these sources form the foundation for understanding
the complex processes of linguistic change.
External Sources
External sources provide context for linguistic reconstruction, often offering insights from history,
archaeology, and interdisciplinary studies.
One significant type of external evidence comes from historical and literary studies. Texts from
ancient historians and writers frequently contain descriptions of languages or linguistic practices. For
example, Tacitus´ Germania (96 AD) describes the Germanic tribes and their language, providing a
crucial link between early Germanic dialects and their historical context. Similarly, literary works,
such as Homer’s Iliad, preserve archaic linguistic forms that help linguists trace the evolution of
Ancient Greek.
• Archaeological Remains
Archaeological remains also contribute to the reconstruction of linguistic history. Inscriptions on
runic stones found in Scandinavia and Northern Europe document early Germanic languages and
provide evidence of their phonological and morphological characteristics. Additionally, weapons,
jewelry, and burial inscriptions often contain short phrases, names, or formulas that reflect the
language and culture of their time. Burial sites, in particular, offer indirect evidence of linguistic
contact and spread by linking material culture to specific linguistic communities.
Present-day linguistic theory offers another valuable external source. Modern theories of linguistic
change, language acquisition, and sociolinguistics guide the interpretation of historical data. For
example, theories of regular sound change, such as Grimm’s Law, explain systematic shifts in
phonology within the Germanic language family. Insights from language acquisition studies shed light
on how children learn and replicate language, helping linguists understand how certain changes
might have been transmitted across generations. Similarly, sociolinguistic theories explain how social
dynamics, such as prestige and contact, drive linguistic innovation.
Internal Sources
Internal sources are primarily linguistic and come directly from written texts and the reconstruction
of spoken language.
• Written Texts
Written texts are among the richest sources of linguistic data. From the 7th century onward, many
languages have a written record that provides insights into their historical development. Medieval
manuscripts, such as Beowulf in Old English, preserve the vocabulary, grammar, and syntax of earlier
stages of the language. Legal documents, personal letters, and religious texts, such as Wulfila’s Gothic
Bible, reflect regional and colloquial speech patterns, offering a more dynamic view of the spoken
language of their time. These texts allow linguists to track changes in spelling, vocabulary, and
grammatical structures over centuries.
Beyond written records, linguists reconstruct spoken language at various linguistic levels. Lexical
reconstruction focuses on identifying shared vocabulary among related languages, often tracing
words back to a common proto-language. Phonological reconstruction uses systematic sound
correspondences, such as those between Latin, Sanskrit, and Ancient Greek, to infer earlier
pronunciations. Morphosyntactic reconstruction examines grammatical structures, such as word
order and inflection, to uncover earlier syntactic patterns. Together, these methods allow linguists
to move beyond the written word and hypothesize about the spoken forms of lost languages.
In conclusion the reconstruction of linguistic history requires a multidisciplinary approach, drawing
on both external and internal sources. External sources, such as historical records, archaeological
findings, and modern linguistic theories, provide essential context for understanding the cultural and
social factors influencing language change. Internal sources, particularly written texts, offer direct
linguistic evidence, enabling detailed analysis and reconstruction at lexical, phonological, and
morphosyntactic levels. By combining these sources, historical linguists can piece together the
evolution of languages, uncovering not only their structures but also the processes that shaped them
over time.
Written texts, though indispensable to historical linguists, present numerous challenges and
limitations. They are often incomplete, biased, or unrepresentative of everyday language, posing
difficulties for accurately reconstructing past linguistic states. As Labov (1994) aptly remarked,
historical linguistics is "the art of making the best use of bad data," highlighting the need for creativity
and rigor when working with imperfect sources.
The availability of written texts from the past is shaped by both chance and deliberate actions. Many
texts have been lost due to material decay, wars, or other historical disruptions. Moreover, before
the invention of the printing press (1492), texts had to be laboriously copied by hand, which limited
their reproduction. Scribes and patrons typically chose to preserve documents considered valuable,
such as religious, legal, or literary works, while discarding everyday writings like letters or mundane
records.
This selective preservation means that the texts available today often reflect a narrow slice of
historical linguistic reality, leaving significant gaps in our understanding of past language use.
Informal speech patterns, regional dialects, and non-prestigious varieties are particularly
underrepresented.
The further back we go in history, the fewer texts we have, and the more limited their variety. Earlier
periods are often represented by a small corpus of formal or religious writings, such as:
This lack of diversity restricts the ability of linguists to reconstruct the full spectrum of a language’s
usage, especially in terms of informal speech, regional dialects, and sociolinguistic variation.
• Letters and Dialogues: While more informal than literary or legal documents, these still reflect
stylized writing conventions and rarely capture spontaneous speech.
• Sermons and Homilies: These texts represent rhetorical and performative language,
influenced by their formal and persuasive purposes.
Writing often lags behind speech in reflecting linguistic innovation. For instance, spoken changes in
pronunciation or grammar may not appear in writing until much later. This disconnect complicates
efforts to reconstruct spoken forms based solely on written evidence.
Texts provide positive evidence—what was actually written or said—but they rarely offer negative
evidence, such as forms that were avoided or no longer used. For instance:
• A text may show that a certain word was used in a specific context but cannot reveal whether
an alternative word or grammatical construction was also possible.
• The absence of a linguistic feature in a text does not necessarily mean that it did not exist; it
may simply not have been recorded.
This limitation requires linguists to rely on inference, comparative methods, and external
corroboration to fill in gaps.
The uniformity principle is the assumption that linguistic processes observed today operated similarly
in the past. This principle is crucial because it allows historical linguists to infer past changes based
on modern linguistic theory and patterns. For example:
• Sound changes, such as lenition (weakening of consonants) or vowel shifts, are consistent
across languages and time periods.
• Sociolinguistic dynamics, such as the adoption of prestige forms, can explain historical shifts
in grammar or vocabulary.
An illustrative example of linguistic reconstruction is the sound change observed in words like Spanish
“humo” ("smoke") from Latin “fūmus”. The question arises: did “f > h” (Latin “f” weakening to “h”)
or “h > f” (an earlier “h” strengthening to “f”)?
• The evidence overwhelmingly supports “f > h” due to the consistent weakening of initial stops
in Romance languages.
• This process can be seen across Spanish vocabulary, where Latin “f” regularly weakened,
while in other Romance languages like Italian (“fumo”), the original sound was preserved.
By analyzing systematic patterns like this, historical linguists reconstruct the pathways of sound
change and clarify linguistic relationships.
Written texts, though invaluable, present significant challenges for historical linguistic research. Their
survival depends on a combination of chance and selective preservation, and their content often
reflects formal, stylized language rather than everyday speech. Texts offer only positive evidence,
limiting their utility for understanding linguistic variation or innovation. However, with the help of
principles like uniformity and systematic methods of reconstruction, linguists can navigate these
limitations, making the most of the "bad data" they have to work with.
Classifying languages into families and subgroups requires systematic methods grounded in linguistic
theory. Two foundational principles guide this process: the Uniformitarian Principle and the
regularity of sound change. These concepts allow linguists to identify shared ancestry and trace
linguistic evolution with precision.
In the field of historical linguistics, understanding how languages change over time is essential for
reconstructing their evolution and relationships. A key concept that helps linguists achieve this is the
Uniformitarian Principle, which asserts that the processes of language change observed today have
been operating in similar ways throughout history. This principle, combined with the idea of regular
sound change, provides the framework for classifying languages and tracing their development over
time.
The Uniformitarian Principle can be summarized as the notion that "knowledge of processes that
operated in the past can be inferred by observing ongoing processes in the present." In the context
of language, this means that language must work now in the same way as it did in the past. This idea
is grounded in the belief that language change is a continuous process, governed by consistent
principles that operate in both the present and the past. Linguists rely on the principle to infer how
languages have evolved, often using contemporary linguistic changes as a model to understand
historical shifts.
For example, by observing the sound changes occurring in modern languages, linguists can
hypothesize that similar changes took place in the past. The ongoing shifts in vowel pronunciation in
languages like English can be compared with historical shifts, such as the Great Vowel Shift in Middle
English, which was a series of changes in the pronunciation of vowels. The Uniformitarian Principle
allows linguists to apply this understanding of vowel shifts to other languages and time periods,
facilitating the reconstruction of older forms of a language that no longer exist.
Language is Arbitrary
A fundamental aspect of the Uniformitarian Principle is the idea that language is arbitrary. This means
that there is no inherent or natural connection between the linguistic symbol (a word) and the
concept it represents. For example, there is nothing about the word "dog" that inherently connects
it to the animal it refers to. Similarly, the word "chien" in French or "Hund" in German also refer to
the same animal, yet the sounds of the words are entirely different. This arbitrariness highlights that
language evolves and changes according to social and cultural conventions, rather than being an
intrinsic reflection of the world.
The arbitrariness of language plays a significant role in historical linguistics, as it allows for the
possibility of substantial linguistic change over time. Words and sounds can evolve without any
inherent logic behind their transformations, driven by factors such as social dynamics, language
contact, and internal innovations. The concept of arbitrariness underscores the flexibility of language
and its ability to adapt, shift, and change over generations.
Regularity of Sound-Change
In historical linguistics, the concept of regularity of sound change is a foundational principle. This
principle asserts that sound changes in a language are not random or sporadic but occur in
predictable, systematic patterns. This regularity is essential for reconstructing languages and
understanding their evolution. Linguists rely on the assumption that any sound change will affect all
words that contain the sound or combination of sounds that is undergoing change. This idea has
profound implications for how linguists trace the development of languages and establish
connections between related languages.
The principle of regular sound change means that when a particular sound undergoes a change, it
will affect all instances of that sound in the language. For example, if a language undergoes a
phonological shift in which a specific vowel sound changes to another, that change will apply to all
words that contain that vowel sound. This is true across the language, regardless of the meaning or
position of the word.
For instance, if a language experiences a vowel shift in which a particular vowel sound like “a”
changes to “o” (e.g., “bat” becomes “bot”), the sound change will apply to all words containing that
vowel, such as “cat” becoming “cot” or “mat” becoming “mot”. This regularity enables linguists to
track and predict sound changes over time, allowing them to reconstruct older forms of languages
based on contemporary ones.
The regularity of sound change is crucial for the field of historical linguistics. Linguists use this
principle to compare related languages and trace their evolution from common ancestral languages.
For example, in the case of the Indo-European language family, systematic sound correspondences
between languages such as Latin, Greek, and Sanskrit can be explained by regular sound changes that
occurred over thousands of years. These changes are not random but follow predictable patterns,
which allow linguists to establish relationships between languages and reconstruct their common
ancestors.
A well-known example is the Grimm’s Law, which describes a series of sound changes in the Germanic
languages, where certain consonants in Proto-Indo-European (PIE) shifted in a regular pattern as they
evolved into Germanic languages. For example, the PIE voiceless stops “p, t, k” became “f, th, h” in
Proto-Germanic (e.g., PIE “pater” ("father") became “fader” in Old English). Grimm’s Law provides a
systematic explanation of how sound changes occurred in the Germanic branch of Indo-European
languages and demonstrates the regularity of phonetic evolution.
While the regularity of sound change is a foundational assumption in historical linguistics, there are
cases where sound changes appear to have exceptions. However, these exceptions are typically not
deviations from the regularity of sound change itself but are often the result of language contact,
social factors, or analogy. For example, certain words may resist a sound change because they are
perceived as prestigious or because they have been borrowed from other languages. These cases do
not represent true exceptions to the rule of regular sound change but rather reflect external
influences on the language.
To illustrate, consider the following examples from Old English (OE) to Modern English (ModE):
• OE “cnafa” /knava/ → ModE “knave” /nejv/
• OE “cniht” /knixt/ → ModE “knight” /najt/
At first glance, it seems there might be an irregular pattern or exception in the way these two words
underwent sound changes, particularly concerning the treatment of the initial /k/ sound. So, what’s
the rule?
The key to understanding this lies in the phonological environment in which the sound change occurs.
Both of these words undergo a shift in their initial consonant cluster, but the specific rule that governs
this change applies to particular environments. Specifically, the rule that deletes the initial /k/ sound
in some words (e.g., “cnafa” → “knave”) does not apply universally across all words. In this case, the
deletion of the /k/ only occurs before certain consonants, such as /n/, which is evident in the
transformation of “cniht” /knixt/ into “knight” /najt/.
Now consider the word OE “cyning” /kyniŋ/ (meaning "king"). In Modern English, this word has
evolved to “king” /kɪŋ/. Why does the initial /k/ remain in this case, instead of being deleted as in
other words like “knave” and “knight”? The rule that deletes initial /k/ does not apply here because
it only operates before /n/. This is why OE “cyning” > ModE “king” retains the /k/ sound, unlike OE
“cniht” > ModE “knight”, which undergoes the /k/-deletion.
Thus, the rule of deleting the initial /k/ is still exceptionless, but it is subject to a specific phonological
environment—it only applies when the /k/ is followed by /n/. This pattern is consistent with the idea
that sound change is regular, but it is governed by the surrounding phonological context. In this way,
historical linguistics maintains the principle that sound changes are systematic and predictable, even
though the application of those changes might be constrained by specific phonological conditions.
These examples demonstrate how sound change operates in a regular and predictable manner, but
the exact rules are often conditioned by their phonological environments. While it may appear that
exceptions exist, closer examination usually reveals that these “exceptions” are not true violations
of the regularity of sound change, but rather the result of specific, rule-governed processes in the
language.
If we assume that sound-change is regular and exceptionless in this way, we can use systematic
comparison of languages to see the relationships between them.
The comparative method is a systematic approach used in historical linguistics to reconstruct features
of a common ancestral language (a proto-language) from its daughter languages by systematically
comparing their semantic, phonological, and morpho-syntactic systems. This involves tracing cognate
words—words in different languages that share a common origin in a proto-language.
Steps
1. Establish Genetic Relatedness: By inspection, identify languages that are likely genetically related,
meaning they descend from a common ancestor. This can be suggested by shared vocabulary,
structural similarities, and historical documentation.
2. Compile Word Lists: Collect words with similar meanings from the languages under study. These
words are chosen for their potential to show systematic relationships, such as common vocabulary
items like kinship terms, numerals, or body parts.
3. Identify Systematic Correspondences: Analyze the word lists for recurring patterns in sounds (e.g.,
where one language has /p/, another consistently has /f/). These correspondences hint at regular
phonological changes from the proto-language.
5. Reconstruct Proto-Sounds: For each correspondence set, hypothesize what the original sound in
the proto-language might have been. Consider plausible phonological changes known from language
evolution, such as lenition, fortition, or assimilation.
6. Reconstruct Proto-Words: Using the proto-sounds, reconstruct the likely forms of specific words
in the proto-language. This involves piecing together sounds according to the patterns observed in
the daughter languages.
7. Reconstruct the Proto-Phonological System: Combine the reconstructed sounds and words to infer
the proto-language's phonological system. Identify the inventory of sounds (phonemes) and the
phonotactic rules (how sounds combine) in the ancestral language.
Illustrative Example
- A systematic correspondence might be: “p” (A) ↔ “v” (B) ↔ “p” (C).
Main Aim
To determine whether languages are genetically related and reconstruct features of the ancestral
language by:
Cognate Words
Cognate words, are those words in different languages that are descended from the same word in a
common ancestor.
For example, the word for “father” in several languages: “pater” (Greek), “pitar” (Sanskrit), “pater”
(Latin), “fadar” (Gothic), “padre” (Spanish and Italian), “père” (French), “pai” (Portuguese) and “Pare”
(Catalan). The consistent similarity across these forms suggests they are cognates deriving from the
same proto-language (e.g., Proto-Indo-European “pǝtḗr”).
However, here are some counterexamples of non-cognate words: “athir” (unrelated in origin to
“pater” - Old Irish), “ataataq” (unrelated - Eskimo). These are false cognates since their resemblance
is coincidental and not due to common ancestry.
In 1786, Sir William Jones, a British philologist and judge, famously observed structural and lexical
similarities between Sanskrit, Greek, Latin, and other languages. His observations laid the
groundwork for the idea of the Indo-European language family and the use of the comparative
method. For example, he observed that the parallels in words for “father” (above) are a classic
example of the similarities that convinced Jones and subsequent linguists of the genetic connection
among these languages.
Outcomes
Through the comparative method, linguists uncover the deep history of languages and their
evolution.
EXTRA 5.8 NOTES OF CAUTION
This highlights a critical caution in the comparative method in historical linguistics: the need to
distinguish between native words and borrowed words (loanwords) to avoid incorrect conclusions
about the genetic relationship between languages.
Many languages borrow words from others due to cultural contact, trade, migration, etc. Some
loanwords have been part of a language for so long that they appear native, but they are not.
The English word “coffee” and the Mandarin word “kafe” sound similar and mean the same thing. At
first glance, one might (wrongly) conclude that this indicates a genetic relationship between English
and Mandarin.
However, the English word “coffee” is not native. It originates from the Arabic word “qahwa”, which
was borrowed into Turkish (“kahve”), then Dutch (“koffie”), and finally into English. Similarly, the
Mandarin word “kafe” is a modern loanword derived from English or another European language.
Thus, the similarity between these words does not reflect a genetic relationship but rather the fact
that both languages borrowed the word from a common source.
• Avoiding False Conclusions: Do not assume that words that sound similar in two languages
are necessarily cognates. Investigate whether the word is part of the native vocabulary or a
recent loan.
• Considering Historical and Cultural Context: Analyze historical factors (e.g., trade routes,
colonization, cultural exchanges) that could explain linguistic borrowing.
• Distinguishing Cognates from Loanwords: Cognates are words that derive from a common
ancestor in genetically related languages. Loanwords are adopted from another language
without indicating a genetic connection.
In conclusion, the example of “coffee” and “kafe” illustrates how loanwords can superficially
resemble cognates but do not indicate a genetic relationship between languages. A careless analysis
might lead one to conclude, for instance, that English and Mandarin are related, when in fact the
similarity is due to cultural and linguistic contact. This is why the comparative method requires careful
analysis supported by multiple lines of evidence.
Chance Resemblances
This issue refers to the possibility of chance resemblances between words in unrelated languages.
While such coincidences are rare, they can occur and must be carefully distinguished from genuine
genetic relationships or borrowings.
This concept refers to the words in different, unrelated languages may resemble each other in form
and meaning purely by coincidence, without sharing a common origin. This happens because
languages have a finite set of sounds and limited ways to combine them, increasing the likelihood of
accidental overlaps.
Some examples:
• English “bad” vs. Persian “bad” ("bad" in both languages): Despite their identical meaning and
similar pronunciation, these words are not cognates. English “bad” comes from Germanic
roots, while Persian “bad” is from Indo-Iranian roots.
• Dutch “elkaar” ("each other") vs. Basque “elkar” ("each other"): The similarity in meaning and
form is coincidental. Dutch is Indo-European, and Basque is a language isolate with no known
relatives.
• English “man” / Korean “man”: No connection; meanings are not even the same.
• German “nass” ("wet") / Zuni “nas”: Resemblance is coincidental.
• Italian “donna” ("woman") / Japanese “onna”: Pure chance, as Italian is Indo-European and
Japanese is from a completely different linguistic family.
• English “black” / Berber “bak’l”: The resemblance is accidental.
1. Analyze Broader Vocabulary: If two languages were related, we would expect to find systematic
similarities across many words, not just isolated resemblances. The examples above fail this test—
further comparison of the languages’ vocabularies reveals no broader connections.
2. Check Historical Linguistic Records: Study the history of sound changes, etymologies, and
borrowings to confirm whether the resemblance is genuine or coincidental.
3. Contextualize Within Language Families: For related languages, resemblances occur in predictable
ways, following regular sound correspondences. Random similarities do not follow such patterns.
In conclucion, chance resemblances, while uncommon, highlight the importance of rigorous linguistic
analysis. Apparent similarities in isolated words (like “bad” in English and Persian or “elkaar” in Dutch
and Basque) do not necessarily indicate genetic relationships. They serve as a reminder that
superficial similarities can be misleading, and a deeper examination of linguistic systems is crucial to
determine actual connections.
5.9 INTERNAL METHOD
The internal reconstruction method is a technique in historical linguistics used to study the history of
a single language, particularly when there are no closely related languages available for comparison.
It relies on patterns and irregularities within the language itself to hypothesize about earlier stages
of that language.
1. Linguistic Probability: Assumes that irregularities in a language today were likely regular patterns
in the past. Languages tend to evolve through regular sound changes, though exceptions arise over
time due to analogical changes, borrowing, or other processes.
2. Uniformity Principle: Assumes that the linguistic mechanisms observed today (e.g., sound changes,
analogy) operated similarly in the past.
Steps
1. Identify a Pattern in the Language: Look for recurring linguistic forms or structures in the present-
day language. Ex: Regular verb conjugations in English (e.g., “walk → walked”).
2. Note Exceptions to the Pattern: Find irregular forms that deviate from the pattern. Ex: Irregular
past-tense verbs like “go → went” or “be → was/were”.
3. Hypothesize Original Regularity: Assume that the exceptions originally followed the regular
pattern. Hypothesize that these irregularities arose due to later changes.
4. Reconstruct the Ancestral Stage: Propose an earlier stage of the language where all forms followed
the regular pattern. Ex: Proto-forms of irregular English verbs might have had regular past-tense
formations that were later altered by analogy, sound changes, or borrowing.
5. Identify Disruptive Changes:Determine the historical processes that caused the deviations from
regularity. Ex: In English, “go → went” is irregular because “went” originally belonged to a different
verb (“wend”), which later replaced the past tense of “go”.
• Self-contained: Can be used when no related languages are available for comparison.
• Insightful: Reveals how irregularities develop over time, showing the dynamic nature of
language change.
Example in Action
In Latin:
- Regular nouns: “rosa” (nom.), “rosae” (gen.)—a simple vowel change marks the case.
Originally, “domus” may have had a distinct genitive form that followed the regular vowel-alternation
pattern but became irregular due to phonological leveling or sound change.
The ancestral stage likely had a distinct genitive (“domos”), which was later lost or merged with the
nominative form.
The study of historical linguistics relies on reconstructing languages' past forms to understand their
origins and evolution. For English, this journey begins with its deep connections to the Proto-Indo-
European (PIE) language, extends through its development in Old and Middle English, and is
documented through various forms of textual evidence. This essay examines the protolexicon of PIE,
the proposed homeland of its speakers, and the significant texts and linguistic materials that trace
the evolution of English.
The protolexicon refers to the reconstructed vocabulary of PIE, the hypothesized ancestor of many
modern languages, including English. By examining the meanings of PIE words, linguists can infer
details about the homeland and culture of its speakers. For example, PIE has terms for seasonal
phenomena such as “winter”, “snow”, and “rain”, which suggest a temperate climate with distinct
seasons. Additionally, PIE vocabulary includes names for trees such as “oak”, “birch”, and “ash”—all
of which grow in temperate regions—but lacks terms for Mediterranean and Asian flora, such as
“olive” and “palm”.
The reconstructed lexicon also reflects an agricultural and pastoral society. PIE speakers had words
for “cattle, sheep, horse, dog,” and “pig”, as well as for farming tools like the “plough”. Familiarity
with rivers and streams is evident through words for “ship” and “salmon” (“lax”), but there are no
terms related to the sea, suggesting a landlocked or semi-inland region. These clues have led linguists
to place the PIE homeland between northern Europe and southern Russia, in the area between the
Vistula and Elbe rivers.
As English evolved from PIE through Proto-Germanic, evidence of its development appears in runic
inscriptions written in the Futhark alphabet. These inscriptions, dating from the 2nd to 12th centuries
CE, provide some of the earliest records of Germanic languages. They offer insights into vocabulary,
phonology, and writing systems before the advent of Old English.
By the time English had developed into its Anglo-Saxon form, extensive written materials began to
appear. Glossaries pairing Latin words with Anglo-Saxon equivalents reflect the bilingual context of
early medieval England and provide valuable data on specialized vocabulary and phonology. For
example, glossed Latin texts—Latin works with English transliterations or notes—help reveal not only
word meanings but also phonological and morphological features of Old English.
Old English reached its literary peak in the 8th to 10th centuries, producing both prose and poetry.
One of the most significant prose works is the Anglo-Saxon Chronicle, a historical account that spans
events from the departure of the Romans (~60 BCE) to decades after the Norman Conquest (1066
CE). This text, commissioned during the reign of King Alfred, is a crucial source for understanding Old
English syntax, vocabulary, and historical context.
In poetry, the epic Beowulf stands out as a masterpiece of Old English literature. Composed between
the 8th and 10th centuries, it combines complex poetic structures with rich vocabulary, offering
insights into the cultural and linguistic landscape of early England.
Another notable Middle English work is the Aȝenbite of Inwit, written in 1340. Translated as The
Remorse of Conscience, this Kentish prose text illustrates regional variations in Middle English,
particularly in vocabulary and grammar. Such texts show how English diversified during this period,
influenced by contact with Norman French and Latin.
Beyond literary works, the history of English is also documented through practical tools such as
grammar books, dictionaries, and orthographies, which became prominent from the 16th century
onward. These resources helped formalize English and chart its evolution into the modern period.
Additionally, the Bible, continuously translated and adapted throughout English history, serves as a
linguistic touchstone that reflects changes in vocabulary, syntax, and style across centuries.