Introduction to Machine
Translation
RANIL M. MONTARIL, MSECE
What is Machine Translation
Machine Translation / Automatic Translation
➢ It is the translation from one natural language
(source language (SL)) to another language (target
language (TL)) using computerized systems and, with
or without human assistance.
The development of machine translation through
four points are: Computational
Linguistics
Artificial
1. Surveys the chronological development of Intelligence
NLP
Application
machine translation, Oriented
Machine
Translation
Linguistics
2. The different approaches developed (linguistic
and computational),
3. The types of machine translation
4. We try to answer an important question which is
how to evaluate a machine translation?
Why MT Matters
This field is one of the oldest applications of computers. Over the years, Machine Translation has
been a focus of investigations by linguists, psychologists, philosophers, computer scientists and
engineers.
A. Social
B. Political
C. Commercial
D. Scientific
E. Intellectual
F. Philosophical
Social/Political
❑MT arises from the socio-political importance of translation in communities where more than
one language is generally spoken. Here the only viable alternative to rather widespread use of
translation is the adoption of a single common ‘lingua franca’, which (despite what one might
first think) is not a particularly attractive alternative, because it involves the dominance of the
chosen language, to the disadvantage of speakers of the other languages, and raises the
prospect of the other languages becoming second-class, and ultimately disappearing.
❑Since the loss of a language often involves the disappearance of a distinctive culture, and a way
of thinking, this is a loss that should matter to everyone. So, translation is necessary for
communication — for ordinary human interaction, and for gathering the information one needs
to play a full part in society.
Commercial
MT is a result of related factors.
❑First, translation itself is commercially important: faced with a choice between a product with
an instruction manual in English, and one whose manual is written in Japanese, most English
speakers will buy the former — and in the case of a repair manual for a piece of manufacturing
machinery or the manual for a safety critical system, this is not just a matter of taste.
❑Secondly, translation is expensive. Translation is a highly skilled job, requiring much more than
mere knowledge of a number of languages, and in some countries at least, translators’ salaries
are comparable to other highly trained professionals
❑Moreover, delays in translation are costly. Estimates vary, but producing high quality
translations of difficult material, a professional translator may average no more than about 4-6
pages of translation (perhaps 2000 words) per day, and it is quite easy for delays in translating
product documentation to erode the market lead time of a new product.
Scientific
❑MT is interesting, because it is an obvious application and testing ground for many ideas in
Computer Science, Artificial Intelligence, and Linguistics, and some of the most important
developments in these fields have begun in MT.
❑To illustrate this: the origins of Prolog, the first widely available logic programming language,
which formed a key part of the Japanese ‘Fifth Generation’ programme of research in the late
1980s, can be found in the ‘Q-Systems’ language, originally developed for MT.
Philosophical
❑MT is interesting, because it represents an attempt to automate an activity that can require the
full range of human knowledge — that is, for any piece of human knowledge, it is possible to
think of a context where the knowledge is required.
❑For example, getting the correct translation of negatively charged electrons and protons into
French depends on knowing that protons are positively charged, so the interpretation cannot be
something like “negatively charged electrons and negatively charged protons”. In this sense, the
extent to which one can automate translation is an indication of the extent to which one can
automate ‘thinking’.
Popular Misconceptions about MT
Popular Conception of MT
History of MT
Although we may trace the origins of machine translation (MT) back to seventeenth century
ideas of universal (and philosophical) languages and of ‘mechanical’ dictionaries, it was not until
the twentieth century that the first practical suggestions could be made. The history of machine
translation can be divided into five (05) periods
First period (1948-1960): The beginning.
❑1949 : Warren Weaver in his Memorandum of 1949 proposed the first ideas
on the use of computers in translation, by adopting the term computer
translation.
❑1952 : The first symposium of machine translation, entitled Conference on
Machine Translation, held in July 1952 at MIT under leadership of Yehoshua Warren Weaver
Bar-Hillel. American scientist and mathematician
❑1954 : The development of the first automatic translator (very basic) by a
group of researchers from Georgetown University in collaboration with IBM,
which translates into more than sixty (60) Russian sentences into English. The
authors claimed that within three to five years, machine translation would not
be a problem.
❑1954 : Victor Yngve published the first journal on MT, entitled « Mechanical
translation devoted to the translation of languages by the aid of machines ».
Victor Huse Yngve was a professor of
linguistics at the University of Chicago and
the Massachusetts Institute of Technology
Second Period (1960-1966) Parsing and
disillusionment
❑Early 1960s This parsing is put forward as the only possible avenue of research
to advance the machine translation. Thus, there are already many parsers
developed from different types of grammars, such as grammar and dependency
grammar Tesnière stratificationnelle Lamb
❑ 1961 : In February of this year that computational linguistics is born, thanks to
weekly lectures organized by David G. Hays at the Rand Corporation in Los
Angeles. These conferences will be included as papers at the First International
Conference on Machine Translation of Languages and Applied Language Analysis Lucien Tesnière
of Teddington in September 1961 with the participation of linguists and influential French linguist.
computer scientists involved in the translation as: Paul Garvin, Sydney M. Lamb,
Kenneth E. Harper, Charles Hockett, Martin Kay and Bernard Vauquois.
❑1964 : the creation of committee ALPAC(Automatic Language Processing
Advisory Committee) with American government to studies the perspectives and
the chances of machine translation • 1966 : ALPAC published his famous rapport
in which it concluded that its works on machine translation is just wasting of
time and money ; the conclusion of this rapport is it had a negative impact on
their search (MT) for a number of years
Third period (1966-1980): New birth and
hope
❑1970 : Start of the project REVERSO by a group of Russian researchers.
❑1970 : Development of System SYSTRAN1 (Russian-English) by Peter Toma,
who was at that time a member of a group search for Georgetown.
❑1976 : Creation of system WEATHER in the project TAUM (machine
translation in the university of Montreal) under the direction of Alai
Colmerauer for the machine translation weather forecasts for the general
public, this system was created by group of researchers
❑1978 : Creation of system ATLAS2 by the Japanese firm FUJITSU, this
translator was based on rules also he is able to translate from Korean to
Japanese and vice versa
Fourth Period (1980-1990): Japanese
invaders
❑1982 : The Japanese firm SHARP markets its Automatic translator DUET (English
- Japanese), this translator was based on rules an approach to translation
transfer
❑1983: as computer giant, NEC develops it’s own system of translation based on
algorithm called PIVOT. Marketed under the name of Honyaku Adaptor II, the
version public the system of translation of NEC is also based on the method of
pivot, by using Interlingua.
❑1986: Development of system PENSEE by OKI3, which is a translator (Japanese-
English) based on rules.
❑1986: The group Hitachi developed his own translation system based on rules
(which is an approach taken by transfer), christened on HICATS (Hitachi
Computer Aided Translation System / Japanese- English).
Fifth Period (since 1990): the Web and
the new vague of translators
❑1993: The project C-STAR (Consortium for Speech Translation Advanced Research) is an
international cooperation. The theme of project is the machine translation of the parole in the
field of tourism (dialogue client travel agent), by videoconference. these project birth the system
C-STAR I which dealt three (03) languages (English, German et Japanese) and made the first
demonstrations transatlantic trilingual in January 1993
❑1998: Marketing the translator REVERSO by the company Softissimo.
❑2000: the Development of system ALPH by Japanese laboratory ATR, this translator (Japanese-
English and Chinese - English) takes an approach based on examples.
Fifth Period (since 1990): the Web and
the new vague of translators- cont.
❑2005: The appearance of the first web site for automatic translation ,like
Google ([Link]
❑2007: METIS-II is a hybrid machine translation system, in which insights
from Statistical, Example based, and Rule-based Machine Translation (SMT,
EBMT, and RBMT respectively) are used.
❑2008 : 23% of internet users, have used the machine translation and 40 %
considering doing so
❑2009: 30% the professionals have used the machine translation and 18%
perform a proofreading.
❑2010: 28% of internet users, have used the machine translation and 50%
planning to do.
Architectures of machine translation
systems
Linguistic Architecture
In the linguistic architecture there are three basic approaches being used for developing MT
systems that differ in their complexity and sophistication. These approaches are:
Bernard Vauquois
a French mathematician and computer
scientist. He was a pioneer of computer
The Vauquois triangle science and machine translation (MT) in
France.
Computational Architecture
1. Rule Based approach:
2. Corpus-based approach:
3. Hybride approach:
Rule Based Approach
Rule-based MT has two approaches: Interlingua and transfer. Rule-Based MT Systems rely on
different levels of linguistic rules for translation.
This MT research paradigm has been named rule-based MT due to the use of linguistic rules of
diverse natures.
For instance, rules are used for lexical transfer, morphology, syntactic analysis, syntactic
generation, etc. In RBMT the translation process consists of:
1. Analyzing input text morphologically, syntactically and semantically.
2. Generating text via structural conversions based on internal structures.
Corpus Based Machine Translation
Corpus-Based Machine Translation, also referred as data driven machine translation, is an
alternative approach for machine translation to overcome the knowledge acquisition problem of
rule-based machine translation.
There are two types of CBMT:
1. Statistical Machine Translation (SMT)
2. Example-Based Machine Translation (EBMT).
Corpus based MT automatically acquires the translation knowledge or models from bilingual
corpora. Since this approach has been designed to work on large sizes of data, it has been
named Corpus-Based MT
Hybride Approach
Some recent work has focused on hybrid approaches that combine the transfer approach with
one of the corpus–based approaches.
This was designed to work with fewer amounts of resources and depend on the learning and
training of transfer rules.
The main idea in this approach is to automatically learn syntactic transfer rules from limited
amounts of word aligned data.
This data contains all the needed information for parsing, transfer, and generation of the
sentences
Types of Machine Translation
1. Machine Translation for Watcher (MT-W)
2. Machine Translation for Revisers (MT-R)
3. Machine Translation for Translators (MT-T)
4. Machine Translation for Authors (MT-A)
Machine Translation for Watcher (MT-W)
This is intended for readers who wanted to gain access to some information written in foreign
language who are also prepared to accept possible bad translation rather than nothing.
This was the type of MT envisaged by the pioneers.
This came in with the need to translate military technological documents.
This was almost the dictionary based translation far away from linguistic based machine
translation.
Machine Translation for Revisers (MT-R)
This type aims at producing raw translation automatically with a quality
comparable to that of the first drafts produced by human.
The translation output can be considered only as brush-up so that the
professional translator freed from that very boring an time consuming task can
be promoted to revisers
Machine Translation for Translators (MT-T)
This aims at helping human translators do their job by providing on-line dictionaries, thesaurus
and translation memory.
This type of machine translation system is usually incorporated into the translation work stations
and the PC based translation tools. “Tools for individual translators have been available since the
beginning of office automation.”
And those systems running on standard platforms and integrated with several text processors
are the ones that attained operational and commercial success
Machine Translation for Authors (MT-A)
This aims at authors wanting to have their texts translated into one or several languages and
accepting to write under control of the system or to help the system disambiguate the utterance
so that satisfactory translation can be obtained without any revision.
This is an “interactive MT, The interaction was however done both during analysis and during
transfer, and not by authors, but by specialists of the system and language(s).”
In short, there have been no operational successes yet in MT-A, but the designs are becoming
increasingly user oriented and geared towards the right kind of potential users, people users,
people needing to produce translations, preferably into several languages
Evaluation of Machine Translation
Systems
1. BLEU (BiLingual Evaluation Understudy)
2. WER (Word Error Rate)
3. PER (Position-independent word Error Rate)
4. TER (Translation Error Rate)
BLEU (BiLingual Evaluation Understudy)
The BLEU metric, proposed by Papineni in 2001 was the first automatic measurement accepted as a
reference for the evaluation of translations. The principle of this method is to calculate the degree of
similarity between candidate (machine) translation and one or more reference translations based on
the particular n-gram precision.
The BLEU score is defined by the following formula:
Where:
• “pn”: the number of n-grams of machine translation is also present in one or more reference
translation, divided by the number of total n-grams of machine translation.
• “wi”: positive weights.
BLEU (BiLingual Evaluation Understudy)
“BP”: Brevity Penalty, which penalizes translations for being “too short". The brevity penalty is
computed over the entire corpus and was chosen to be a decaying exponential in “r/c”, where
“c” is the length of the candidate translation and “r” is the effective length of the reference
translation.
WER (Word Error Rate)
The WER metric, Proposed by Popovic and Ney in 2007. Originally used in Automatic Speech
Recognition, compares a sentence hypothesis refers to a sentence based on the Levenshtein
distance.
It is also used in machine translation to evaluate the quality of a translation hypothesis in
relation to a reference translation.
For this, the idea is to calculate the minimum number of edits (insertion, deletion or substitution
of the word) to be performed on hypothesis translation to make it identical to the reference
translation. The number of editss to be performed, noted “dL(ref, hyp)” is then divided by the
size of the reference translation, denoted “Nref” as shown in the following formula
WER (Word Error Rate)
Where:
• dL(ref, hyp): is the Levenshtein distance between the reference translation “ref” and the
hypothesis tanslation “hyp”.
A shortcoming of the WER is the fact that it does not allow reordering of words, whereas the
word order of the hypothesis can be different from word order of the reference even though it is
correct translation.
PER (Position-independent word Error
Rate)
The PER metric, proposed by Tillman in 1997. compare the words of machine translation with those
of the reference regardless of their sequence in the sentence.
The PER score is defined by the following formula:
Where:
• dper: calculates the difference between the occurrences of words in machine translation and the
translation of reference.
A shortcoming of the PER is the fact that the word order can be important in some cases.
TER (Translation Error Rate)
The TER metric, proposed by Snover in 2006. Is defined as the minimum number of edits needed to
change a hypothesis so that it exactly matches one of the references. The possible edits in TER include
insertion, deletion, and substitution of single words, and an edit which moves sequences of
contiguous words. Normalized by the average length of the references. Since we are concerned with
the minimum number of edits needed to modify the hypothesis, we only measure the number of
edits to the closest reference.
The TER score is defined by the following formula:
Where:
• Nb (op) : is the minimum number of edits;
• Avreg Nref: the average size in words references.
???