See discussions, stats, and author profiles for this publication at: [Link]
net/publication/281857330
Improving the quality of Machine Translation using Rule Based Tense
Synthesizer for Hindi
Conference Paper · June 2015
CITATIONS READS
2 388
4 authors:
Shashi Pal Singh Ajai Kumar
Centre for Development of Advanced Computing Centre for Development of Advanced Computing
51 PUBLICATIONS 474 CITATIONS 79 PUBLICATIONS 885 CITATIONS
SEE PROFILE SEE PROFILE
Hemant Darbari Anshika Gupta
Centre for Development of Advanced Computing Banasthali University
75 PUBLICATIONS 966 CITATIONS 1 PUBLICATION 2 CITATIONS
SEE PROFILE SEE PROFILE
All content following this page was uploaded by Shashi Pal Singh on 18 September 2015.
The user has requested enhancement of the downloaded file.
Improving the quality of Machine Translation using
Rule Based Tense Synthesizer for Hindi
Shashi Pal Singh*1, Ajai Kumar*2, Dr. Hemant Darbari*3, Anshika Gupta#4,
*
AAI, Center for development of Advanced Computing, Pune, India,
1
shashis@[Link]
2
ajai@[Link]
3
darbari@[Link]
#
Banasthali Vidyapith, Banasthali, India
4
[Link]@[Link]
Abstract—Translation of English documents into Hindi language should be ȡ °ȯ (char ladke); noun-adjective agreement ex:
is becoming an integral part for facilitating communication. The
major population of India where 366 million people uses Hindi as Ȫȡ °ȧ (gauraa ladki) is incorrect, here °ȧ (ladki) is
primary language, for them providing information in Hindi is an noun, gender is female and person is 3rd, accordingly the
important task the translation of English documents may be adjective used should be Ȫȣ (gauri). Thus, the correct
manually or automatically. When translation done manually the
chances are rare for errors, but when we use any translation sentence should be Ȫȣ °ȧ (gauri ladki).
machines or engine the outputs received from these machines for There are some errors which occur due to phrasal
Hindi have lots of grammatical mistakes or errors. One of the
major issues observed is of tense.
ambiguity like ȡ ȯȣ ]ȱɉ ȡ ȡȡ ¡Ȱ (ram meri ankho ka
taaraa hai); here ]ȱɉ ȡ ȡȡ (ankho ka taaraa) is a phrase
Thus, we propose solution to build a rule based tense synthesizer which should be considered as a single token, if it is used for
that would recognise the subject, verb and auxiliary verb,
analyse the tense, then modify the verb and auxiliary verb
any further processing. For example: if we use the above
according to the subject and put the sentence in the correct tense. phrase in reference with the female like Ȣȡ ȯȣ ]ȱɉ ȡ ȡȡ
This system could be integrated with Machine Translation ¡Ȱ (sita meri ankho ka taaraa hai), it should be used the same
engines to boost up the quality of Hindi translation.
way, it should not be modified according to the nouns like
Keywords— Word sense disambiguation, ontology, phrasal Ȣȡ ȯȣ ]ȱɉ ȧ ȡȣ ¡Ȱ (sita meri ankho ka taari hai), that
ambiguity, POS Tag, Tokenization, lemmatize, inflectional, would be incorrect. Another type of ambiguity is Phrasal
tense markers, translation, hybrid morphological analyser, verbs. Consider the following examples:
ambiguity. They brought up the child in luxury. (ȡȡ) (palna)
I. INTRODUCTION They brought up the table to the first floor. (a ȡȡ)
India is a diverse nation with varied regional languages. In (upar lana)
order to communicate with people that follow different They brought up the issue in the house. (Ǖȡ `ȡȡ)
language, we need to know their languages. But knowing so (muddha oothana)
many different languages is not practical [1]. Thus, machine Here, different translations of brought up need to be used
translation engines were designed. Anglabharti (anubharati), according to the context.
Anusaaraka, MaTra, Sampark, Anuvadak, are designed to get
translation for Indian languages [2]. Translations have 4 major Another kind or type is structural ambiguity
aspects morphology, Lexicalization, syntax and semantics. Flying Planes can be dangerous.
Direct translators focus more on morphology and syntax ¡ȡ_ ¡ȡ« `°ȡ ȡ ¡Ȫ ȡ ¡Ȱ
whereas indirect translators focus on semantics along with the
morphology and syntax [3, 4]. havaee jahaj udna khatrnak ho sakta hai
In this paper we shall focus on improving the syntactic part `°ȡ ¡Ǖ] ¡ȡ_ ¡ȡ« ȡ ¡Ȫ ȡ ¡Ȱ
or structural part of the translation based on tenses. We see Udta huaa havaee jahaj khatrnak ho sakta hai
that the Hindi translations outputs majorly suffer from a lot of Here, more than one translation could be produced depending
grammatical errors which includes the following. on the part of speech tagging for the words [5].
Noun-Verb Agreement with Number example: ǒȡ k Case markers play an important role in determining the
correct parts of speech. As, Hindi is a free word-order
Ĥȡ ȡ_ ¡Ȱ (vikas aur pratap bhai hai) is incorrect it would be language, part of sentences can be moved around without
ǒȡ k Ĥȡ ȡ_ ¡ɇ (vikas aur pratap bhai hain), Noun- making difference in the core meaning provided that case
markers are correctly handled. For e.g.
number agreement ex: ȡ °ȡ (char ladka) is incorrect it
“ȡ Ȫ ȡ ȯ ȡȡ”
978-1-4799-8047-5/15/$31.00 2015
c IEEE 415
Ravan ko ram ne mara It is also found that MT does not translate non-standard
“ȡ ȯ ȡ Ȫ ȡȡ” language with the same accuracy as the standard languages.
Ram ne ravan ko mara Also, there occurs ontology because of the lack of world
knowledge. For example, a machine faces ambiguity when
The identification of ȡ (ram) as a subject and ȡ (ravan) translating sentences like “I saw a man/star/molecule with
as an object comes from the case markers ȯ (ne) and Ȫ (ko) microscope/telescope/binoculars”.
respectively. Therefore case markers along with the
morphology play important role in determining the correct
translation. III. PROPOSED SYSTEM
Now the most important thing which includes parts of all This paper focuses on the first aspect i.e. Grammar. We
the aspects mentioned above is Tenses. The subject, its gender, propose to build a rule based Tense Synthesizer for correcting
number, person should go according to the verb and the the Hindi received from machine translation engines. Rules
auxiliary verb present in the sentence [8]. Consider a very were designed based on the analysis done on the outputs
simple sentence of future perfect. I will have sung. The received from translators like Google, TDIL etc. along with
translation received from [7] is “ɇ ȡ ¡Ȱ ȡfȡ” (main gaya the grammar of English and Hindi as in [16-18].
The system consists of hybrid morphological analyser and
hai jaega), which is incorrect. However, it should be “ɇ ȡ rule based tense analyser and corrector.
Ǖȡ ¡Ȫaȱȡ” (main gaa chuka hounga) or “ɇ ȡ Ǖȡ ¡ȪaȱȢ” The hybrid morphological analyser is combination of rules
(main gaa chukka houngi). based, corpus based and Hidden Markov model tagging
process which gives more accurate result then simple
II. LITERATURE REVIEW morphological analyser.
Today, a tonne of data have been digitalized and is present The machine translated Hindi is given as an input to the
on cloud. But this data is available in different languages. To system. The pre-processing part of the system consist of
have access to these data, we need to translate them into the various functionality such as cleaning of the input sentence,
language that we know. Translation and its evaluation are very tokenization of the sentence, identifying the type of sentences
time consuming, expensive, subjective and highly prejudiced it belongs, phrase identification, clause boundary
task.[9],[10],[5] Also it requires people who have expertise in identification. Then tokenized or phrased are passes to hybrid
different languages. Thus, there was a need for finding a better morphology analyser which give better and accurate POS
way of analysing and utilizing large amount of this data. tagging.
This problem was solved by the development of machine The hybrid morphological analyser uses the rules and the
translation engines, which could computerize this task of lexical knowledgebase to recognise the part of speech, number,
conversion of data from one language to another. This data person and gender for each token. For example
could then reach the audience that needs it and could be used
for the betterment of the society [11]. ǒȡ \Íȡ °ȡ ¡Ǘȱ
Vikas achha ladka ho
Translation engines could be bilingual or multilingual. ǒȡ /N-M-S \Íȡ /JJ °ȡ/NN ¡ȱǗ /VAUX {apply rules}
They could have followed direct approach or indirect
ǒȡ \Íȡ °ȡ ¡Ȱ
(interlingua, transfer) approach. At present various MT
engines based on Rule, Statistical, Example, Dictionary and Vikas achha ladka hai
hybrid approaches have been designed for translation of text
from one language to another [4]. It then passes for tense correction to the tense analyser and
corrector. It recognises the tense and corrects it by modifying
Following are the issues that still persist in the existing MT the verb (lemmatization + tense markers) and auxiliary verb
Engines. The translated text contains errors relating to Part of according to number, gender of the subject. For example some
Speech (POS) Tagging, lemmatization, noun- of the rules we apply on tenses are:
number/adjective agreement, tenses, case markers, phrase
ambiguity etc. [12, 13]. Present Continuous tense {rules}:
In some cases, when there occurs ambiguous words, which Verb (root) + ¡ȡ(raha) /¡ȯ (rahe) /¡ȣ(rahi) + Present
have more than one meaning, then recognising the correct Tense
sense of the word in reference to the context and assigning the Example: ɇ ǒ + ¡ȡ+ ¡ȲǕ
POS tag becomes difficult. e.g. ¡ȡȱ f ] ]Ȣ ¡Ȱ ( vhan Main pi+ raha+ hoon
ek aam aadmi hai), here the word ] (aam) becomes Past tense {rules}:
ambiguous as it has two meanings, first it can be the Verb (root) + ʜȡ/ʜȯ/ʜȢ
noun(name of a fruit) and second, it can be the [Link] Example: ɇ ǒ+ȡ
comes under Word Sense Disambiguation [14, 15].
Main pi+ ya
Future Continuous tense {rules}:
416 2015 IEEE International Advance Computing Conference (IACC)
Verb (root) + ȡ (ta) / ȯ (te) / Ȣ (ti) + Ǿȱȡ (karunga) /ȯ ȡ
B. Algorithm
(karega) / Ʌ ȯ (karange)/ Ȫȯ (karoge)
Algorithm for Rule Based Tense Synthesizer
Example: ɇ ǒȡ ¡Ǘȱȡ Step1: Take Hindi sentence as input.
Main pita rahunga Step2: moves sentence to Pre-Processing module
The system is built on 190 rules and has achieved an Step3: Clean the input sentence
accuracy of 70-75%. Step4: Tokenize the sentence.
An effort to develop a prototype has been developed, which Step5: Identify the Phrase in given sentence (Adjective,
could be used with the MT engines to correct and synthesise Adverb and Noun)
the grammatical errors and boost up the quality of Hindi Step6: Identify the Clause in given sentence (Adverb
Translation received from various MT engines. clauses, Adjective Clauses, Noun clauses)
Step7: Hybrid morphologically analyse the sentence using
A. Architecture POS, Rule based and HMM).
Step8: Find the subject of the sentence.
Step 8.1 Find the person of the subject, if third person go
Input (Hindi Text) to step 8.2 else Goto 8.3
Step 8.2: Find the number and gender
Step 8.3: Find the auxiliary verb, check which array list it
Pre Processing belongs to and set variables accordingly.// Different array
lists have auxiliary verbs for different tenses.
Cleaning Tokenization Phrase Identification Step 8.4: Find the main verb
Step 8.5: Lemmatize the main verb
Step 8.6: Based on the variables initialised, recognize the
tense.
Clause Identification Step 8.7: Apply Rules and modify the verb and auxiliary
verb according to the tense recognised.
Step9: Print the output.
Hybrid Morphological C. Example
1: °ȯ/N-M-P ȯȡ/NN ȡɅ/VM Ǖȯ/VAUX ¡ɉȢ/VM
POS Number Sub: °ȯ/N-M-P
Gender Verb aux – Ǖȯ, ¡ɉȢ
Verb main – ȡɅ
Rule Based Case
Adjective
Lemmatize verb main, we receive the stem verb ȡ
Tense recognised “future perfect”
Verb
Rule:
N-M-P + Stem verb + Ǖȯ + ¡ɉȯ
HMM
(Hidden Markov model tagging) Output: °ȯ ȡ Ǖȯ ¡ɉȯ
2: ]/PP ȡȯ/VM ¡Ȫȯ/VAUX (Here, to
overcome the ambiguity of possessive pronoun as to
which gender it belongs to, the system will give output
Tense Synthesizer
for male and female)
Time of Action
Rule Based Sub: ]/PP
Synthesizer
Verb aux – ¡Ȫȯ
Verb main – ȡȯ
Analyser & Prediction Lemmatize verb main, we receive the stem verb ȡ
Tense recognised “future continuous”
Correction
Rules:
]+stem+ȯ ¡Ʌ ȯ
]+stem+Ȣ ¡Ʌ Ȣ
Output (Correct Hindi Text)
Fig.1. Architecture
2015 IEEE International Advance Computing Conference (IACC) 417
Output: ] ȡȯ ¡Ʌ ȯ and ] ȡȢ ¡Ʌ
Ȣ Accuracy = (n1*(Score of A))+ n2*(Score of B)+ n3*(Score
of C)+ n4*(Score of D)+ n5*(Score of E))/ Total number of
3: Ǖ/PP ȡɅ/VM Ǔȡ/VA AUX ¡ȪȢ/AUX Sentences(T)
(Exception of some Auxiliary verb) T=n1+n2+n3+n4+n5
Sub - Ǖ (2nd person) Here,
n1 represents number of senteences tagged ‘A’;
Auxiliary verb – Ǔȡ, ¡ȪȢ
n2 represents number of senteences tagged ‘B’
Verb main – ȡɅ n3 represents number of senteences tagged ‘C’
Lemmatize verb main, we receive thee stem verb ȡ n4 represents number of senteences tagged ‘D’
n5 represents number of senteences tagged ‘E’
Tense recognised “Simple past”
Accuracy of the system was found
f to be 70-75(approx.) %.
Rule:
Ǖ + ȯ + stem +Ǔȡ + ¡Ȫȡ V. CONC CLUSION
Output: Ǖ ȯ ȡ Ǔȡ ¡Ȫȡ Tense synthesizer is capable of analysing and correcting the
tenses of Hindi sentences up to certain extent. It follows a rule
IV. EVALUATION based approach and has been trained
t on 190 rules currently
For evaluating the accuracy of thee system, Human and more rules are required to handle
h some complex sentences.
Evaluation is used. Testing is performed onn 100 sentences of We have achieved an accuraacy of 70-75% for the tense
present, past, future tense. A collection of sentences were handling. Similarly if we work on Verb, phrases and adjective
taken on English and they were translated into Hindi through synthesizer more accuracy can be
b achieved. For example
Google, Bing and TDIL. Then theses traanslated Hindi was If the English input comes too machine for translation: The
passed through Tense synthesizer. The ouutputs were graded old man and woman the input anda Hindi output is ambiguous,
using the following scores. until we handle the phrases. In this case if old-man (Ǘ±ȡ
]Ȣ) ( budha aadmi) is phrasee then the correct output would
be Ǘ±ȡ ]Ȣ k k (budhaa aadmi aur aurat) otherwise it
Number of senten
nces would be Ǘ±ȯ ]Ȣ k k. (B
Budhe aadmi aur aurat)
As Indian languages are hiighly inflectional, this system
100% could be further improved byy designing more rules after
analysing more data. Further, accuracy can be improved by
50%
%
incorporating named entity reecognition better methods of
lemmatization for Hindi as disccussed in [13] and using HMM
0% approach as in [12], rule based approach as in for pos tagging
A B C D E in the morphological analyser. Also various alignment tools
Number of like Giza, Mossess could be inttegrated for sentence alignment
72% 8% 10% 3
3% 7%
sentences for better result. More post editting tools may be developed or
incorporated for the user to makke auto changes
Fig.2. Evalution VI. REFEERENCES
[1] Sneha Tripathi andd Juran Krishna Sarkhel,
“Approaches to macchine translation”, Annals of
TABLE [Link] BOARD
Library and Informatiion Studies, vol. 57, pp. 388-
Grade Translation Quality Sccore 393, December 2010.
A Excellent 100 [2] G.V. Garje1 and G.K.. Kharate, “Survey of machine
translation systems in India”, Int. J. Natural
B Very Good 90 Language Computingg, vol. 2, no. 4, pp. 47-67,
C Average 80 October 2013.
D Good 60 [3] V.A. Poroshin. Sem mantic analysis of natural
language. Available:
E Poor 40
[Link]
Using the following formula, the accuraccy was calculated /Fall/Seminar/[Link]
418 2015 IEEE International Advance Computing Conference (IACC)
[4] W. John Hutchins and Harold L. Somers, An [14] Preeti Yadav and Sandeep Vishwakarma, “Mining
Introduction to Machine Translation. London: Association rules based approach to Word Sense
Academic Press, 1992. Disambiguation for Hindi language”, Int. J.
Emerging Technology and Advanced Engineering,
[5] S. R. Priyanga and A. Azhagu Sindhu, “Rule based
vol. 3, no. 5, pp. 470-473, May 2013.
statistical hybrid machine translation”, Int. J. Science
and Modern Engineering (IJISME), vol. 3, no. 5, pp. [15] Kamlesh Dutta, Nupur Prakash and Saroj Kausik.
7-10, April 2013. Resolving Pronominal Anaphora in Hindi Using
Hobbs [Online]. Available:
[6] Machine Translation from Technology Development
[Link]
for Indian Languages (TDIL), Department of
Electronics & Information Technology (DeitY), [16] English Grammar, Tenses. Available:
Ministry of Communications & Information [Link]
Technology, Government of India. Available: english_grammar_tenses.pdf
[Link]
[17] Hindi Lessons. Available:
[Link]/components/com_mtsystem/CommonUI/home
[Link]
[Link]
[18] Shashi Pal Singh, Ajai Kumar, Hemant Darbari,
[7] R. Ananthakrishnan, M. Kavitha, J. Hegde Jayprasad,
Kanak Mohnot and Neha Bansal, “The Hindi named
Chandra Shekhar, Ritesh Shah, Sawani Bade and M.
entity recognizer using hybrid morphological
Sasikumar. “MaTra: A Practical Approach to Fully-
analyzer framework”, Int. J. Computer Applications
Automatic Indicative English-Hindi Machine
in Engineering Sciences, vol. IV, no. II, pp. 22-27,
Translation”, in Proceedings of the 1st National
June 2014.
Symposium on Modeling and Shallow Parsing of
Indian Languages, 2-4 April 2066, IIT Bombay, [19] [Link]
India. 5
[8] Aditi Kalyani, Hemant Kumud, Shashi Pal Singh,
Ajai Kumar and Hemant Darbari, “Evaluation and
ranking of machine translated output in Hindi
language using precision and recall oriented
metrics”, Int. J. Advanced Computer Research, vol.
4, no. 1, issue 14, pp. 54-58, March 2014
[9] P. Gupta, N. Joshi and I. Mathur, “Automatic
Ranking of MT Outputs using Approximations”, Int.
J. Computer Applications, vol. 81, no. 17, pp. 27-31,
November 2013.
[10] Based Machine Translator and Synthesizer System
English Language Essay, Available:
[Link]
language/based-machine-translator-and-synthesizer-
[Link]
[11] Nisheeth Joshi, Hemant Darbari and Iti Mathur,
“HMM Based POS Tagger for Hindi”, in the
Proceeding of International Conference on Artificial
Intelligence, Soft Computing (AISC-2013), pp. 341-
349, 2013.
[12] Snigdha Paul, Nisheeth Joshi and Iti Mathur,
“Development of a Hindi Lemmatizer”, Int. J.
Computational Linguistics and Natural Language
Processing, vol. 2, no. 5, pp. 380-384, May 2013.
[13] Parul Rastogi and S.K. Dwivedi, “Performance
comparison of Word Sense Disambiguation (WSD),
algorithm on Hindi language supporting search
engines”, Int. J. Computer Science Issues, vol. 8, no.
2, pp. 375-379, March 2011.
2015 IEEE International Advance Computing Conference (IACC) 419
View publication stats