0% found this document useful (0 votes)
23 views23 pages

Chapter 5

Uploaded by

Goku
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
0% found this document useful (0 votes)
23 views23 pages

Chapter 5

Uploaded by

Goku
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
After finishing this chapter, you will be able to: } Know the phonetical foundations in Sanskrit Language > Understand the basic structure of Sanskrit Language Grammar | » Get introduced to some computational aspects in Astadhyayi > Appreciate the importance of verbs in Sanskrit Language > Appreciate the potent Sanskrit offers in Natural Language Processing, Panini’s work on Sanskrit grammar has been acclaimed to be one of the best in the field of linguistics and phonetics. It has attracted wide attention. The accompanying figure is a Birch bark manuscript from Kashmir of the Rupavatara, a grammatical textbook based on the Sanskrit grammar of Panini. It was composed by Dharmakirti, a Buddhist monk from Ceylon. The ‘manuscript was transcribed in 1663. Source: hetps://[Link]!wiki/File:Birch_bark_MS_from_Kashmir_of_the_Rupavatra_ Wellcome. 0032691 jpg = 114 _Introdveton to Indian Knowledge SystemC ‘An Ecosystem for Sans! nt languages (Sanskrit Sanskrit is one of three an Greek, and Latin). While ancient Greek and Latin have attained a state of being classical, the Sanskrit language continues to be in use, albeit in specific domains by researchers. Hi in recent years, there is a growing recognitio of Sanskrit and several in ; taken globally to integrate technology aids with nguage. One of them is building n of the importance jatives have been the Sanskrit an ecosystem of language processing tools for Sanskrit. The efforts began in the late 1980s with several projects digitizing Sanskrit texts. Over a decade 2 central registry of digitized Indic texts was created in some Universities in Germany. Another progress in this direction was the creation of digital dictionaries of Sanskrit. The Cologne Digital Sanskrit Lexicon project digitized the Monier Willams’ Sanskrit-English Dictionary. Furthermore, the Digital Dictionaries of South Asia project at the University of Chicago digitized Apte’s and MacDonell's Sanskrit-English dictionaries. The next step in tis direction was the processing of Sanskrit text using machine coding. Academy of Sanskrit Research, Melkote developed a verbal cognition generator for Bhandarkar's Sanskrit primer in the early 1990s. By the turn of the century, the Ministry of Communication and Information Technology, Government of india set up the Technology Languages program (TOIL). has been a spurt in the efforts international Develo Further to this, t evident from multiple collaborative efforts in this direction developed Anusdraka’, The Akshar Bharati gro that employed techniques ed on Pénini's Astadhyayi. Beginning in 20 f Sanskrit computational tools was developed rough several research projects. These activities y (INU), Delhi as been the most effective ugh some basic comm 9S can elab d compl concepts and Applic for our commun andled v % ideas. Scientific discovery, the advancement of Jkrit Language Processing University of Hyderabed, IT Kenpo and Rashtriya Sanskrit Vidyapeeth, Tirgpatt Global, other Sanskrit linguistic researchers et allaborting on these Some of the fallout of these are developin Sensi including the following «+ Web-based Sanskrit reader program at Brown University called Kramapstha 2 digital edition of Whitney's roots that served as the source of verbal stems forthe ilectional generation sofware The international Digital Sanskrit Library Integration projectin the Classics Department at Brown University Sanskrit Library Phonetic encoding that allows all sounds represented in Vedic texts to be represented digitally Sanskrit Heritage platform (SH), centred around an electronic version of the Sanskrit- French Sanskrit Heritage dictionary Sanskrit Reader, allowing segmentation (andhi analysis), tagging, and parsing In 2008, the Indian Government funded a major consortium project to develop various tools for analysis of Sanskrit text and a Sanskrit-Hindi machine translation system. Sanskrit scholars and computational linguists collaborated to develop several tools for enhancing the Sanskrit language processing ecosystem. This includes Sandhi and Samasa analyser, morphological analyser, generator, sandhi joiner and splitter, and fullledged parser. These tools are now being deployed in some machine translation systems for testing and further refinements Source: Adopted from, Goyal, P., Huet, G., Kulkami, A,, Scharf, P, and Bunker, R. (2012). "A Distributed tform for Sanskrit Processing’, Proceedings of Papers, pp. 1011-1028, COLING Mumbai, December ication since time immemorial. th gestures, language becomes inevitable sowledge, and collaborative working require a common method of commun\ca¥er ine tenga this roe ina eviized society. As technology and computing prowess improves, we aFe ys o develop better applications using Artificial Intelligence (AN) techniaues, This requires able Nevelop efficient Natural Language Processing (NLP) capabilities. A systematic OT eguages and their capabilities for the evolving requirements has become important for the langice and technology community. In this chapter we shal see the developments in the fie sf yanguage in ancient India, 5.1 COMPONENTS OF A LANGUAGE ow many of us can read and comprehend Shakespeare's works or native Indian {ete Heth as Ramayana? Even if we know the language in which these texts are written, it IS difficult icause the way the language is used undergoes changes.§ pecmtuage 1s a tool used by everyone ina community and # One of the Vedangas Kt i karana focuses on linguistics itis very difficult to maintain it unchanged. We also notice 4 a tide there are differences in the same language spoken by and Phonetics aspects of San people from different regions. For these reasons, the, itaghya EO a ee Tees neompreheraiis fr te Weta a ra people in the future though they may be speaking the same and the best avalable descriptive fanguage. Studying the structure of the language helps us model of a languages not lose the underlying principles that govern a language and ensures that the received wisdom from the ancestors is not Jost. Italso maintains continuity in language processing. ‘Communication is key to trade, science and technology, and societal progress. It hinges on our ability to effectively process language as itis central to all human transactions and pursuits Language processing has two dimensions: receptive and productive. The receptive part of a language deals With the ability of an individual to receive language inputs from multiple sources and process them to decipher the intended message and comprehend them. On the other hand, the productive part of the language is to transmit back to others for their consumption. The focus in the former is on listening and reading, whereas it is on speaking and writing in the latter. Viewed from another perspective, sound (listening and speaking) and script (reading and writing) are the essential elements of a language. Therefore, language processing can be represented in a 2x2 framework as agua ge shown in Figure 5.1. Linguistics addresses all these aspects of a language. Phonetics will cater to the receptive and productive aspects of the sound and a syntactical structure will cater to the scriptural aspects of a language. Linguistics is a branch of language research that provides a scientific study of a language. It iS a systematic study of language to understand speech sounds, grammatical structures, and mea ing It helps us analyse the language form and meaning and identify systematic methods “integral to the Tanguage to derive the word forms a nd their meaning using structured rules and syntax. The earliest approach to a systematic treatment of linguistics is attributed to the Indian FIGURE 5.1 Components of a ‘oncepts and Applications to Indian Knowledge Syster ho lve inthe 6th Century BCE, His work on Sanslert grammar grammarian, Panint tum opus on linguistics. In this chapter, we shall see more dets ae jum opus on lin as Astadhydyris a magn! 5.2. PANINI’S WORK ON SANSKRIT GRAMMAR in the Indian cont e the pi /edas and their meanings was given ext, since the preservation of the Vedas anc ae ae ae a wy Sas an important discipline to be studied by everyone. Therena ames importance, linguistics was ref siigas, Siksa, Nirukta and Vyakarana focused on phonetical and ae a vogee hve Brel disused these ln Chapter 2 of the Book Deeg aspects of ARS. erin Sanskrit can comprehend ancien Ween the epics, eee aie ee ie carpas ina lossless fashion. Several fa] farig aaa borrowed the Partha sed ghonedi structure, language organization, and vocabalaes aa Panini’s work aaa gettimar Therefore, iis important to knove this in somalia aman woe aera raat, kno as Agiidkyi sl process of refiteuaet and syntactical suc Se ence a eae aa eation of human intelligence. - ae athe china is “a culmination of a_long grammatical tradition. Panini brought hig originality into action and proposed a structure to the grammar. Te is the best available descriptive model of a language. The greatness of this work is evident from the fact that fe has overshadowed all previous attempts on Sanskrit grammar, Moreover, several commentaries — written on its another proof of its prominence in linguistics. In the 4th century BCE, Katyayana composed a commentary (vartika) on the Paninan work, which served to: provide further explanation of the work and clarified certain aspects of Astadhyayi. Another great work on Astadhyayi is the Mahabhasya, a commentary on the Astadhyayi, authored by Patafijali in the 2nd century BCE. These three are considered the three great sages (tri-muni) of Sanskrit grammar. : While framing a syntax and linguistic framework, the challenge is to find an efficient set of rules to accommodate all the observed patterns and variations in the use of a language by the society, If there are robust rules that govern a language ¢ Pinini composed 3,983 rules to leaving several exceptions, it is not useful, At the other accommodate all the patterns and extreme, to accommodate all the possible variations, the Te spoanikt language: | number of rules framed may become very large, again * Te dotinguching teen wend making it less useful. The greatness of the Sansleit gramme Sat ae laid out by Papin bes In fe asiienec i ee eternal in its appeal framework for linguistics that nicely balanced both these — aspects, Panini did not write a set of rules and insisted that s to use the language. This is possible when we develop a for example. Panini composed 3,983 rules (known as Patterns and variations in the Sanskrit language and jm in eight chapters (therefore, the name Astédhyayi). Each chapter is | into 4 quarters, thereby making it 32 quarters in all ck ye f PSnini and its distinguishing features make Sanskrit a po d eternal in its appeal. These are summarized below: everyone must follow these computer language such as C or Python, ) to mostly accommodate the cd th basic approach of he entire vocabulary of the Sans rit language could be created using the 3,983 rules. “wv exceptions need special handling. The rules are aphorisms (known as siitras), share easy to commit to memory. What it implies is that ifsomeone is familiar with Linguistics 117 __ ee er et these rules and how it needs to be applied, then it amounts to gaining unambjgvov mastery of Sanskrit language ie. The educational systems in India until the introduction of the Macaulayan system of ff education in the 19th century CE ensured this mastery for the students. ‘+ Language processing and word generation are strictly rule-based and derivative ature. Proper application of the rules must result in a valid word (a form of the Taguage), There is no need to make any additional assumptions to derive any word + The derivation of words using the rul les could be done using a step-by-step process: What it implies is that technically speaking, st tarting from a base (a verbal root or wfominal Toot), the logic can search if any of the 3,983 rules could be applied to the current transformation of the root. If the rule can modify, the current structure it will perfo rm the operation. The procedure stops when none of the rules could be further applic .d. The result is the final word. This rule-based recursive structure to language is highly amenable to computer-based processing. The entire process !s logical, unambiguous, and rational. '¢ The entire scheme for word generation follows a highly modular approach. Two basic components form a word. Each word is formed out of a base (verbal or nominal root); ‘To this, one or more suffixes are added to generate the word. On account of this method of word generation, the Sanskrit language inherently has a high degree of ‘patterns’ of words and word structures that makes it very efficient to develop vocabulary. + Since the entire language is rule-based, the vocabulary is not fixed or static for the language. The rules can be used for generating new words, as long as the rules are not violated. Therefore, the language is dynamic, can construct new words as demand ‘rises, and can maintain its relevance. This has implications for lexicographic studies in the language. + Astidhyayi deploys several interesting data structures and computational elements that make it unique among linguistic studies. ‘These distinguishing aspects of Panini's grammar are discussed details in Sections 5.5, 5.6 and 5.8 of this chapter. Since language has the phonetic (sound) and syntactical (script) dimensions, we shall see Panini’s approach to these in the Sanskrit language. 5.3. PHONETICS IN SANSKRIT Phonetics is the study of sounds in a language, particularly the production of sound in a language and how it communicates the language corresponding to the scripts of the language. It also addresses the issue of how the sound is perceived in the language. Phonetics in the Sanskrit language has been addressed in some detail, since this is vital because the ancient Indian knowledge tradition is oral. The entire transmission of the Vedas from time immemorial has survived several thousand years on account of scientific methods of oral rendering, This has been possible on account of a well-developed science of phonetics. This also ensured an impeccable textual transmission superior to the classical texts of other cultures. This is perhaps the reason for UNESCO recognizing Vedas in the form of oral knowledge as a heritage for preservation. The science of the study of sound, known as Siksa, forms one of the six of the Vedaigas. In Chapter 2, we have a brief introduction to Siksa. In the earliest Indian traditions, Pratisakhyas sand Applications eton to Indian Krowdve = isare produced. Reveda-Pratisakhya and the Taittirtyaq e issue of how sounds are Pro o these texts, the sound pri enon ‘on the subject. According to Primarily. are the earliest works on one se nest due to the movement of the breath in the jael the junetion of the Tey a diferent locations resulting in various sound pattem sin the oral cavity As technology and computing prowess » According to Sanskrit grammar, two pho improves, we are able to develop better are considered to be homogenous ff thet ications using Artificial intelligence (Al) produced with the same articulatory chniques. This requires us to develop and at the same place of pronunciation, * Natural Language Processing (NLP) > The entire scheme for word generat follows a highly modular approach. Two | reation of components form a wor st available suffix (pratyaya). Each word is formed out of a base (verbal 6 tominal root). To this, one or more si are added to generate the word. One of the distinguishing aspects of # Paninian approach to linguistics and Sanskti fadhyayi is considered a fir human intelig criptive model of a language. 3,983 rules (known as mmodate most of the patterns skrit language and apters, According to earliest Indian primarily arises at # the sound grammar is the use of several featur sri tw cher a Junction of the throat which in modern parlance map to certal e chest due to the movement of the om, breath in the body and manifests -omputational concepts. Cavity at diferent locations resuting in various” 20INIS system of applying grammatiel sound panes. ue conditions to derive words works exactly like 2 rule-based engine, Every conceivable word in the Sanskrit language can be derived using a set of rules in a strict algorithmic fashion. Astadhyayl employs recursive logic to process ammatical applications. One such example the logic of samasa, a method of creating ‘compound wards from a group of nouns On account of the karaka concepts and linguistic arrangements, the sentence formation has a robust structure and leaves very little scope for ambiguity in the Sanskrit language. Most words (both the verb forms and noun forms) originate from the dhatus. In Sanskrit, we often find several synonyms for a word. REVIEW QUESTIONS 1. Briefly sketch the organization of Astadhyayi. How many rules are there in As Zz 3 4 10. 1. 12. 13, Linguistics 135 Each synonym for 2 word is often derived from a dhatu, Prefixes, known as upa-sargas, are appended to the verb forms in order to create additional words. There are 22 upa-sargas and these could be prefixed to a ver’ form. By adding the upa: sargas, it is possible to express the meaning in many ways. Studies show that certain inherent advantages in the structure of the Sanskrit language and the grammar makes it an attractive candidate for NLP and artificial intelligence-related work. \dhyayi? What are the unique aspects of the Sanskrit grammar propounded by Panini? How Is the issue of phonetics addressed in the Sanskrit language? What are the phonetical aspects, pertaining to vowels specified in Sanskrit grammar? ‘What do you understand by the terms ‘Prakrti' and ‘Pratyaya’? What is the relevance of these in Sanskrit grammar? Write short notes on the following: (@)_ The Sanskrit language is derivational in nature (b) Sanskrit grammar is rule-driven (Q) Sanskrit grammar is modular in structure ‘What are the ways by which one can generate noun forms in Sanskrit? Similarly, forms can be generated, explain how verb What do you understand by the terms ‘subanta’ and ‘titanta’? What do you understand by the term ‘mnemonics’? How is it used in Astadhyayi? Briefly explain how compound words are generated in the Sanskrit language, What is the relevance of the ‘Karaka concept? Using an example identify the Karakas in your example. Can you also relate it to the vibhaktis? Write a one-page note explaining why knowledge of dhatus in Sanskrit is very valuable. What happens when an upa-sarga is prefixed to a verb form? Give some exam your argument. ples in support of Comment on the statement, “The Sanskrit language has potential for use in NLP and Al applications",

You might also like