0% found this document useful (0 votes)
3 views45 pages

NLP Applications in Healthcare Systems

The document discusses the applications of Natural Language Processing (NLP) in medicine, highlighting its importance for analyzing clinical texts and electronic health records. It covers various NLP tasks, components, and methods, emphasizing the challenges posed by medical language and the need for specialized NLP systems. Key topics include low-level and high-level NLP components, such as tokenization, part-of-speech tagging, and named entity recognition, along with their relevance in improving healthcare quality and patient outcomes.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views45 pages

NLP Applications in Healthcare Systems

The document discusses the applications of Natural Language Processing (NLP) in medicine, highlighting its importance for analyzing clinical texts and electronic health records. It covers various NLP tasks, components, and methods, emphasizing the challenges posed by medical language and the need for specialized NLP systems. Key topics include low-level and high-level NLP components, such as tokenization, part-of-speech tagging, and named entity recognition, along with their relevance in improving healthcare quality and patient outcomes.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Module 4: NLP in Medicine

CO4:illustrate NLP applications for healthcare systems.

Example:-
Alexa,
Apple Siri,
Microsoft Cortana etc

Content:-
• NLP tasks in Medicine, Low-level NLP components, High level NLP
components, NLP Methods
• Clinical NLP resources and Tools, NLP Applications in Healthcare. Model
Interpretability using Explainable AI for NLP applications.
BY Prof. Ichhanshu Jaiswal (VCET,VASAI)
Common Use of NLP:
• Text/Document classifications
• Sentiments Analysis ( Emotion Analysis )
• Information Retrieval ( Extract Name , Place, etc from document )
• Language Translator
• Speech/Text Based chatbots
• Knowledge Graphs
• Text Summarizations
• Topic Modelling ( From Text decide on whom its talking about )
• Text Correction
• Speech to text conversion BY Prof. Ichhanshu Jaiswal (VCET,VASAI)
Language Diversity ??
There are more than
5k language in the
world so how to get
dataset for less
spoken language
BY Prof. Ichhanshu Jaiswal (VCET,VASAI)
What is NLP?
• NLP refer to any language used by people for communication other than machine
language or computer programming language.
• Natural language processing (NLP), a field of artificial intelligence and
computational linguistics, related to automated analysis of natural language.
Why NLP in Medicine ?
• In medical journals there are millions of text articles published in a single year that
can offer new treatment options .
• In clinical domain large amount of EHR data is required to process, to improve
healthcare quality through disease prediction, Expert System and
Recommendation system .
• Doctor use abbreviation/short name while writing prescription, so its challenge for
NLP as because medical sublanguages differ largely from general English language
Example:- The chest is clear, No abnormality found in urine, BP recorded 80/130 are
Medical language , NLP will help computer to understand such terms
BY Prof. Ichhanshu Jaiswal (VCET,VASAI)
• NLP is a set of techniques that extract meaning from text. These techniques
determine the meaning of a word, phrase, sentence, or document by recognizing the
grammatical rules of the language.
• NLP techniques can identify and extract elements such as proper names, locations,
actions, or events to find the relationships among them across documents.
• Translating unstructured content from a corpus of information, into a meaningful KB
is the task of NLP
• NLP breaks down the text to provide meaningful understanding.
• The focus of NLP is on determining the underlying grammatical and semantic patterns
that occur within a language.
• EXAMPLE: In the travel industry the word “fall” refers to a season of the year but In
a medical context it refers to a patient falling. So NLP looks not just at the domain,
but also at the levels of meaning that each of the following areas provide to our
understanding.
BY Prof. Ichhanshu Jaiswal (VCET,VASAI)
Example:

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


• NLP is an interdisciplinary field that applies statistical and rules-based
modelling of natural languages to automate the capability to interpret the
meaning of language.
• Steps in Natural Language processing in general are:
1. Language Identification & tokenization
2. Phonology ( Sound Understanding )
3. Morphology ( Structure of word ) PreExam, PostExam
4. Lexical Analysis ( Identify word ) Tank ( water or military )
5. Syntax Analysis. ( Relationship between word ) Old men and women
6. Semantic Analysis ( Meaning from structure ) Car hit pole while moving
7. Discourse Analysis ( Before & After ) He wanted that……. What is that ??
8. Pragmatic Analysis ( Context Understanding) Police Came… why ??

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


1. Language Identification & tokenization
• In any analysis of incoming text, the first process is to identify which
language the text is written in and then to separate the string of
characters into words ( tokenization ) take place.
2. Phonology
• Phonology is the study of the physical sounds of a language and how
those sounds are uttered in a particular language.
• This area is important for speech recognition and speech synthesis but
is not important for interpreting written text.
• Example: A person who is angry may use the same words as a person
who is confused; however, differences in intonation will convey
differences in emotion

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


3. Morphology:
• Morphology refers to the structure of a word
• Morphology gives us the stem of a word and its additional elements
of meaning.
• Example: “Pre-Exam” here I am talking of activity happening before
exam.
• Here we identify all kind of prefix, postfix, singular, plural etc to find
out exact meaning of word.
• There is a huge difference in meaning if someone uses the verb
“come” versus the verb “came.

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


4. Lexical Analysis ( Identify word )
• Lexical analysis is a vocabulary that includes its words and
expressions.
• It depicts analyze, identify and describe the structure of words.
• It includes dividing a text into paragraphs, sentences and word
• Individual words are analysed into their components, and nonword
tokens such as punctuations are separated from the words.
• Example: “Tank was full” Tank ???? water tank or military tank

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


5. Syntax analysis/Parsing (Identify Relationship between words)
• Syntax focus about the proper ordering of words which can affect its
meaning.
• This involves analysis of the words in a sentence by following the
grammatical structure of the sentence.
• The words are transformed into the structure to show hows the
word are related to each other.
• Example: Old man and women were taken to safe place??

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


6. Semantic Analysis/Parsing ( Meaning from structure )
• Semantic analysis is a structure created by the syntactic analyzer
which assigns meanings.
• Generated answer is how close to ideal answer ( Example: HOT
ICECREAM has no meaning)
• Semantics focuses only on the literal meaning of words, phrases, and
sentences.
• This only abstracts the dictionary meaning or the real meaning from
the given context.
• The structures assigned by the syntactic analyzer always have
assigned meaning
• Example: Car hit the pole while it was moving
BY Prof. Ichhanshu Jaiswal (VCET,VASAI)
[Link] Integration(Consider Before & After Sentence )
• It means to understand context of sentence. The meaning of any
single sentence sometimes depends upon before and after sentences.
• For example, the word "that" in the sentence "He wanted that"
depends upon the prior discourse context/sentence.

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


[Link] Analysis ( Understanding the context )
• Pragmatic Analysis deals with understanding meaning of sentence in
various situation
• Context help machine to identify true meaning of sentence
• Pragmatic analysis helps users to discover this intended effect by
applying a set of rules that characterize cooperative dialogues.
• Example: The Police are coming ??

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


Why NLP Task in Medicine ?
• NLP algorithms require special development for medical tasks because medical
sublanguages differ largely from general English
• Example: Clinical notes are entered by physicians who have limited time, so they
frequently use domain-specific abbreviations, omit information that can be
assumed by context, and have language problems such as misspellings or
incorrect word usage.
• It is very uncommon for physicians at different hospital sites to develop their own
local jargon for devices, techniques, or other items.
• One major challenge for medical NLP systems are barriers faced from data
availability and consistency.
• Many medical NLP systems need to access medical documents from EHR clinic
information systems.
• This can often be problematic because access to patient records is confidential,
requires the approval of institutional Review Boards (IRBs), and may require data
de-identification.
BY Prof. Ichhanshu Jaiswal (VCET,VASAI)
NLP task in medicine:-
• NLP task in medicine is divided into 2 parts
1. Low level NLP components
( Tokenization, SBD, POS, Parsing )

2. High level NLP components


(Negative Selection, Relationship Extraction, NER, WSD, SRL, IE )

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


Low level NLP components
1. Tokenization:
• Tokenization is an initial step of automated processing of a text.
• Full stop, white space, semi-colon, colon are mostly used to separate
token.
• Some of the difficulties that occur with tokenization from ambiguous
punctuation such as the colon in “2:30am” or full stop in “M.D.”
• The biomedical literature will also have certain technical terms such as
“Blood Pressure” and “Blood-Pressure,” which add additional
difficulties in tokenization as these are 2 token or 1 token.
• For this reason, a simple tokenizer for general English text will typically
not work well in biomedical text

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


2. Sentence Boundary Detection(SBD)
• This task aims to detect where sentences start and where it end.
• A simple SBD system can identify sentence boundaries using a small set of rules.
• Capital word and full stop helps in identify SBD.
• However, the task can be complicated by the fact that punctuation marks such
as question marks, semicolons, and white space are often ambiguous and need
more complex logic in special cases.
• AI methods such as decision trees, neural networks, and Hidden Markov Models
(HMMs) are frequently used for SBD.
• Biomedical literature a are full of abbreviations (e.g., “q.i.d.,” “p.r.n.”), acronyms
(e.g., “OD,” “OS”), and symbolic constructions (e.g., “blood pressure: 130/67”)
that add difficulty to SBD.
• For medical NLP systems, one frequent approach for SBD includes the use of
domain lexical resources such as the National Library of Medicine’s (NLM’s)
SPECIALIST Lexicon and annotated domain corpora to ensure satisfactory SBD
performance.
BY Prof. Ichhanshu Jaiswal (VCET,VASAI)
3. Part-of-Speech Tagging (POS) [Meaning of word for sentence understanding]
• The part of speech, indicates how the word functions in meaning as well as
grammatically within the sentence based on both definition as well as local
context.
• POS tagging is an essential step of NLP systems for proper understanding of
sentence so that error don’t propagate in syntactic output.
• POS taggers trained merely on general English do not usually performance on
medical text.
• A number of POS taggers have been developed specifically for the medical
domain.
1. Trigrams’n’Tags (TnT) tagger. trained on a relatively small set of clinical notes
2. MedPost tagger a POS tagger based on an HMM and trained on manually
tagged sentences in medical text.

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


POS tagging Example

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


4. Parsing: Used to check grammar of sentence
It gives structural
representation to
input data by
creating parse tree
as per grammar of
that language

The sentence
like “Door
open the”, for
example,
would be
rejected by
parser as it is
wrong
grammar. BY Prof. Ichhanshu Jaiswal (VCET,VASAI)
BY Prof. Ichhanshu Jaiswal (VCET,VASAI)
4. Shallow Parsing
• Shallow parsing (also chunking or light parsing) is an analysis of a
sentence which first identifies constituent parts of sentences (nouns,
verbs, adjectives, etc.) and then links them to higher order
units(parse tree) that have discrete grammatical meanings (noun
groups or phrases, verb groups, etc.).
• In the medical domain, shallow parsing is used in a wide range of
tasks such as drug–drug interaction (DDI) detection, medical
problem assertion detection, biological entity relation extraction,
and medical information extraction (IE).
• Several shallow parsers have been built for medical text processing,
such as the SPECIALIST minimal commitment parser which produces
high-level syntactic information rather than the traditional full
syntactic information for better noun phrase discovery in medical text

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


5 Deep Parser:
• Deep parsing is the process to produce an ordered, rooted tree that
represents the syntactic structure of a string according to some formal
grammar
• Full syntactic parsing of text can provide a large amount of deep linguistic
information such as sentence voice, phrase type, and POS tags, which are
shown to perform considerably better than surface-oriented features
(e.g., pattern matching) for many NLP tasks.
• General parser trained on English corpus have limited performance on
medical text.
• NLP experts have investigated several methods to adapt parsers trained
on general English to new target domains.
• New entries can be imported from domain resources to existing parser
lexicons using morphological clues, heuristic mapping, and direct expansion
to make them suitable for parsing medical knowledge.
BY Prof. Ichhanshu Jaiswal (VCET,VASAI)
High level NLP components
1. Negation Detection: ( -ve word convey meaningful information in Medical )
• Many medical documents such as discharge summaries contain large amounts of
important information of patients that can be used for a wide range of secondary
applications like Mediclaim, insurance etc.
• Report may contain sentence like “They have not noticed any abnormal behavior”
OR “no significant complications of bleeding.”
• Due to this, Negation detection is a critical component in medical NLP systems
as here NOT may be used to show positive symptoms.
• Consider two sentences: “The child is not tired” and “The child is not very tired.”
In the first sentence, the word “not” scope over “tired,” while in the second
sentence, the word “very” redirects the scope of “not” to itself and away from
“tired.” so identifying scope of negation is challenge
• Negation detection systems can detect negation solely by rules developed on a
handcrafted list of negation phrases that appear before or after a term of interest
or through the use of an ontology of negative medical concepts generated from
a standard medical dictionary.
BY Prof. Ichhanshu Jaiswal (VCET,VASAI)
2. Relationship Extraction:
• Relation extraction aims to determine or discover relationships between
entities (e.g., drugs, diseases, findings, genes) in medical texts.
• A large variety of relations exists in medical field, such as reactions
between drugs, genes, associations between diseases and symptoms, and
relations between patient problems and treatments.
• Rule-based approaches for relation extraction work, by exploiting the
linguistic patterns exhibited by relations.
• Rules can be manually defined by domain experts or derived from
annotated corpus.
• Machine learning-based systems rely on machine learning techniques
along with a variety of features based on the nature of the relationship
such as lexical, syntactic, semantic, and dependency features

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


3 Name Entity Recognition:
• Named entity recognition (NER) aims to identify and classify elements into
named entities, which are predefined categories such as names (e.g., drugs,
genes, person), findings, diseases, and medications.
• Example: MRI,XRAY, CTSCAN belong to entity reports while fever, malaria,
dengue belong to entity disease.
• The biomedical literature is full of terms particular to the biomedical domain that
are typically not detected by conventional general English NLP systems.
• In order to extract relations between entities, it is crucial for the system to be able
to detect unknown nouns or named entities.
• To achieve this Rule-based systems which require a significant manual effort,
and these rules may not easily extend to a new domain can be used for NER
• We can also use machine learning system can also be used this approaches
requires large annotation corpus for model training but it can adapt to new
domain easily

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


4. Word Sense Disambiguation (WSD):
• Ambiguity is a problem inherent to natural language, where a term
can have more than one meaning depending upon the context or use
of the term in a particular text
• Word sense disambiguation (WSD) is the process of understanding
which sense of a term, including single words, abbreviations, or
acronyms, is being used in a particular context.
• This problem is more extensive in the medical domain due to the
high use of abbreviations and acronyms in medical documents.
• Approaches to WSD generally rely on a particular domain knowledge
source, such as the Unified Medical Language System (UMLS) for
sense collection, sample collection, and model training.

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


5. Semantic Role Labeling ( SRL ) [ Help in asking Question from data]
• Semantic role labeling (SRL) is the task of detecting semantic roles
associated with predicates/verbs as it helps to answers questions such
as “who,” “when,” “what,” “where,” and “why.”
• Example: Ravi placed the ball beside the couch,” here the predicate is
the verb “place.” The semantic roles associated with “place” include
“placer”—Ravi; “thing placed”—ball; and “location”—beside the couch.
• SRL can be used for Information Extraction, question answering (QA),
text summarization.
• In order to build an SRL system to process medical text, domain
resources such as semantic annotated corpus and semantic frames
often must be created.
• Domain-specific features often boost the SRL performance in medical
domain .

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


[Link] Extraction (IE):
• IE is a task that involves extracting problem-specific information
from the text of interest and then transforming this information into
structured form.
• For example: 1. Vaccination reactions can be extracted from medical
reports 2. Relationships between genes and diseases from the
biomedical literature
• These systems were built mostly using pattern matching techniques
or machine learning systems
• Deep parsing, NER, and WSD are often part of an IE system
• In the medical domain, a variety of IE systems have been built for
various tasks as well as many NLP tools for IE, such the Medical
Language Extraction and Encoding System & clinical Text Analysis
and Knowledge Extraction System (cTAKES).

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


NLP Methods:
• NLP methods include symbolic (Grammer Rules-based) methods,
statistics-focused methods, and Machine learning methods.
• Symbolic methods are built based on rules of that language, while
statistical methods and Machine learning methods require training to
build models.
• This includes:-
1. Support Vector Machine ( ML Approach )
2. Maximum Entropy Modeling ( Statistics Approach )
3. N-Gram Modeling ( Statistics Approach )
4. Hidden Marcov Model ( Statistics Approach )

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


1. Support Vector Machine(SVM)
• SVM can be used for Text Classification.
• Example : We can classify Emails into spam or non-spam, news articles into
different categories like Politics, Stock Market, Sports, etc. using SVM.
• SVM can do linear or nonlinear classification, regression, and even outlier
detection tasks. • Multiple separation lines can be drawn, but how to get best line
?? This is not possible by just single line
• So we can draw 2 parallel line above and below this line, this
will helps in better classification
• These 2 parallel lines are called as margin
• The points on these margin lines are called as support vectors,
because they support in O/P.
• In 2-dimension separation line its called line but in higher
dimension its called as hyperplane
• SVM kernel comes into picture when we convert low dimension
data into high dimension for better classification.
BY Prof. Ichhanshu Jaiswal (VCET,VASAI)
• It works by looking for an optimal hyperplane that maximizes the distance
between the hyperplane and the nearest samples from each of the two
classes.
• For an N-feature case, the input may be transformed mathematically using
a kernel function into higher dimension to allow linear separation of the
data points
• The prediction accuracy of SVM is generally high because of the sound
mathematical theory behind it and the robustness of the method.
• It generally works well when training examples contain errors, as well,
because of its use of a separation process.
• SVM is computationally expensive as its training process requires to
raise/increase dimension of data points.
• In the medical domain, SVM has been shown to perform well on many
classification tasks such as smoking status classification, disease detection
etc
BY Prof. Ichhanshu Jaiswal (VCET,VASAI)
2. Maximum Entropy Modelling (MEM)
• The principal of maximum entropy modeling (MEM) is simple and realistic.
• It is a statistical learning method that models all that is known and assumes
nothing about that which is unknown.
• It is a probability-based model, in which probability of occurrence of each
content is readily available in corpus, and output is content with max
probability.
• For Example: You have employed some expert to translate English into Tamil,
this person has in his mind which Tamil words to use for a particular English
word.
• The translation made by him purely depends how good is his vocabulary and
there may be case for certain English word he could not convert to Tamil as P (
Tamil/English) for this is 0 as he has no knowledge of this.
• In case of multiple Tamil word for a particular English word, one which has max
probability/accuracy is delivered depending on the vocabulary of expert.
BY Prof. Ichhanshu Jaiswal (VCET,VASAI)
3. N Gram Modelling (Train Model to predict next
word)
• It is stastical based approach for NLP
• N-gram is a sequence of the N-words in the modeling of NLP.
• Consider an example of the statement for modeling. “I love reading
history books and watching documentaries”.
• In one-gram or unigram, there is a one-word sequence. As for the above
statement, in one gram it can be “I”, “love”, “history”, “books”, “and”,
“watching”, “documentaries”.
• In two-gram or the bi-gram, there is the two-word sequence i.e. “I
love”, “love reading”, or “history books”.
• In the three-gram or the tri-gram, there are the three words sequences
i.e. “I love reading”, “ reading history books,” “and watching
documentaries.

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


• For N-1 words, the N-gram modeling predicts most occurred words
that can follow the sequences.
• The model is the probabilistic language model which is trained on the
collection of the text.
• This model is useful in applications i.e. speech recognition, and
machine translations.
• So, the N-gram language model is about finding probability
distributions over the sequences of the word.

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


4. Hidden Marcov Model (HMM)
• HMM is another statistical NLP method. Markov models are built on
the Markov assumption that the next state depends on current state
only that why marcov model are called as memoryless model
• Example:- Here we know that Ravi is Happy/Sad and depending on
this we predict weather condition at this location. This is HMM

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


• States of Marcov Chain are hidden but we can observe some
variables that are dependent on that states and this is called HMM.
• For example, in a speech recognition system, the sound we hear is
the output of hidden states, such as vocal chords, the size of the
person’s throat, the position of the person’s tongue, and many other
factors.
• Each sound of a word is generated from changes of these hidden
factors.
• It has been used in medicine to describe the effect of alcoholism
treatment on the likelihood of healthy/unhealthy populations,

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


Clinical NLP Resources & Tools
• Here we mainly focus on several NLP resources and tools available in
the biomedical domain.
• They are as below:-
1. Unified Medical Language System (UMLS)
2. Corpora
3. SPECIALIST NLP Tool
4. MetaMap
5. SemRep

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


1. Unified Medical Language System (UMLS)
• UMLS was developed by and is maintained by the National Library of
Medicines(NLM) to provide health care professionals and researchers
with a biomedical domain knowledge resource.
• UMLS is a structured KB that connects different biomedical sources
and enables biomedical research application development.
• UMLS contains three knowledge sources as given below

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


i. Metathesaurus: It is a large, multi-purpose, and multi-lingual
vocabulary database that contains information about biomedical
and health related concepts, their various names, and the
relationships among them. It is built from the electronic record of
many different medical books, terms used in patient care, health
services billing, public health statistics, biomedical literature,
clinical data, and health services research.
ii. Semantic Network: It is knowledge structure that depicts how
concepts are related to one another and illustrates how they are
interconnect. Semantic networks use artificial intelligence (AI)
programming to mine data, connect concepts and establish
relation between them.
Example:-

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


iii. SPECIALIST Lexicon: Is an English dictionary including over 200,000 biomedical
terms as well as common English words. It also contains syntactic, morphological,
and orthographic information of each term or word.
For example, each lexical record contains base forms of the term; the part of
speech; a unified identifier; spelling variants; and inflection for nouns, verbs, and
adjectives. BY Prof. Ichhanshu Jaiswal (VCET,VASAI)
2. Corpora: A medical KB
• Medical corpus is a domain-specific corpus made up of texts in
English related to medical science, collected from the Internet .It
follows standard given by NLM
• The MEDLINE (one such corpus) is a collection of biomedical
abstracts. It is maintained by the NLM and contains over 21 million
reference from 1946 to the present
• The corpus has been annotated with various levels of linguistic and
semantic information covering POS, syntactic, term, event, relation,
and coreference annotation
• Most research groups created their own clinical text corpus and
annotations for specific NLP tasks locally.

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


3. SPECIALIST NLP Tools
• They are computer programs developed by the NLM to aid in dealing with
different biomedical NLP tasks.
• Tools include lexical tools such as lexical variant generator (LVG), normalized
string generator (Norm), word index generator (WordInd), dTagger POS tagger,
subterm mapping tools (STMTs), and others.
• LVG contains a series of commands to perform lexical transformation of text.
• Norm provides a normalization process for those terms included in the SPECIALIST
Lexicon. Norm can help to find similar terms and to map terms to UMLS concepts.
• WordInd creates a sequence of alphanumeric characters by reading the text,
helping UMLS to produce the word index for the Metathesaurus.
• dTagger is a POS tagger built specifically with SPCIALIST Lexicon. dTagger was
trained on MedPost corpus, a set of annotated MEDLINE abstracts, and tokenizes
text into multiword terms
• STMT was built to provide comprehensive subterm-related features, including all
subterms, the longest prefix subterm, and synonymous subterm substitutions.

BY Prof. Ichhanshu Jaiswal (VCET,VASAI)


[Link]
• MetaMap is a program developed by the NLM to map biomedical
text to the UMLS Metathesaurus.
• MetaMap provides various options, including data option (choose
specific vocabularies and data model); processing options (such as
author-defined acronyms/abbreviations, negation detection,
WSD/ambiguity) and output options (human readable, machine
output, and XML).
• Released application programming interfaces (APIs) provide options
to integrate MetaMap into other programs
• It is an effective tool to map biomedical terms and has been widely
used in applications of the clinical domain, such as detection of
clinical findings.
BY Prof. Ichhanshu Jaiswal (VCET,VASAI)

You might also like