Process Model Generation From Natural Language Text: Abstract. Business Process Modeling Has Become An Important Tool For
Process Model Generation From Natural Language Text: Abstract. Business Process Modeling Has Become An Important Tool For
1 Introduction
Business process management is a discipline which seeks to increase the effi-
ciency and effectiveness of companies by holistically analyzing and improving
business processes across departmental boundaries. In order to be able to ana-
lyze a process, a thorough understanding of it is required first. The necessary
level of insight can be obtained by creating a formal model for a given business
process.
The required knowledge for constructing process models has to be made ex-
plicit by actors participating in the process [1]. However, these actors are usually
not qualified to create formal models themselves [2]. For this reason, modeling
experts are employed to iteratively formalize and validate process models in col-
laboration with the domain experts. This traditional procedure of extracting
process models involves interviews, meetings, or workshops [3]. It entails con-
siderable time and costs due to ambiguities or misunderstandings between the
involved participants [4]. Therefore, the initial elicitation of conceptual models
is considered to be a knowledge acquisition bottleneck [5]. According to Herbst
[1] the acquisition of the as-is model in a workflow project requires 60% of the
total time spent. Accordingly, substantial savings are possible by providing ap-
propriate tool support to speed up the acquisition phase.
H. Mouratidis and C. Rolland (Eds.): CAiSE 2011, LNCS 6741, pp. 482–496, 2011.
c Springer-Verlag Berlin Heidelberg 2011
Process Model Generation from Natural Language Text 483
2 Background
the documents upon completeness and registers the claim. Then, the Handling
department picks up the claim and checks the insurance. Then, an assessment
is performed. If the assessment is positive, a garage is phoned to authorise the
repairs and the payment is scheduled (in this order). Otherwise, the claim is re-
jected. In any case (whether the outcome is positive or negative), a letter is sent
to the customer and the process is considered to be complete.” Such information
is usually provided by people working in the process and then formalized as a
model by system analysts [2].
For our model generation approach, we will employ methods from compu-
tational linguistics and natural language processing. This branch of artificial
intelligence deals with analyzing and extracting useful information from natural
language texts or speech. For our approach, three concepts are of vital impor-
tance: syntax parsing, which is the determination of a syntax tree and the gram-
matical relations between the parts of the sentence; semantic analysis, which
is the extraction of the meaning of words or phrases; and anaphora resolution,
which involves the identification of the concepts which are references using pro-
nouns (“we”,“he”,“it”) and certain articles (“this”, “that”). For syntax parsing
and semantic analysis, there are standard tools available.
The Stanford Parser is a syntax parsing tool for determining a syntax tree.
This tree shows the dependencies between the words of the sentence through the
tree structure [10]. Additionally, each word and phrase is labeled with an appro-
priate part-of-speech and phrase tag. The tags of the Stanford Parser are the
same which can be found in the Penn Tree Bank [11]. The Stanford Parser also
produces 55 different Stanford Dependencies [12]. These dependencies reflect the
grammatical relationships between the words. Such grammatical relations pro-
vide an abstraction layer to the pure syntax tree. They also contain information
about the syntactic role of all elements.
Process Model Generation from Natural Language Text 485
There are also tools available for semantic analysis. They provide semantic
relations on different levels of detail. We use FrameNet [13] and the lexical
database WordNet[14]. WordNet provides various links to synonyms, homonyms,
and hypernyms for a particular class of meaning associated with a synonym-set.
FrameNet defines semantic relations that are expected for specific words. These
relations are useful, e.g., to recognize that a verb “send” would usually go with
a particular object being sent. Syntax parsers and semantic analysis are used in
our transformation approach, augmented with anaphora resolution.
3 Transformation Approach
The most important issue we are facing when trying to build a system for gener-
ating models is the complexity of natural language. We collected issues related
to the structure of natural language texts from the scientific literature and an-
alyzed the test data, which is described in section 4. Thereby, we were able to
identify four broad categories of issues which we have to solve in order to ana-
lyze natural language process descriptions successfully (see Table 1). Syntactic
Leeway relates to the fact that there is a mismatch between the semantic and
syntactic layer of a text. Atomicity deals with the question of how to construct
a proper phrase-activity mapping. Relevance has to check whether parts of the
text might be irrelevant for the generated process model. Finally, Referencing
addresses the question of how to resolve relative references between words and
between sentences.
through model generation. This approach was also taken by most of the other
works which built a similar system [17,22,18,23]. The data structure used by the
approach of the University of Rio de Janeiro [23] was taken from the CREWS
project [15]. The authors argue that it is suited well for this task as a scenario
description corresponds to the description of a process model. Therefore, we also
use the CREWS scenario metamodel as starting point. However, we modified
several parts as, e.g., we explicitly represent connections between the elements
using the class “Flow”. Additionally, we explicitly considered traceability as a
requirement. Thus, attributes relating an object to a sentence or a word are
added to the World Model. The four main elements of our World Model are
Actor, Resource, Action, and Flow. This World Model will be used throughout
all phases of our transformation procedure to capture syntactic and semantic
analysis results. Each phase is allowed to access, modify and add data.
The rest of this section is dedicated to analyzing and discussing the issues
collected in Table 1. We will then seize the developed suggestions and reference
these issues during the description of our transformation approach. Section 3.1
discusses sentence level analysis for finding actions. Section 3.2 investigates text
level analysis for enriching the data stored in the world model. Finally, Sec-
tion 3.3 describes the generation of a BPMN model. While we focus on the
general procedure here, we documented details of all algorithms in [25].
The first step of our transformation procedure is a sentence level analysis. The
extraction procedure consists of the steps that are outlined as a BPMN model
in Figure 2. This overview also shows the different components upon which our
transformation procedure builds and their usage of Data Sources.
The text is processed in several stages. First, a tokenization splits up the text
into individual sentences. The challenge here is to distinguish a period used for
an abbreviation (e.g. [Link].) from a period marking the end of a sentence.
Afterwards, each sentence is parsed by the Stanford Parser using the factored
model for English [11]. We utilize the factored model and not the pure proba-
bilistic context free grammar, because it provides better results in determining
the dependencies between markers as “if” or “then”, which are important for
the process model generation. Next, complex sentences are split into individual
phrases. This is accomplished by scanning for sentence tags on the top level of
the Parse Tree and within nested prepositional, adverbial, and noun phrases.
Once the sentence is broken down into individual constituent phrases, actions
can be extracted. First, we determine whether the parsedSentence is in active
or passive voice by searching for the appropriate grammatical relations (Issue
1.1). Then, all Actors and Actions are extracted by analyzing the grammatical
relations. To overcome the problem of example sentences mentioned earlier (Issue
3.2) the actions are also filtered. This filtering method simply checks whether the
sentence contains a word of a stop word list called example indicators. Then, we
extract all objects from the phrase and each Action is combined with each Object.
The same is done with all Actors. This procedure is necessary as an Action
is supposed to be atomic according to the BPMN specification [8] and Issue
2.1. Therefore, a new Action has to be created for each piece of information as
illustrated in the following example sentences. In each sentence the conjunction
relation which causes the extraction of several Actors, Actions or Resources is
highlighted. As a last step, all extracted Actions are added to the World Model.
◦ “Likewise the old supplier creates and sends the final billing to the cus-
tomer.” (Action)
◦ “It is given either by a sales representative or by a pre-sales employee
in case of a more technical presentation.” (Actor)
◦ “At this point, the Assistant Registry Manager puts the receipt and
copied documents into an envelope and posts it to the party.” (Resource)
The last step of the text level analysis is the generation of Flows. A flow
describes how activities are interacting with each other. Therefore, during the
process model generation such Flows can be translated to BPMN connecting
objects. When creating the Flows we build upon the assumption that a process
is described sequentially and upon the information gathered in the previous steps.
The word, which connected the items is important to determine how to proceed
and what type of gateway we have to create. So far we support a distinction
between “or”, “and/or”, and “and”. Other conjunctions are skipped.
In the last phase of our approach the information contained in the World Model
is transformed into its BPMN representation. We follow a nine step procedure, as
depicted in Figure 3.3. The first 4 steps: creation of nodes, building of Sequence
Flows, removal of dummy elements, the finishing of open ends, and the processing
of meta activities are used to create an initial and complete model. Optionally,
the model can be augmented by creating Black Box Pools and Data Objects.
Finally, the model is laid out to achieve a human-readable representation.
Fig. 4. Structural overview of the steps of the Process Model Generation phase
The first step required for the model creation is the construction of all nodes
of the model. After the Flow Object was generated, we create a Lane Element
representing the Actor initiating the Action. If no Lane was determined for an
Action, it is added to the last Lane which was created successfully as we assume
that the process is described in a sequential manner. The second step required
during the model creation is the construction of all edges. Due to the defini-
tion of Flows within our World Model, this transformation is straight-forward.
Whenever a Flow which is not of the type “Sequence” is encountered, a Gate-
way appropriate for the type of the flow is created. An exception to that is the
490 F. Friedrich, J. Mendling, and F. Puhlmann
type “Exception”. If the World Model contains a flow of this type, an exception
intermediate event is attached to the task which serves as a source and this In-
termediate Event is connected to the target instead of the node itself. We then
skip dummy actions, which were inserted between gateways directly following
each other.
Step four is concerned with open ends. So far, no Start and End Events
were created. This is accomplished in this step. The procedure is also straight
forward. We create a preceding Start event to all Tasks which do not have any
predecessors (in-flow = 0) and succeeding End Events to all Tasks which do
not have any successors (out-flow = 0). Additionally, Gateways whose in- and
out-flow is one receive an additional branch ending in an End Event.
The last step in the model creation phase handles Meta-Activities (Issue 3.3).
We search and remove redundant nodes directly adjacent to Start or End Events.
This is required as several texts contain sentences like “[...] the process flow at
the customer also ends.” or “The process of “winning” a new customer ends
here.” If such sentences are not filtered, we might find tasks labeled “process
ends” right in front of an end event or “start workflow” following a start event.
We remove nodes whose verb is contained in the hypernym tree of “end” or
“start” in WordNet if they are adjacent to a Start or End Event.
The execution of these five steps yields a full BPMN model. As the elements of
this model do not contain any position information yet, our generation procedure
concludes with an automated layout algorithm. We utilize a simple grid layout
approach similar to [26], enhanced with standard layout graph layout algorithms
as Sugiyama [27] and the topology-shape-metric approach [28]. For the example
text of the claims handling process from Section 2 we generated the model given
in Figure 5. The question of how far this result can be considered to be accurate
is discussed in the following section.
to create pairs of nodes and edges. We use the greedy heuristic as it showed the
best performance without considerable accuracy trade-offs. After the mapping
is created, a Graph Edit Distance value can be calculated given:
◦ Ni - set of nodes in model i
◦ Ei - set of edges of model i
◦ Ni - the set of nodes in model i which were not mapped
◦ Ei - the set of edges in model i which were not mapped
◦ M - The mapping between the nodes of model 1 and 2
An indicator for the difference between the models can be calculated as:
|M|
∗ i=1 1 − sim(Mi ) if|M | > 0
m = (1)
1.0 otherwise
As a last step weights for the importance of the differences (wmap ), the un-
mapped Nodes (wuN ), and the unmapped Edges (wuE ) have to be defined. For
our experiments we gave the difference a slightly higher importance and assigned
wmap = 0.4 and wuN = wuE = 0.3. The overall graph edit distance then becomes:
Table 3. Result of the application of the evaluation metrics to the test data set
and edges within the generated models. We can see that the transformation
procedure tends to produce models which are on average 9-15% larger in size
then what a human would create. This can be partially explained by noise and
meta sentences which were not filtered appropriately. On the other hand, humans
tend to abstract during the process of modeling. Therefore, we often find more
detail of the text also in the generated model. The results are highly encouraging
as our approach is able to correctly recreate 77% of the model in average. On
a model level up to 96% of similarity can be reached, which means that only
minor corrections by a human modeler are required.
During the detailed analysis we determined different sources of failure, which
resulted in a decreased metric value. These are noise, different levels of abstrac-
tions, and processing problems within our system. Noise includes sentences or
phrases that are not part of the process description, as for instance “This ob-
ject consists of data elements such as the customers name and address and the
assigned power gauge.” While such information can be important for the under-
standing of a process, it leads to unwanted Activities within the generated model.
To tackle this problem, further filtering mechanisms are required. Low similarity
also results from difference in the level of granularity. To solve this problem, we
could apply automated abstraction techniques like [34] on the generated model.
Finally, the employed natural language processing components failed during the
analysis. At stages, the Stanford Parser failed at correctly classifying verbs. For
instance, the parser classified “the second activity checks and configures” as a
noun phrase, such that the verbs “check” and “configure” cannot be extracted
into Actions. Furthermore, important verbs related to business processes are
not contained in FrameNet, as “report”. Therefore, no message flow is created
between report activities and a Black Box Pool. We expect this problem to
be solved in the future as the FrameNet database grows. With WordNet, for
instance, there is a problem with times like “2:00 pm”, where pm as an abbre-
viation for “Prime Minister” is classified as an Actor. To solve this problem a
reliable sense disambiguation has to be conducted. Nevertheless, overall good
results were achieved by using WordNet as a general purpose Ontology.
5 Related Work
6 Conclusion
texts making use of indentions, or texts which are of low quality is not possible
at the moment and presents opportunities for further research.
While the evaluation conducted in this thesis evinced encouraging results
different lines of research could be pursued in order to enhance the quality or
scope of our process model generation procedure. As shown the occurrence of
meta-sentences or noise in general is one of the severest problems affecting the
generation results. Therefore, we could improve the quality of our results by
adding further rules and heuristics to identify such noise. Another major source
of problems was the syntax parser we employed. As an alternative, semantic
parsers like [39] could be investigated.
References
1. Herbst, J., Karagiannis, D.: An inductive approach to the acquisition and adapta-
tion of workflow models. In: Proceedings of the IJCAI, pp. 52–57 (1999)
2. Frederiks, P., Van der Weide, T.: Information modeling: the process and the
required competencies of its participants. Data & Knowledge Engineering 58(1),
4–20 (2006)
3. Scheer, A.: ARIS-business process modeling. Springer, Heidelberg (2000)
4. Reijers, H., Limam, S., Van Der Aalst, W.: Product-based workflow design. Journal
of Management Information Systems 20(1), 229–262 (2003)
5. Gruber, T.: Automated knowledge acquisition for strategic knowledge. Machine
Learning 4(3), 293–336 (1989)
6. Blumberg, R., Atre, S.: The problem with unstructured data. DM Review 13, 42–49
(2003)
7. White, M.: Information overlook. EContent(26:7) (2003)
8. OMG, eds.: Business Process Model and Notation (BPMN) Version 2.0 (June 2010)
9. Freund, J., Rücker, B., Henninger, T.: Praxishandbuch BPMN. Hanser (2010)
10. Melčuk, I.: Dependency syntax: theory and practice, New York (1988)
11. Marcus, M., Marcinkiewicz, M., Santorini, B.: Building a large annotated corpus
of English: The Penn Treebank. Computational Linguistics 19(2), 330 (1993)
12. de Marneffe, M., Manning, C.: The Stanford typed dependencies representation.
In: Workshop on Cross-Framework and Cross-Domain Parser Evaluation, pp. 1–8
(2008)
13. Baker, C., Fillmore, C., Lowe, J.: The berkeley framenet project. In: 17th Int. Conf.
on Computational Linguistics, pp. 86–90 (1998)
14. Miller, G.A.: Wordnet: A lexical database for english. CACM 38(11), 39–41 (1995)
15. Achour, C.B.: Guiding scenario authoring. In: 8th European-Japanese Conference
on Information Modelling and Knowledge Bases, pp. 152–171. IOS Press, Amster-
dam (1998)
16. Li, J., Wang, H., Zhang, Z., Zhao, J.: A policy-based process mining framework:
mining business policy texts for discovering process models. ISEB 8(2), 169–188
17. Yue, T., Briand, L., Labiche, Y.: An Automated Approach to Transform Use Cases
into Activity Diagrams. Modelling Foundations and Appl., 337–353 (2010)
18. Fliedl, G., Kop, C., Mayr, H., Salbrechter, A., Vöhringer, J., Weber, G., Winkler,
C.: Deriving static and dynamic concepts from software requirements using sophis-
ticated tagging. Data & Knowledge Engineering 61(3), 433–448 (2007)
19. Kop, C., Mayr, H.: Conceptual predesign–bridging the gap between requirements
and conceptual design. In: 3rd Int. Conf. on Requirements Eng. p. 90 (1998)
496 F. Friedrich, J. Mendling, and F. Puhlmann
20. Ghose, A., Koliadis, G., Chueng, A.: Process Discovery from Model and Text Arte-
facts. In: 2007 IEEE Congress on Services, pp. 167–174 (2007)
21. Ghose, A.K., Koliadis, G., Chueng, A.: Rapid business process discovery (R-BPD).
In: Parent, C., Schewe, K.-D., Storey, V.C., Thalheim, B. (eds.) ER 2007. LNCS,
vol. 4801, pp. 391–406. Springer, Heidelberg (2007)
22. Sinha, A., Paradkar, A., Kumanan, P., Boguraev, B.: An Analysis Engine for
Dependable Elicitation on Natural Language Use Case Description and its Ap-
plication to Industrial Use Cases. Technical report, IBM (2008)
23. de AR Gonçalves, J.C., Santoro, F.M., Baião, F.A.: A case study on designing pro-
cesses based on collaborative and mining approaches. In: Int. Conf. on Computer
Supported Cooperative Work in Design, Shanghai, China (2010)
24. Fliedl, G., Kop, C., Mayr, H.: From textual scenarios to a conceptual schema. Data
& Knowledge Engineering 55(1), 20–37 (2005)
25. Friedrich, F.: Automated generation of business process models from natural
language input. Master’s thesis, Humboldt-Universität zu Berlin (November 2010)
26. Kitzmann, I., Konig, C., Lubke, D., Singer, L.: A Simple Algorithm for Automatic
Layout of BPMN Processes. In: IEEE Conf. CEC, pp. 391–398 (2009)
27. Seemann, J.: Extending the sugiyama algorithm for drawing UML class diagrams:
Towards automatic layout of object-oriented software diagrams. In: Graph Drawing,
pp. 415–424. Springer, Heidelberg (1997)
28. Eiglsperger, M., Kaufmann, M., Siebenhaller, M.: A topology-shape-metrics
approach for the automatic layout of UML class diagrams. In: Proceedings of the
2003 ACM Symposium on Software Visualization, p. 189. ACM, New York (2003)
29. White, S., Miers, D.: BPMN Modeling and Reference Guide: Understanding and
Using BPMN. Future Strategies Inc. (2008)
30. Holschke, O.: Impact of granularity on adjustment behavior in adaptive reuse of
business process models. In: Hull, R., Mendling, J., Tai, S. (eds.) BPM 2010. LNCS,
vol. 6336, pp. 112–127. Springer, Heidelberg (2010)
31. Reijers, H.: Design and control of workflow processes: business process management
for the service industry. Eindhoven University Press (2003)
32. Dijkman, R., Dumas, M., van Dongen, B., Käärik, R., Mendling, J.: Similarity of
business process models: Metrics and evaluation. Inf. Sys. 36, 498–516 (2010)
33. Dijkman, R., Dumas, M., Garcıa-Banuelos, L., Käärik, R.: Graph Matching
Algorithms for Business Process Model Similarity Search. In: Dayal, U., Eder, J.,
Koehler, J., Reijers, H.A. (eds.) BPM 2009. LNCS, vol. 5701, pp. 48–63. Springer,
Heidelberg (2009)
34. Polyvyanyy, A., Smirnov, S., Weske, M.: On application of structural decomposi-
tion for process model abstraction. In: 2nd Int. Conf. BPSC, pp. 110–122 (March
2009)
35. Kop, C., Vöhringer, J., Hölbling, M., Horn, T., Irrasch, C., Mayr, H.: Tool Sup-
ported Extraction of Behavior Models. In: Proc. 4th Int. Conf. ISTA (2005)
36. Yue, T., Briand, L., Labiche, Y.: Automatically Deriving a UML Analysis Model
from a Use Case Model. Technical report, Carleton University (2009)
37. Wang, H.J., Zhao, J.L., Zhang, L.J.: Policy-Driven Process Mapping (PDPM):
Discovering process models from business policies. DSS 48(1), 267–281 (2009)
38. Sinha, A., Paradkar, A.: Use Cases to Process Specifications in Business Process
Modeling Notation. In: 2010 IEEE Int. Conf. on Web Services, pp. 473–480 (2010)
39. Shi, L., Mihalcea, R.: Putting Pieces Together: Combining FrameNet, VerbNet and
WordNet for Robust Semantic Parsing. In: Proceedings of the 6th Int. Conf. on
Computational Linguistics and Intelligent Text Processing, p. 100 (2005)