0% fanden dieses Dokument nützlich (0 Abstimmungen)
3 Ansichten278 Seiten

Dissertation

Die Dissertation von Michael Werner behandelt die Automatisierung der Analyse von Geschäftsprozessen im Rahmen von Finanzprüfungen unter Verwendung von Process Mining-Techniken. Sie präsentiert einen Multilevel Process Mining (MLPM) Algorithmus, der speziell für Jahresabschlussprüfungen entwickelt wurde und die Analyse von Geschäftsprozessen verbessert, um die Effizienz und Effektivität der Prüfungen zu steigern. Die Arbeit zeigt auf, dass trotz der Automatisierung bestimmter Prüfungsverfahren zukünftige Bilanzskandale nicht vollständig verhindert werden können, jedoch Ressourcen freigesetzt werden, um risikobehaftete Transaktionen besser zu prüfen.

Hochgeladen von

9dv4jq5g8w
Copyright
© All Rights Reserved
Wir nehmen die Rechte an Inhalten ernst. Wenn Sie vermuten, dass dies Ihr Inhalt ist, beanspruchen Sie ihn hier.
Verfügbare Formate
Als PDF, TXT herunterladen oder online auf Scribd lesen
0% fanden dieses Dokument nützlich (0 Abstimmungen)
3 Ansichten278 Seiten

Dissertation

Die Dissertation von Michael Werner behandelt die Automatisierung der Analyse von Geschäftsprozessen im Rahmen von Finanzprüfungen unter Verwendung von Process Mining-Techniken. Sie präsentiert einen Multilevel Process Mining (MLPM) Algorithmus, der speziell für Jahresabschlussprüfungen entwickelt wurde und die Analyse von Geschäftsprozessen verbessert, um die Effizienz und Effektivität der Prüfungen zu steigern. Die Arbeit zeigt auf, dass trotz der Automatisierung bestimmter Prüfungsverfahren zukünftige Bilanzskandale nicht vollständig verhindert werden können, jedoch Ressourcen freigesetzt werden, um risikobehaftete Transaktionen besser zu prüfen.

Hochgeladen von

9dv4jq5g8w
Copyright
© All Rights Reserved
Wir nehmen die Rechte an Inhalten ernst. Wenn Sie vermuten, dass dies Ihr Inhalt ist, beanspruchen Sie ihn hier.
Verfügbare Formate
Als PDF, TXT herunterladen oder online auf Scribd lesen

BUSINESS PROCESS ANALYSIS AUTOMATION

FOR FINANCIAL AUDITS

A design science-oriented approach to support internal and external auditors in pro-


cess audits by using process mining techniques

Kumulative Dissertation

von

Michael Werner

zur Erlangung des akademischen Grades


eines Doktors der Wirtschafts- und Sozialwissenschaften
(Dr. rer. pol.)
der Fakultät für Betriebswirtschaft
der Universität Hamburg

Hamburg

Oktober 2014

Tag der Einreichung: 7. Oktober 2014


Tag der Annahme: 15. April 2015
Tag der Disputation: 15. April 2015

Vorsitzender: Prof. Dr. Mark Heitmann


Erstgutachter: Prof. Dr. Markus Nüttgens
Zweitgutachter: Prof. Dr. Tilo Böhmann

I
Foreword

This thesis presents the results of a challenging research endeavor that lasted over
several years and stretched out to numerous different countries that were visited for
discussion, reflection and knowledge exchange with international colleagues. The
thesis presents innovative solutions to automate the analysis of business processes
in the context of financial audits by using process mining techniques. It is set up as a
cumulative dissertation that consists of ten published scientific papers. 28 high qual-
ity reviews from international fellow researchers have been received in the past years
for the submitted papers. These reviews have contributed significantly to the pro-
gress of the presented research and several reliable relationships to experts in the
domain of process mining and compliance checking have been established across the
world. The preparation of the research was extremely work intensive and could not
have been accomplished without the support of my family, friends and colleagues. I
would like to thank my wonderful wife for her everlasting support. Without her mo-
tivation and encouragement I would never have been able to put the necessary ded-
ication into the presented work as it has been the case. She guided me like a warming
light through the ups and downs of this academic voyage. I would also like to thank
my family and especially my parents who have always provided good advice and as-
sistance. I would like to devote special thanks to my sister who reviewed my papers
sometimes even in night work to ensure that my use of the English language is ac-
ceptable for international standards. I would also like to thank my PhD supervisors
Prof. Dr. Markus Nüttgens and Prof. Dr. Horst Zündorf for providing advice during my
PhD studies and especially Prof. Dr. Nüttgens who has always supported my partici-
pation in scientific conferences and exchange with international experts. This thesis
would not have been written if the related research project had not been initiated by
Prof. Dr. Nick Gehrke who also helped me to manage the sometimes difficult entry
into the academic career and community. I would like to thank him for his assistance
and fruitful discussions. Many persons have contributed to the success of the re-
search that is presented in this thesis. My PhD colleagues Niels Müller-Wickop and
Martin Schultz have been valuable companions in our common research project as
well as my colleagues at our department at the Business School of the University of
Hamburg and from the Department of Informatics at the Nordakademie. I would like
to devote special thanks to Boris Böttcher for reviewing the final thesis and many
thanks also to the many unknown reviewers who dedicated big amounts of their time
to prepare constructive feedbacks that could be used to improve the results that are
presented in this thesis.

For my parents Sylvia and Hans-Joachim

II
I. Abstract German
Enterprise Resource Planning (ERP) Systeme sind in modernen Organisationen heut-
zutage integraler Bestandteil zur Unterstützung und Automatisierung von Geschäfts-
prozessen. Unternehmen veröffentlichen Geschäftsberichte, um verschiedene Inte-
ressensgruppen über ihre wirtschaftliche und finanzielle Lage zu informieren. Eine
wesentliche Datenquelle für die Erstellung dieser Berichte sind die Daten, die in ERP
Systemen erzeugt werden. Aufgrund ihrer wesentlichen Rolle für das Wirtschaftssys-
tem werden Jahresabschlussberichte von Wirtschaftsprüfern geprüft. Bilanzskandale
der vergangenen Jahre haben gezeigt, dass Wirtschaftsprüfer nicht in der Lage waren,
diese zu verhindern oder zumindest Verstöße frühzeitig aufzudecken. Ein wichtiger
Bestandteil der Prüfung von Jahresabschlüssen ist die Prüfung von Geschäftsprozes-
sen und relevanten internen Kontrollen. Die Prüfung von Geschäftsprozessen ist in
der Annahme begründet, dass wohlkontrollierte Geschäftsprozesse zu vollständigen
und richtigen Buchungseinträgen auf den Finanzkonten führen. Trotz der zunehmen-
den Integration von Informationstechnologie für die Automatisierung von Geschäfts-
prozessen verwenden Wirtschaftsprüfer weiterhin vorwiegend traditionelle und ma-
nuelle Prüfungsprozeduren, um die Prüfungen durchzuführen. Diese Prozeduren sind
zeitintensiv und fehleranfällig. Das Ergebnis ist ein Ungleichgewicht zwischen auto-
matisierter Transaktionsverarbeitung auf Seiten der Unternehmen und manuellen
Prüfungsprozeduren auf Seiten der Wirtschaftsprüfer, welches zu ineffizienten oder
ineffektiven Prüfungen führt. Der Einsatz von automatisierten Prüfungsprozeduren
würde dieses Ungleichgewicht reduzieren. Diese Dissertation folgt einem gestal-
tungsorientierten Forschungsansatz und stellt unterschiedliche Artefakte vor, die es
erlauben, Geschäftsprozessmodelle automatisiert mit Hilfe von Process Mining Tech-
niken zu generieren. Solche Techniken analysieren die Ereignisdaten, die im Zuge der
Transaktionsverarbeitung erzeugt werden. Diese Arbeit beschreibt als wesentliches
Ergebnis einen Multilevel Process Mining (MLPM) Algorithmus, der speziell für den
Einsatz in Jahresabschlussprüfungen entwickelt wurde. Verschiedene weitere Arte-
fakte, die in dieser Arbeit präsentiert werden, sind jedoch auch für andere Einsatz-
szenarien nützlich. Der vorgestellte Algorithmus vereinigt die Kontrollfluss- und Da-
tenflussperspektive. Er arbeitet auf verschiedenen Abstraktionsebenen, erzeugt prä-
zise und passende Prozessmodelle, verarbeitet nicht-beschriftete und nicht-lineare
Ereignisdaten aus ERP Systemen als Eingabedaten und verwendet Datenbeziehun-
gen, um den Kontrollfluss abzuleiten. Sein Einsatz kann die Analyse von Geschäfts-
prozessen verbessern, die ein wichtiger Bestandteil von Prozessprüfungen ist. Die Au-
tomatisierung bestimmter Prozeduren für die Prüfung von Geschäftsprozessen wird
zukünftige Bilanzskandale wahrscheinlich nicht verhindern können. Aber der Einsatz
des vorgestellten Algorithmus kann Prozessprüfungen verbessern und Prüfungsres-
sourcen freisetzen, die derzeit für die Prüfung von Standardgeschäftsvorfällen ver-
wendet werden, um von Standardverfahren abweichende Transaktionen zu prüfen,
die in der Regel ein wesentlich höheres Risiko aufweisen als Standardtransaktionen.

III
II. Abstract English
Enterprise resource planning (ERP) systems are key components in modern organiza-
tions to support and automate the operation of business processes. Companies pub-
lish financial reports to inform stakeholders about the economic and financial perfor-
mance of the organization. A major data source for preparing these reports is the
data that is produced by ERP systems. Due to their important role in the economic
system financial reports are audited by public accountants. Accounting scandals in
recent years have shown that auditors have not been able to prevent these scandals
or at least to indicate any violations before the actual collapse. An important part of
financial audits is the audit of business processes and related internal controls. The
rationale for auditing business processes is the assumption that well-controlled busi-
ness processes lead to complete and correct postings on the financial accounts. De-
spite the increasing integration of information technology for the automation of busi-
ness processes public accountants primarily still use traditional and mostly manual
audit procedures to carry out their process audits. These procedures are time-con-
suming and error-prone. The result is an imbalance between automated transaction
processing on the companies’ side and manual audit procedures on the auditors’ side
leading to inefficient or ineffective audits. The application of automated audit proce-
dures would reduce this imbalance. This thesis follows a design science-oriented re-
search approach and introduces several artifacts that can be used to create business
process models by using process mining techniques. Such techniques analyze the
event log data that is recorded during the processing of business transactions. This
thesis presents a Multilevel Process Mining (MLPM) algorithm that has been espe-
cially tailored for the use in financial audits. Several other presented artifacts are also
useful for other application areas. The algorithm integrates the control flow and data
flow perspective. It operates on different abstraction levels, creates precise and fit-
ting process models, accepts unlabeled and non-linear event logs from ERP systems
as input, and considers data relationships to infer the control flow. Its application can
improve the analysis of business processes which is an important part in process au-
dits. The automation of certain process audit procedures will most likely not prevent
accounting scandals in the future. But it can be used to improve process audits and
to set free resources from auditing standard business transactions that can then be
spent on the auditing of non-standard transactions that commonly exhibit a much
higher risk than standard transactions.

IV
III. Content
Foreword ...................................................................................................................... II
I. Abstract German ..................................................................................................... III
II. Abstract English ......................................................................................................IV
III. Content ....................................................................................................................V
IV. List of Figures .........................................................................................................VII
V. List of Tables ........................................................................................................... IX
VI. List of Abbreviations ................................................................................................ X
1 Introduction ............................................................................................................. 1
2 Research Area and Objectives ................................................................................. 4
2.1 Financial Audits .............................................................................................. 4
2.2 Contemporary Audit Approaches and Tool Support...................................... 4
2.3 Imbalance between Automated Transaction Processing and Manual Audit
Procedures...................................................................................................... 5
2.4 Research Questions and Objectives ............................................................... 6
3 Research Structure and Methodology ..................................................................... 7
3.1 Epistemological and Ontological Orientation ................................................ 7
3.2 Research Scope and Subjects ......................................................................... 8
3.2.1 Segmentation Framework ................................................................ 8
3.2.2 Research Scope .............................................................................. 10
3.2.3 Research Subjects .......................................................................... 12
3.3 Research and Thesis Structure ..................................................................... 13
3.4 Research Methodology ................................................................................ 15
3.4.1 Analysis Methods ........................................................................... 15
3.4.2 Design Methods ............................................................................. 16
3.4.3 Evaluation Methods ....................................................................... 16
3.4.4 Method Overview .......................................................................... 17
3.5 Cumulative Publications ............................................................................... 20
4 Analysis .................................................................................................................. 23
4.1 Related Scientific Work ................................................................................ 23
4.1.1 Business Process Management ...................................................... 23
4.1.2 Process Mining ............................................................................... 24
4.2 Requirements from the Application Domain ............................................... 29
4.2.1 The Role of Business Processes in Financial Audits ....................... 29
4.2.2 Data Structure ................................................................................ 31
4.2.3 Requirements Summary................................................................. 33
5 Design..................................................................................................................... 36
5.1 Conceptual Models for Specifying the Problem Domain and Solution........ 36
5.2 CPN Specification for Integrating the Data Perspective .............................. 38
5.3 Complexity Reduction Algorithm ................................................................. 41

V
5.4 Data-dependent Sequencing Algorithm....................................................... 44
5.5 Multilevel Process Mining Algorithm ........................................................... 46
5.6 Software Prototype ...................................................................................... 48
6 Evaluation .............................................................................................................. 51
6.1 Experimental Setup ...................................................................................... 51
6.2 Simulation..................................................................................................... 52
6.3 Results Analysis ............................................................................................ 53
6.4 Requirements Fulfillment ............................................................................. 59
7 Diffusion ................................................................................................................. 59
8 Summary and Outlook ........................................................................................... 61
8.1 Summary....................................................................................................... 61
8.2 Limitations .................................................................................................... 63
8.3 Outlook and Future Research ...................................................................... 64
9 Bibliography ........................................................................................................... 65
10 Appendix A: Publications ....................................................................................... 75
10.1 Who Is Afraid of the Big Bad Wolf - Structuring Large Design Science
Research Projects ....................................................................................... 76
10.2 Process Mining ........................................................................................... 96
10.3 Potentiale und Grenzen automatisierter Prozessprüfungen durch
Prozessrekonstruktionen ......................................................................... 117
10.4 Einsatzmöglichkeiten von Process Mining für die Analyse von
Geschäftsprozessen im Rahmen der Jahresabschlussprüfung ................ 136
10.5 Business Process Mining and Reconstruction for Financial Audits ......... 153
10.6 Colored Petri Nets for Integrating the Data Perspective in Process
Audits ...................................................................................................... 169
10.7 Tackling Complexity: Process Reconstruction and Graph Transformation
for Financial Audits .................................................................................. 178
10.8 Improving Structure: Logical Sequencing of Mined Process Models ...... 194
10.9 Towards Automated Analysis of Business Processes for Financial
Audits ...................................................................................................... 210
10.10 Multilevel Process Mining for Financial Audits........................................ 226
11 Appendix B: Publication List from Literature Review .......................................... 251
12 Appendix C: Final Declaration .............................................................................. 267

VI
IV. List of Figures
Figure 1 Imbalance Between Automated Transaction Processing and Manual Audit
Procedures in Financial Audits .................................................................. 6
Figure 2 Segmentation Framework (adapted from Werner et al., 2014, p. 8) ............ 9
Figure 3 Research Scope (adapted from Werner et al., 2014, p. 12) ......................... 10
Figure 4 Horizontal Abstraction Levels (Weske 2012 p. 76)....................................... 12
Figure 5 Research Structure ....................................................................................... 14
Figure 6 Research Phases, Artifacts and Thesis Structure.......................................... 14
Figure 7 Integrating the Application Domain and Knowledge Base in DSR (adapted
from Hevner et al., 2004, p. 80) .............................................................. 23
Figure 8 Distribution of Worldwide Process Mining Related Scientific Publications . 25
Figure 9 Distribution of Process Mining Related Scientific Publications in Europe ... 26
Figure 10 Audit Process (adapted from Werner and Gehrke, 2011, p. 105) ............. 29
Figure 11 Example Model of a Purchase Process (Werner and Gehrke, 2015, p.
823) .......................................................................................................... 30
Figure 12 Audit Procedures (adapted from Werner and Gehrke, 2011, p. 105) ....... 30
Figure 13 ERM for Accounting Data Structure (Werner, 2013, p. 389) ..................... 32
Figure 14 Source Data Structure for a Purchase Process Instance (adapted from
Gehrke and Müller-Wickop, 2010, p. 7) .................................................. 33
Figure 15 Design in the Information Systems Research Framework (adapted from
Hevner et al., 2004, p. 80) ....................................................................... 36
Figure 16 Metaphor Illustrating the Relationship between Business Processes,
Financial Accounts and Application Controls (adapted from Werner et
al., 2012a, p. 5356) .................................................................................. 37
Figure 17 Conceptual Solution ................................................................................... 38
Figure 18 Simple Example of a Purchase Process Using CPN (Werner, 2013, p.
392) .......................................................................................................... 41
Figure 19 Example A: Normal Process Instance (Werner et al., 2012b, p. 4) ............ 42
Figure 20 Example B: Large Process Instance (Werner et al., 2012b, p. 5) ................ 42
Figure 21 Example C: Monster Process Instance (Werner, 2012, p. 211)22 ............... 42
Figure 22 Aggregated Process Instance A (Werner et al., 2012b, p. 8)...................... 44
Figure 23 Logical Dependency (Werner and Nüttgens, 2014, p. 3892) ..................... 45
Figure 24 Clearing Deadlock (Werner and Nüttgens, 2014, p. 3892) ........................ 45
Figure 25 Deadlock Resolution (Werner and Nüttgens, 2014, p. 3893) .................... 46
Figure 26 Example of a Mined Process Model (adapted from Werner and Gehrke,
2015, p. 828) ............................................................................................ 48

VII
Figure 27 Prototype Structure .................................................................................... 49
Figure 28 Prototype GUI Main Screen ........................................................................ 50
Figure 29 Prototype GUI Configuration Screen .......................................................... 50
Figure 30 Evaluation in the Information Systems Research Framework (adapted
from Hevner et al., 2004, p. 80) .............................................................. 51
Figure 31 Experimental Setup .................................................................................... 52
Figure 32 Example Process Instance before Simulation ............................................ 53
Figure 33 Example Process Instance after Simulation ............................................... 53
Figure 34 Data Set 1 Distribution of Number of Net Elements over the Number of
Instances (Werner et al., 2013, p. 382) ................................................... 54
Figure 35 Data Set 2 Distribution of Number of Net Elements over the Number of
Instances (Werner et al., 2013, p. 382) ................................................... 54
Figure 36 Data Set 3 Distribution of Number of Net Elements over the Number of
Instances (Werner et al., 2013, p. 382) ................................................... 55
Figure 37 Data Set 1 Distribution of the Number of Instances for Different
Transaction Code Combinations (Werner et al., 2013, p. 383) ............... 55
Figure 38 Data Set 2 Distribution of the Number of Instances for Different
Transaction Code Combinations (Werner et al., 2013, p. 384) ............... 56
Figure 39 Data Set 3 Distribution of the Number of Instances for Different
Transaction Code Combinations (Werner et al., 2013, p. 384) ............... 56
Figure 40 Data Set 1 Distribution of the Number of Instances over Number of
Accounts (Werner et al., 2013, p. 386).................................................... 57
Figure 41 Data Set 2 Distribution of the Number of Instances over Number of
Accounts (Werner et al., 2013, p. 386).................................................... 57
Figure 42 Data Set 3 Distribution of the Number of Instances over Number of
Accounts (Werner et al., 2013, p. 386).................................................... 57
Figure 43 Frequency Distribution for Data Set 3 (Werner and Gehrke, 2015, p.
829) .......................................................................................................... 58
Figure 44 Scatter Diagram for Data Set 326 (Werner and Gehrke, 2015, p. 829)....... 58
Figure 45 Frequency Distribution for Data Set 4 (Werner and Gehrke, 2015, p.
829) .......................................................................................................... 58
Figure 46 Scatter Diagram for Data Set 426 (Werner and Gehrke, 2015, p. 829)....... 58

VIII
V. List of Tables
Table 1 Research Methods ......................................................................................... 17
Table 2 Overview of Research Artifacts, Research Methods, Diffusion Types and
Research Segments.................................................................................. 18
Table 3 Overview of Cumulative Publications ............................................................ 22
Table 4 Distribution of Process Mining Related Scientific Publications in the World 26
Table 5 Distribution of Process Mining Related Scientific Publications in Europe .... 27
Table 6 Process Mining Tools ..................................................................................... 28
Table 7 Event Log Structure (Gehrke and Werner, 2013, p. 935) .............................. 31
Table 8 CPN Specification for Process Mining in Financial Audits (Werner, 2013, p.
390) .......................................................................................................... 39
Table 9 Net Characteristics for Process Instance Examples ....................................... 42
Table 10 Prototype Modules ...................................................................................... 48
Table 11 Evaluation Data Sets (adapted from Werner et al., 2013, p. 381) .............. 54
Table 12 Overview of Net Size Distribution Characteristics (Werner et al., 2013, p.
383) .......................................................................................................... 55
Table 13 Overview of Account Distribution Characteristics (Werner et al., 2013, p.
385) .......................................................................................................... 56
Table 14 Requirements Overview .............................................................................. 59
Table 15 Overview Conferences and Presented Papers ............................................ 60

IX
VI. List of Abbreviations
ACE Automated Controls Evaluator
ACL Audit Command Language
ASC Accounting Standards Codification
AS Auditing Standard
BI Business Intelligence
BPM Business Process Management
BPMN Business Process Model and Notation
CAATs Computer-Assisted Audit Techniques
CORE Computing Research & Education
CPN Colored Petri Nets
DAAD Deutscher Akademischer Austauschdienst
DFG Deutsche Forschungsgemeinschaft
DRSC Deutsches Rechnungslegungs Standards Committee
DSR Design Science Research
EPCs Event Driven Process Chains
ERA Excellence in Research for Australia
ERM Entity-Relationship-Model
ERP Enterprise Resource Planning
FASB Financial Accounting Standards Board
FPN Financial Petri Net
GAAP Generally Accepted Accounting Principles
GUI Graphical User Interface
HTD Human-Technical Dimension
IAASB International Auditing and Assurance Standards Board
IASB International Accounting Standards Board
IDE Integrated Development Environment
IDEA Interactive Data Extraction and Analysis
IDW Institut der Wirtschaftsprüfer in Deutschland e.V.
IDW-PS IDW-Prüfungsstandard
IFAC International Federation of Accountants
IFRS International Financial Reporting Standards

X
ISA International Standards on Auditing
MLPM Multilevel Process Mining
PN Petri Nets
PCAOB Public Company Accounting Oversight Board
RCD Research Contribution Dimension
RQD Research Question Dimension
SOX Sarbanes-Oxley-Act
SQL Structured Query Language
US-GAAP United States Generally Accepted Accounting Principles
VHB Verband der Hochschullehrer für Betriebswirtschaft e.V.

XI
INTRODUCTION

1 Introduction
Enterprise resource planning (ERP) systems are key components in modern organizations.
They are a specific type of enterprise information systems that are primarily used to sup-
port and automate the operation of business processes. As they process business transac-
tions they also create data that provides information about the economic and financial per-
formance of an organization. Companies use this data source to prepare financial reports.
Such reports get published as financial statements in periodic time intervals. These reports
play a critical role for the smooth functioning of our economic system because they are an
important prerequisite for stakeholders to direct their decisions. Due to their informative
significance governments and regulatory institutions have issued laws and regulations that
intend to safeguard the correctness of published financial statements (e.g. Deutscher Bun-
destag, 2013, para. 316–324; United States Congress, 2012a, 2012b). They have entrusted
public accountants to carry out audits for ensuring that accounting standards are adhered
to and that the published information is free of material misstatements.
Accounting scandals in recent years (e.g. Enron 2001, MCI WorldCom 2002, Parmalat 2003,
Lehman Brothers 2008, Fannie Mae 2008, Satyam 2009, Olympus 2011 or HRE 2011) have
shown that auditors have not been able to prevent these scandals or at least to indicate
any violations before the actual collapse. This raises the question of how auditors can be
supported to improve their audits. An important part of financial audits is the audit of busi-
ness processes and related internal controls (IDW, 2009; IFAC, 2012; United States Con-
gress, 2002). The rationale for auditing business processes is the assumption that well-con-
trolled business processes lead to complete and correct postings on the financial accounts.
Companies use information systems to support or automate the operation of their business
processes. These produce increasing data volumes and become more and more complex
with the increasing integration of information technology. Despite this technological pro-
gress public accountants primarily still use traditional and mostly manual audit procedures
to carry out their process audits. These procedures become inefficient or ineffective in au-
dit environments that are characterized by a high integration of information systems for
the automation of transaction processing. The result is an imbalance between automated
transaction processing on the companies’ side and manual audit procedures on the audi-
tors’ side.
A solution to moderate this imbalance would be the application of automated audit proce-
dures. Business processes are currently audited with the use of manual audit procedures
like interviews and inspection of selected documents. These procedures are time-consum-
ing and error-prone. Interviews become ineffective if the interviewee does not have a com-
prehensive knowledge of the relevant process activities. This can be the case if certain ac-
tivities are automatically executed by an information system without any human interac-
tion. It can also be doubted that the manual inspection of a small sample of transactions is
adequate if millions or even billions of transactions are executed in a company every day
as, for example, in the telecommunication industry. The application of automated audit
procedures is always possible when the mere number of processed transactions makes
manual audit procedures inappropriate. Whenever this is the case information systems
must be involved that support or automate the processing. Otherwise the number of trans-

1
INTRODUCTION

actions could still be effectively audited with manual audit procedures. ERP systems pro-
duce data that is stored in the systems’ databases. The stored data includes the journal
entries and other logging information that can be used to reconstruct relationships be-
tween the stored data entries and their corresponding originally executed process activi-
ties. The content of the stored data in ERP systems that is related to financial transactions
is generally suitable for automated audit procedures if such techniques are able to analyze
and exploit the available data source.
Data analysis techniques that can be used for this purpose have to fit the specific require-
ments that are relevant for financial audits. This thesis deals with the question as to how
data analysis techniques can be used in the context of financial audits to automate the
analysis of business processes. The analysis of relevant business processes is the first step
in a process audit to get an understanding about its structure, the data volumes that are
processed and how the business processes relate to the financial accounts. The research
that is presented in this thesis is embedded in a broader research project that lasted several
years with the objective to develop tools and methods for the automation of process audits
in financial audits.1 The presented research focusses on the discovery and analysis of pro-
cess models by using process mining techniques. The source data are the journal entries
that are stored in ERP systems. The research is complementary to research efforts made
by other researchers that cover the adequate representation of information in process
models for the purpose of financial audits (Müller-Wickop, 2014) and the integration of
internal controls (Schultz, 2015).
The main challenges for the development of the presented results are the specific require-
ments from the application domain that have to be accounted for as well as the complexity
and volume of the source data. Process Mining is a research field that has emerged in the
late 1990s. Research on process mining has matured in the past decade with the develop-
ment of powerful general purpose mining algorithms that use simple deterministic (van der
Aalst et al., 2004), heuristic (Günther and van der Aalst, 2007; Weijters et al., 2006) or ge-
netic (de Medeiros, 2006) approaches. A fundamental challenge in process mining is to
balance competing model criteria (van der Aalst et al., 2012, sec. C6; van der Aalst, 2011a,
chap. 5.4.3) such as fitness, precision, simplicity and generalizability (Rozinat, 2007; Rozinat
et al., 2008). Auditors require highly fitting and precise process models to avoid false neg-
ative and false positive audit results. The source data that is available in ERP systems and
that can be used for financial audits is different compared to traditional event logs that are
used for process mining because it is unlabeled and not linear (Werner and Nüttgens,
2014). A further challenge is the integration of the data perspective to model the relation-
ship between process activities and financial accounts (Werner, 2013) and the ability to
inspect created models at different levels of abstraction to trace a data entry from its point
of origin to the final output on the financial accounts. The data perspective has widely been
neglected in the academic community so far (de Leoni and van der Aalst, 2013; Stocker,
2012) and contemporary general purpose mining algorithms are not designed to provide
process models at different abstraction levels.
The main research output that is described in this thesis is a special purpose process mining
algorithm that can be used in the context of financial audits. It integrates the control flow
1
Virtual Accounting Worlds sponsored by the German Federal Ministry of Education and Research (grant
number 01IS10041) (University of Hamburg, 2014).

2
INTRODUCTION

and data flow perspective, operates on different abstraction levels, creates precise and fit-
ting process models, accepts unlabeled and non-linear event logs as input, and considers
data relationships to infer the control flow. The application of the introduced artifacts will
probably not prevent accounting scandals. But it would enable public accountants to audit
standard transactions very efficiently and it would free resources that can be spent on au-
diting non-standard and high risk transactions and thus improve the overall audit process
and results.
The research presented in this thesis follows a design science research approach (DSR). This
approach was chosen because of the proximity of the investigated research questions to
practical problems and the intention to develop artifacts that have a high contribution and
relevance for the application domain. DSR commonly consists of four different research
phases: analysis, design, evaluation and diffusion (Österle et al., 2010). The structure of this
thesis follows these phases. It starts with a discussion of the research area and objectives
in chapter 2 and the research structure and methodology in chapter 3. The different re-
search phases were conducted in iterative cycles resulting in different research artifacts
that were created in the overall research effort. The outputs of one iteration served as the
input for the next iteration until a satisfying mining algorithm was designed. The analysis,
design and evaluation results of each cycle are described in the corresponding chapters 4,
5 and 6. Chapter 4 deals with the analysis phase. It provides an overview of related scientific
work in chapter 4.1 and the results of the requirement analysis in chapter 4.2. Chapter 5
describes the design phase and the different research artifacts that have been developed.
Each main artifact is described in one of the chapters 5.1 to 5.6. Chapter 6 illustrates the
different evaluation approaches and results that have been achieved with the description
of the general experimental setup in chapter 6.1, the discussion of conducted simulations
in chapter 6.2 and the presentation of the analysis results on mined process models in
chapter 6.3. Chapter 6.4 summarizes how the requirements identified in the analysis phase
have been satisfied. This thesis has been prepared as a cumulative dissertation. The re-
search results have been published in ten scientific articles that form the core of this thesis.
They are included in appendix A. Chapter 7 illustrates how these publications contributed
to the diffusion of the research results as the last phase of design science-oriented re-
search. The thesis closes with a summary in chapter 8.1, a discussion on identified limita-
tions in chapter 8.2 and an outlook to future research in chapter 8.3.

3
RESEARCH AREA AND OBJECTIVES

2 Research Area and Objectives


2.1 Financial Audits
Financial audits are important for the smooth functioning of economic markets. They are a
control mechanism to prevent the publication of false financial information. The reliability
of published financial statements is crucial for stakeholders like shareholders, creditors, tax
and regulatory authorities, employees, clients, financial analysts or competitors (Küting
and Reuter, 2004) to direct their decisions. Organizations are legally obliged to prepare
their financial statements fairly and truthfully. The requirements are specified in national
laws such as the Handelsgesetzbuch (Deutscher Bundestag, 2013) in Germany or the Secu-
rities Act and the Securities Exchange Act (United States Congress, 2012a, 2012b) in the
United States of America. Companies are required to apply generally accepted accounting
principles (GAAP) and accounting standards when preparing their financial statements. This
ensures that the published statements display correct and comparable information to the
addressees. National governments have entrusted specific entities to issue accounting
standards. Such standard setting bodies are the International Accounting Standards Board
(IASB) on the international level or the Deutsches Rechnungslegungs Standards Committee
(DRSC) in Germany and the Financial Accounting Standards Board (FASB) in the United
States of America on the national level. They issue the International Financial Reporting
Standards (IFRS), the Deutsche Rechnungslegungs Standards (DRS) and the FASB Account-
ing Standards Codification (ASC). In order to protect addressees from misinformation the
financial statements are audited by public accountants. The obligation to engage registered
public accountants for the auditing of financial statements is generally mandated by law.
Public accountants act as referees ensuring the adherence to laws, regulations and ac-
counting standards by auditing financial statements. They assess if the audited statements
give a fair and true view of the financial situation of a company and if the statements are
free of material misstatements. Standards on auditing provide guidance for conducting fi-
nancial audits. They are issued by regulatory bodies such as the International Auditing and
Assurance Standards Board (IAASB) for the International Standards on Auditing (ISA), the
Institut der Wirtschaftsprüfer (IDW) for the IDW-Prüfungsstandards (IDW-PS) or the Public
Company Accounting Oversight Board (PCAOB) for the Auditing Standards (AS). Laws, reg-
ulations and standards differ between countries. But a convergence has been taken place
over recent years between internationally significant accounting frameworks (Schipper,
2005).

2.2 Contemporary Audit Approaches and Tool Support


Auditing standards require the application of a risk based audit approach that takes into
account the internal control framework over relevant business processes and underlying
information systems. The ISA 315 states that:
“The objective of the auditor is to identify and assess the risks of material mis-
statement, whether due to fraud or error (…) through understanding the entity
and its environment, including the entity’s internal control (…)” (IFAC, 2012,
sec. 3).
Business processes play a significant role in financial audits. The basis for the consideration
of business processes is the assumption that well-controlled business processes lead to

4
RESEARCH AREA AND OBJECTIVES

correct entries in the financial accounts. It is much more efficient to investigate the struc-
ture of business processes and related embedded controls than to inspect single transac-
tions. ISA 315 explicitly refers to business processes and related information systems: “The
auditor shall obtain an understanding of the information system, including the related busi-
ness processes, relevant to financial reporting (…)” (IFAC, 2012, sec. 18). The focus on busi-
ness processes and internal controls is a consequence of several accounting scandals at the
beginning of the new millennium that resulted in legislation such as the Sarbanes-Oxley-
Act (SOX) (United States Congress, 2002). Similar requirements can be found in national
audit standards such as the IDW-PS 261 (IDW, 2009) or the Auditing Standard No. 12
(PCAOB, 2010).
Primary audit procedures to test internal controls and relevant business processes are in-
quiries of staff members, observation, inspection of documents or reports and the tracing
of transactions through the information system relevant to financial reporting (IFAC, 2012,
sec. A74). ISA 330 also states that computer-assisted audit techniques (CAATs) may be used
to obtain additional evidence. CAATs are software tools that assist the auditor. Braun and
Davis discuss that one possible interpretation of CAATs refers to “any use of technology to
assist in the completion of an audit” (Braun and Davis, 2003, p. 726). This is a very wide
definition and it would mean that simple project management and documentation tools
could also be considered as CAATs. It is therefore appropriate to limit the use of the term
to “tools and techniques employed to audit computer applications and to tools and tech-
nique that extract and analyze data from computer applications” (Braun and Davis, 2003,
p. 726). Examples of commercial CAATs are ACL (Audit Command Language) or IDEA (Inter-
active Data Extraction and Analysis). Accounting firms like PwC also use proprietary soft-
ware like the PwC SAP ACE (Automated Controls Evaluator). Such tools support certain au-
dit procedures and provide functionality for data queries, sample extractions and statistical
analysis or specific purposes like user access or segregation of duties analysis. The usage of
CAATs is relatively low (Bierstaker et al., 2014) and they only support certain specific tasks
in financial audits. None of currently available CAATs support process audits in a holistic
manner.

2.3 Imbalance between Automated Transaction Processing and Manual Au-


dit Procedures
Accounting scandals in recent years (e.g. Enron 2001, MCI WorldCom 2002, Parmalat 2003,
Lehman Brothers 2008, Fannie Mae 2008, Satyam 2009, Olympus 2011 or HRE 2011) have
shocked the economic markets. Public accountants have not been able to prevent one of
these scandals or were able at least to reveal any violation before the actual collapse. Com-
panies use information systems to automate the processing of business transactions result-
ing in increasingly complex processes and enormous data volumes that exhibit character-
istics of Big Data (Chen et al., 2012) in terms of velocity of accumulations and volume.

5
RESEARCH AREA AND OBJECTIVES

The task of auditing business processes is getting more and more challenging with the in-
creasing integration of information systems for the automation of transaction processing
and the growing amount of produced data. Traditional audit procedures like interviews and
inspections of selected documents become inefficient or even ineffective in such audit en-
vironments (Werner and Gehrke, 2011). Interview partners may no longer have overall in-
formation about a business process if parts of it are operated in an automated way and
nontransparent in the information system without any human interaction. Furthermore it
is questionable if the in-
spection of relatively few Company Auditor
samples is a sufficient au- Perspective Perspective
dit procedure when mil- Discrepancy
Automated Manual
lions of transactions are Transaction Audit
processed. Software Processing Procedures
tools are rarely used and
– Prepares financial reports – Issues an opinion in
only for very specific using the data that is created regard to the truth and
tasks. This situation leads in information systems during fairness of the financial
transaction processing statements
to an imbalance between
– Processing is automated and – Uses primarily traditional
automated transaction creates increasingly large data and manual audit
procedures
processing on the compa- amounts

nies’ side and manual au-


dit procedures on the au- Imbalance between automated processing on the
companies’ side and manual audit procedures on
ditors’ side which is illus- the auditors’ side
trated in Figure 1. The Inefficient or even ineffective audits
outcomes are inefficient
or even ineffective au- Figure 1 Imbalance Between Automated Transaction Pro-
dits. cessing and Manual Audit Procedures in Financial Audits

2.4 Research Questions and Objectives


A solution to reduce the imbalance between automated transaction processing on the
companies’ side and manual audit procedures on the auditors’ side would be the applica-
tion of automated audit procedures that exploit the data that is created and recorded in
the course of transaction processing. The advantage of this data is that it is recorded auto-
matically and independently from any involved person and therefore provides a much
more reliable data source compared to the data sources that are used in traditional audit
procedures (Jans et al., 2010). Automated audit procedures would have to be able to oper-
ate with the recorded data and comply with the specific requirements that generally have
to be accounted for in financial audits. The overall research question that has to be an-
swered for this aim can be formulated as followed:
Overall research question: How can data analysis techniques be used to auto-
mate the audit of business processes in the context of
financial audits?
The overall research question is very broad and has been addressed by different research-
ers that focused on different research aspects to find solutions to this question in a com-
mon research effort. This thesis deals with the aspect of creating and analyzing business
models by exploiting the source data which is stored in information systems that support

6
RESEARCH AREA AND OBJECTIVES

or automate the processing of business transactions. The creation of process models from
the source data is a core requirement for the development of automated audit procedures
that support the audit of business processes in a holistic manner. The specific research
question that is addressed in this thesis can be phrased as followed:
Specific research question: How can reliable process models be automatically re-
constructed by analyzing data stored in information
systems that process financially relevant transactions?
The objective of the described research is the development of research artifacts in form of
construct, models, methods and instantiations that can be used to create and analyze pro-
cess models on the basis of financially relevant data recorded in the source systems. Pro-
cess mining is a business intelligence approach that provides powerful mining algorithms
that are able to create process models from recorded event logs (van der Aalst, 2011a). It
therefore serves as a starting point and important knowledge base for the research pre-
sented in this thesis.

3 Research Structure and Methodology


3.1 Epistemological and Ontological Orientation
The research approach and methodology for each research effort has to fit the intended
purpose. Different paradigms and approaches prevail in the information systems research
community. A fierce controversy has taken place in the international information systems
research community in recent years between advocates of different research paradigms
(Baskerville et al., 2010; Österle et al., 2010). Two differing research paradigms in infor-
mation systems research are positivism and interpretivism. They differ in terms of the on-
tological and epistemological assumptions (Niehaves, 2007). Interpretivist research focus-
ses on the perception of humans and the individuality of cognition. The basic ontological
assumption is that a single objective world might indeed not exist and that hence objective
cognition is impossible. Positivist research prevails in natural sciences and it is character-
ized by the ontological assumption that a real world exists that can be described unambig-
uously in combination with the epistemological assumption that objective cognition is
hence possible. Social or behavioral science research tends to be descriptive in nature and
to follow an interpretivist paradigm. Behavioral science research and interpretivist-ori-
ented research methods are especially prevalent in the Anglo-Saxon research community
(Palvia et al., 2004, 2003; Schauer, 2011; Wilde and Hess, 2007). Design science research is
a research approach that has gained increased attention over the last decade in the inter-
national scientific community as an important research approach (Vaishnavi and Kuechler,
2004). It is normative in nature and DSR researchers tend to follow a positivist research
paradigm. DSR is prevalent in German speaking countries (Wilde and Hess, 2007) where
information systems research has traditionally been closely related to the natural and en-
gineering sciences (Schauer, 2011) but also an important research approach in the Anglo-
Saxon research community (Baskerville et al., 2010). DSR is characterized by the duality of
the epistemological and design objective (Riege et al., 2009). It intends to create research
artifacts in forms of constructs, models, methods (March and Smith, 1995) and theories
(Gregor, 2006; Gregor and Hevner, 2013) that on the one hand contribute to the scientific

7
RESEARCH STRUCTURE AND METHODOLOGY

knowledge base and that are on the other hand useful for the application domain (Hevner
et al., 2004).
The research presented in this thesis intends to develop solutions to improve the audit of
business processes in financial audits. The research questions are closely related to chal-
lenges in practice. DSR is therefore perceived as an adequate research approach to develop
the intended research outputs. Several scholars have pointed out that focusing on a single
research approach or research method might lead to paradigmatic bias (Mingers, 2001)
and that an interaction between DSR and social science research is necessary to achieve
progress in the scientific information systems discipline (Gregor and Baskerville, 2012). We
therefore relied on a variety of research methods with differing paradigmatic orientation
to prevent any methodological or paradigmatic bias and to benefit from observing the phe-
nomena under investigation from different scientific angles. Detailed information on the
research structure and methodology is described in chapters 3.3 and 3.4.

3.2 Research Scope and Subjects

3.2.1 Segmentation Framework


The presented research was conducted in the research project Virtual Accounting Worlds.
This project was sponsored by the German Federal Ministry of Education and Research
(grant number 01IS10041) and followed the objective to develop tools and methods to
support auditors in process audits in the context of financial audits (University of Hamburg,
2014). The overall research task was divided into three different research sub-scopes that
were assigned to three different researchers. A segmentation framework was developed
to confine and assign the different work tasks to the involved researchers (Werner et al.,
2014).
The framework itself was developed as part of the research project. It is used in this section
to describe the research scope that is covered by this thesis. The framework is extensively
described in (Werner et al., 2014). It consists of three dimensions that are suitable to divide
the overall research task in well-defined and manageable research segments.
The first dimension refers to the human-technical interaction. The main subjects of interest
in information systems research are information technologies and the man/machine inter-
action. The distinguishing characteristic of the information systems research discipline is
the investigation of phenomena that emerge from the interaction of humans and technol-
ogy in socio-technical systems (Gregor, 2006). Categorizations for different levels of infor-
mation technology and human interaction can be found in various models. A common
model from the field of information management is presented by Krcmar (Krcmar, 2010).
The model was adapted from (Wollnik, 1988) and distinguishes between three different
levels of management tasks. The lowest level contains tasks for the management of the
technical infrastructure that are necessary for the use of information and communication
technology at the higher levels. The second level deals with the management of infor-
mation systems and includes the management of data, processes, applications and their
lifecycles. At the highest level reside the tasks for the management of the information econ-
omy. The main objective of the tasks at this level is the management of the resource infor-
mation, its supply, demand and usage. The two lower levels are mainly concerned with the

8
RESEARCH STRUCTURE AND METHODOLOGY

technical aspects of information management. The human aspect is considered at the high-
est level where the requirements for the lower levels are defined on the basis of the needs
of human information recipients and users of the applications, processes, data and tech-
nology that are located on the lower levels. The model of information management ad-
dresses both elements that are subject to research in information system science: infor-
mation systems and human interaction in socio-technical systems. It is therefore useful for
the categorization of research content because each researched artifact in design science-
oriented research can be characterized if it addresses one or more of the different levels.
The original model represents applications, data and processes at the same level. Promi-
nent conferences in the information systems science area like the Business Process Man-
agement conference (BPM, 2014), comprehensive publications (Weske, 2012) and exten-
sive reviews (van der Aalst, 2013, 2012) show that business processes are a key component
in information systems research. We therefore considered it appropriate to divide the level
of information systems into the two levels of software applications and processes. The pro-
cess level is the connecting layer where process participants use components from the
lower application level to satisfy information demands from the higher level.
The resulting four levels of the Human-Technical Dimension (HTD) are very broad. Although
they can be used to distinguish research content on a technical vs. human-interaction di-
mension they are not sufficient to divide the research content into manageable segments
(Werner et al., 2014). Hevner et al. present a research framework for information systems
research that provides an illustration of how different concepts that are relevant for re-
search projects relate to each other (Hevner et al. 2004). It describes the relationships be-
tween the main research activities for design science (build and evaluate) and behavioral
science (develop and justify) research, the environment and the knowledge base. The en-
vironment or application domain defines the problem space. The phenomena of interest
for design science-oriented research should be derived from the environment. The
knowledge base represents the pool of already existing scientific expertise. Each research
project should take into account already
existing knowledge to execute the re-
search activities, provide useful artifacts
to the application domain and add addi-
tional generalized knowledge to the
knowledge base. The objective to contrib-
ute to the application domain on the one
hand and the scientific knowledge base
on the other hand characterizes design
science-oriented research projects (Riege
et al., 2009). The distinction between the
research contribution target domains can
be used as a second Research Contribu-
tion Dimension (RCD) for the categoriza-
tion of research content.
Each research project addresses an over-
Figure 2 Segmentation Framework
all research question. The research ques-
(adapted from Werner et al., 2014, p. 8)
tions in large research projects are, as a

9
RESEARCH STRUCTURE AND METHODOLOGY

rule, complex. Otherwise it would be debatable if such a project has the characteristics of
a large project in the first place. Complex research questions can usually be divided into
detailed lower-level research questions. These research questions can be used as a catego-
rization criterion for a third Research Question Dimension (RQD). Figure 2 shows the seg-
mentation framework with all three dimensions. It illustrates how separate segments
emerge based on the different dimension categories. Each segment can be referenced by
using its x- (RQD), y- (HTD) and z-coordinates (RCD) in the cube. The reference model shown
in Figure 2 has to be instantiated to be useful for a specific research project which is de-
scribed for the research at hand in in the next chapter.

3.2.2 Research Scope


Figure 3 shows the instantiated framework that was used to divide the overall research
task into manageable research segments and to confine the scope for the research pre-
sented in this thesis. The primary application domain for the research work at hand is fi-
nancial audits (compare chapter 2.1). It represents the first category in the Research Con-
tribution Dimension.
The thesis focusses on presenting solu-
tions for producing reliable process mod-
els automatically by analyzing data stored
in information systems that process fi-
nancially relevant transactions. Process
models are abstractions from reality that
are used to graphically represent busi-
ness processes. A business process is of a
set of related activities that are per-
formed in an organizational and technical
environment to realize a business goal
(Reichert and Weber, 2012; van der Aalst
and Stahl, 2011; Weske, 2012; Workflow
Management Coalition, 1999). “A busi-
ness process model consists of a set of ac-
tivity models and execution constraints
between them.” (Weske, 2012, p. 7).
Modeling is the activity that leads to the Figure 3 Research Scope (adapted from
definition of models (Hansen and Neu- Werner et al., 2014, p. 12)
mann, 2009). Process models are usually
defined by using modeling languages like Business Process Model and Notation (BPMN),
Event Driven Process Chains (EPCs) or Petri Nets. Business process modeling is a fundamen-
tal part of Business Process Management (BPM). It is therefore perceived as one important
scientific knowledge base.
The fundamental idea to develop the solutions presented in this thesis is the application of
process mining techniques. Process mining allows creating process models from recorded
event log data (van der Aalst, 2011a). It is a Business Intelligence (BI) approach that bridges
the gap between BPM and BI (van der Aalst, 2011b). BI can be seen as an amalgamation of
business economics, operations research, data mining, statistics and data warehousing

10
RESEARCH STRUCTURE AND METHODOLOGY

(Müller and Lenz, 2013) and it is traditionally used to support decision making processes
(Turban et al., 2007; Vercellis, 2009). The aim of process mining as an BI approach is to
discover, analyze and enhance process models (van der Aalst et al., 2012). The segmenta-
tion framework was used to define and structure the research task for the overall research
project. This thesis focusses on process mining techniques but it also relates to other data
analysis techniques that were used to statistically analyze the source data and produced
process models. In order not to be too restricted and also to account for the other research
tasks from the overall research project that are not part of this thesis BI was chosen as the
more general term as a relevant scientific knowledge base. BPM and BI in combination form
the second category on the Research Contribution Dimension as illustrated in Figure 3.
The categories for the third Research Question Dimension were derived from breaking the
overall research question (compare chapter 2.4) down into three detailed sub-questions.
The detailed, lower-level research questions were labeled with the keywords reconstruc-
tion, assessment and visualization. They were formulated by logically reasoning what kind
of questions on a more specific level had to be answered to find solutions for the overall
research question. They were formulated as follows:
(1) Reconstruction: How can reliable process models be automatically recon-
structed by analyzing data stored in information systems
that process financially relevant transactions?
(2) Assessment: How can process models be automatically assessed from
an audit perspective by integrating control data that is
stored in the source systems?
(3) Visualization: How can process models be graphically represented to dis-
play information that is relevant to auditors and that can
be applied in real audit environments?
Figure 3 shows the precise scope of the research work that is presented in this thesis. The
relevant research segments are highlighted in red. It is necessary to consider the applica-
tion level (segments (3,2,1) and (3,2,2)) because software applications like ERP systems that
process financially relevant transactions provide the source data that is necessary to recon-
struct process models. The focus of the research lies on the reconstruction (segments
(3,3,1) and (3,3,2)) and analysis (segment (2,3,1)) of process models for financial audits.
The research scope on the horizontal process-level stretches into the vertical assessment
category because the analysis of reconstructed process models already provides infor-
mation that is useful for the assessment of business processes in the context of financial
audits. The usage-level is concerned with the demand and supply of information and its
consumption. This level is relevant to take primarily formal requirements into account that
are important for the adequacy of the produced models (segments (3,4,1) and (3,4,2)). Fig-
ure 3 shows that contributions were planned for the application domain (segments (3,2,1),
(2,3,1), (3,3,1), (3,4,1)) and the knowledge base (segments (3,2,2), (3,3,2), (3,4,2)).
The infrastructure-level (segments (x,1,z)) was out of scope because it is not necessary to
consider this level in order to achieve the research objective as all necessary source data is
provided on the application-level. Aspects of process visualization from a usability and end-
user perspective (segments (1,y,z)) are not covered in this thesis and neither are aspects

11
RESEARCH STRUCTURE AND METHODOLOGY

for the assessment of business processes that relate to internal controls (segments (2,y,z)
except (2,3,1)).

3.2.3 Research Subjects


The research subjects are the data that is produced by information systems during the
course of processing financially relevant transactions and the created process models.
The presented research focusses on data produced by ERP systems because they are the
predominant type of information systems that are used to support and automate the pro-
cessing of business transactions (Konradin Mediengruppe, 2011). The produced data from
ERP systems is generally suitable for process mining purposes (van der Aalst et al., 2012).
The term model has many different mean-
ings depending on the context and scien- M3: Meta-Metamodel
tific discipline. A model according to the
Instance of describes
interpretation prevailing in information
systems science is a simplified image of a M2: Metamodel Notation
selected part of reality (Heinrich et al.,
2004) that serves a specific purpose Instance of describes
expresses
(Becker et al., 2012). Model and reality are
related to each other (model-relation) M1: Model
which means that it is possible to conclude Instance of describes
from observed model characteristics on
reality and vice versa (isomorphism-rela- M0: Instance
tion) (Heinrich et al., 2004).
Figure 4 illustrates that models in the con- Figure 4 Horizontal Abstraction Levels
text of BPM can be differentiated depend- (Weske 2012 p. 76)
ing on the level of horizontal abstraction
(Weske, 2012). Models (M1) are created by using constructs that are defined by metamod-
els (M2). These are associated with notations that often exhibit a graphical nature. The
Petri Net metamodel (M2), for example, defines that Petri Nets (M1) consist of places and
transitions that form a directed bipartite graph. “The complete set of concepts and associ-
ations between concepts is called metamodel.“ (Weske, 2012, pp. 76–77). Constructs of a
metamodel (M2) can themselves be defined by a meta-metamodel (M3).
Process models reside on the model-level (M1) and are instances of metamodels (M2) that
are expressed using an associated modeling language like BPMN, EPCs or Petri Nets. A pro-
cess model is an abstraction of a business process and represents a set of similar executed
entities of a business process. A process instance model (M0) is an abstraction of a single
executed entity of a process model.
“The instance level reflects the concrete entities that are involved in business
processes. Executed activities, concrete data values, and resources and persons
are represented at the instance level.” (Weske, 2012, p. 75).
Process models and process instance models are especially relevant for the investigated
research question. Process models provide an overview of the structure of a business pro-
cess. Process instance models provide detailed information for each single execution of a

12
RESEARCH STRUCTURE AND METHODOLOGY

business process that is relevant from an audit perspective to inspect individual transac-
tions and associated data values. The two types of models are on the one hand the primary
outputs that are produced by the designed research artifact (compare chapter 5) but on
the other hand also the subjects of investigation in the evaluation phase (chapter 6).

3.3 Research and Thesis Structure


Different research frameworks exist that provide guidance on how to conduct design sci-
ence-oriented research.2 Österle et al. suggest a generally accepted framework that divides
DSR into the phases analysis, design, evaluation and diffusion (Österle et al., 2010). The
research and thesis structure follows these four different phases. Figure 5 provides an over-
view of the main artifacts that were developed in the design phase and the research meth-
ods that were used in the different research phases. It also shows the diffusion types3 for
the diffusion phase.

2
Peffers et al. present a research methodology that consists of six phases: (1) identify problem and moti-
vate, (2) define objectives of a solution, (3) design and development, (4) demonstration, (5) evaluation
and (6) communication (Peffers et al., 2007, 2006). Their research framework is composed of process el-
ements that have been identified by different scholars working in the information systems (Cole et al.,
2005; Hevner et al., 2004; Nunamaker et al., 1991; Takeda et al., 1990; Walls et al., 1992) and engineer-
ing (Archer, 1984; Eekels and Roozenburg, 1991) discipline. Österle et al. suggest four phases for DSR: (1)
analysis, (2) design, (3) evaluation and (4) diffusion (Österle et al., 2010). Gregor and Baskerville examine
the research process from a philosophy of science perspective with the objective to provide a framework
for the combination of design science and social science research. The presented research process con-
sists of the phases (A) construct and test artefacts, (B) formulate prescriptive knowledge and theory, (C)
study artefact(s) in use, (D) test knowledge of artefacts in use and (E) formulate descriptive knowledge
(Gregor and Baskerville, 2012). All authors explicitly emphasize the iterative relationship between the
different research steps in each model. Alturki et al. present the most detailed model that consists of 14
research steps (Alturki et al., 2011).
3
The diffusion types are presented in Figure 5 to provide a complete overview of the research structure.
They cannot be considered as research methods and are therefore not described in the following chapter
3.4 but instead in chapter 7.

13
RESEARCH STRUCTURE AND METHODOLOGY

Design
Designed Artifacts
- Conceptual Models
- CPN Specification
- Complexity Reduction Algorithm
- Data-dependent Sequencing
Algorithm
- Multilevel Process Mining Algorithm
- Software Prototype
Analysis Design Methods: Evaluation
- Modeling
- Method Engineering
Analysis Methods: - Prototyping Evaluation Methods:
- Document and Data Analysis - Simulation
- Literature Review - Laboratory Experiment
- Modeling - Quantitative Analysis

Diffusion

Diffusion Types:
- Scientific Publications
- Conference Presentations

Figure 5 Research Structure


The different phases were undergone in iterative loops where the research outputs of one
research cycle formed the input of the following one. Figure 6 shows how the different
designed artifacts depend on each other and how the research structure refers to the struc-
ture of this thesis.
Chapter

Analysis 4

Conceptual Models 5.1

CPN Specification 5.2

Complexity Reduction Algorithm 5.3


Design
Input for

Data-dependent Sequencing Algorithm 5.4

Multilevel Process Mining Algorithm 5.5

Software Prototype 5.6

Evaluation 6

Diffusion 7

Research Progress

Figure 6 Research Phases, Artifacts and Thesis Structure

14
RESEARCH STRUCTURE AND METHODOLOGY

3.4 Research Methodology


The term methodology is used ambiguously in information systems science.4 We use the
term in the sense of a combination of research methods that are used for a specific re-
search project. The presented research follows as design science-oriented research ap-
proach. An important aspect when following a specific research approach is the prevention
of paradigmatic bias that might occur by focusing on a single research paradigm. We there-
fore used different research methods with diverse paradigmatic orientation to prevent
such a bias and to benefit from triangulation (Jick, 1979) and an appropriate mix between
qualitative and quantitative research methods (Venkatesh et al., 2013). Figure 5 in chapter
3.3 illustrates the variety of research methods that were used in the different research
phases. The following chapters briefly describe these methods.

3.4.1 Analysis Methods

Document and Data Analysis


The primary analysis methods include a literature review and the analysis of source data
and documents (Recker, 2012) by inspecting the data, data structure and system documen-
tation from test and productive ERP systems.

Literature Review
The literature review referred to relevant scientific publications but also to laws, regula-
tions and standards that specify the formal requirements that have to be taken into ac-
count in the context of financial audits. It was conducted following the guidelines published
by various scholars (Fettke, 2006; Rowley and Slack, 2004; vom Brocke et al., 2009; Webster
and Watson, 2002). The collection of empirical data regarding requirements directly from
the application domain was not in scope of the research presented in this thesis but was
conducted in the related research project (Müller-Wickop et al., 2013; Müller-Wickop and
Schultz, 2013a; Schultz et al., 2012; University of Hamburg, 2014). The empirical investiga-
tion included the application of qualitative research methods such as structured interviews
(Gubrium and Holstein, 2002) and the quantitative research methods in the form of surveys
(Fowler, 1984). The achieved results were incorporated in this work as part of the literature
analysis to achieve an adequate mix between requirements identified using a positivist per-
spective (objective laws, regulations and standards) and an interpretive perspective (re-
sults from interviews and surveys).

Modeling
Models are, besides constructs, methods, instantiations and theories (Gregor, 2006), one
of the main research artifacts in DSR (March and Smith, 1995). Two types of models are
used in this thesis: conceptual models to define the problem and solution space like entity-

4
Mingers identifies three widespread and different interpretations for the term methodology (Mingers,
2001).The most general meaning refers to the study of methods. The most specific meaning is related to
a particular research study and it refers to the actual research method(s) that are used in a specific piece
of research. The third interpretation is a generalization of the second. A methodology is referred to as a
specific combination of particular methods that is deliberately designed a priori in the sense of a blue-
print that occurs many times in practice.

15
RESEARCH STRUCTURE AND METHODOLOGY

relationship-models (ERM) in the analysis phase and process models as output of the de-
signed methods and software prototype in the design and evaluation phases. Modeling or
conceptual modeling is the research method that was used to create different models.
Modeling is an engineering process that creates simplified images of reality by using an
inductive approach relying on observations or a deductive approach relying on theories
(Wilde and Hess, 2006).

3.4.2 Design Methods

Method Engineering
The primary research methods for designing the research artifacts were method engineer-
ing and prototyping. Method Engineering is a commonly used research method in infor-
mation systems research (Österle et al., 2010; Wilde and Hess, 2007, 2006) for the system-
atic design of methods (Brinkkemper, 1996). A method in this context consists of different
parts (method fragments) that can be combined and reused (Harmsen et al., 1994). A new
method can be engineered by combining existing method fragments in a new manner or
by developing completely new method fragments.

Prototyping
Prototyping is a software programming approach originating from the area of software en-
gineering (Naumann and Jenkins, 1982). Prototyping consists of the phases: requirement
identification, development, implementation, revision and enhancement. The aim is to de-
velop a software artifact that implements the intended core functionality in iterative cycles.
It is a common research method in design science-oriented research (Österle et al., 2010;
Schauer, 2011; Wilde and Hess, 2007, 2006). This research method is particularly suitable
for the research at hand because it allows to design, implement and evaluate different re-
search artifacts that relate to each other as shown in Figure 6.

3.4.3 Evaluation Methods


DSR is often criticized for a lack of research rigor (Österle et al., 2010) and several scholars
highlight the importance of rigorous evaluation in design science-oriented research (Riege
et al., 2009; Venable et al., 2012). The duality of the design and the epistemological objec-
tive (Riege et al., 2009) require the consideration of evaluation methods that can be used
to verify the adequateness of the research outputs in regard to the identified research gap
and to validate if the developed solutions are able to satisfy the demands from the appli-
cation domain. Several scholars provide guidelines and frameworks for the selection of ap-
propriate research methods.
The primary evaluation methods that were used in the presented research include simula-
tion, lab experiment and quantitative analyses. Simulations and lab experiments are suita-
ble for artificial ex-post evaluations (Venable et al., 2012). Quantitative analyses using de-
scriptive statistics (Bhattacherjee, 2012) were added to describe and analyze the output
data that was generated during the lab experiments. The design and evaluation methods
are closely related. The implementation of the designed artifacts in a prototype already
serves as a proof of concept from an evaluation perspective because it allows to verify if

16
RESEARCH STRUCTURE AND METHODOLOGY

the theoretical constructs, models and methods can actually be executed. It also provides
the foundation for simulations and lab experiments.

Simulation
Simulations are goal-oriented experiments to gain information on models that are difficult
to represent as formal models due to their inherent complexity. The models that were cre-
ated by the software prototype were made of up to several hundred thousand net ele-
ments (Werner et al., 2012b). Simulations were used to verify if the produced process mod-
els satisfied the identified requirements for example in respect to specific model charac-
teristics like soundness (van der Aalst, 2011a; Weske, 2012).

Laboratory Experiment
The prototype was used in laboratory experiments. The conducted experiments differ from
traditional experiments that are commonly used in behavioral science-oriented research.
In such experiments one or more independent variables are manipulated by the researcher.
Subjects are randomly assigned to different treatment levels and the results of the treat-
ments on outcomes are observed (Bhattacherjee, 2012, chap. 10). The subject in the design
science-oriented experiments is the prototype itself and the outcomes that are observed
are the produced process models (Riege et al., 2009). The variables are different configu-
ration parameters.

Quantitative Analysis
The output that was produced by the designed artifacts were finally analyzed with the use
of descriptive statistics (Bhattacherjee, 2012; Schira, 2009). These analysis methods were
used to inspect the voluminous output data on an aggregate level and to gain insights into
the distributions and relationship between different model characteristics.

3.4.4 Method Overview


Table 1 provides an overview of the used research methods, paradigmatic orientation and
literature references that were primarily used to apply the listed methods. The table shows
that a mix between positivist and interpretive oriented research methods has been
achieved by using diverse quantitative and qualitative analysis, design and evaluation
methods.
Table 1 Research Methods
Paradigmatic
Research Method Type Primary References
Orientation
Document and Quantitative / (Bhattacherjee, 2012; Österle et al.,
Positivist
Data Analysis Qualitative 2010; Recker, 2012)
(Fettke, 2006; Rowley and Slack,
Positivist / Quantitative /
Literature Review 2004; vom Brocke et al., 2009;
interpretivist Qualitative
Webster and Watson, 2002)
(Österle et al., 2010; van der Aalst and
Stahl, 2011; Wilde and Hess, 2007,
Modeling Interpretivist Qualitative
2006)

17
RESEARCH STRUCTURE AND METHODOLOGY

(Brinkkemper, 1996; Harmsen et


Method
Positivist Qualitative al., 1994; Österle et al., 2010; Wilde
Engineering
and Hess, 2007, 2006)
(Naumann and Jenkins, 1982; Wilde
Prototyping Positivist Qualitative
and Hess, 2007, 2006)
(Riege et al., 2009; van der Aalst
Simulation Positivist Quantitative and Stahl, 2011; Wilde and Hess,
2007, 2006)
(Bhattacherjee, 2012; Österle et al.,
Laboratory
Positivist Quantitative 2010; Riege et al., 2009; Wilde and
Experiment
Hess, 2007, 2006)
(Bhattacherjee, 2012; Gujarati and
Quantitative
Positivist Quantitative Porter, 2009; Österle et al., 2010;
Analysis
Recker, 2012)

Table 2 lists the different artifacts, related research methods and diffusion types in the left
column. The diagrams in the right column illustrate how the artifacts relate to the research
segments described in chapter 3.2.2.

Table 2 Overview of Research Artifacts, Research Methods, Diffusion Types and Re-
search Segments
Research Artifacts, Research Methods and
Related Research Segments
Diffusion Types
Artifact Conceptual Models
Segments All

Research Methods
Document and Data Analysis
Analysis
Literature Review
Design Modeling
Evaluation -

Book Chapters 3 and 4


Diffusion Type
Conference Paper 5

18
RESEARCH STRUCTURE AND METHODOLOGY

Artifact CPN Specification


Segments (3,4,2)

Research Methods
Document and Data Analysis
Analysis
Literature Review
Modeling
Design
Prototyping
Laboratory Experiment
Evaluation
Simulation

Diffusion Type Conference Paper 6

Complexity Reduction Algo-


Artifacts
rithm
Segments (3,3,2)

Research Methods
Analysis Document and Data Analysis
Method Engineering
Design
Prototyping
Laboratory Experiment
Evaluation Simulation
Quantitative Analysis

Diffusion Type Conference Paper 7

Data-dependent Sequencing
Artifacts
Algorithm
Segments (3,2,2)

Research Methods
Analysis Document and Data Analysis
Method Engineering
Design
Prototyping
Laboratory Experiment
Evaluation Simulation
Quantitative Analysis

Diffusion Type Conference Paper 8

19
RESEARCH STRUCTURE AND METHODOLOGY

Multilevel Process Mining Al-


Artifact
gorithm
Segments (3,3,1),(3,3,2), (2,3,1)

Research Methods
Analysis Document and Data Analysis
Method Engineering
Design
Prototyping
Laboratory Experiment
Evaluation Simulation
Quantitative Analysis
Conference Paper 9
Diffusion Type
Journal Paper 10

Artifact Software Prototype


Segments (3,2,1),(3,3,1), (2,3,1),(3,4,1)

Research Methods
Analysis -
Design Prototyping
Laboratory Experiment
Evaluation
Simulation
Conference Paper 9
Diffusion Type
Journal Paper 10

3.5 Cumulative Publications


This thesis has been prepared as a cumulative dissertation. The research results that are
presented in this thesis have been published in different scientific articles. Table 3 provides
information on the different publications and how they relate to the different chapters in
this thesis. The first column shows the number of the paper that is used for reference pur-
poses in this thesis. The second column provides a reference to the appendix that contains
the complete originally published papers.5 The third column shows to which chapter the
publication primarily refers to. The papers generally follow the publication guidelines for
DSR (Gregor and Hevner, 2013). Each paper is a complete scientific publication and does
not refer exclusively to a single chapter but also to aspects that are mentioned in various
chapters of this thesis. The assignment to specific chapters is only intended to illustrate
which core aspects are described in the different publications and how they generally refer
to the structure of this thesis.

5
The papers have been formatted in a consistent way. References to page numbers in this thesis refer to
the originally published papers.

20
RESEARCH STRUCTURE AND METHODOLOGY

Column four shows the full title of each paper and the fifth column provides information
on the publication outlet. Two papers were published as book chapters, six in conference
proceedings and two in journals.6 The sixth column provides the reference to the entry in
the bibliography. Column seven lists the acceptance rate for scientific papers that were
provided by the conference chairs. Columns eight to eleven provide information on the
ranking of the publication outlets according to the VHB Jourqual 2.1 (Verband der
Hochschullehrer für Betriebswirtschaft e.V., 2011), WKWI (WKWI, 2008), ERA 2010 (Aus-
tralian Research Council, 2014) and CORE 2013 (CORE, 2014) rankings.
Column twelve describes the review procedure and lists how many different reviews from
fellow researchers were received during the review process. The used abbreviations have
the following meaning:
DB: Double blinded
B: Blinded
The number of authors that contributed to each paper is listed in column 13. Column 14
shows how many dissertation points can be assigned to each paper using the formula 2 /
(number of authors +1). Column 15 shows the level of authorship from the author of this
thesis that was commonly agreed among all co-authors. Two papers were prepared in sin-
gle authorship. The average authorship for the remaining papers is 91 %. None of the listed
papers is part of any other dissertation thesis.

6
Paper ten was initially submitted for the IEEE TSC journal in August 2013. It passed the first review round in
December 2013, the second in July 2015 before being published in December 2015. The included version
refers to the preprint version at the time of submission of this dissertation.

21
RESEARCH STRUCTURE AND METHODOLOGY

Table 3 Overview of Cumulative Publications

Part of other Dis-


Rankings

Primary Related

dure / Reviews
Number of Au-
Review Proce-

Dissertation
Acceptance

Authorship

sertations
Appendix

Chapters

WKWI 2008

CORE 2013
VHB JQ 2.1

Points
ERA 2010

thors
Rate
# Title Publication Outlet Reference

Who Is Afraid of the Big Bad Wolf - Struc-


22nd European Conference on Infor- (Werner et al., B
1 10.1 3 turing Large Design Science Research Pro- 34 % A A A DB (4) 4 0.40 85% No
mation Systems (ECIS) 2014) (7.37)
jects
(Gehrke and Wer- 7) E
2 10.2 4.1 Process Mining wisu - das wirtschaftsstudium - - - B (-) 2 0.67 95% No
ner, 2013) (2.86)
Potentiale und Grenzen automatisierter
4.2 (Werner and 7)
3 10.3 Prozessprüfungen durch Prozessrekon- Forschung für die Wirtschaft - - - - B (-) 2 0.67 95% No
5.1 Gehrke, 2011)
struktionen
Einsatzmöglichkeiten von Process Mining
4.2 7)
4 10.4 für die Analyse von Geschäftsprozessen Forschung für die Wirtschaft (Werner, 2012) - - - - B (-) 1 1.00 100% No
5.1
im Rahmen der Jahresabschlussprüfung
Business Process Mining and Reconstruc- 45th Hawaii International Confer- (Werner et al., C
5 10.5 5.1 54 % B A A DB (6) 3 0.50 90% No
tion for Financial Audits ence on System Sciences (HICSS) 2012a) (6.44)
Colored Petri Nets for Integrating the Data 32nd International Conference on B8)
6 10.6 5.2 (Werner, 2013) 32 %8) B8) A8) A8) DB (3) 1 1.00 100% No
Perspective in Process Audits Conceptual Modeling (ER) (7.59)
Tackling Complexity: Process Reconstruc-
33rd International Conference on In- (Werner et al., A10)
7 10.7 5.3 tion and Graph Transformation for Finan- 29 %9) A10) A10) A*10) DB (3) 5 0.33 80% No
formation Systems (ICIS) 2012b) (8.48)
cial Audits
Improving Structure: Logical Sequencing 47th Hawaii International Confer- (Werner and C
8 10.8 5.4 56 % B A A DB (4) 2 0.67 95% No
of Process Models ence on System Sciences (HICSS) Nüttgens, 2014) (6.44)
Towards Automated Analysis of Business 11th International Conference on (Werner et al., C
9 10.9 6 26 %11) A C C DB (5) 3 0.50 90% No
Processes for Financial Audits Wirtschaftsinformatik (WI) 2013) (6.73)
5.5 Multilevel Process Mining for Financial IEEE Transactions on Services Com- (Werner and 12)
10 10.10 -13) A -13) -13) B (3)14) 2 0.67 95% No
6 Audits puting Gehrke, 2015)

7
Invited for publication
8
Submitted as full paper (14 pages), accepted as short paper (8 pages), presented at the ER 2013 and published in the conference proceedings, overall acceptance rate 32 %
9
Acceptance rate for Completed Research Paper 28.92 %, Research-in-Progress 29.2 %
10
Submitted and accepted as Research-in-Progress paper (12 pages), presented at the ICIS 2014 and published in the conference proceedings
11
Acceptance rate for the relevant track 21.6 %
12
The journal does not communicate acceptance rates for special or regular issues.
13
Not explicitly listed in the VHB JQ2.1, ERA 2010 or CORE 2013 rankings. The ISI impact factor for this journal in 2014 was 3.049.
14
Refers to the first review round, three more reviews were received during the second review round.

22
ANALYSIS

4 Analysis
The analysis phase is the first phase of the DSR cycle (Österle et al., 2010). Figure 7 shows
a framework that provides guidance on how to carry out design science-oriented research
(Hevner and Chatterjee, 2010; Hevner et al., 2004). It is useful to identify which aspects are
important for the analysis phase. A key aspect is the analysis of the application domain for
the identification and confinement of the research question. Taking requirements from the
application domain into account ensures that the investigated question is indeed relevant
for a specific practical purpose. A second aspect is the analysis and consideration of already
existing knowledge from the scientific knowledge base which is important to ensure the
research rigor by relying on knowledge that has already been evaluated and accepted in
the scientific community. The following sub-chapters illustrate how knowledge from the
knowledge base (chapter 4.1) and requirements from the application domain have been
taken into account (chapter 4.2).

Environment IS Research Knowledge Base


Relevance Rigor

People Develop / Build Foundations


• Theories
• Artifacts
Business Applicable
Needs Knowledge
Organizations

Justify / Evaluate Methodologies


• Quantitative
Technology Research Methods
• Qualitative Research
Methods

Application in the Additions to the


Appropriate Environment Knowledge Base

Figure 7 Integrating the Application Domain and Knowledge Base in DSR (adapted from
Hevner et al., 2004, p. 80)

4.1 Related Scientific Work

4.1.1 Business Process Management


Business process management (BPM) is an important and mature research area that is rel-
evant to the presented research.
“Business process management includes concepts, methods, and techniques to
support the design, administration, configuration, enactment, and analysis of
business processes” (Weske, 2012, p. 5).
It “is the discipline that combines knowledge from information technology and
knowledge from management sciences and applies this to operational business
processes” (van der Aalst, 2013, p. 1)

23
ANALYSIS

Several up-to-date textbooks provide fundamental and extensive knowledge on BPM (Du-
mas et al., 2013; Weske, 2012). Of particular interest are those publications related to the
modeling of business processes using Petri Nets (van der Aalst and Stahl, 2011) because
Petri Nets represent the most frequently used modeling language in the field of process
mining (Tiwari et al., 2008). Van der Aalst provides a comprehensive survey on business
process management and identifies 20 different use cases and six key concerns (van der
Aalst, 2013). His survey relies on the analysis of 289 scientific publications published at the
International Conference on Business Process Management (van der Aalst, 2012). The most
relevant use case for the research at hand is the ‘discover model from event data’ which
strongly relates to the research domain of process mining.

4.1.2 Process Mining


Process mining is a relatively new research area that emerged in the late 1990s initiated by
the work from Cook and Wolf about discovering software processes (Cook and Wolf, 1998a,
1998b, 1999). The discovery of process models was also already subject of research in var-
ious established scientific disciplines like concurrency theory, inductive inference, stochas-
tic, data mining, machine learning or computational intelligences (van der Aalst, 2011a).
Agrawal et al. and Maxeiner et al. were among the first who used workflow logs to create
process models (Agrawal et al., 1998; Maxeiner et al., 2001). Schimm, Herbst and Karagi-
annis first used process mining in the context of workflow and business process manage-
ment (Herbst, 2003, 2000a, 2000b; Herbst and Karagiannis, 1999, 1998; Schimm, 2001a,
2001b). Substantial and extensive research work on process mining has been published by
researchers in the past decade. Van der Aalst provides a comprehensive summary of basic
and advanced concepts in process mining that have been discovered in recent years (van
der Aalst, 2011a). The Process Mining Manifesto is an important publication for the re-
search in this area. It illustrates contemporary chances and challenges for process mining
(van der Aalst et al., 2012). An overview of the state-of-the-art in process mining until 2008
is summarized in (Tiwari et al., 2008).
Figure 8 shows the distribution of scientific publications in the world until the end of 2013
that relate to process mining. The publications were identified by conducting an extensive
literature review following the guidelines suggested by several scholars (Fettke, 2006; Row-
ley and Slack, 2004; vom Brocke et al., 2009; Webster and Watson, 2002). Scientific publi-
cations were identified by using the search term ‘process mining’ for the following data-
bases:
• ProQuest
• AISel
• EBSCO
• Springer Link
• ScienceDirect
• ACM

24
ANALYSIS

The search was restricted to the title and peer-reviewed publications when possible. The
search also included a review of related articles in international top information systems
journals (AIS Senior Scholars' Basket of Journals (AIS, 2014)):
• European Journal of Information Systems
• Information Systems Journal
• Information Systems Research
• Journal of the Association for Information Systems
• Journal of Information Technology
• Journal of Management Information Systems
• Journal of Strategic Information Systems
• Management Information Systems Quarterly

These journals were searched for hits in the title, abstract and keywords. The literature
review finally covered 236 distinct articles15. The mapping to countries was achieved by
investigating the heritage of the authors.16 To prevent duplicate mappings only the first
author of a publication was used to assign a publication to a specific country. Table 4 lists
the publications per country. It can be seen that most publications were prepared by au-
thors from the Netherlands followed by China and Germany.

Figure 8 Distribution of Worldwide Process Mining Related Scientific Publications17

15
The complete list of publications from this literature review is provided in appendix B.
16
For the case that different information on the heritage of the author was found possibly because the au-
thor had switched his or her location in the academic career the most current information was used.
17
The diagrams were prepared using Google Geocharts (Google, 2014)

25
ANALYSIS

Table 4 Distribution of Process Mining Related Scientific Publications in the World


Country Publications Country Publications
Netherlands 68 Iran 2
China 27 Turkey 2
Germany 22 Australia 2
Belgium 14 Algeria 2
USA 11 Mexico 2
Spain 11 Poland 2
Italy 10 Portugal 2
South Korea 9 Thailand 1
Great Britain 8 Canada 1
Taiwan 7 Colombia 1
Brazil 6 Chile 1
France 4 Macedonia 1
Japan 3 Hungary 1
Norway 3 Tunisia 1
Austria 3 Vietnam 1
Egypt 2 Rumania 1
Denmark 2 Israel 1
Russia 2

Figure 9 and Table 5 show the distribution of process mining in Europe on the city level. It
shows that Eindhoven can be considered the epicenter of research on process mining with
64 publications which represent 43 % of the publications in Europe and 27 % worldwide.

Figure 9 Distribution of Process Mining Related Scientific Publications in Europe

26
ANALYSIS

Table 5 Distribution of Process Mining Related Scientific Publications in Europe


Country City Publi-cati- Country City Publi-
ons cations
Belgium Ghent 3 France Clermont-Fer- 1
Belgium Diepenbeek 4 rand
Belgium Leuven 7 Italy Milan 1
Germany Ingolstadt 1 Italy Bologna 2
Germany Potsdam 2 Italy Rende 1
Germany Ulm 1 Italy Bari 2
Germany Paderborn 1 Italy Pavia 1
Germany Dresden 4 Italy Ferrara 3
Germany Oldenburg 1 Macedonia Skopje 1
Germany Magdeburg 1 Netherlands Eindhoven 64
Germany Hamburg 4 Norway Groningen 4
Germany Frankfurt 1 Norway Trondheim 2
Germany Münster 2 Norway Bergen 1
Germany Freiburg 4 Poland Poznan 1
Denmark Lyngby 1 Poland Warsaw 1
Denmark Aarhus 1 Portugal Lisbon 1
Great Britain Birmingham 1 Portugal Santa Maria 1
Great Britain Sheffield 1 da Feira
Great Britain Bedford 5 Romania Cluj-Napoca 1
Great Britain Ipswich 1 Spain Barcelona 7
France Paris 2 Spain Ciudad Real 4
France Marseilles 1 Hungary Veszprem 1

The identified articles represented the starting point for the scanning of the scientific
knowledge base for the research at hand. They are just a subset of scientific articles that
relate to the field of process mining. Many relevant articles do not include the search term
‘process mining’ in the title and were therefore not identified in the initial review. The sci-
entific knowledge base was extended during the iterative research cycles. Relevant related
scientific work is referenced explicitly in the individual publications (compare appendix A).
The following paragraphs provide an overview of the most important publications for the
presented research. Fundamental aspects in regard of process mining have been published
in (Gehrke and Werner, 2013). Process mining is used for process discovery, process en-
hancement, conformance and compliance checking (Gehrke and Werner, 2013; van der
Aalst et al., 2012). Process discovery, conformance and compliance checking are especially
relevant in the context of process audits.
Several powerful general purpose mining algorithms have been developed in recent years
that use simple deterministic (van der Aalst et al., 2004), heuristic (Günther and van der
Aalst, 2007; Weijters et al., 2006) or genetic (de Medeiros, 2006) approaches. An important
aspect for the evaluation of the appropriateness of mining algorithms are quality criteria
that can be used to assess mined process models (Rozinat, 2007; Rozinat et al., 2008). Sev-
eral publications deal with process mining in the context of compliance checking. Compli-

27
ANALYSIS

ance checking refers to the question if the observed process behavior complies with rele-
vant rules. Ramezani et al. identify 55 control flow oriented compliance rules formalized in
terms of Petri Net patterns that can be used to check mined models if they comply with
these patterns (Ramezani et al., 2012). Van der Werf et al. scrutinize how information from
the organizational contexts can be connected to recorded data in event logs to check com-
pliance rules that are independent from individual process instances (van der Werf et al.,
2012). Accorsi and Lehmann incorporate the data perspective to identify information leaks
in business process models (Accorsi and Lehmann, 2012). Caron et al. suggest a rule-based
compliance checking and risk management approach (Caron et al., 2013). Van der Aalst et
al. introduce a conceptual model for online auditing using process mining techniques (van
der Aalst et al., 2011) whereas Jans et al. discuss opportunities, challenges and limitations
for using process mining in the context of audits (Jans, 2012; Jans et al., 2010). They also
present several case studies (Jans et al., 2011, 2008). Conformance checking aims to iden-
tify deviant behavior in a process. It requires the existence of a model that is used for com-
parison. Rozinat and van der Aalst introduce an approach to compare mined process mod-
els to a reference model (Rozinat and van der Aalst, 2008). They use concepts like fitness
and appropriateness to identify deviations. A similar approach is used by Adriansyah et al.
who introduce a cost-based fitness analysis (Adriansyah et al., 2011). Van der Aalst and de
Medeiros present a two-step approach (van der Aalst and de Medeiros, 2005). A reference
model is mined in a first step and subsequent executions recorded in the event log are then
used to identify any deviations compared to the previously mined model. Bezerra and
Wainer analyze the event log to detect anomalies (Bezerra and Wainer, 2013). Yang and
Hwan present a framework and case study to identify fraud in the healthcare sector (Yang
and Hwang, 2006). Van der Aalst presents a case study using data from a Dutch govern-
mental institution (van der Aalst, 2005).
Several academic and commercial process mining software tools exist. The most common
are listed in Table 6. A major academic and open-source tool is ProM (Process Mining
Group, 2015). It provides plug-ins for many different mining algorithms, as well as analysis,
conversion and export modules. Disco (fluxicon, 2015) is a commercial application that ben-
efits from intuitive and easy usability. It also provides integrated functionality for the filter-
ing and loading of event logs. ProM and Disco are the software tools that were primarily
used for evaluation and comparison purposes in the presented research.
Table 6 Process Mining Tools
Product Name Link
ARIS Process Performance Manager [Link]
celonis business intelligence [Link]
Disco [Link]
Genet/Petrify [Link]
Interstage Business Process Manager [Link]
QPR ProcessAnalysizer [Link]
ProM [Link]
ProcessGold [Link]
Rbminer/Dbminer [Link]
ReflectOne [Link]
ServiceMosaic [Link]

28
ANALYSIS

4.2 Requirements from the Application Domain

4.2.1 The Role of Business Processes in Financial Audits


The aim of the research that is described in this thesis is the development of analysis tech-
niques that can be used as a foundation to automate the audit of business processes in the
context of financial audits. A prerequisite to assess the requirements from the application
domain that should be taken into account is an understanding of the role of business pro-
cesses in financial audits. National and international laws and audit standards are an im-
portant information source for this purpose. ISA 315, IDW-PS 261 (IDW, 2009) and the AS
No. 12 (PCAOB, 2010) require a risk based approach that takes business processes, internal
controls and related information systems into account. A financial audit is commonly made
up of different phases that are illustrated in Figure 10.
Information Collection and Risk Assessment

Design Effectiveness Testing of Internal Controls

Operating Effectiveness Testing of Internal Controls

Substantive Audit Procedures

Reporting

Figure 10 Audit Process (adapted from Werner and Gehrke, 2011, p. 105)
The process starts with the collection of necessary information on the audited entity and
its business environment for the risk assessment, the identification of possible causes of
risk and the determination of necessary materiality levels. It follows the testing of design
effectiveness of internal controls. Internal controls are defined as
“the process designed, implemented and maintained by those charged with
governance, management and other personnel to provide reasonable assur-
ance about the achievement of an entity’s objectives with regard to reliability
of financial reporting, effectiveness and efficiency of operations, and compli-
ance with applicable laws and regulations.” (IFAC, 2012, para. 4(c))
With regard to business process management and process modeling internal controls can
commonly be considered as activities that influence the processing of business transac-
tions. The objective of the test of the design effectiveness is to ensure that the internal
controls are appropriately designed to achieve the intended control objectives. This assess-
ment is only possible if the auditor has information about how the controls relate to the
business transactions that actually create the postings on the financial accounts. The rela-
tionship between business processes and financial accounts is illustrated in Figure 11. The
model represents a simple purchase process. The rectangles represent activities that are
executed in an ERP system. They create postings on financial accounts. According to the
double-entry bookkeeping system each posting consists of at least one credit and one debit
entry that balance each other. Not every activity creates an entry in the financial accounts.
If a company orders goods from a supplier this does not result in a posting, because the
liability only occurs when the goods are received. The delivery of the ordered goods leads

29
ANALYSIS

to entries in the raw materials and the goods received / invoices received account. When
the invoice is received the goods received / invoices received is cleared and a correspond-
ing entry is posted on the trade payables account. When the invoice is finally paid the entry
on the trade payables account is cleared with a corresponding entry on the bank account.
This example illustrates the relationship between business processes and financial ac-
counts that are important from an audit perspective

Goods Received /
Raw Materials Trade Payables Bank Account
Invoices Received

10,000 10,000 10,000 10,000 10,000 10,000

Order Goods Receive Goods Receive Invoices Pay Invoices

Figure 11 Example Model of a Purchase Process (Werner and Gehrke, 2015, p. 823)
The next step in a financial audit is the testing of the operating effectiveness of the identi-
fied controls to ensure that the controls were indeed effective in the audited period. The
audit process continues with substantive testing procedures that include traditional ana-
lytical procedures, physical examinations of inventory or the assessment of confirmations
of balances from suppliers. The audit process terminates with the final audit report.
Auditors use different types of audit procedures to test the design and operating effective-
ness of internal controls. The primary audit procedures to test internal controls are inquiry,
observation, inspection and re-performance (IFAC, 2012, para. A73). They differ in terms of
audit effort and gained audit reliance as illustrated in Figure 12.

Re-performance
Audit Reliance

Inspection

Observation

Inquiry

Audit Effort

Figure 12 Audit Procedures (adapted from Werner and Gehrke, 2011, p. 105)
The audit procedures illustrated in Figure 12 are all manual procedures. They are time-
consuming and error-prone and may become inefficient or even ineffective if information
systems are used on the company’s side to automate the operation of business processes
and if very large amounts of data are processed (Werner and Gehrke, 2011). A key prereq-
uisite for the effectiveness of inquiries is the assumption that the interview partner has
sufficient information about relevant business processes and related internal controls and
that the auditor is able to receive all necessary information in an interview. This might not

30
ANALYSIS

be the case anymore if process activities are partially or completely processed by an infor-
mation system without any human interaction. Observations, inspections and re-perfor-
mances are carried out on the basis of sampling. With an increase of processed transactions
the sample size has to be increased. When millions or even billions of transactions are pro-
cessed during an audit period it can be assumed that they become very inefficient when
the sample size is increased accordingly or that they even become ineffective if an appro-
priate sample size cannot be selected because of limited audit resources.

4.2.2 Data Structure


Manual audit procedures can be substituted by automatic audit procedures on the basis of
process mining techniques if the data that is stored in information systems during the
course of processing is used. The original source data can commonly not be used directly
for process mining purposes because it does not exhibit the data structure that is required
by established process mining algorithms (van der Aalst et al., 2012) .
Mining algorithms use event logs as input that exhibit a specific structure (Günther and
Verbeek, 2012). An event log is basically a table which contains all recorded events that
relate to executed business activities. Each event is mapped to a case that corresponds to
a single execution of a business process which is called process instance. The sequence of
recorded events in a case is called a trace. Cases and events are characterized by classifiers
and attributes. Classifiers ensure the distinctness of cases and events by mapping unique
names to each case and event. Attributes store additional information that can be used for
analysis purposes (Gehrke and Werner, 2013). An example of an event log is given in Table
7.
Table 7 Event Log Structure (Gehrke and Werner, 2013, p. 935)
Case ID Event ID Timestamp Activity
1 1000 01.01.2013 Order Goods
1001 10.01.2013 Receive Goods
1002 13.01.2013 Receive Invoice
1003 20.01.2013 Pay Invoice
2 1004 02.01.2013 Order Goods
1005 01.01.2013 Receive Goods
… … …

The data entries that are recorded in ERP systems differ from the structure illustrated in
Table 7. The necessary data is commonly stored in various database tables. This data ex-
hibits a specific structure that relates to the systematic of double-entry bookkeeping. The
Entity-Relationship-Model (ERM) in Figure 13 illustrates the relationship between transac-
tions that are executed in ERP systems, data entries and financial accounts. The illustrated
entity attributes follow the naming of data labels used in SAP ERP systems. The execution
of a transaction that is labeled with a transaction code creates one or more posting docu-
ments.18 These documents contain two or more journal entry items that are posted to a
specific account. If the account is enabled for open-item-accounting each open item has to

18
If the execution is not financially relevant no posting document is created.

31
ANALYSIS

be cleared by a clearing posting. If this is the case a cleared item carries a reference to a
posting document that creates the clearing items.

Transaction 1 0...N Posting Document 1 2...N Journal Entry Item 0...N 1 Financial Account
creates contains posted on
TransactionCode DocumentNr DocumentNr AccountNr
UserName PositionNr AccountType
PostingDate AccountNr Balance
TransactionCode 0...1 0...N Amount
is cleared
PostingText CreditOrDebit
ClearingDocNo

Figure 13 ERM for Accounting Data Structure (Werner, 2013, p. 389)


This data structure is different compared to the one presented in Table 7. Events in a tra-
ditional event log exhibit a strict linear order within each case. No concurrency is allowed
on the process instance level. This strict order is a fundamental prerequisite for traditional
process mining algorithms to infer the control flow. It is not given for financially relevant
data as illustrated in Figure 13. Due to the 1 to 2…N cardinality of ‘the contains’ relationship
and the existence of the ‘is cleared’ relationship with the 0…N to 0…1 cardinality the result-
ing data structure is not linear. The source data is also not labeled which means that events
are not mapped to cases. Gehrke and Müller-Wickop introduce an algorithm that maps
events to cases according to the data structure shown in Figure 13. Figure 14 shows an
example of the source data structure of a process instance extracted from an SAP system
using the approach and the notation introduced by (Gehrke and Müller-Wickop, 2010). A
rounded rectangle represents an activity. A simple rectangle represents a cleared journal
entry item that is involved in open-item-accounting. Hexagons represent those items that
are not involved in open-item-accounting. Gray rectangles and hexagons represent debit
and white ones represent credit postings. Each rectangle and hexagon displays values for
the journal entry item number, the account it was posted on and the posted amount. The
activity symbols display the transaction code that was used to create the posting docu-
ments. Activities and related journal entries belonging to the same document numbers are
clustered in groups. The example shows that the recorded events which are represented
by the activity symbols do not exhibit a strict linear order but show a concurrent behavior
on the process instance level. It also becomes obvious that no direct information on the
causal direction of activities can be derived from the presented model.

32
ANALYSIS

5000004384

003
0000310000
001 13907.14
0000310000 005
10880.29 0000310000
14776.34

MB01
002
0000191100
10880.29
006
004 0000191100
0000191100 14776.34
13907.14
0100008540

FB1S
0100008541

0100008542
FB1S

FB1S

5100004301

003
002 0000191100
0000191100 13907.14
10880.29 004
0000191100
0100008537
14776.34 2000000217
FB1S
MR1M
004 003
0000154000 0000113101
005
0000154000
306.86 76067.06 5100004300
001
5934.56 5000004383
0000160000
001
45498.33
0000160000
F110 006 004
32921.32 0100008539
002 0000154000 008 0000191100
0000160000 4294.08 0000191100 8916.93
007 003 8042.62
78419.65
0000230051 0000191100 FB1S
001 0.01
0000276000 8916.93 002
003
2045.73 MR1M 0000191100
0000310000
3722.20
008 8916.93
005
0000230051 MB01
0000191100
0.00
8042.62 006 005
0100008536 0000191100 0000310000
004 002 7945.48 7945.48
0000191100 0000191100 007 001
7945.48 3722.20 FB1S 0000310000 0000310000
8042.62 3722.20

0100008538

FB1S

Figure 14 Source Data Structure for a Purchase Process Instance (adapted from Gehrke
and Müller-Wickop, 2010, p. 7)

4.2.3 Requirements Summary


The previous two chapters illustrated the role of business processes in financial audits and
the data structure of the source data that can be used for process mining purposes. A pro-
cess mining algorithm that is used to discover and analyze process models in financial au-
dits should meet the different criteria that stem from the application domain and the used
ERP systems. Müller-Wickop and Schultz et al. conducted empirical investigations using ex-
pert interviews and surveys to identify key concepts and information requirements for pro-
cess audits (Müller-Wickop and Schultz, 2013a; Schultz et al., 2012). They show that the
process flow is a central concept and highly relevant from an audit perspective in practice.
Two further important concepts for external auditors are financial statements and materi-
ality19. Only those business transactions are inspected in a financial audit that can have a
material effect on the financial statements. It is therefore necessary to receive information
on the value flow that is created by the audited business process to decide if it needs to be
audited from a materiality perspective or if it can be neglected. These requirements corre-
spond to the logical considerations on the basis of the review of relevant audit standards
in chapter 4.2.1. They can be summarized as follows:

19
Materiality is defined in ISA 320 as followed: “Misstatements, including omissions, are considered to be
material if they, individually or in the aggregate, could reasonably be expected to influence the economic
decisions of users taken on the basis of the financial statements.” (IFAC, 2009, para. 2).

33
ANALYSIS

Requirement I: The mined process models should provide information on the


control flow of process activities.
Requirement II: The mined process models should provide information on the
relationship of business processes and financial accounts.
Requirements I and II can be satisfied by integrating the control flow and the data flow
perspective.
Another critical requirement in financial audits is the preservation of the audit trail. The
audit trail is a fundamental concept in financial accounting. It is a path in an information
system that allows tracing a transaction from the point of origin to the final output. It is
used to verify the accuracy and validity of journal entries (Romney and Steinbart, 2008).
Translating this requirement into the context of process mining implies that a mining algo-
rithm may not alter the original data during the mining process.
Requirement III: The mining algorithm should not alter the original source data
for creating process models.
If process mining is used in financial audits the generated models are used to discover in-
compliant behavior. The provided process models should therefore only represent infor-
mation on process behavior that was actually observed in the source data. A fundamental
challenge for process mining is the balancing between competing quality criteria (van der
Aalst et al., 2012). Process mining is generally used to reduce complexity by visual repre-
sentation and abstraction. Abstraction means that certain details are not represented in
the created process models. This can lead to the phenomenon of miss-fitting process mod-
els. Models can be under- or over-fitting. A model is under-fitting if it allows execution
paths in the process model that are not represented in the event log and over-fitting if they
do not allow for any additional behavior that is not included in the event log. Rozinat et al.
identify four quality criteria for the evaluation of mining algorithms: fitness, precision, gen-
eralization and structure (Rozinat et al., 2008). Fitness indicates if a model is able to repre-
sent all cases in the event log. Precision is the complementary criterion. It indicates if a
process model does not allow additional behavior that was not observed in the log. Gener-
alization addresses the capability of a model to express more behavior than recorded in
the log. Structure refers to the graphical representation of a business process and depends
on the graphical components of the target language. Other scholars use the closely related
criterion of simplicity instead of structure (van der Aalst et al., 2012). Mined process models
in financial audits should be as precise and fitting as possible. If the produced process mod-
els are over-fitting certain behavior recorded in the event log is not represented in the pro-
cess model. The auditor would therefore assume that no incompliant behavior has oc-
curred. But in reality it is just not represented in the miss-fitting process model which would
eventually lead to false positive audit results. If the process model is too general, process
behavior is illustrated that actually did not occur. This would lead to false negative audit
results and unnecessary investigations by the auditor (Werner and Gehrke, 2015).
Requirement IV: The mined process models should be as precise as possible.
Requirement V: The mined process models should be as fitting as possible.
A suitable mining algorithm must also be able to present the detailed source data to the
auditor for investigation purposes. Detailed information is necessary to inspect individual

34
ANALYSIS

data values to identify or confirm compliance violations. But on the other hand the mining
algorithm should also be able to present information at an adequate abstraction level to
provide an overview of the control and data flow. It should therefore provide information
on the investigated business process on different abstraction levels.
Requirement VI: The mining algorithm should be able to produce process models
at different abstraction levels.
Additional requirements originate from the structure of the available source data as dis-
cussed in chapter 4.2.2. The source data from ERP systems is commonly not labeled and
the activities on process instance level may exhibit concurrent behavior.
Requirement VII: The mining algorithm should be able to use unlabeled event logs
as input.
Requirement VIII: The mining algorithm should be able to use event logs with non-
linear traces as input.

35
DESIGN

5 Design
The second phase in the DSR cycle is the design of research artifacts such as constructs,
models, methods, instantiations (March and Smith, 1995) and theories (Gregor, 2006). The
design phase lies at the heart of design science-oriented research as illustrated in Figure
15. The aim is to use the results from the analysis phase concerning the requirements and
business needs on the one hand and the scientific knowledge from the knowledge base on
the other hand to develop novel solutions. They should be useful for the application do-
main and should also contribute to the general body of knowledge as shown with the feed-
back loops in the lower part of Figure 15. Chapter 3.3 has illustrated how the different ar-
tifacts relate to each other and that they have been developed in iterative cycles. The used
design methods have been described in chapter 3.4. The following sub-chapters describe
the different research artifacts as the main research output of the presented research.
Environment IS Research Knowledge Base
Relevance Rigor
People Develop / Build Foundations
• Theories
Business
• Artifacts Applicable
Needs Knowledge
Organizations

Justify / Evaluate Methodologies


• Quantitative
Technology Research Methods
• Qualitative Research
Methods

Application in the Additions to the


Appropriate Environment Knowledge Base

Figure 15 Design in the Information Systems Research Framework (adapted from He-
vner et al., 2004, p. 80)

5.1 Conceptual Models for Specifying the Problem Domain and Solution
The key problem that motivates this research is the imbalance between automated trans-
action processing on the companies’ side and manual audit procedures on the auditors’
side leading to inefficient or even ineffective process audits (compare chapter 2.3). The
idea to reduce this imbalance is the development and application of automated data anal-
ysis techniques such as process mining. But how should this work from a conceptual per-
spective?
We use a metaphor to describe the problem domain and possible solution. Business pro-
cesses are a set of related activities to achieve a business goal (Weske, 2012). A business
process can be interpreted as a flow or river. The execution of activities creates data entries
that flow through the involved information systems. Financially relevant information finally
ends up as journal entries on the financial accounts. This interpretation is conceptually il-
lustrated in Figure 16. It shows different process flows that represent different business

36
DESIGN

processes. These processes create data which is the water in our metaphor that feeds the
financial accounts represented as the lake and final destination.
The figure also shows several application controls. Application controls are internal controls
that are implemented in the information systems. They support or automate the pro-
cessing of business transactions and regulate the transaction processing or respectively the
flow of the water in our metaphor.
We illustrate the relationship between the process activities and application controls with
an example. A typical purchase process starts with a purchase order (activity 1) that creates
an order document. Purchase orders can usually only be issued up to a certain amount
without further approval (control A). The ordered goods are eventually delivered (activity
2). A two-way-match (control B) ensures that the received goods can only be entered in the
ERP system if the ordered goods match with the purchase order (Chuprunov, 2012). Finally
the invoice for the delivered goods arrive (activity 4). A three-way-match ensures that the
billed goods match with the purchase order and the delivered goods (control C)
(Chuprunov, 2012). Activity 3 shows that not all application controls relate to all possible
process flows or sub-processes. The two-way-match is usually not used for delivered ser-
vices because they are not physically delivered and no corresponding delivery document
exists.

Process Activities
# Description
1 Purchase order
2 Goods receipt for purchase order
3 Services receipt
4 Incoming invoice
Application Controls
# Description
System based approval for pur-
A
chase order
B Two-way-match
C Three-way-match

Figure 16 Metaphor Illustrating the Relationship between Business Processes, Financial


Accounts and Application Controls (adapted from Werner et al., 2012a, p. 5356)
A public accountant has to provide an opinion during a financial audit about whether the
financial statements present a fair and true view of the financial situation of the company.
If we rely on the presented metaphor this would mean that the auditor has to ensure that
the lake only contains water from specific sources with a defined quality. Following con-
temporary audit approaches (compare chapter 2.2) the auditor would manually take sam-
ples from different places in the lake to verify the water quality and origin. This would be
the equivalent of manual audit procedures like the inspection of selected documents or
interviewing individuals. Based on experience and professional judgment he would maybe
also choose several rivers for inspection and verify manually if control mechanisms are in
place which regulate the flow and quality of the water that would be the equivalent of

37
DESIGN

manual internal controls tests. An auditor has to rely on the information received from
interviews and in best cases from available process descriptions but without actually know-
ing precisely which rivers and concurrent flows indeed exist and how much water they carry
into the lake.
One of the most important criteria for an auditor is to know the relationship between the
business processes and financial accounts (compare chapter 4.2.3.). The first step to sup-
port the auditor is the provision of data analysis techniques that provide information on
the interaction of processes and accounts. Using process mining techniques enables the
auditor to get a visual representation of relevant business processes. This is the foundation
for the auditor to receive an all-embracing understanding of the relevant business pro-
cesses and how they affect the financial accounts. When this information is available appli-
cation controls can be taken into account in a second step. Application controls can be
viewed as filters from an audit perspective. If the design and operating effectiveness of all
necessary application controls for a specific business process are verified such a process
does not need any further attention because the controls ensure that only valid and correct
financial entries end up in the financial accounts from the controlled business processes.
Audit resources can then be spent on those processes that are not controlled by application
controls and that usually exhibit a higher au-
dit risk than standard processes that are well- Process Perspective Control Perspective

controlled. Automated
Process Mining
The fundamental idea for automating pro- Techniques + Control Testing
Techniques
cess audits is the combination of process
mining and automated application controls
testing techniques. This is conceptually illus- Combined Perspective
trated in Figure 17. The automation of the
Automated Business
process discovery and analysis is the aim of
Process Audit
the research presented in this thesis (com-
Techniques
pare chapter 3.2.2). It is the first step and a
fundamental prerequisite for developing au-
tomated audit procedures. Figure 17 Conceptual Solution

5.2 CPN Specification for Integrating the Data Perspective


The expressive power of a mining algorithm partially depends on the used modeling lan-
guage. The majority of process mining algorithms use Petri Nets for modeling mined pro-
cess models (Tiwari et al., 2008). The usage of Petri nets allows the application of mature
mining and analysis methods that are already implemented in academic software tools
such as ProM (Process Mining Group, 2015). Petri nets are suitable for the modeling of
business processes and offer a formal but also graphical notation which is also comprehen-
sible to non-experts (van der Aalst and Stahl, 2011). They provide a sound mathematical
foundation which allows the simulation and verification of Petri net models (Jensen and
Kristensen, 2009).
It is possible to satisfy the requirement I from chapter 4.2.3 that postulates that the mined
models should provide information on the control flow of process activities by modeling
mined process models as Petri Nets. Petri Nets support all common control flow patterns

38
DESIGN

(van der Aalst et al., 2003) and are therefore suitable to model the control flows in business
processes.
Requirement II from chapter 4.2.3 postulates that the mined process models should pro-
vide information on the relationship of business processes and financial accounts which
can be achieved by integrating the data perspective in process models. The data perspec-
tive has generally not been considered extensively yet for process mining (de Leoni and van
der Aalst, 2013; Stocker, 2012). The majority of mining algorithm use low-level Petri Nets.
Data objects can be included by using Colored Petri Nets (CPN). Such objects can be mod-
eled as colored tokens and places (Werner, 2013).
A Colored Petri Net can formally be expressed as a tuple CPN = ( , , , Σ, , , , , ) (Jen-
sen and Kristensen, 2009). Table 8 provides a specification for integrating the data perspec-
tive using CPN for process mining in financial audits.20

Table 8 CPN Specification for Process Mining in Financial Audits (Werner, 2013, p. 390)
is a finite set of transitions
The transitions represent the activities that were executed in the process. They dis-
Transactioncode

play the name of the activity. Further information, for example the transaction code
Activity Name

name, can be added.


is a finite set of places
Places in the FPN represent financial accounts and control places.
Control Places Control places determine the control flow in a process model. For
every process model one source place is modeled that connects to the
start transactions. The sink place marks the termination of the pro-
cess. The control places between the start and end transition deter-
Source Sequence Sink mine the execution sequence of the process model. A control place
Place Place Place belongs to the set of control places CP.

Account Places
Account
debit side of a Account
credit side of The account places represent financial accounts
balance sheet balance sheet that are affected by the execution of activities in a
account account process. The symbol color indicates the meaning of
Account
credit side of a Account
debit side of a an account. Account places belong to the set of ac-
profit and loss profit and loss
account account
count places AP.

∈ ×
∪ × is a set of arcs also called flow relation
Control arcs connect control places with transitions. They model the control flow
Control in the model.
Arc
Posting arcs illustrate the relationship between activities represented as transi-
Posting tions in the model and financial accounts that are modeled as account places.
Arc
Clearing arcs are used to model that an activity cleared an entry on the corre-

tical abbreviation for two arcs (p,t) and (t,p).


Clearing sponding account. Clearing arcs are double-headed arcs and are used as an syntac-
Arc

20
CPN following this specification are called Financial Petri Nets (FPN) in the related publication.

39
DESIGN

is a set of non-empty color sets


The color set contains the possible values that are posted or
colset VALUES = double
colset ACCOUNTS = string
cleared.
The color set contains all account numbers.
colset ACOUNTTYPE = boolean
The color set contains {1,0} indicating if the represented ac-
count is a balance sheet or a profit and loss account.
colset CREDorDEB = boolean
The color contains {1,0} indicating if the account place is a
representation of the debit or credit side of an account.
colset EXECUTIONS = int
The color contains {1,…,n} indicating how often a path was

ACCOUNTPLACES is the color set as a product of VALUES *


chosen in the FPN.

colset ACCOUNTPLACES ACCOUNTS * ACCOUNTTYPE * CREDorDEB


* is a finite set of typed variables such that Type[,] ∈ for all variables , ∈ *.

and = {}.
In FPN models arc inscriptions are modeled as constants. Variables are therefore not necessary

1: → is a color set function that assigns a color set to each place.


678 9 : ;< 4 ∈
(4) 5
= 7 68: ;< 4 ∈
The color set function in FPN assigns dif-
ferent color sets to places depending if
they belong to the group of control or
account places:
>: → ?@ A* is a guard function that assigns a guard to each transition B such that
CDE [>(B)] = FGGHEIJ.
= K is the set of expressions that is provided by the used inscription language. FPN do not ex-
plicitly include guards because they do not model dynamic behavior of transitions that depends

for FPN is therefore defined as (L) = LMNO <PM QRR L ∈ .


on specific input but illustrate the processing of already executed processes. The guard function

?: → ?@ A* is an arc expression function that assigns an arc expression to each arc I such
that CDE [?(I)] = 1(D)ST , where p is the place connected to the arc I.
The arc expressions in a FPN are constants. The arc expression function assigns to each posting
and clearing arc a set of constants that denote the posted or cleared value, the account num-
ber, account type and an indicator if it is a credit or debit posting. For each control flow arc the
number of execution times is assigned indicating how often this path was chosen in the process

{VQR ∈ 97 :, QWW ∈ 678 :, QWWLX4 ∈ 678 Y , WPMZ ∈ K [PM[ \}


model.

];Lℎ X4O [ (Q)] = (4)_` = 678 9 : ;< 4 ∈


(Q) U
{Oa ∈ = 7 68: ];Lℎ X4O [ (Q)] = (4)_` = = 7 68: ;< 4 ∈

b: → ?@ A∅ is an initialization function that assigns an initialization expression to each


place D such that CDE [b(D)] = 1(D)ST

ef Oa ∈ = 7 68: ];Lℎ X4O [ (4)] = (4)_` = = 7 68: ;< 4 = gPNMWO


The initialization function of a FPN assigns initialization expressions to each place as follows:

(4) d
∅_` PLℎOM];gO
Only the source place is initialized in a FPN. The initialization expression for 4 = gPNMWO gener-
ates e tokens in the initial marking hi (4), one for each connected start transition. The inscrip-
tion of each token is a member of the set = 7 68:.

40
DESIGN

Figure 18 shows an example of a simple purchase process that can be modeled by using
the specification illustrated in Table 8.
100_D 200_C 200_D 400_D
Raw Materials Good Receipt / Invoices Receipt Good Receipt / Invoices Receipt Bank Account

[50,000] [50,000]
[50,000] [50,000]
[50,000]

MB01 MIRO F110

[1] Post Goods [1] [1] Enter Incoming [1] [1] Automatic [1]
[1]
Receipt for PO Invoice Payment
Source S1 S2 Sink

[50,000]

[50,000] [50,000]

300_C 300_D
Creditor Account Creditor Account

Figure 18 Simple Example of a Purchase Process Using CPN (Werner, 2013, p. 392)

5.3 Complexity Reduction Algorithm


A common challenge in process mining is the complexity of mined process models. Ex-
tremely complex models are called spaghetti and lasagna models due to their graphical
appearance (van der Aalst, 2011c). Source data derived from ERP systems related to finan-
cially relevant transactions is complex (Werner et al., 2013, 2012b). Figure 19 shows a typ-
ical process instance. It consists of 12 transitions and 115 net elements21. The average pro-
cess instance in the used data sets from three different companies consisted of 7 to 9 net
elements (Werner et al., 2013). But the vast majority of the mined models just represented
trivial instances that consist of one transition. The instance in Figure 19 can therefore be
perceived as a typical process instance for a purchase process. Figure 20 and Figure 21 show
two additional examples of process instance models that were derived from data provided
by a company operating in the manufacturing industry. The model in Figure 19 can still be
interpreted by simple observation. But this is not the case anymore for the models from
Figure 20 and Figure 21. These represent instances that consist of 358 (example B) and
2,966 (example C)22 transitions. Table 9 shows some characteristics of the three example
instances.

21
The sum of transitions, places and arcs in the CPN.
22
We called them monster instances due to their large size and necessary computation time for mining, and
the fact that they do not show the same graphical characteristics as spaghetti or lasagna processes but
exhibit a more organic shape.

41
DESIGN

Figure 19 Example A: Normal Process Instance (Werner et al., 2012b, p. 4)

Figure 20 Example B: Large Process In- Figure 21 Example C: Monster Process In-
stance (Werner et al., 2012b, p. 5) stance (Werner, 2012, p. 211)22

Table 9 Net Characteristics for Process Instance Examples


Example Instance A B C
Number of transitions 12 358 2,966
Number of places 43 1,804 8,936
Number of arcs 60 2,549 14,867
Sum of net elements 115 4,711 29,735

The complexity of mined models can be reduced by using graph transformation techniques
(Werner et al., 2012b). The complexity of a graph can be defined in many different ways
depending on the relevant perspective und purpose (Neel and Orrison, 2006). A graph is

42
DESIGN

of edges with E ⊆ V2 (Diestel, 2010). We consider the complexity of the graph simply as a
generally defined as a pair of disjoint sets G = (V,E) .V is the set of vertices and E is the set

function of its number of edges and vertices which is the same as the number of net ele-
ments as shown for example in Table 9.
The mined models consist of transitions that represent the process activities and account
places (compare chapter 5.2). The account places represent the financial accounts. An ac-
count place is modeled for each journal entry item that was created by an activity. Although
the complexity of mined models can be extremely high it is striking that the number of
different accounts in such models is comparatively small. The mean value of different ac-
counts per instance ranged from 2.33 to 2.95 (Werner et al., 2013). To reduce the complex-
ity of mined models it would therefore be a promising approach to aggregate account
places that carry the same account number using graph transformation techniques (Heckel,
2006; Rozenberg, 1997).
An important requirement described in chapter 4.2.3 is the preservation of the audit trail
and the necessity to keep the source data unchanged during the mining process (require-
ment III). This requirement can be articulated in more detail in regard to the intended graph
transformation (Werner et al., 2012b):
Requirement III a): The set of firing sequences has to stay constant.
Requirement III b): Different arc types may not be merged.
Requirement III c): The arc inscriptions representing the value of posted journal en-
tries have to be preserved.
The aggregation algorithm shown in Listing 1 uses CPN with tuple N = (P, T, F, C, cd, W, m0)23
as input and aggregates places in accordance with requirements III a) to c).
Listing 1 Aggregation Algorithm (Werner et al., 2012b, p. 7)
PItem ⊆ P
PItemAgg= ∅
set of all places representing journal entry items in the net
initially empty set for aggregated places
FArcs set of all arcs in the net

While PItem≠ ∅
Aggregate Places

Take pi ∈ PItem
Select all pj ∈ PItem with cd(pi)=cd(pj)
Merge arcs for each pi and pj
Add pi to PItemAgg
Remove pi and pj from PItem
Set PItem=PItemAgg

arcs FIncArcsI ⊆ FArcs for pi


Merge Arcs

arcs FIncArcsJ ⊆ FArcs for pj


Get incoming

ai(tm,pi) ∈ FIncArcsI and arc aj(tn,pj) ∈ FIncArcsJ


Get incoming
For each arc
If the arc type of ai = arc type of aj
And if tm = tn then add W(aj) to W(ai)and /Case 1
remove aj from FArcs
Else set aj(tm,pj) /Case 2,3 and 4

, C to Σ, cd to , W to and m0 to hi (4).
23
The used CPN tuple elements refer to the tuple elements specified in chapter 5.2 where F is equivalent to

43
DESIGN

Get outgoing arcs FOutArcsI ⊆ FArcs for pi


Get outgoing arcs FOutArcsJ ⊆ FArcs for pj
For each arc ai(pi,tm) ∈ FOutArcsI and arc aj(pj,tn) ∈ FOutArcsJ
If the arc type of ai = arc type of aj
And if tm = tn then add W(aj) to W(ai)and /Case 1
remove aj from FArcs
Else set aj(pj,tm) /Case 2,3 and 4

Figure 22 shows the result when the algorithm is used to reduce the complexity of example
A (Figure 19). The average of the number of net elements per instance could be reduced
by 23.2 % and 23.4 % for two different data sets originating from a company operating in
the retail business and a test SAP system (Werner et al., 2012b).
[100008542]
[100008542]
100008542
FB1S
[5100004301] [5000004384]
D021845 [100008539]
[5100004301] [5000004384]
[100008539]
5100004301 100008539 5000004384
MR1M FB1S MB01
BOLLINGER D021845 [100008537] BOLLINGER

[14776.34] [100008537] [14776.34]


100008537

[45498.33] [8042.62] FB1S [8042.62] [10880.29, 13907.14, 14776.34] [10880.29, 13907.14, 14776.34]
[2000000217] [10880.29, 13907.14, 14776.34]
[5934.56] D021845 [100008540]
[2000000217]
154000 160000 154000 191100 [8916.93] [100008540] 191100 310000
2000000217 100008540 [8916.93]
F110 [10880.29] FB1S
[306.86] [45498.33, 32921.32] [10880.29]
MONCHANIN D021845 [100008538]
[4294.08] [7945.48] [100008538]
100008538 [7945.48] [3722.20, 8916.93, 7945.48, 8042.62] [3722.20, 8916.93, 7945.48, 8042.62]
[32921.32] [3722.20, 8916.93, 7945.48, 8042.62]
[78419.65] [76067.06] [2045.73] [3722.20] [3722.20]
FB1S
[100008536]
160000 113101 276000 230051 230051 D021845 5000004383
5100004300 [13907.14] [100008536] [13907.14]
MB01
MR1M 100008536
[0.00] [0.01] BOLLINGER
BOLLINGER FB1S
[5000004383]
[5100004300] D021845 [100008541]
[100008541] [5000004383]
[5100004300] 100008541
FB1S
D021845

Figure 22 Aggregated Process Instance A (Werner et al., 2012b, p. 8)

5.4 Data-dependent Sequencing Algorithm


The source data in ERP systems that can be used to mine business processes that relate to
financial accounts is unlabeled and recorded events do not exhibit a strict linear order on
the process instance level (compare requirements VII and VIII from chapter 4.2.3). Gehrke
and Müller-Wickop present an algorithm that can be used to map events to cases satisfying
requirement VII (Gehrke and Müller-Wickop, 2010).
Traditional mining algorithms use the temporal ordering of events to infer the control flow.
This is not possible if the events are not strictly ordered. Instead of relying on the temporal
ordering it is possible to derive the control flow by analyzing the data dependency between
events in a case. The control flow can be determined by analyzing which activity cleared
items that were created by other activities. This is illustrated in Figure 23.

44
DESIGN

0001900113
1 1

[15029.81] [15029.81]
0002811000 0012490379 0007904673 0005035200

Post
Payment
[15029.81] with Clearing [17.47]

2010/09/03 2010/09/03
0005004040
0001900111
[15062.42]
[15.14]

0001900113

[15029.81]

Figure 23 Logical Dependency (Werner and Nüttgens, 2014, p. 3892)


The activity Payment in Figure 23 posted an item with the value of 15,029.81 on the account
0001900113. This item was cleared by the activity Post with Clearing. The logical sequence
for these activities can therefore be derived as Payment → Post with Clearing. This ap-
proach can be used for all activities in a process instance to infer the complete control flow.
Start nodes can be determined by identifying those activities that do not have any incoming
control arcs. End nodes do not have any outgoing control arcs (Werner and Nüttgens,
2014).
A problem with this approach is the occurrence of clearing deadlocks. This constellation is
shown in Figure 24. The shown Clear Postings activity did not create any other posted item
and would therefore be defined as an end node. This is a problem because the clearing
activity is actually not an end node but just an intermediary node in the process model.

0002810200 0002810200

1 1 1
[914.68] [914.68] [914.68] [914.68]

0050157332 0095359370 0015980342

Post Post
Clear Postings
Received Goods Received Invoice

2010/02/16 2010/09/03 2010/09/03


[68.08] [910.99]
[846.60] [3.69]

0001400100 0004000070 0002811000 0004000070

Figure 24 Clearing Deadlock (Werner and Nüttgens, 2014, p. 3892)


Clearing deadlocks can be removed by using graph transformation techniques (Rozenberg,
1997). Clearing activities can be identified by selecting all activities that did not post any
items but only cleared items from other activities. Clearing activities commonly clear items
from two or more other activities. At least one of these activities must have posted an
additional item that was cleared by another activity different from the clearing activity for
a deadlock to appear. Otherwise the clearing activity would represent a valid end node and

45
DESIGN

such a constellation would not be considered a clearing deadlock (Werner and Nüttgens,
2014). The procedure to remove a clearing deadlock is graphically illustrated in Figure 25.
Removing clearing deadlocks reduces the complexity and improves the readability of mined
models but it violates requirement III described in chapter 4.2.3 because it alters the origi-
nal data (Werner and Nüttgens, 2014). Applications in practice are necessary to evaluate if
the alteration is acceptable or if it may not be used for specific purposes.

Activity B Activity C Activity B Activity C

Clearing Clearing
Activity X Activity X

Activity A Activity A

a) Deadlock Constellation b) Deadlock Resolution

Figure 25 Deadlock Resolution (Werner and Nüttgens, 2014, p. 3893)


In summary the control flow of non-linear event logs can be inferred by:
(1) Defining causal dependencies
(2) Removing clearing deadlocks
(3) Defining start and end nodes

5.5 Multilevel Process Mining Algorithm


Chapters 5.1 to 5.4 showed different constructs, models and methods that solve specific
problems for mining business process models in the context of financial audits. Listing 2
shows the final Multilevel Process Mining Algorithm (MLPM) that was designed by integrat-
ing the results from iterative research cycles. The individual research results have been de-
scribed in the previous chapters and form the foundation for the algorithm laid out in the
Listing 2.

Listing 2 Multilevel Process Mining Algorithm (Werner and Gehrke, 2015, p. 824)
1. Mine Cases (Section 1)
2. D set of all posting document numbers
3. J set of all journal entry item numbers

Di = ∅ initially empty set of document numbers belonging to case i ∈ ID


4. ID set of all case IDs

Ji = ∅ initially empty set of journal entry item numbers for case i ∈ ID


5.

While D ≠ ∅
6.

Remove d ∈ D from D and insert d into Di


7.

Insert all j ∈ J posted by d into Ji and remove j from J


8.

Insert all d ∈ D that cleared j ∈ Ji into Di and remove d from D


9.

Repeat 54. and 55. for all d ∈ Di and j ∈ Ji


10.
11.

46
DESIGN

IG = ∅
12. Reconstruct Instance Graphs (Section 2)
13. initially empty set of instances graphs

Ni ≠ ∅, Ei⊆ Ni× Ni is the set of arcs, Li the set of task labels


14. IGi instance graph (Ni,Ei,Li,li) for case i with the set of nodes

and li: Ni→ Li is a labeling function mapping nodes onto Li


15.

17. For all i ∈ ID


16.

Create n ∈ Ni for each d ∈ Di with li(n)= transaction code of d


Create e(nj,nk) ∈ Ei for each dj and dk ∈ Di if dk cleared an item
18.

j ∈ Ji that was posted by dj


19.
20.
21. Insert IGi into IG

IM = ∅
22. Reconstruct Instance Models (Section 3)

instance model (Ti,Pi,Ai,Σi,Vi,Ci,Gi,Ei,Ii)for case i ∈ ID


23. initially empty set of instances models

25. For all i ∈ ID


24. IMi

For each e(nj,nk)∈ Ei create p ∈ Pi, a(tj,p), a(p,tk)


26. Set Ti = Ni

For each d ∈ Di
27.

Create p ∈ Pi for each j ∈ Ji that was posted by d


28.

Create a(t,p) ∈ Ai for each j ∈ Ji that was posted by d and


29.

a(t,p), a(p,t) ∈ Ai for each j ∈ Ji that was cleared by d


30.

Aggregate all places pk and pj ∈ Pi if Ci(pk) = Ci(pj)


31.

Aggregate all transitions tk and tj ∈ Ti if li(tk) = li(tj)


32.
33.
34. Insert IMi into IM

PM = ∅
35. Mine Process Models (Section 4)

37. Compute causal matrix CM(IMi) for all i ∈ ID


36. initially empty set of process models

38. While IM ≠ ∅

For each IMk ∈ IM


39. Remove IMj from IM and insert IMj into PM
40.

Merge IMj and IMk with Tjk=Tj∪Tk, Pjk=Pj∪Pk, Ajk=Aj∪Ak, Σjk=Σj∪ Σk


41. If CM(IMj) = CM(IMk)

Aggregate all places pl and pm ∈ Pjk if Cjk(pl) = Cjk(pm)


42.

Aggregate all tl and tm ∈ Tjk if ljk(tl) = ljk(tm)


43.
44.
45. Remove IMk from IM

A complete description of the algorithm is available in (Werner and Gehrke, 2015). The
algorithm is divided into the sections Mine Cases, Reconstruct Instance Graphs, Reconstruct
Instance Models, and Mine Process Models. The algorithm produces models at different
levels of abstraction (compare chapter 3.2.3) and therefore fulfills requirement VI de-
scribed in 4.2.3. It takes unlabeled and non-linear event logs as input and produces CPN as
specified in chapter 5.2. It integrates method fragments to label the event log by matching
events to cases in Section 1, infers the control flow based on data dependency in Section 2,
integrates the data perspective in Section 3 and reduces the complexity of mined models
in Section 3 and 4.
Figure 26 shows an example of a mined process model using the MLPM algorithm. It illus-
trates the control flow by showing the sequence of activities and the data flow by illustrat-
ing the amounts that were posted by the different activities on the financial accounts. The
process model represents eight process instances. The MLPM algorithm can also be used
to create the related instance models to enable the auditor to inspect individual instances
on a detailed level.

47
DESIGN

0005035200
[5,208.68]

0004002000
0002811000
[117.28] [35.55] 0001900101
[36,937.40]
0004000070
[35.55]
[308.86] [7,426.98] 0005004900 [648,155.17]
[35.55] 0001900103 0003160100
[36,922.40] [36,922.40] [4,429.73]
[12] [12]
0004000070
[49,864.85] [9,364.44] 0005004900 0001900311
0001900313
(B) Enter
Incoming Invoice
[35.55]
0001421100
0002811000 [366,410.06] [8]
[365,748.82] (E) Post with
[365,748.82] Clearing [36,922.40] 0001900103
(D) Payment
[593,846.02] [648,190.72] [648,190.72] [17] [17]

0004000010 0004000110
0001900111
[83] [1,078.42]
[83] [475.10]
[8] (A) Post Goods [245,171.63]
[8] [96] [96] [83]
Receipt [643,402.01] (C) Clear 0001900113
Account [83]
0001900113
0002810200
0002810200
[245,008.85] [245,008.85]
[245,008.85]
[645,022.01] [645,022.01] 0001900313
[643,402.01] 0005004040
[25] [25] [365,748.82]
[60.07]

Figure 26 Example of a Mined Process Model


(adapted from Werner and Gehrke, 2015, p. 828)24

5.6 Software Prototype


The research artifacts described in 5.2 to 5.5 Table 10 Prototype Modules
were implemented in a software prototype. Module Functionality
The conceptual structure of the prototype is Core Graph Modeling
shown in Figure 27. The prototype was imple- M1 Case Matching and Labeling
mented using the JAVA programming language M2 Reconstruction
(Oracle, 2014a) and the integrated develop- M3 Aggregation
ment environment (IDE) NetBeans (Oracle, M4 Graphical User Interface
2014b). The software prototype uses event M5 SQL Database Management
logs as input. The event logs are provided by a M6 Graph Database Management
separate extraction module that extracts rele- M7 File Writing
vant data tables from the source ERP systems. M8 Utilities
These are stored in a relational SQL database. M9 Configuration
The software consists of 10 software packages as listed in Table 10. The core package pro-
vides the functionality for graph modeling. Separate modules can be used for case match-
ing and labeling, reconstruction of models based on the matched and labeled event data
and aggregation for complexity reduction. Additional modules are used for managing the
connection to the databases, the creation of output files and auxiliary purposes. The mod-
ular structure facilitates the integration of additional modules. The prototype uses two dif-
ferent kinds of databases. The original event log data is stored in a traditional relational
SQL database (H2 Database Engine, 2014). The mined graphs are stored in a graph database
Neo4j (Neo Technology Inc., 2014) for performance reasons.
The prototype is able to create a variety of different output files that can be used as input
for other software tools. Several independent software applications were used for specific
purposes in the research cycle. The software yEd (yWorks GmbH, 2015) provides powerful
layout algorithms and was used to graphically represent the mined models. Renew is a CPN

24
The arc inscriptions for posting and clearing arcs only display the assigned constant for the posted or
cleared value. The inscriptions for the account type, account number and credit or debit indicator are
omitted for better readability. The same is the case for the inscriptions of the connected account places
that only show the account number.

48
DESIGN

tool (University of Hamburg, 2015) that was used to simulate the execution of mined mod-
els to verify if the created models behaved like expected. Statistical analyses were carried
out using Stata (StataCorp LP, 2014). ProM (Process Mining Group, 2015) and Disco (flux-
icon, 2015) were used for evaluation purpose to compare mining results created by the
prototype with those created by general purpose mining algorithms.

Figure 27 Prototype Structure


Figure 28 and Figure 29 show the prototype’s graphical user interface (GUI). It is split into
two screens. The main screen can be used to mine processes and to produce the different
kind of output files. The created files are shown on the Visualization panel and can be
opened using the yEd, Renew, ProM or Disco software. Two Display panels provide infor-
mation on selected entries. The configuration screen provides configuration options on
basic and advanced parameters.

49
DESIGN

Figure 28 Prototype GUI Main Screen

Figure 29 Prototype GUI Configuration Screen

50
EVALUATION

6 Evaluation
Design science-oriented research is often criticized for its lack of scientific rigor due to in-
sufficient evaluation efforts (Österle et al., 2010). Evaluation is an essential part of DSR as
highlighted in Figure 30. The research presented in this paper was set up to ensure that
sufficient evaluation can be carried out. The instantiation of the designed artifacts in a soft-
ware prototype set the foundation to observe the behavior of the designed artifacts when
exposed to test and real life data and to analyze the output that was created by the proto-
type. The following sub-chapters describe the evaluation efforts and results that could be
achieved using different evaluation methods (compare chapter 3.4).

Environment IS Research Knowledge Base


Relevance Rigor
People Develop / Build Foundations
• Theories
Business
• Artifacts Applicable
Needs Knowledge
Organizations

Justify / Evaluate Methodologies


• Quantitative
Technology Research Methods
• Qualitative Research
Methods

Application in the Additions to the


Appropriate Environment Knowledge Base

Figure 30 Evaluation in the Information Systems Research Framework (adapted from


Hevner et al., 2004, p. 80)

6.1 Experimental Setup


A laboratory experiment was used to test the designed artifacts and to create output for
observation and analysis purposes. Several experiments were conducted during the itera-
tive research cycles but their basic structure was similar. It is illustrated in Figure 31. The
first step was the extraction of the event data from the source systems using the extraction
module. The extracted data was checked for first and second order defects (Kemper et al.,
2010, pp. 29–32). The event log data served as input for the mining module. The imple-
mented methods were used to create process models. The created models were observed
using the yEd software to check if they fulfilled the desired requirements. The manual in-
spection was especially important in the first research cycles to check the proper operation
of the implemented methods. The created models were further checked in respect to cer-
tain characteristics in simulations by using the Renew CPN tool (compare chapter 6.2). The
output from the prototype was additionally compared to the results that could be achieved
by using general purpose mining algorithms in ProM and Disco (Werner and Nüttgens,
2014). The created models and selected model characteristics were finally used for quanti-
tative analyses (compare chapter 6.3).

51
EVALUATION

Input ERP Data Event Log Process Models Process Models Event Log Process Models

Activity Data Extraction Mining Observation Verification Comparison Quantitative Analysis

Qualitative Quantitative Quantitative


Output Event Log Process Models Process Models
Data Data Data

Software
Extraction Module Mining Module yEd Renew Disco / ProM Stata
Components

Figure 31 Experimental Setup

6.2 Simulation
Simulations can be used to gain information if it is difficult to represent the subject of in-
vestigation in a formal mathematical model due to its inherent complexity. The examples
from chapter 5.3 illustrate that the complexity of the mined models tended to be very high.
The Renew software tool was used to test on a sample basis if the prototype created sound
process models according the definition used by (van der Aalst, 2011a, p. 39; Weske, 2012,
pp. 326–329) with regard to:
• proper completion
• option to complete
• absence of dead transitions
• safeness
These criteria are formulated for low-level Petri Nets. The colored account places as de-
fined in the CPN specification in chapter 5.2 were excluded for completion and safeness
testing due to the fact that account places can hold multiple tokens that represent the
journal entry items and that by definition remain on the account places even after the pro-
cess is finished.
Figure 32 shows the CPN model from Figure 19 in Renew. Figure 32 shows the model in the
initial state and Figure 33 the same model after the simulation has been executed. Every
transition has fired and produced the tokens on the account places that represent the
posted journal items.

52
EVALUATION

Figure 32 Example Process Instance before Simulation

Figure 33 Example Process Instance after Simulation

6.3 Results Analysis


The produced process instances were used as a source data for descriptive analyses. Se-
lected model characteristics like the number of transitions, model complexity or the num-
ber of represented instances in a process model were computed and used as input for sta-
tistical analyses (Werner et al., 2013). Four different source data sets were used during the
overall research process. The first set was extracted from the SAP IDES test system. The
test system is available for universities participating in the SAP University Alliance Program
(SAP, 2015). Three data sets were gathered from companies operating in the retail, manu-
facturing and media industries. Table 11 provides an overview of the volume of the used
data sets.

53
EVALUATION

Table 11 Evaluation Data Sets (adapted from Werner et al., 2013, p. 381)
1 2 3 4
Data Set
SAP IDES Retail Manufacturing Media
Number of journal entries 115,060 92,487 1,764,773 156,604
Number of journal entry items 419,106 222,901 7,395,434 559,506
Number of process instances 81,171 40,130 1,035,805 18,975
Number of processes 361 307 841 516
Covered period 17 years 1 year 1 year 1 year

The source data was used to create process instance and process models. Figure 34 to Fig-
ure 36 show the distributions of the number of process instances over the number of net
elements with logarithmic scaling on the x- and y-axes. They illustrate that only very few
instances consist of very many net elements. The vast majority of instances consist of rela-
tively few net elements. Table 12 provides an overview of specific characteristic values of
the distributions.

Figure 34 Data Set 1 Distribution of Number of Net Elements over the Number of In-
stances (Werner et al., 2013, p. 382)

Figure 35 Data Set 2 Distribution of Number of Net Elements over the Number of In-
stances (Werner et al., 2013, p. 382)

54
EVALUATION

Figure 36 Data Set 3 Distribution of Number of Net Elements over the Number of In-
stances (Werner et al., 2013, p. 382)

Table 12 Overview of Net Size Distribution Characteristics (Werner et al., 2013, p. 383)
#1 #2 #3
Data Set
SAP IDES Retail Manufacturing
Mean value of net ele-
15.31 21.28 20.30
ments per instance
Median value of net ele-
9 9 7
ments per instance
Maximum of net ele-
32,519 275,870 4,769,379
ments per instance
Standard deviation of
127.83 1,380.41 4,688.30
number of net elements

Figure 37 to Figure 39 show the distributions of transaction code combinations over the
number of instances. The y-axes follow a logarithmic scaling. Each number on the x-axis
represents a transaction code combination. The majority of instances only contain very few
combinations. In combination with the results from analyzing the distribution of net sizes
it can be assumed that the majority of instances are very limited in size and reveal the same
transaction code combinations.

Figure 37 Data Set 1 Distribution of the Number of Instances for Different Transaction
Code Combinations (Werner et al., 2013, p. 383)

55
EVALUATION

Figure 38 Data Set 2 Distribution of the Number of Instances for Different Transaction
Code Combinations (Werner et al., 2013, p. 384)

Figure 39 Data Set 3 Distribution of the Number of Instances for Different Transaction
Code Combinations (Werner et al., 2013, p. 384)
Figure 40 to Figure 42 show the distributions of accounts in the instance models. Table 3
provides an overview of characteristic values of these distributions. The maximum number
of used accounts in a process instance model is relatively low compared to the maximum
of possible net sizes listed in Table 12. For the interpretation of the illustrated distributions
and their characteristic values it is helpful to take into account that only one instance in
data set two uses the maximum number of 596 accounts and one in data set three the
maximum of 399 accounts. All other instances do not use more than 63 accounts in data
set two and no more accounts than 45 in data set three. The observation of the distribu-
tions reveals that journal entry items are posted to relatively few accounts in a specific
process instance. This is reasonable because a specific process generally only uses a subset
of the available set of accounts.
Table 13 Overview of Account Distribution Characteristics (Werner et al., 2013, p. 385)
1 2 3
Data Set
SAP IDES Retail Manufacturing
Mean value 2.95 2.30 2.33
Median value 2 2 2
Maximum value 36 596 399
Standard deviation 1.61 3.37 0.92

56
EVALUATION

Figure 40 Data Set 1 Distribution of the Number of Instances over Number of Accounts
(Werner et al., 2013, p. 386)

Figure 41 Data Set 2 Distribution of the Number of Instances over Number of Accounts
(Werner et al., 2013, p. 386)

Figure 42 Data Set 3 Distribution of the Number of Instances over Number of Accounts
(Werner et al., 2013, p. 386)

achieved fitness and precision. The metrics <lmno and 4lmno were used to measure fitness
The mined process models from data sets 2 to 4 were further inspected in terms of

and precision (Werner and Gehrke, 2015) with:

|/q | q ∈ lrs ∧ ∈ luv 0| |lrs |w|/q | q ∈ lrs ∧ ∉ luv 0|


<lmno . |luv |
4lmno . |lrs |

57
EVALUATION

The average values over the data sets 2 to 4 for these metrics were:
<lmno = 1
4lmno = 0.8125
Figure 43 and Figure 45 show the frequency distributions of the mined process models de-
pending on their model complexity measured as the number of included transitions. They
show that the process models are distributed similarly to a normal distribution. Figure 44
and Figure 46 present scatter diagrams for data sets 3 and 4. They illustrate the distribution
of the number of represented instance models in a process model depending on the model
size. The diagrams show that the vast majority of process instances actually belong to very
simple process models that contain only a few transitions.
20

1.0e+06

100000

Number of Represented Instances


15

10000
Frequency

1000
10

100
5

10

1
0

0 5 10 15 0 5 10 15
Number of Transitions Model Complexity

Figure 43 Frequency Distribution for Data Figure 44 Scatter Diagram for Data Set 326
Set 3 (Werner and Gehrke, 2015, p. 829) (Werner and Gehrke, 2015, p. 829)
20

1.0e+06

100000
Number of Represented Instances
15

10000
Frequency

1000
10

100
5

10

1
0

0 5 10 15 20 25 0 5 10 15 20 25
Number of Transitions Model Complexity

Figure 45 Frequency Distribution for Data Figure 46 Scatter Diagram for Data Set 426
Set 4 (Werner and Gehrke, 2015, p. 829) (Werner and Gehrke, 2015, p. 829)

25
Process models including loops were excluded from the calculation.
26
The dependent variables in Figure 44 and Figure 46 use a logarithmic scaling.

58
EVALUATION

6.4 Requirements Fulfillment


Several requirements have been identified as being relevant for the development of a pro-
cess mining algorithm for financial audits (compare chapter 4.2.3). Table 14 provides an
overview of the different requirements and how they were met.
Table 14 Requirements Overview
Requirement Accounted for by Fulfillment
The mined process models should pro- Using CPN to model the control
I vide information on the control flow of flow. Yes
process activities.
The mined process models should pro- Using CPN to integrate the data
vide information on the relationship of flow perspective.
II Yes
business processes and financial ac-
counts.
The mining algorithm should not alter Simulation results show that
the original source data for creating the algorithm does not alter
III process models. the original source data and Yes27
that it creates sound process
models.
The mined process models should be Analysis results show that a
IV as precise as possible. high but not perfect precision Partially
is achieved.
The mined process models should be Analysis results show that a
V Yes
as fitting as possible. perfect fitness is achieved.
The mining algorithm should be able to The algorithm is able to create
produce process models at different models on different abstrac-
abstraction levels. tion levels (instance graphs
VI Yes
representing the source data
structure, process instance
models and process models).
The mining algorithm should be able to The algorithm preprocesses
VII use unlabeled event logs as input. the input data and assigns a Yes
case to each event.
The mining algorithm should be able to The algorithm is able to take
use event logs with non-linear traces event logs with non-linear
VIII as input. traces as input and infers the Yes
control flow by using data de-
pendencies.

7 Diffusion
The last phase in DSR is the diffusion of the achieved research results into the application
domain and the scientific knowledge base. Scientists use a variety of communication types
to distribute achieved research results amongst scholarly colleagues and practitioners. Pub-
lication types encompass conference articles, presentations, scientific- or practice-oriented

27
The MLPM algorithm can be used with our without enabled clearing deadlock resolution. If the resolution
is disabled no alteration of the original source data takes place.

59
DIFFUSION

reports, dissertations or habilitations, theses, textbooks, websites, guidelines, standards,


lectures, seminars, research proposals, implementations, spin-offs and so forth.
The publication of scientific articles as book chapters, in conference proceedings or journals
in combination with the presentation of these articles in scientific conferences constitute
the primary diffusion types that were used in the course of the presented research. Ten
different publications were published that are listed in chapter 3.5. The publications fol-
lowed the iterative research cycle (compare Figure 5 in chapter 3.3). All publications went
through a rigorous review process. 28 independent and formal reviews from scientists were
received for the listed publications in the course of submission, revision and acceptance of
the published papers. The review and feedback from scientific experts in the relevant re-
search area ensured a high quality of the research outputs and provided an external eval-
uation of the achieved results that could be used to improve the research outputs.
The research from the various publications was presented at scientific conferences. Table
15 lists the conferences and presented papers.
Table 15 Overview Conferences and Presented Papers
EMISA 2011 Discussion and presentation of the research exposé at the doctorial con-
sortium of the 4th International Workshop on Enterprise Modeling and
Information Systems Architectures in Hamburg
HICSS 2012 Presentation of a the research paper Business Process Mining and Re-
construction for Financial Audits (Werner et al., 2012a) at the 45th Hawaii
International Conference on System Sciences in Wailea
BPM 2012 Discussion and presentation of the research exposé at the doctorial con-
sortium of the 10th Business Process Management conference in Tallinn
ICIS 2012 Presentation of the Research-in-Progress paper Tackling Complexity:
Process Reconstruction and Graph Transformation for Financial Audits
(Werner et al., 2012b) at the 33rd International Conference on Infor-
mation Systems in Orlando sponsored by a scholarship of the German
Academic Exchange Service (DAAD)
WI 2013 Presentation of the paper Towards Automated Analysis of Business Pro-
cesses for Financial Audits (Werner et al., 2013) at the 11th International
Conference on Wirtschaftsinformatik in Leipzig
ER 2013 Presentation of the paper Colored Petri Nets for Integrating the Data
Perspective in Process Audits (Werner, 2013) at the 32nd International
Conference on Conceptual Modeling (ER 2013) in Hong Kong
HICSS 2014 Presentation of the paper Improving Structure: Logical Sequencing of
Process Models (Werner and Nüttgens, 2014) at the 47th Hawaii Interna-
tional Conference on System Sciences in Waikoloa sponsored by a schol-
arship of the German Academic Exchange Service (DAAD)
ECIS 2014 Presentation of the paper Who is afraid of the Big Bad Wolf - Structuring
Large Design Science Research Projects (Werner et al., 2014) at the 22nd
European Conference on Information Systems in Tel Aviv

60
DIFFUSION

Research results were also continuously published on the research project website (Uni-
versity of Hamburg, 2014) and presented in workshops to the project partners from the
industry to achieve a diffusion into the application domain. The presented research results
also served as input for the development of a commercial software application that is cur-
rently under development. The application of the presented academic prototype in field
experiments is planned for future research.
The positive feedback from scientists in the form of reviews for published papers and dis-
cussions at several conferences has shown that the aim has been achieved to distribute the
research results into the academic knowledge base. Distribution into the application do-
main was initiated via the cooperation with project partners from industry and contacts
with commercial software companies such as fluxicon, the software development company
for the process mining tool Disco (fluxicon, 2015).
The presented research results have further been used as a scientific foundation to prepare
and submit two research proposals to the German Research Foundation (DFG).

8 Summary and Outlook


8.1 Summary
Financial audits play a significant role in our economic system as an important control
mechanism to safeguard the correctness and reliability of published financial information.
Process audits are an important part of financial audits but public accountants face difficul-
ties in conducting these audits with the ongoing integration of information systems for the
operation of business processes, increased process complexities and growing amounts of
processed data. Manual audit procedures that are traditionally used to carry out process
audits become inefficient or even ineffective in environments that exhibit such character-
istics.
The aim of the research that is presented in this thesis is the development of data analysis
techniques that can be used to support auditors in process audits and that are able to pro-
duce reliable process models by using the data which is stored in information systems that
process financially relevant transactions. It has been carried out using a design science-
oriented research approach due to the proximity of the research question to practical prob-
lems and the objective to create research artifacts that contribute both to the scientific
knowledge base and the application domain. This aim has been achieved by the develop-
ment of several research artifacts and in particular a special purpose mining algorithm that
takes the requirements from the application domain and knowledge from the scientific
knowledge base into account.
The research followed the design science phases analysis, design, evaluation and diffusion.
These phases were iterated in several cycles leading to different developed methods that
have been formalized as algorithms as well as constructs and models for describing the
problem domain and solution space. The main output of the research is a Multilevel Process
Mining algorithm that produces a variety of models. The algorithm integrates the control
flow and data flow perspective. It operates on different abstraction levels, creates precise
and fitting process models, accepts unlabeled and non-linear event logs from ERP systems

61
SUMMARY AND OUTLOOK

as input, and considers data relationships to infer the control flow. It is an innovative solu-
tion especially designed for the application in financial audits. Gregor and Hevner provide
a framework that can be used to categorize the knowledge contribution of designed arti-
facts in DSR (Gregor and Hevner, 2013). They use the four domains routine design, improve-
ment, exaptation and invention to categorize the knowledge contribution. The presented
research work can be assigned to the exaptation quadrant because the main objective is to
provide a solution for a new application area by partly using already existing knowledge.
But it also affects the improvement quadrant by introducing new methods to model the
control flow and data flow simultaneously in mined models, to label event logs from ERP
systems and to infer the control flow relying on data dependencies between events rec-
orded in non-linear event logs. Gregor and Hevner further differentiate between three lev-
els of contribution types that range from abstract, complete and mature knowledge on the
highest level to more specific, limited and less mature knowledge on the lowest level. The
research results presented in this thesis are mainly located on the second level providing
constructs and methods for the mining of process models and on the first level presenting
an instantiated software artifact. The results that can be achieved by analyzing the mining
outcomes can also be input for the third and highest knowledge contribution level.
Although the MLPM algorithm is meant to be a special purpose mining algorithm many
research results that were achieved on the route to the final algorithm are also generally
applicable. The integration of the data perspective is an important aspect for process min-
ing and it has not been addressed yet sufficiently in the academic arena (de Leoni and van
der Aalst, 2013; Stocker, 2012). The provided CPN specification shows a solution how the
data perspective can be integrated which can also be adapted for other application areas.
Contemporary process mining algorithms require labeled event logs and strict linearly or-
dered events. The MLPM algorithm accepts unlabeled and non-linear event logs as input.
Labeling events is an important aspect in process mining and not well researched (Ferreira
and Gillblad, 2009). Non-linear event logs are also present in other application areas like
process mining in online discussion forums (Wang et al., 2014).
The research artifacts have been implemented in an academic software prototype that can
be used in real scenarios and that is able to take test and real life data from SAP ERP systems
as input. The researched artifacts have been evaluated extensively in simulations with test
and real data. The produced models have been analyzed using descriptive statistics. They
provide a novel empirical data base for observing real business processes in organizations.
Knowledge on business processes gained on the basis of data that is created by using pro-
cess mining techniques is still scarce. The provided information on the mined models can
be seen as a first step to broaden this knowledge base.
The MLPM algorithm is supposed to be used to support public accountants in process au-
dits. The application of automated data analysis techniques alone will most likely not pre-
vent accounting scandals. But it can be used as a tool to improve process audits. Its appli-
cation enables auditors to receive all-embracing and reliable information on the audited
business processes and their relationships to the financial accounts. It can be used as a
foundation to automate the analysis and audit of standard processes that generally exhibit
a lower risk than non-standard process that usually exhibit a higher risk. Efficient and ef-
fective automated analyses of standard processes set free audit resources that can then be
spent on non-standard transactions.

62
SUMMARY AND OUTLOOK

8.2 Limitations
The MLPM algorithm is able to discover process models in accordance with the identified
requirements to a large extent as discussed in chapter 6.4. Requirement IV could just par-
tially be met. The mined models are not absolutely precise. It has to be validated in further
research if the achieved level of precision is sufficient in practice.
Several additional limitations have to be taken into account when discussing the achieved
research results. The mined process models do not represent sound workflow nets accord-
ing to commonly used definitions (van der Aalst, 2011a, p. 39; Weske, 2012, pp. 326–329).
The MLPM algorithm can be used to mine precise and fitting process models based on the
available source data from ERP systems. Formally well-structured process models are not
critical from this point of view. Sound process models could be achieved by neglecting the
data perspective by not modeling the account places. The remaining models would then
represent sound workflow nets but without representing the data flow perspective.
All test and real data sets were extracted from SAP ERP systems. It can therefore not be
concluded that the research results also hold true for other data sources. But the MLPM
algorithm exploits the general structure of accounting entries as described in chapter 4.2.2.
This structure is independent from the implemented data structures of a particular ERP
system.
The resolution of deadlocks as discussed in chapter 5.4 changes the original source data
and therefore violates requirement III described in chapter 4.2.3. The deadlock resolution
can be enabled or disabled in the implemented prototype. Field experiments could provide
further information on the feasibility of the deadlock resolution in real settings.
Some mined process models showed loops. These loops can occur when a transaction has
cleared a journal item that was posted by the same transaction or by a transaction located
in the subsequent execution path. This constellation leads to a deadlock in the process
model. Such a deadlock is not critical for the interpretation of the model from an audit
perspective but generally not desired for the modeling of correct process models. A solu-
tion could be the prevention of aggregating transitions carrying the same label if this would
result in a loop.
The mining algorithm produces precise and fitting process models at the cost of lacking
generalization. It is therefore not applicable for scenarios with highly variable business pro-
cesses. In the worst case scenario all process instances show a different behavior. The min-
ing algorithm would then produce a process model for each process instance. The data
presented in Table 11 shows that this will most likely not occur in the context of financial
audits. Business processes are usually standardized to a certain degree when they are sup-
ported by ERP systems. The data shows that the number of process models ranges from
307 for the smallest data set to 841 models for the largest. This may still seem to be a big
number, but many of the process models represent trivial processes. 63 process models in
data set 3 only consist of one transition. The activities in these processes were mostly car-
ried out by using a single general purpose transaction. They are of little interest from a
process perspective. The process models that reflect the major business processes are
those that contain many transitions and represent a high number of process instances. Data
set 3 contains 135 process models consisting of 5 transitions. But just two of them already

63
SUMMARY AND OUTLOOK

represent 62% of the instances of this category. It can therefore be assumed that the ma-
jority of models for more complex processes only represent very infrequent behavior and
can be tested traditionally by inspecting individual journal entries.

8.3 Outlook and Future Research


The presented research results support the automated discovery and analysis of process
models in process audits and lay the foundation for the development of completely auto-
mated audit procedures as conceptually illustrated in chapter 5.1. The next research step
would be the integration of application controls.
Several aspects that are important from an audit perspective have not been considered
sufficiently yet. Materiality is a key concept in financial audits. Research on the role of busi-
ness processes from a materiality perspective is currently pending. The presented MLPM
algorithm is able to produce process models. But it does not provide an overall process
map as used as a metaphor in chapter 5.1. Methods for mining process maps are currently
under development but have not been evaluated yet.
The prototype uses data from SAP systems. Data from other source systems have not been
considered extensively yet. The development of extensions to integrate other data sources
is intended for future research. The used event log data can be used as input for academic
and commercial software tools like ProM and Disco. But due to the non-linear ordering of
events in these logs the non-linear traces have to be sliced into linear traces as described
in (Müller-Wickop and Schultz, 2013b). But this slicing leads to undesired side effects like
the duplication of events and related data values. The integration of the designed solutions
into academic software tools like ProM is planned for future research.

64
BIBLIOGRAPHY

9 Bibliography
Accorsi, R., Lehmann, A., 2012. Automatic Information Flow Analysis of Business Process
Models, in: Barros, A., Gal, A., Kindler, E. (Eds.), Business Process Management, Lec-
ture Notes in Computer Science. Springer, pp. 172–187.
Adriansyah, A., van Dongen, B.F., van der Aalst, W.M.P., 2011. Conformance Checking Us-
ing Cost-based Fitness Analysis, in: 15th IEEE International Enterprise Distributed
Object Computing Conference. pp. 55–64.
Agrawal, R., Gunopulos, D., Leymann, F., 1998. Mining Process Models from Workflow
Logs, in: Proc. Sixth Int’l Conf. Extending Database Technology. pp. 469–483.
AIS, 2014. Senior Scholars’ Basket of Journals [WWW Document]. URL [Link]
[Link]/?SeniorScholarBasket (accessed 7.16.14).
Alturki, A., Gable, G.G., Bandara, W., 2011. A Design Science Research Roadmap, in: Ser-
vice-Oriented Perspectives in Design Science Research. Springer, pp. 107–123.
Archer, L.B., 1984. Systematic Method for Designers, in: Developments in Design Method-
ology. John Wiley, London, pp. 57–82.
Australian Research Council, 2014. The Excellence in Research for Australia (ERA) - Aus-
tralian Research Council (ARC) [WWW Document]. URL [Link]
(accessed 8.21.14).
Baskerville, R., Lyytinen, K., Sambamurthy, V., Straub, D., 2010. A Response to the Design-
oriented Information Systems Research Memorandum. European Journal of Infor-
mation Systems 20, 11–15.
Becker, J., Probandt, W., Vering, O., 2012. Grundsätze ordnungsmäßiger Modellierung
Konzeption und Praxisbeispiel für ein effizientes Prozessmanagement. Springer Ga-
bler, Berlin; Heidelberg.
Bezerra, F., Wainer, J., 2013. Algorithms for Anomaly Detection of Traces in Logs of Pro-
cess Aware Information Systems. Information Systems 38, 33–44.
Bhattacherjee, A., 2012. Social Science Research: Principles, Methods, and Practices. A.
Bhattacherjee, Tampa, Fla.
Bierstaker, J., Janvrin, D., Lowe, D.J., 2014. What Factors Influence Auditors’ Use of Com-
puter-assisted Audit Techniques. Advances in Accounting 30, 67–74.
BPM, 2014. BPM 2014 [WWW Document]. URL [Link] (accessed
10.29.13).
Braun, R.L., Davis, H.E., 2003. Computer-assisted Audit Tools and Techniques: Analysis
and Perspectives. Managerial Auditing Journal 18, 725–731.
Brinkkemper, S., 1996. Method Engineering: Engineering of Information Systems Develop-
ment Methods and Tools. Information and Software Technology 38, 275–280.
Caron, F., Vanthienen, J., Baesens, B., 2013. Comprehensive Rule-Based Compliance
Checking and Risk Management with Process Mining. Decision Support Systems 54,
1357–1369.

65
BIBLIOGRAPHY

Chen, H., Chiang, R.H.L., Storey, V.C., 2012. Business Intelligence and Analytics: From Big
Data to Big Impact. MIS Quarterly 36, 1165–1188.
Chuprunov, M., 2012. Handbuch SAP-Revision: internes Kontrollsystem und GRC. Galileo
Press, Bonn.
Cole, R., Purao, S., Rossi, M., Sein, M.K., 2005. Being Proactive: Where Action Research
Meets Design Research, in: Proceedings of the 26th International Conference on In-
formation Systems.
Cook, J.E., Wolf, A.L., 1999. Software Process Validation: Quantitatively Measuring the
Correspondence of a Process to a Model. ACM Transactions on Software Engineer-
ing and Methodology 8, 147–176.
Cook, J.E., Wolf, A.L., 1998a. Discovering Models of Software Processes from Event-based
Data. ACM Transactions on Software Engineering and Methodology 7, 215–249.
Cook, J.E., Wolf, A.L., 1998b. Event-based Detection of Concurrency. ACM SIGSOFT Soft-
ware Engineering Notes 23, 35–45.
CORE, 2014. Computing Research & Education & Conference Rankings [WWW Docu-
ment]. URL [Link] (accessed 8.21.14).
de Leoni, M., van der Aalst, W.M.P., 2013. Data-Aware Process Mining: Discovering Deci-
sions in Processes Using Alignments, in: 28th Annual ACM Symposium on Applied
Computing. Coimbra, Portugal, pp. 1454–1461.
de Medeiros, A.K.A., 2006. Genetic Process Mining. Eindhoven University of Technology,
Eindhoven.
Deutscher Bundestag, 2013. Handelsgesetzbuch.
Diestel, R., 2010. Graph Theory, 4th ed. Springer, Heidelberg; New York.
Dumas, M., La Rosa, M., Mendling, J., Reijers, H.A., 2013. Fundamentals of Business Pro-
cess Management. Springer.
Eekels, J., Roozenburg, N.F.M., 1991. A Methodological Comparison of the Structures of
Scientific Research and Engineering Design: Their Similarities and Differences. De-
sign Studies 12, 197–203.
Ferreira, D., Gillblad, D., 2009. Discovering Process Models from Unlabelled Event Logs.
Business Process Management 143–158.
Fettke, P., 2006. State-of-the-Art des State-of-the-Art. Wirtschaftsinformatik 48, 257–266.
fluxicon, 2015. Process Mining and Process Analysis - Fluxicon [WWW Document]. URL
[Link] (accessed 2.1.13).
Fowler, J.F.J., 1984. Survey Research Methods, Auflage: 4th. ed. SAGE Publications, Inc.
Gehrke, N., Müller-Wickop, N., 2010. Basic Principles of Financial Process Mining A Jour-
ney through Financial Data in Accounting Information Systems, in: Proceedings of
the 16th Americas Conference on Information Systems, Lima, Peru.
Gehrke, N., Werner, M., 2013. Process Mining. wisu - das wirtschaftsstudium 934–943.

66
BIBLIOGRAPHY

Google, 2014. Visualization: Geochart - Google Charts — Google Developers [WWW Doc-
ument]. URL [Link]
chart?hl=de (accessed 7.16.14).
Gregor, S., 2006. The nature of theory in information systems. MIS Quarterly 30, 611–642.
Gregor, S., Baskerville, R., 2012. The Fusion of Design Science and Social Science Research.
ISF 2012.
Gregor, S., Hevner, A.R., 2013. Positioning and Presenting Design Science Research for
Maximum Impact. MIS Quarterly 37, 337–355.
Gubrium, J.F., Holstein, J.A., 2002. Handbook of Interview Research: Context and Method.
SAGE.
Gujarati, D.N., Porter, D.C., 2009. Basic Econometrics. McGraw-Hill Irwin, Boston.
Günther, C., van der Aalst, W.M.P., 2007. Fuzzy Mining – Adaptive Process Simplification
Based on Multi-perspective Metrics. Business Process Management 328–343.
Günther, C.W., Verbeek, E.H., 2012. XES Standard Definition. Eindhoven University of
Technology, Eindhoven.
H2 Database Engine, 2014. H2 Database Engine [WWW Document]. URL
[Link] (accessed 7.25.14).
Hansen, H.R., Neumann, G., 2009. Wirtschaftsinformatik 1 Grundlagen und Anwendun-
gen, 10th ed. Lucius & Lucius, Stuttgart.
Harmsen, A.F., Brinkkemper, J.N., Oei, H., 1994. Situational Method Engineering for Infor-
mation System Project Approaches. University of Twente, Department of Computer
Science.
Heckel, R., 2006. Graph Transformation in a Nutshell. Electronic Notes in Theoretical
Computer Science 148, 187–198.
Heinrich, L.J., Heinzl, A., Roithmayr, F., 2004. Wirtschaftsinformatik-Lexikon. Oldenbourg,
München.
Herbst, J., 2003. Ein induktiver Ansatz zur Akquisition und Adaption von Workflow-Model-
len. Tenea Verlag Ltd.
Herbst, J., 2000a. Dealing with Concurrency in Workflow Induction, in: European Concur-
rent Engineering Conference. SCS Europe.
Herbst, J., 2000b. A Machine Learning Approach to Workflow Management, in: López de
Mántaras, R., Plaza, E. (Eds.), Proceedings ofMachine Learning: 11th European Con-
ference Machine Learning. Springer Berlin Heidelberg, Berlin, Heidelberg, pp. 183–
194.
Herbst, J., Karagiannis, D., 1999. An Inductive Approach to the Acquisition and Adaptation
of Workflow Models, in: Proceedings of the IJCAI. pp. 52–57.
Herbst, J., Karagiannis, D., 1998. Integrating Machine Learning and Workflow Manage-
ment to Support Acquisition and Adaptation of Workflow Models, in: Proceedings

67
BIBLIOGRAPHY

9th International Workshop on Database and Expert Systems Applications


(DEXA’98). pp. 745–752.
Hevner, A., Chatterjee, S., 2010. Design Science Research in Information Systems, in: De-
sign Research in Information Systems. Springer US, Boston, MA, pp. 9–22.
Hevner, A.R., March, S.T., Park, J., Ram, S., 2004. Design Science in Information Systems
Research. MIS Quarterly 28, 75–105.
IDW, 2009. IDW PS 261 Feststellung und Beurteilung von Fehlerrisiken und Reaktionen
des Abschlussprüfers auf die beurteilten Fehlerrisiken.
IFAC, 2012. ISA 315 (Revised), Identifying and Assessing the Risks of Material Misstate-
ment through Understanding the Entity and Its Environment.
IFAC, 2009. ISA 320 Materiality in Planning and Performing an Audit.
Jans, M., 2012. Process Mining in Auditing: From Current Limitations to Future Chal-
lenges, in: Daniel, F., Barkaoui, K., Dustdar, S. (Eds.), Business Process Management
Workshops. Springer, Berlin, Heidelberg, pp. 394–397.
Jans, M., Alles, M., Vasarhelyi, M., 2011. Process Mining of Event Logs in Internal Auditing:
A Case Study, in: 2nd International Symposium on Accounting Information Systems.
Jans, M., Alles, M., Vasarhelyi, M., 2010. Process Mining of Event Logs in Auditing: Oppor-
tunities and Challenges. Working paper. Hasselt University. Belgium.
Jans, M., Lybaert, N., Vanhoof, K., Van Der Werf, J.M., 2008. Business Process Mining for
Internal Fraud Risk Reduction: Results of a Case Study, in: Proceedings of the Inter-
national Research Symposium on Accounting Information Systems, Paris.
Jensen, K., Kristensen, L.M., 2009. Coloured Petri Nets. Springer.
Jick, T.D., 1979. Mixing Qualitative and Quantitative Methods: Triangulation in Action. Ad-
ministrative Science Quarterly 24, 602–611.
Kemper, H.-G., Mehanna, W., Baars, H., 2010. Business Intelligence - Grundlagen und
praktische Anwendungen. Vieweg + Teubner, Wiesbaden.
Konradin Mediengruppe, 2011. ERP Studie 2011.
Krcmar, H., 2010. Informationsmanagement. Springer, Berlin; Heidelberg.
Küting, K., Reuter, M., 2004. Bilanzierung im Spannungsfeld unterschiedlicher Adressaten.
DSWR 9, 230–233.
March, S.T., Smith, G.F., 1995. Design and Natural Science Research on Information Tech-
nology. Decision Support Systems 15, 251–266.
Maxeiner, M.K., Küspert, K., Leymann, F., 2001. Data Mining von Workflow-Protokollen
zur teilautomatisierten Konstruktion von Prozessmodellen, in: Datenbanksysteme in
Büro, Technik Und Wissenschaft (BTW), 9. GI-Fachtagung,. Springer-Verlag, pp. 75–
84.
Mingers, J., 2001. Combining IS Research Methods: Towards a Pluralist Methodology. In-
formation Systems Research 12, 240–259.

68
BIBLIOGRAPHY

Müller, R.M., Lenz, H.-J., 2013. Business Intelligence, [Link]. Springer, Berlin, Hei-
delberg.
Müller-Wickop, N., 2014. Integration von Wertflüssen in Geschäftsprozessmodellierungs-
sprachen: Ein gestaltungsorientierter Ansatz zur Unterstützung von Revisoren bei
der wertflussorientierten Planung und Durchführung von Prozessprüfungen. Ham-
burg.
Müller-Wickop, N., Schultz, M., 2013a. Modelling Concepts For Process Audits - Empiri-
cally Grounded Extension Of BPMN, in: Proceedings of the 21st European Confer-
ence on Information Systems, Utrecht.
Müller-Wickop, N., Schultz, M., 2013b. ERP Event Log Preprocessing: Timestamps vs. Ac-
counting Logic, in: Proceedings of the 8th International Conference on Design Sci-
ence Research in Information Systems and Technology, Springer Berlin Heidelberg,
Berlin, Heidelberg, pp. 105–119.
Müller-Wickop, N., Schultz, M., Peris, M., 2013. Towards Key Concepts for Process Audits
– A Multi-Method Research Approach, in: Proceedings of the 10th International
Conference on Enterprise Systems, Accounting and Logistics, Utrecht.
Naumann, J.D., Jenkins, A.M., 1982. Prototyping: the New Paradigm for Systems Develop-
ment. MIS Quarterly 29–44.
Neel, D.L., Orrison, M.E., 2006. The Linear Complexity of a Graph. The Electronic Journal
of Combinatorics 13, 1–19.
Neo Technology Inc., 2014. Neo4j - The World’s Leading Graph Database [WWW Docu-
ment]. URL [Link] (accessed 7.25.14).
Niehaves, B., 2007. On Epistemological Diversity in Design Science: New Vistas for a De-
sign-oriented IS Research, in: Proceedings of the 28th International Conference on
Information Systems, Montreal, pp. 1–13.
Nunamaker, J.F., Chen, M., Purdin, T.D.M., 1991. Systems Development in Information
Systems Research. Journal of Management Information Systems 7, 89–106.
Oracle, 2014a. Java [WWW Document]. URL [Link] (accessed
7.25.14).
Oracle, 2014b. Welcome to NetBeans [WWW Document]. URL [Link] (ac-
cessed 7.25.14).
Österle, H., Becker, J., Frank, U., Hess, T., Karagiannis, D., Krcmar, H., Loos, P., Mertens, P.,
Oberweis, A., Sinz, E.J., 2010. Memorandum on design-oriented information sys-
tems research. European Journal of Information Systems 20, 7–10.
Palvia, P., Leary, D., Mao, E., Midha, V., Pinjani, P., Salam, A.F., 2004. Research methodol-
ogies in MIS: an update. Communications of the Association for Information Sys-
tems (Volume 14, 2004) 526, 542.
Palvia, P., Mao, E., Salam, A.F., Soliman, K.S., 2003. Management Information System Re-
search: What’s There in a Methodology? Communications of the Association for In-
formation Systems 11, 289–309.

69
BIBLIOGRAPHY

PCAOB, 2010. Auditing Standard No. 12 Identifying and Assessing Risks of Material Mis-
statement.
Peffers, K., Tuunanen, T., Gengler, C.E., Rossi, M., Hui, W., Virtanen, V., Bragge, J., 2006.
The design science research process: a model for producing and presenting infor-
mation systems research, in: Proceedings of the 1st International Conference on De-
sign Science Research in Information Systems and Technology (DESRIST), pp. 83–
106.
Peffers, K., Tuunanen, T., Rothenberger, M.A., Chatterjee, S., 2007. A Design Science Re-
search Methodology for Information Systems Research. Journal of Management In-
formation Systems 24, 45–77.
Process Mining Group, 2015. ProM [WWW Document]. URL [Link]
[Link]/prom/start (accessed 4.8.13).
Ramezani, E., Fahland, D., van der Aalst, W.M.P., 2012. Where Did I Misbehave? Diagnos-
tic Information in Compliance Checking, in: Barros, A., Gal, A., Kindler, E. (Eds.), Busi-
ness Process Management, Lecture Notes in Computer Science. Springer, pp. 262–
278.
Recker, J., 2012. Scientific Research in Information Systems: a Beginner’s Guide. Springer,
New York.
Reichert, M., Weber, B., 2012. Enabling Flexibility in Process-aware Information Systems
Challenges, Methods. Springer, Berlin; New York.
Riege, C., Saat, J., Bucher, T., 2009. Systematisierung von Evaluationsmethoden in der ge-
staltungsorientierten Wirtschaftsinformatik. Wissenschaftstheorie und gestaltungs-
orientierte Wirtschaftsinformatik 69–86.
Romney, M.B., Steinbart, P.J., 2008. Accounting Information Systems, 11th Revised edi-
tion. ed. Prentice Hall.
Rowley, J., Slack, F., 2004. Conducting a Literature Review. Management Research News
27, 31–39.
Rozenberg, G., 1997. Handbook of Graph Grammars and Computing by Graph Transfor-
mation, illustrated edition. ed. World Scientific Pub Co, Singapore.
Rozinat, A., 2007. Towards an Evaluation Framework for Process Mining Algorithms (BPM
Center Report). [Link].
Rozinat, A., de Medeiros, A.K.A., Günther, C.W., Weijters, A., van der Aalst, W.M.P., 2008.
The Need for a Process Mining Evaluation Framework in Research and Practice, in:
Business Process Management Workshops. pp. 84–89.
Rozinat, A., van der Aalst, W.M.P., 2008. Conformance Checking of Processes Based on
Monitoring Real Behavior. Information Systems 33, 64–95.
SAP, 2015. SAP-UCC [WWW Document]. URL [Link] (accessed 5.3.12).
Schauer, C., 2011. Die Wirtschaftsinformatik im internationalen Wettbewerb. Gabler,
Wiesbaden.

70
BIBLIOGRAPHY

Schimm, G., 2001a. Process Mining linearer Prozessmodelle - Ein Ansatz zur Automatisier-
ten Akquisition von Prozesswissen, in: Proceedings 1. Konferenz Professionelles
Wissensmanagement.
Schimm, G., 2001b. Process Mining Elektronischer Geschäftsprozesse, in: Proceedings
Elektronische Geschäftsprozesse.
Schipper, K., 2005. The introduction of International Accounting Standards in Europe: Im-
plications for International Convergence. European Accounting Review 14, 101–126.
Schira, J., 2009. Statistische Methoden der VWL und BWL Theorie und Praxis. Pearson Stu-
dium, München, Boston.
Schultz, M., 2015. Business Process Compliance from an Audit Perspective - A Design Sci-
ence Research Approach for an Integrated Processing of Business Process Models
and Internal Control in the Context of Process Audits. Hamburg.
Schultz, M., Müller-Wickop, N., Nüttgens, M., 2012. Key Information Requirements for
Process Audits - an Expert Perspective, in: Proceedings of the 5th International
Workshop on Enterprise Modelling and Information Systems Architectures, Vienna.
StataCorp LP, 2014. Stata | Data Analysis and Statistical Software [WWW Document]. URL
[Link] (accessed 7.25.14).
Stocker, T., 2012. Data Flow-oriented Process Mining to Support Security Audits, in: Ser-
vice-Oriented Computing-ICSOC 2011 Workshops. pp. 171–176.
Takeda, H., Veerkamp, P., Yoshikawa, H., 1990. Modeling Design Process. AI magazine 11,
37–48.
Tiwari, A., Turner, C.J., Majeed, B., 2008. A Review of Business Process Mining: State-of-
the-art and Future Trends. Business Process Management Journal 14, 5–22.
Turban, E., Aronson, J.E., Liang, T.-P., Sharda, R., 2007. Decision Support and Business In-
telligence Systems. Pearson Education International, Upper Saddle River, London.
United States Congress, 2012a. Securities Act of 1933.
United States Congress, 2012b. Securities Exchange Act of 1934.
United States Congress, 2002. Sarbanes-Oxley Act, Public Law 107–204.
University of Hamburg, 2015. Renew - The Reference Net Workshop [WWW Document].
URL [Link] (accessed 7.14.15).
University of Hamburg, 2014. Virtual Accounting Worlds [WWW Document]. URL
[Link] (accessed 7.2.14).
Vaishnavi, V., Kuechler, W., 2004. Design Science Research in Information Systems [WWW
Document]. URL [Link]
van der Aalst, W.M.P., 2013. Business Process Management: A Comprehensive Survey.
ISRN Software Engineering 2013.
van der Aalst, W.M.P., 2012. A Decade of Business Process Management Conferences:
Personal Reflections on a Developing Discipline, in: Business Process Management.
Springer, pp. 1–16.

71
BIBLIOGRAPHY

van der Aalst, W.M.P., 2011a. Process Mining: Discovery, Conformance and Enhancement
of Business Processes, 1st Edition. ed. Springer, Berlin Heidelberg.
van der Aalst, W.M.P., 2011b. Using Process Mining to Bridge the Gap Between BI and
BPM. Computer 44, 77–80.
van der Aalst, W.M.P., 2011c. Process Mining: Discovering and Improving Spaghetti and
Lasagna Processes, in: IEEE Symposium on Computational Intelligence and Data
Mining (CIDM). pp. 1–7.
van der Aalst, W.M.P., 2005. Business Alignment: Using Process Mining as a Tool for Delta
Analysis and Conformance Testing. Requirements Engineering Journal Vol. 10, pp.
198–211.
van der Aalst, W.M.P., Andriansyah, A., de Medeiros, A.K., Arcieri, F., Baier, T., Blickle, T.,
Bose, J.C., van den Brand, P., Brandtjen, R., Buijs, J., 2012. Process Mining Mani-
festo, in: BPM 2011 Workshops Proceedings. pp. 169–194.
van der Aalst, W.M.P., de Medeiros, A.K.A., 2005. Process Mining and Security: Detecting
Anomalous Process Executions and Checking Process Conformance. Electronic
Notes in Theoretical Computer Science 121, 3–21.
van der Aalst, W.M.P., Hofstede, A.H.M. ter, Kiepuszewski, B., Barros, A.P., 2003. Work-
flow Patterns. Distributed and Parallel Databases 14, 5–51.
van der Aalst, W.M.P., Stahl, C., 2011. Modeling Business Processes : a Petri Net-oriented
Approach. MIT Press, Cambridge, Mass.
van der Aalst, W.M.P., van Hee, K., van der Werf, J.M., Kumar, A., Verdonk, M., 2011. Con-
ceptual Model for Online Auditing. Decision Support Systems 50, 636–647.
van der Aalst, W.M.P., Weijters, T., Maruster, L., 2004. Workflow Mining: Discovering Pro-
cess Models from Event Logs. IEEE Transactions on Knowledge and Data Engineering
16, 1128–1142.
van der Werf, J.M.E.M., Verbeek, H.M.W., van der Aalst, W.M.P., 2012a. Context-Aware
Compliance Checking, in: Barros, A., Gal, A., Kindler, E. (Eds.), Business Process Man-
agement, Lecture Notes in Computer Science. Springer, pp. 98–113.
van der Werf, J.M.E.M., Verbeek, H.M.W., van der Aalst, W.M.P., 2012b. Context-Aware
Compliance Checking, in: Barros, A., Gal, A., Kindler, E. (Eds.), Business Process Man-
agement, Lecture Notes in Computer Science. Springer, pp. 98–113.
Venable, J., Pries-Heje, J., Baskerville, R., 2012. A Comprehensive Framework for Evalua-
tion in Design Science Research. Design Science Research in Information Systems.
Advances in Theory and Practice 423–438.
Venkatesh, V., Brown, S.A., Bala, H., 2013. Bridging the qualitative-quantitative divide:
Guidelines for conducting mixed methods research in information systems. MIS
Quarterly 37, 21–54.
Verband der Hochschullehrer für Betriebswirtschaft e.V., 2011. VHB-Jourqual2.1.
Vercellis, C., 2009. Business Intelligence. Wiley, Chichester.

72
BIBLIOGRAPHY

vom Brocke, J., Simons, A., Niehaves, B., Reimer, K., Plattfaut, R., Cleven, A., 2009. Recon-
structing the Giant - on the Importance of Rigour in Documenting the Literature
Search Process, in: Proceedings of the 17th European Conference on Information
Systems, pp. 2–13.
Walls, J., Widmeyer, G., El Sawy, O., 1992. Building an Information System Design Theory
for Vigilant EIS. Information Systems Research 3, 36–59.
Wang, G.A., Wang, H.J., Li, J., Fan, W., 2014. Mining Knowledge Sharing Processes in
Online Discussion Forums, in: Proceedings of the 47th Hawaii International Confer-
ence on System Science, IEEE, pp. 3898–3907.
Webster, J., Watson, R.T., 2002. Analyzing the Past to Prepare for the Future. MIS Quar-
terly 26, xiii – xxiii.
Weijters, A., van der Aalst, W.M.P., de Medeiros, A.K.A., 2006. Process Mining with the
Heuristics Miner-algorithm. Technische Universiteit Eindhoven, Tech. Rep. WP 166.
Werner, M., 2013. Colored Petri Nets for Integrating the Data Perspective in Process Au-
dits, in: Proceedings of the 32nd International Conference on Conceptual Modeling
(ER 2013), Springer-Verlag, Hong Kong, China, pp. 387–394.
Werner, M., 2012. Einsatzmöglichkeiten von Process Mining für die Analyse von Ge-
schäftsprozessen im Rahmen der Jahresabschlussprüfung, in: Plate, G. (Ed.), For-
schung Für Die Wirtschaft. Cuvillier Verlag, Göttingen, pp. 199–214.
Werner, M., Gehrke, N., 2015. Multilevel Process Mining for Financial Audits. IEEE
Transactions on Services Computing 8, 820–832.
Werner, M., Gehrke, N., 2011. Potentiale und Grenzen automatisierter Prozessprüfungen
durch Prozessrekonstruktionen, in: Plate, G. (Ed.), Forschung für die Wirtschaft. Sha-
ker Verlag, Aachen, pp. 99–120.
Werner, M., Gehrke, N., Nüttgens, M., 2013. Towards Automated Analysis of Business
Processes for Financial Audits, in: Proceedings of the 11th International Conference
on Wirtschaftsinformatik (WI 2013), Leipzig, pp. 375–389.
Werner, M., Gehrke, N., Nüttgens, M., 2012a. Business Process Mining and Reconstruc-
tion for Financial Audits, in: Proceedings of the 45th Hawaii International Confer-
ence on System Sciences (HICSS 2012), Maui, pp. 5350–5359.
Werner, M., Nüttgens, M., 2014. Improving Structure - Logical Sequencing of Process
Models, in: Proceedings of the 47th Hawaii International Conference on System Sci-
ences (HICSS 2014), Big Island, pp. 3888–3897.
Werner, M., Schultz, M., Müller-Wickop, N., Gehrke, N., Nüttgens, M., 2012b. Tackling
Complexity: Process Reconstruction and Graph Transformation for Financial Audits
(Research in Progress), in: Proceedings of 33rd International Conference on Infor-
mation Systems (ICIS 2012), Orlando, pp. 1–12.
Werner, M., Schultz, M., Müller-Wickop, N., Nüttgens, M., 2014. Who Is Afraid of the Big
Bad Wolf - Structuring Large Design Science Research Projects, in: Proceedings of
the 22nd European Conference on Information Systems (ECIS 2014), Tel Aviv, Israel,
pp. 1–16.

73
BIBLIOGRAPHY

Weske, M., 2012. Business Process Management Concepts, Languages, Architectures.


Springer, Berlin; New York.
Wilde, T., Hess, T., 2007. Forschungsmethoden der Wirtschaftsinformatik. Wirtschaftsin-
formatik 49, 280–287.
Wilde, T., Hess, T., 2006. Methodenspektrum der Wirtschaftsinformatik: Überblick und
Portfoliobildung. Arbeitspapiere des Instituts für Wirtschaftsinformatik und Neue
Medien, LMU München 2.
WKWI, 2008. WI-Orientierungslisten. Wirtschaftsinformatik 50, 155–163.
Wollnik, M., 1988. Ein Referenzmodell des Informationsmanagements. Information Ma-
nagement 3, 34–43.
Workflow Management Coalition, 1999. Terminology & Glossary. Technical Report.
Yang, W.-S., Hwang, S.-Y., 2006. A Process-mining Framework for the Detection of
Healthcare Fraud and Abuse. Expert Systems with Applications 31, 56–68.
yWorks GmbH, 2015. yEd - Graph Editor [WWW Document]. URL
[Link] (accessed 7.15.15).

74
APPENDIX A: PUBLICATIONS

10 Appendix A: Publications

75
APPENDIX A: PUBLICATIONS

10.1 Who Is Afraid of the Big Bad Wolf - Structuring Large Design Science Re-
search Projects

Number 1
Who Is Afraid of the Big Bad Wolf - Structur-
Title
ing Large Design Science Research Projects
Appendix 10.1
Primary Related Chapters 3
Type Conference Paper
22nd European Conference on Information Sys-
Conference
tems (ECIS 2014)
Reference (Werner et al., 2014)
Acceptance Rate 34 %
VHB JQ 2.1 Ranking B (7.37)
WKWI Ranking A
ERA 2010 A
CORE 2013 A
Review Procedure Double Blinded
Number of Reviews 4
1. Michael Werner
2. Martin Schultz
Authors
3. Niels Müller-Wickop
4. Markus Nüttgens
Dissertation Points 0.40
Authorship
Overall 85%
Design 85%
Realization 85%
Writing 85%
Status Published
Part of other Dissertations No
[Link]
Link
[Link]

76
APPENDIX A: PUBLICATIONS

Who Is Afraid of the Big Bad Wolf - Structuring Large Design Science Research
Projects

WERNER, MICHAEL
University of Hamburg, Germany
[Link]@[Link]

SCHULTZ, MARTIN
University of Hamburg, Germany
[Link]@[Link]

MÜLLER-WICKOP, NIELS
University of Hamburg, Germany
[Link]-wickop@[Link]

NÜTTGENS, MARKUS
University of Hamburg, Germany
[Link]@[Link]

Abstract: Design Science Research is an important research approach that has


gained increased attention over the past decade in the international scientific
community. Recommendations and methodologies for design science-oriented
research published in recent years provide good guidance for scientists to con-
duct and publish design science-oriented research. But little attention has yet
been paid to the setup and structuring of large research projects. Such projects
sometimes look like dangerous and greedy big bad wolves that threaten the
innocent project participants. Large research projects are special because of
the volume and complexity of the research scope and the interaction of partic-
ipating project members. We look for solutions to tame the wolf and present a
framework to confine and structure the content of large research projects into
well-defined and individually manageable research segments. This framework
is suitable for projects characterized by manifold research tasks and the in-
volvement of many project participants. This article presents the designed
framework and illustrates benefits, limitations and practical implications. Its ap-
plication is illustrated by means of a research project case study that was car-
ried out over a three year period with the participation of several research in-
stitutions and partner companies.

Keywords: Design Science, Research Process, Research Framework, Case Study

77
APPENDIX A: PUBLICATIONS

1 Introduction
Design Science Research (DSR) has gained increased attention over the last decade in the
international scientific community as an important research approach (Vaishnavi and
Kuechler, 2004). DSR is especially prevalent in German speaking countries (Wilde and Hess,
2007) where information systems research has traditionally been closely related to the nat-
ural and engineering sciences in contrast to research communities in Anglo-Saxon countries
which show a tendency to positivist, behavioristic research methods (Schauer, 2011). Sev-
eral publications on the role and interaction of design science, natural science (March and
Smith, 1995) and social science (Gregor and Baskerville, 2012) have led to more clarity
about the differences, relationships, and interactions between different research ap-
proaches in the scientific domain. Scientists have made valuable contributions on how to
conduct DSR in a structured and rigorous manner (Hevner et al., 2004; Peffers et al., 2007;
Österle et al., 2010; Hevner and Chatterjee, 2010) and for positioning DSR results in the
academic arena (Gregor and Hevner, 2013). But little attention has yet been paid to the
question of how large research projects can actually be structured and set up. Large re-
search projects sometimes appear like hungry and dangerous wolves that threaten the
frightened project participants because they do not know how to manage the voluminous
research tasks appropriately to achieve the intended research goals. Projects get out of
control, out of budget, out of time and consume valuable research resources without
achieving the intended research objectives. The involved researchers feel eaten up, be-
come frustrated and burn out. We describe a framework that can help to improve the re-
search process by giving guidance as to how the research content and tasks can be struc-
tured and divided into manageable components.
Design science-oriented research in information systems research is defined by different
steps (Peffers et al., 2007; Gregor and Baskerville, 2012) that commonly at least include the
phases analysis, design, evaluation and diffusion (Österle et al., 2010). The completion of
all phases is a resource- and time-consuming task. In comparison to purely descriptive sci-
ence, DSR is characterized by the duality of the epistemological and the design objective
(Riege et al., 2009). The objective of creating artifacts that are valuable for practical pur-
poses (Hevner et al., 2004) requires the involvement and participation of project members
from the application domain. Evaluation is an essential component in design science-ori-
ented research projects (Riege et al., 2009; Venable et al., 2012). Many evaluation methods
like simulations, lab or field experiments require instantiated artifacts that can be used in
natural or artificial evaluation environments. The instantiation of designed artifacts is com-
monly a time-consuming task that requires different skills than the design phase. The ne-
cessity to include all necessary research phases and to conduct extensive evaluation for
rigorous research results as well as the aim to create valuable research artifacts that on the
one hand contribute to the scientific knowledge base but that can on the other hand also
be applied in practice, commonly lead to large research projects that include the participa-
tion of multiple researchers, research assistants, different scientific institutions and partner
companies from the application domain. The interaction of multiple researchers and par-
ticipation of different interest groups contribute to the complexity of these projects.
Our research is guided by the question as to how the research content and tasks in large
research projects can be structured to facilitate parallel research work by simultaneously

78
APPENDIX A: PUBLICATIONS

encouraging mutual research efforts and synergies. We introduce a framework that can be
used to structure large design science-oriented research projects. The framework allows
the participants to divide the research content of a research project into small segments
that can then be addressed individually to identify necessary research artifacts and appro-
priate research methods. We focus on design science-oriented research projects in the in-
formation systems science discipline because of the idiosyncratic characteristics of design
science-oriented research. The duality of the epistemological and design objective in these
projects leads to research activities that are different compared to other research ap-
proaches.
The overall objective of presenting our framework is to enable researchers to confine and
structure the content of research projects in such a way that research segments are well
separated from each other to allow distributed and parallel research work on different seg-
ments but also to enable the research group to keep track of the whole research process
and to benefit from the interrelations of participating researchers and the mutual research
efforts. The framework focusses on setting the scope of a research project and for dividing
its content into small and well-defined parts that can be addressed by different research-
ers. It should not be regarded as a project management framework. Project management
is a complex process and recommendations and guidance on project management is for
example available in relevant national and international standards (International Organiza-
tion for Standardization, 2012; Deutsches Institut für Normung, 2009; Project Management
Institute, 2013; Great Britain and Office of Government Commerce, 2009). The presented
framework should be seen as a useful tool in the setup and operational phase of a project
to promote the assignment of research tasks to the involved researchers. Our research re-
sults are derived from a large design science-oriented research project that was carried out
in cooperation with four private companies and two research institutions over a three year
period. It included the participation of several senior and junior researchers, research as-
sistants and representatives of the participating companies. The presented framework was
developed and used during this research project. The application of the framework and the
insights from its practical usage - which is included as a case study - illustrates the benefits
and limitation of the presented framework.

2 Methodology and Research Structure


The work presented in this paper follows a DSR approach. DSR addresses research prob-
lems that originate from the application domain and its aim is to develop useful artifacts
that can be applied in practice (Hevner et al., 2004). The application domain of our research
work is the field of information systems research itself. We relied on a DSR approach in
order to create an artifact that can improve the research process in large research projects
and hence the proximity of the research question to practical problems in daily scientific
practice. The presented research work follows the research phases described by Österle et
al. which consist of analysis, design, evaluation and diffusion (Österle et al., 2010). The anal-
ysis of the problem domain is described in section four within the case study. The main
focus of this paper is the presentation of the designed model for structuring and confining
research projects which is addressed in section five. Section six continues with a discussion
on lessons learned based on the described case study, identified benefits, limitations and
evaluation aspects. The diffusion of our research results as the last phase in DSR is already

79
APPENDIX A: PUBLICATIONS

addressed by the publication of this article. A major challenge in scientific endeavors is the
selection of an adequate research method (Galliers and Land, 1987). A variety of well-es-
tablished research methods is available for researchers in information systems science (Pal-
via et al., 2003; Palvia et al., 2004; Wilde and Hess, 2007). The selection of a research
method should follow the intended research objective. The objective of our research lies
in the development of a framework that can be used by scientists to structure and confine
the research content in large research projects. We reviewed relevant frameworks in the
information systems science discipline and amalgamated different models to construct a
new framework useful for the purpose at hand. This approach is comparable to method
engineering (Brinkkemper, 1996). Method engineering is used to construct methods based
on already existing methods and method fragments. But instead of merging existing
method fragments we used a model engineering approach by referring to model compo-
nents that have already been proven useful in the information system research discipline
for creating a novel solution.
Research projects are complex undertakings that include the interaction of many individu-
als with differing or sometimes even opposing motivations. The development of a frame-
work for structuring research projects needs to consider the complex social settings in re-
search projects. We used the developed framework in a real research project which is illus-
trated as a case study in this paper. Case study research is a common research method in
information systems research (Chen and Hirschheim, 2004). It is a special form of qualita-
tive-empirical research methods (Wilde and Hess, 2006) that involves the close examina-
tion of people, topics and issues (Hays, 2004). It is especially suited to investigate complex
phenomena in their natural environments and can be used for behavioral or design-ori-
ented research (Wilde and Hess, 2006). Case study research is commonly criticized for the
lack of generalizability due to the uniqueness of the investigated case. But Yin points out
that similar concerns can, for example, also be applied in the contexts of single experiments
(Yin, 2008). Furthermore the objective of our research is of exploratory nature and not em-
pirical evaluation.

3 Related Work
Scholars have highlighted the need for structured and commonly agreed research pro-
cesses in the DSR community (Leist and Rosemann, 2011). Peffers et al. present a research
methodology28 for DSR in the information systems community (Peffers et al., 2006; Peffers
et al., 2007). The authors present a research framework that consists of six phases: (1) iden-
tify problem and motivate, (2) define objectives of a solution, (3) design and development,
(4) demonstration, (5) evaluation and (6) communication. Their framework is based on a
review of existing scientific publications on the research process in the information systems
and related research disciplines. The research framework is composed of process elements
that have been identified by different scholars working in the information systems (Takeda
et al., 1990; Nunamaker et al., 1991; Walls et al., 1992; Hevner et al., 2004; Cole et al., 2005)
and engineering (Archer, 1984; Eekels and Roozenburg, 1991) discipline. It is interesting

28
The term methodology is interpreted ambiguously in information systems science (Mingers, 2001).
Peffers et al. refer to a methodology as a combination of methods independent of a single research pro-
ject and intend to present a methodology that serves as a commonly accepted framework for carrying
out design science research.

80
APPENDIX A: PUBLICATIONS

that the continental European research community is mostly neglected by the authors alt-
hough the design science discipline has a long lasting tradition especially in German speak-
ing countries (Winter, 2008). Österle et al. as representatives of this community suggest
four phases for DSR: (1) analysis, (2) design, (3) evaluation and (4) diffusion (Österle et al.,
2010). Gregor and Baskerville examine the research process from a philosophy of science
perspective with the objective to provide a framework for the combination of design sci-
ence and social science research. The presented research process consists of the phases (A)
construct and test artefacts, (B) formulate prescriptive knowledge and theory, (C) study
artefact(s) in use, (D) test knowledge of artefacts in use and (E) formulate descriptive
knowledge (Gregor and Baskerville, 2012). All authors explicitly emphasize the iterative re-
lationship between the different research steps in each model. Alturki et al. present a more
detailed model that consists of 14 research steps. The authors also present an extensive
summary of relevant literature (Alturki et al., 2011).
It is interesting to note that researchers from the information systems discipline rarely refer
to project management literature. Available knowledge on project management that has
been standardized in international (International Organization for Standardization, 2012)
or national (Project Management Institute, 2013; Deutsches Institut für Normung, 2009;
Great Britain and Office of Government Commerce, 2009) guidelines has rarely been con-
sidered yet, although research work exhibits all the characteristics that are also associated
with projects (vom Brocke and Lippe, 2010). Vom Brocke and Lippe are among the few
authors that build the bridge between research processes and project management. They
point out the need to tailor existing project management guidelines for research projects
and identify eight characteristics that distinguish design science-oriented research projects
from traditional project types (vom Brocke and Lippe, 2010). We are not aware of quanti-
tative research on project management in the field of information systems science.
The aforementioned literature provides useful insights into the research processes in de-
sign science-oriented research. But little has yet been published that provides guidance as
to the effective structure and set up of large DSR projects. To do exactly that is what we
intend with the presentation of the research results in this article.

4 Case Study
The research project that is described as a case study in this article was concerned with
research questions from the field of financial audits. Companies prepare financial state-
ments to provide interested parties with financially relevant information. The correctness
and reliability of this information are a key requirement for stakeholders to direct their
decisions. National laws and regulations mandate the audit of financial statements by an
independent third party to prevent the distribution of false financial information because
of its paramount role for the well-functioning of economic markets. These audits are car-
ried out by public accountants. Accounting scandals in recent years have shown that audi-
tors have not been able to prevent these scandals or at least indicate any violations before
the actual collapse. A common problem in financial audits is an imbalance between auto-
mated transaction processing of partly huge data volumes on the companies’ side and tra-
ditional and manual audit procedures on the auditors’ side (Werner et al., 2012). Compa-
nies use information systems to support and automate the operation of their business pro-
cesses. Auditors primarily rely on traditional audit procedures like interviews and manual

81
APPENDIX A: PUBLICATIONS

inspections of available documents to achieve the necessary audit comfort. But these pro-
cedures become inefficient or even ineffective in environments where the processing of
transactions is highly automated and includes the handling of large data volumes (Werner
and Gehrke, 2011). A solution to decrease this imbalance would be the application of au-
tomated audit procedures. Business processes play a significant role in financial audits. The
audit of business processes and internal controls that affect the processing of transactions
are an important part of financial audits (International Federation of Accountants, 2012).
The rationale for considering business processes is the assumption that well-controlled pro-
cesses will lead to complete and correct entries on the financial accounts. The objective of
the described research project was the development of methods and tools that support
the auditor by automating parts of the procedures that are necessary to conduct process
audits, and to thereby make these audits more efficient and effective. The fundamental
idea was to use innovative data analysis techniques for automating the discovery of process
models and to automatically assess the design and operating effectiveness of internal con-
trols by analyzing relevant control data. The overall research question was formulated as
follows:
• How can data analysis techniques be used to automate the audit of business pro-
cesses in the context of financial audits?
Process mining provides powerful methods and tools that can be used to reconstruct pro-
cess models based on the analysis of recorded event logs (van der Aalst, 2011). For design-
ing the desired tools and methods the following more detailed research questions had to
be answered.
(1) How can reliable process models be automatically reconstructed by analyzing data
stored in information systems that process financially relevant transactions?
(2) How can process models be automatically assessed from an audit perspective by
integrating control data that is stored in the source systems?
(3) How can process models be graphically represented to display information that is
relevant to auditors and that can be applied in real audit environments?
These questions needed to be answered in order to develop research artifacts that are able
to close the research gap and that are also valuable for the application domain. A main
component of the project was the development of a software prototype. The design of the
prototype was on the one hand desired for the creation of a valuable artifact for the appli-
cation domain but also for evaluation purposes in the research process. The project mem-
bers consisted of two research institutions and four partner organizations. A software com-
pany was responsible for the programming and instantiation of the software tool. A large
auditing company supported the requirement analysis, design and testing phase and pro-
vided necessary data. A small auditing and consulting company was included to reflect the
requirements from small and medium sized companies. A public association for board
members contributed as a project partner to diffuse the research results into the broader
application domain and by providing information about aspects required by the board
level. The project members consisted of three PhD students and two professors from the
information systems area, three software developers, several research assistants and
about 20 contact persons from partner companies who were contacted during the research

82
APPENDIX A: PUBLICATIONS

project. Concerns arose at the beginning of the research work about how the overall pro-
ject should be structured. Each project participant had a different motive to take part on
the project. The partner companies needed a software artifact for practical use. The junior
researchers were eager to advance in their PhD studies and the senior researchers were
concerned about resources and had to keep in mind the overall progress of the involved
scientific institutions. There was the risk that research work on specific tasks would be con-
ducted redundantly and other research tasks be neglected due to uncoordinated research
activities and deviating motivations. A framework was necessary to identify which research
outputs would be critical for the success of the research project, how these related to each
other, which scientific approach would be adequate for the development and evaluation
of research artifacts and which researcher would fit best to accomplish different research
tasks according to available expertise and skills. It was also necessary to decide on the roles
and responsibilities for the communication with the project partners for specific research
aspects such as the requirement analysis, software development and evaluation ap-
proaches.

5 Segmentation Framework
This section deals with the description of the framework for structuring large design sci-
ence-oriented research projects that was developed in the project mentioned above. The
objective of the presented research was the development of a tool that allows the confining
and separation of the research content and tasks. The first step to design a useful frame-
work is the identification of distinguishing features that can be used to categorize different
research contents. The main subjects of interest in information systems research are infor-
mation technologies and the man/machine interaction. Gregor points out that the distin-
guishing characteristic of the information systems area is not only the consideration of both
worlds - technology and humans - but also the investigation of phenomena that emerge
from their interaction in socio-technical systems (Gregor, 2006). A framework for the struc-
turing of research content should therefore take into account the technological and hu-
man-interaction aspects. Categorizations for different levels of information technology and
human interaction can be found in various models. A common model from the field of in-
formation management is presented by Krcmar (Krcmar, 2010). He distinguishes between
three different levels of management tasks. These are accompanied by independent lead-
ership tasks which are relevant for all levels. Figure 1 shows a graphical representation of
the model. The lowest level contains tasks for the management of the technical infrastruc-
ture that is necessary for the use of information and communication technology at the
higher levels. The second level deals with the management of information systems and
includes the management of data, processes, applications and their life-cycles. At the high-
est level reside the tasks for the management of the information economy. The main ob-
jective of the tasks at this level is the management of the resource information, its supply,
demand and usage. The lower two levels are mainly concerned with the technical aspects
of information management. The human aspect is considered at the highest level where
the requirements for the lower levels are defined on the basis of the needs of human in-
formation recipients and users of the applications, processes, data and technology that is
provided by the lower levels. The model of information management addresses both ele-
ments that are subject to research in information system science: information systems and

83
APPENDIX A: PUBLICATIONS

human interaction in socio-technical systems. It is therefore a useful starting point for the
categorization of research content because each researched artifact in design science-ori-
ented research can be characterized if it addresses one or more of the different levels, in-
frastructure, applications and usage. The original model represents applications, data and
processes at the same level.

Management of the Supply


Information management

Demand
Information Economy Usage
Leadership Tasks of

Data
Management of the Processes
Information Systems Application-Life-
Cycle

Management of the Storage


Information and Operation
Communication
Communication Technology Technology Bundle

Figure 1 Model of Information Management Figure 2 Human-Technical Di-


(adapted from Krcmar 2010 p.50) 29 mension
Prominent conferences in the information systems science area like the Business Process
Management conference (BPM, 2014), comprehensive publications (Weske, 2012) and ex-
tensive reviews (van der Aalst, 2012; van der Aalst, 2013) show that business processes are
a key component in information systems research. We therefore considered it appropriate
to divide the level of information systems into the two levels of software applications and
processes. The process level is the connecting layer were process participants use compo-
nents from the lower application level to satisfy information demands from the higher
level. The requirement to include this additional level also became obvious in our research
project because it was not sufficient to consider relevant software applications and how
information was consumed at the usage-level but also how relevant information was cre-
ated within business processes and how these relate to the financial statements that serve
as information input for stakeholders at the highest level. The resulting four levels as shown
in Figure 2 represent the first dimension for categorizing research content.
The presented levels are very broad. Although they can be used to distinguish research
content on a technical vs. human-interaction dimension experiences from our case study
showed that they are not sufficient to divide the research content into manageable seg-
ments. The usage of the four levels as the single separation criterion would mean that re-
searchers would only be concerned about the assigned level which is not a suitable solution

29
The model refers to the reference model that was originally introduced by Wollnik. The three levels of the
information management model relate to the levels of information usage, information and communica-
tion systems and infrastructure of the information processing and communication in Wollnik’s reference
model (Wollnik, 1998).

84
APPENDIX A: PUBLICATIONS

because interdependencies between the different levels would be neglected. Furthermore


it became obvious in our research project that such a broad separation was not adequate
in practical settings when considering the motivation and skill sets of individual research-
ers.
Hevner et al. present a research framework for information systems research that provides
an illustration how different concepts that are relevant for research projects relate to each
other (Hevner et al., 2004). The framework is presented in Figure 3. It shows the relation-
ships between the main research activities for design science (build and evaluate) and be-
havioral science (develop and justify) research, the environment and the knowledge base.
The environment or application domain defines the problem space. The phenomena of in-
terest for design science-oriented research should be derived from the environment. The
knowledge base represents the already existing pool of research results that have been
explored. Each research project should consider already existing knowledge to execute the
research activities, provide useful artifacts to the application domain and add additional
generalized knowledge to the knowledge base.

Environment IS Research Knowledge Base


Relevance Rigor
People Develop / Build Foundations
• Theories
Business
• Artifacts Applicable
Needs Knowledge
Organizations

Justify / Evaluate Methodologies


• Quantitative
Technology Research Methods
• Qualitative Research
Methods

Application in the Additions to the


Appropriate Environment Knowledge Base

Figure 3 Information Systems Research Framework


(adapted from Hevner et al. 2004 p. 80)
DSR projects are characterized by the duality of the epistemological and design objective
which is illustrated in the model through the feedback loops from the research activities to
the environment and the knowledge base. The distinction between the research contribu-
tion target domains can be used as a second dimension for the categorization of research
content. Figure 4 shows the integration of the application domain and knowledge base as
a second dimension.

85
APPENDIX A: PUBLICATIONS

Figure 4 Integration of the Research Con- Figure 5 Integration of the Research


tribution Dimension Question Dimension
Each research project addresses an overall research question. The research questions in
large research projects are, as a rule, complex. Otherwise it would be questionable if such
a project indeed has the characteristics of a large project. Complex research questions can
usually be divided into detailed lower-level research questions. These research questions
can be used as a categorization criterion for a third dimension as shown in Figure 5. The
application of all three dimensions as shown in Figure 5 illustrates how separate segments
emerge based on the different dimension categories. Each segment can now be considered
separately. And groups consisting of segments from different dimensions can be assigned
to different researchers. This ensures that researchers do not just focus on an isolated seg-
ment but that they consider requirements from and interrelationships with different di-
mensions. The division into research segments can be seen as a ‘divide and conquer’ ap-
proach. The overall research project is broken into manageable components that can be
conquered individually. Figure 6 illustrates a single segment of the overall model.

Artifact

Research Methods
- Analysis
- Design
- Evaluation

Diffusion Type

Figure. 6 Research Segment Table 1 Segment Description


Artifacts in design science-oriented research should be developed by using suitable re-
search methods. For each segment it is now possible to identify and describe the relevant

86
APPENDIX A: PUBLICATIONS

artifact(s), the research methods for the analysis, design and evaluation as well as the dif-
fusion type. A template for such a description is illustrated in Table 1. Evaluation methods
for a single artifact can for example be chosen by relying on available frameworks (Venable
et al., 2012), whereas the type of diffusion can for example be determined by referring to
the knowledge contribution level of design science-oriented research (Gregor and Hevner,
2013).
Each segment should be associated with one (or more) artifact. March and Smith define
four types of research artifacts: constructs, methods, models and instantiations (March and
Smith, 1995). Gregor argues that design theories should also be regarded as an important
outcome of design science-oriented research (Gregor, 2006). Many research methods exist
for analysis, design and evaluation purposes. Table 2 illustrates exemplary research meth-
ods that have commonly been cited by renowned scholars. The superscripts a) to e) disclose
the origin of the listed methods that are described in Footnote 30.
Artifacts Research Methods Diffusion Types
• Construct a), e) Analysis • Conference
• Literature review c), survey a), expert in- presentation
• Method a), e)
terview a), c), case study a), data analysis a) • Journal publi-
• Model a), e)
Design cation
• Instantiation a), e)
• Modeling a) ,conceptual c) and reference • Workshop
• Design theory f) presentation
modeling a), method engineering a), argu-
ment-based, concept-based and formal-
deductive analysis b), case study b), c), pro-
totyping a), b), qualitative analysis b), c),
quantitative analysis b), action re-
search b), lab experiment b), c), field ex-
periment b), c), field study c), survey c),
secondary data c)
Evaluation
• Action research d), focus groups d), case
study d), participant observation d), eth-
nography d), survey d), mathematical or
logical proof d), criteria-based evalua-
tion d), lab experiment a), d), field experi-
ment a), d), simulation a), d), pilot applica-
tion a), expert reviews a)
Table 2 Overview of Research Artifacts, Methods and Diffusion Types30
The benefit of applying the segmentation framework becomes obvious when it is illustrated
with examples. Figure 7 shows an instantiation of the reference framework for the de-
scribed case study. The research question domain was separated into three categories that

30
a) (Österle et al., 2010), b) (Wilde and Hess, 2007), c) (Palvia et al., 2003) d) (Venable et al., 2012), e)
(March and Smith, 1995), f) (Gregor, 2006)

87
APPENDIX A: PUBLICATIONS

represent three different research questions. These were derived from the overall research
question and objective (compare section 4). The primary application domain for the re-
search project was financial audits. Business process management and business intelli-
gence were perceived as being the most important scientific disciplines that form the
knowledge base for conducting research in the project.
One segment in the framework is highlighted. It is located on the process-level because the
objective of this segment was the creation of process models that are useful in the context
of financial audits. The research task of this segment was the development of a mining
algorithm that is able to reconstruct process models that comply with the requirements of
the application domain. The designed algorithm exploits the specific structure of financially
relevant transaction data to create process models (Gehrke and Müller-Wickop, 2010) and
uses the data perspective to model the relationship between financial accounts and pro-
cess activities (Werner, 2013). The requirements for developing this algorithm were de-
rived from interviews with experts that showed that the data perspective is of utmost im-
portance in the field of financial audits for illustrating the relationship between business
processes and the financial accounts (Müller-Wickop et al., 2013). The technical require-
ments were investigated by analyzing the relevant data structure in ERP systems. The min-
ing algorithm was designed by using components and research results from already existing
mining algorithms following a method engineering approach (Brinkkemper, 1996).

Financial Process Min-


Artifact
ing Algorithm
Research Methods
Structured Interview
Analysis
Data Analysis
Method Engineering
Design
Prototyping
Simulation
Evaluation
Case Study
Diffusion
Journal
Type

Figure 7 Example Research Segment Table 3 Example Description


The mining algorithm was implemented as a prototype and evaluated by using test data for
simulations (Werner et al., 2013). The evaluation also included a case study in a real world
scenario. Table 3 summarizes the research methods and diffusion type for the designed
artifact
Figure 8 shows a second example and illustrates two research segments. It demonstrates
the relationship between segments that are related to the same level on the technical vs.
human-interaction dimension but address different knowledge contribution categories on
the knowledge contribution dimension. Table 4 lists the relevant artifacts and summarizes
the research methods and diffusion types. A data extraction module was designed for the
extraction of the event log data from the source ERP systems. The inspection of the ex-
tracted event logs revealed that they are not suitable for traditional mining algorithms.
Financially relevant process instances do not exhibit strict linear control flows as assumed

88
APPENDIX A: PUBLICATIONS

by traditional process mining algorithms, their execution behavior includes divergent and
convergent behavior on the business process instance level (Werner and Nüttgens, 2014).
It was necessary to develop a pre-processing algorithm that transforms non-linear event
logs into linear event logs to be able to compare the mining results of the designed Financial
Process Mining Algorithm with other mining algorithms (Mueller-Wickop and Schultz,
2013). The extraction module is an artifact relevant for the application domain because it
allows the extraction of event log data for specific ERP systems in real world scenarios. The
pre-processing algorithm is an artifact that is not specific to the application domain but can
be applied in a variety of application scenarios for transforming non-linear event logs into
linear event logs and can therefore be considered as a generalized contribution to the pro-
cess mining knowledge base. Both artifacts needed to consider the ERP systems that are
used in organizations for processing business transactions. The designed artifacts needed
to be able to interact with these information systems and provided the input for the pro-
cess mining algorithm described in the first example. It was therefore sensible to locate
these artifacts on the application level of the human-technical dimension.

Pre-Processing Al-
Artifact
gorithm
Research Methods
Analysis Experiment
Method Engineer-
Design
ing
Evaluation Simulation
Diffusion Type Conference
Data Extraction
Artifact
Module
Research Methods
Analysis Data Analysis
Design Prototyping
Evaluation Case Study
Diffusion Type Conference

Figure 8 Example Multiple Research Segments Table 4 Example Description


The framework is further useful for the identification of dependencies between individual
segments and the grouping of interdependent segments. Each group can be assigned to
the involved research project participants based on motivational preferences and skill sets.
Figure 9 shows the research scopes that were assigned to the three PhD students that were
involved in the research project. The figure illustrates different aspects. First it is notable
that not all possible segments are covered. This was intended and can be attributed to the
research objective and scope of the research project. Research on the infrastructure-level
was explicitly excluded from the project scope because it was not crucial for the achieve-
ment of the overall research objective. Not covered segments provide opportunities for
additional and subsequent research. The illustration shows a second aspect. It can be seen
that overlaps occurred in some segments. These overlaps were important. A basic require-

89
APPENDIX A: PUBLICATIONS

ment in large research projects is the clear separation of research tasks. But it is also im-
portant to have overlaps that create a common knowledge base that is fundamental for all
research tasks, from preventing completely isolated research efforts and the emergence of
‘Chinese Walls’ within the project.

Research Scope Research Scope Research Scope


Process Reconstruction Process Assessment Process Visualization
Figure 9 Research Scopes
A certain degree of overlap is necessary to encourage communication, mutual research
effort and the creation of research synergies. Too much overlap leads to redundant and
unproductive research work. Finding a balance between separation and overlap is difficult
but crucial for the success of the research project.

6 Discussion
The previous section describes a framework for the structuring of large research projects.
It was applied in a research project that is presented as a case study in this article. Its ap-
plication proved to be useful for the described case. Not all research objectives that were
identified in the initiation phase of the project could be achieved but the main research
questions were answered with the design and evaluation of relevant research artifacts and
their instantiations. The research results were published in 21 peer-reviewed publications
and the project was finished on time and budget with an instantiated software prototype
that included the majority of the developed methods. The provided framework played a
significant role in the research process for the coordination of the research activities of the
participating researchers. The application of the framework showed that several aspects
are crucial for the successful application. A major obstacle was the development of a mu-
tually agreed understanding of terms and definitions regarding the research artifacts and
used research methods among all involved parties. The understanding of specific con-
structs deviated quite significantly between researchers and practitioners but also within
the researcher group. The mutual research in specific research segments was beneficial to
create a common agreement on fundamental concepts and terms. Another crucial success
factor was the development of, agreement on and also implementation of the research and
publication plans that can be developed on the basis of the applied framework. But if these
plans are not followed strictly by all individuals the risk of rivalry and counterproductive

90
APPENDIX A: PUBLICATIONS

activities increases. The implementation of designed artifacts in a software prototype was


an important part of the project. A lack of commonly agreed documentation standards led
to additional programming efforts when software implementations initially developed for
one segment had to be used for another one. A suitable way to prevent such detrimental
developments are mandatory continuous project meetings on a formal level for the discus-
sion of the project progress but also on an informal level to encourage the mutual research
efforts among different research segments and also to minimize communication barriers
between the involved researchers. Dependencies between segments should be considered
when planning the sequence of research activities and to identify synergies that can be
realized if an analysis method for one segment can for example also be used for identifying
requirements for another segment. A crucial outcome from our project was the insight that
is counterproductive to assign segments to project participants that do not fit to the latters’
motivation and skill profile of a participant. It is for example disadvantageous for the pro-
ject progress to assign research tasks that require software implementation to project par-
ticipants that lack relevant programming skills or the willingness to acquire these. (Brooks,
1987) stressed that a critical success factor in software development projects is the selec-
tion of top designers. A similar success factor for large DSR projects is the selection of pro-
ject participants that fit to the identified research segments or segment groups in terms of
motivation and skills.
Although we believe that the presented framework has the potential to facilitate the struc-
turing and management of large research projects a major limitation is its limited evalua-
tion. A case study has been chosen to demonstrate the applicability and usefulness of the
presented framework. Such studies investigate cases for the purpose of illumination and
understanding (Hays, 2004). But single case studies have the disadvantage that it is ques-
tionable if the creation of generalized knowledge is possible based just on one single ob-
servation. Extensive evaluation to strengthen the reliability of the created research results
in the context of research processes is difficult due to the relative long run-time of research
projects and the limited possibility to receive data for evaluation purposes. Although the
limited evaluation might be seen as a constraint for the generalizability of the presented
results it should be kept in mind that the purpose of this article is of exploratory nature.
Our intention is to provide researchers a useful tool that assists in structuring research pro-
jects and therefore improving research processes and outcomes. The presented framework
is a designed model that is not only applicable to an individual situation but that can be
used as a reference model (Heinrich et al., 2004) to structure the content of any DSR pro-
ject. The model can be seen as an improvement of existing tools for project management
in the scientific area. Like in software engineering there is no single solution that fits all
possible project scenarios. Defining and grouping research tasks and assigning them to suit-
able project participants is a big challenge. The provided framework can be used as a tool
to facilitate this task.

7 Summary and Outlook


DSR is an important research approach in the international information systems science
community. A variety of publications exist that provide guidance on how to conduct design
science-oriented research. DSR differs from social science research due to its normative

91
APPENDIX A: PUBLICATIONS

nature and the duality of the research objective. DSR does not only aim at generating gen-
eralized knowledge as a contribution to the scientific knowledge base but also intends to
develop artifacts that are useful for the application domain. This duality leads to research
activities like the instantiation of designed artifacts and evaluation types that are idiosyn-
cratic to DSR compared to other research approaches. Large DSR projects can appear like
dangerous wolves hungry to eat up the helpless project participants and research re-
sources. This paper presents a framework that can be used to structure the content of re-
search projects by dividing the overall research tasks into manageable segments and
thereby tames the wolves. The use of the framework allows the coordination of parallel
and mutual research work of the participating researchers necessary to conduct large and
complex research projects successfully. The applicability of the framework and the benefits
that can be gained by its application have been described by means of a case study. Re-
search on the management of research projects in information systems science is still
scarce. Further research efforts will be made to evaluate the presented framework in fu-
ture research. Many aspects of project management in the information systems science
discipline, e.g. concerning the portfolio management of research projects have not been
investigated yet and empirical research is almost absent. We hope that this field of research
will be addressed more closely also by other researchers.

8 References
Van der Aalst, W.M. (2012). A decade of business process management conferences: per-
sonal reflections on a developing discipline. In Business Process Management.
Springer, pp. 1–16.
Van der Aalst, W.M. (2013). Business Process Management: A Comprehensive Survey.
ISRN Software Engineering, 2013.
Van der Aalst, W.M.P. (2011). Process Mining: Discovery, Conformance and Enhancement
of Business Processes. 1st Edition. Berlin Heidelberg: Springer.
Alturki, A., Gable, G.G. and Bandara, W. (2011). A design science research roadmap. In
Service-Oriented Perspectives in Design Science Research. Springer, pp. 107–123.
Archer, L.B. (1984). Systematic method for designers. In Developments in Design Method-
ology. London: John Wiley, pp. 57–82.
BPM. (2014). BPM 2014. Available from: [Link]
Brinkkemper, S. (1996). Method engineering: engineering of information systems devel-
opment methods and tools. Information and Software Technology, 38(4), pp.275–
280.
Vom Brocke, J. and Lippe, S. (2010). Taking a project management perspective on design
science research. In Global Perspectives on Design Science Research. Springer, pp.
31–44.
Brooks. (1987). No Silver Bullet Essence and Accidents of Software Engineering. Com-
puter, 20(4), pp.10–19.

92
APPENDIX A: PUBLICATIONS

Chen, W. and Hirschheim, R. (2004). A paradigmatic and methodological examination of


information systems research from 1991 to 2001. Information Systems Journal,
14(3), pp.197–235.
Cole, R. et al. (2005). Being Proactive: Where Action Research Meets Design Research. In
Proceedings of the 26h International Conference on Information Systems. ICIS.
Deutsches Institut für Normung. (2009). DIN SPEC 69901 - Projektmanagement. Berlin:
Beuth Verlag GmbH.
Eekels, J. and Roozenburg, N.F.M. (1991). A methodological comparison of the structures
of scientific research and engineering design: Their similarities and differences. De-
sign Studies, 12(4), pp.197–203.
Galliers, R.D. and Land, F.F. (1987). Viewpoint: choosing appropriate information systems
research methodologies. Communications of the ACM, 30(11), pp.901–902.
Gehrke, N. and Müller-Wickop, N. (2010). Basic Principles of Financial Process Mining A
Journey through Financial Data in Accounting Information Systems. In Proceedings
of the 16th Americas Conference on Information Systems. Americas Conference on
Information Systems. Lima, Peru.
Great Britain and Office of Government Commerce. (2009). Managing successful projects
with PRINCE2. London: TSO.
Gregor, S. (2006). The nature of theory in information systems. MIS Quarterly, 30(3),
pp.611–642.
Gregor, S. and Baskerville, R. (2012). The Fusion of Design Science and Social Science Re-
search. ISF 2012.
Gregor, S. and Hevner, A.R. (2013). Positioning and Presenting Design Science Research
for Maximum Impact. MIS Quarterly, 37(2), pp.337–355.
Hays, P.A. (2004). Case study research. In Foundations for research methods of inquiry in
education and the social sciences. Mahwah, N.J.: L. Erlbaum Associates, pp. 217–
234.
Heinrich, L.J., Heinzl, A. and Roithmayr, F. (2004). Wirtschaftsinformatik-Lexikon. Mün-
chen: Oldenbourg.
Hevner, A. and Chatterjee, S. (2010). Design Science Research in Information Systems. In
Design Research in Information Systems. Boston, MA: Springer US, pp. 9–22.
Hevner, A.R. et al. (2004). Design Science in Information Systems Research. MIS Quarterly,
28(1), pp.75–105.
International Federation of Accountants. (2012). ISA 315 (Revised), Identifying and As-
sessing the Risks of Material Misstatement through Understanding the Entity and Its
Environment.
International Organization for Standardization. (2012). ISO 21500 2011 Guidance on pro-
ject management.
Krcmar, H. (2010). Informationsmanagement. Berlin; Heidelberg: Springer.

93
APPENDIX A: PUBLICATIONS

Leist, S. and Rosemann, M. (2011). Research process management. In Proceedings of the


22nd Australasian Conference on Information Systems (ACIS) 2011-Identifying the
Information Systems Discipline. pp. 1–11.
March, S.T. and Smith, G.F. (1995). Design and natural science research on information
technology. Decision support systems, 15(4), pp.251–266.
Mingers, J. (2001). Combining IS research methods: towards a pluralist methodology. In-
formation systems research, 12(3), pp.240–259.
Mueller-Wickop, N. and Schultz, M. (2013). ERP Event Log Preprocessing: Timestamps vs.
Accounting Logic. In Design Science at the Intersection of Physical and Virtual De-
sign. 8th International Conference on Design Science Research in Information Sys-
tems and Technology. Berlin, Heidelberg: Springer Berlin Heidelberg, pp. 105–119.
Müller-Wickop, N., Schultz, M. and Peris, M. (2013). Towards Key Concepts for Process
Audits – A Multi-Method Research Approach. In Proceedings of the 10th Interna-
tional Conference on Enterprise Systems, Accounting and Logistics. ICESAL. Utrecht.
Nunamaker, J.F., Chen, M. and Purdin, T.D.M. (1991). Systems Development in Infor-
mation Systems Research. Journal of Management Information Systems, 7(3),
pp.89–106.
Österle, H. et al. (2010). Memorandum on design-oriented information systems research.
European Journal of Information Systems, 20(1), pp.7–10.
Palvia, P. et al. (2003). Management Information System Research: What’s There in a
Methodolgy? Communications of the Association for Information Systems, 11,
pp.289–309.
Palvia, P. et al. (2004). Research methodologies in MIS: an update. Communications of the
Association for Information Systems (Volume 14, 2004), 526(542), p.542.
Peffers, K. et al. (2007). A Design Science Research Methodology for Information Systems
Research. J. Manage. Inf. Syst., 24(3), pp.45–77.
Peffers, K. et al. (2006). The design science research process: a model for producing and
presenting information systems research. In Proceedings of the first international
conference on design science research in information systems and technology
(DESRIST 2006). pp. 83–106.
Project Management Institute. (2013). A guide to the project management body of
knowledge: (PMBOK Guide). Newtown Square: Project Management Institute.
Riege, C., Saat, J. and Bucher, T. (2009). Systematisierung von Evaluationsmethoden in der
gestaltungsorientierten Wirtschaftsinformatik. Wissenschaftstheorie und gestal-
tungsorientierte Wirtschaftsinformatik, pp.69–86.
Schauer, C. (2011). Die Wirtschaftsinformatik im internationalen Wettbewerb. Wiesba-
den: Gabler.
Takeda, H., Veerkamp, P. and Yoshikawa, H. (1990). Modeling design process. AI maga-
zine, 11(4), p.37.

94
APPENDIX A: PUBLICATIONS

Vaishnavi, V. and Kuechler, W. (2004). Design Science Research in Information Systems.


[online]. Available from: [Link]
systems/.
Venable, J., Pries-Heje, J. and Baskerville, R. (2012). A comprehensive framework for eval-
uation in design science research. Design Science Research in Information Systems.
Advances in Theory and Practice, pp.423–438.
Walls, J., Widmeyer, G. and El Sawy, O. (1992). Building an information system design the-
ory for vigilant EIS. Information Systems Research, 3(1), pp.36–59.
Werner, M. (2013). Colored Petri Nets for Integrating the Data Perspective in Process Au-
dits. In Proceedings of 32nd International Conference on Conceptual Modeling (ER
2013). 32nd International Conference on Conceptual Modeling (ER 2013). Hong
Kong, China: Springer-Verlag, pp. 387–394.
Werner, M. and Gehrke, N. (2011). Potentiale und Grenzen automatisierter Prozessprü-
fungen durch Prozessrekonstruktionen. In G. Plate, ed. Forschung für die Wirtschaft.
Aachen: Shaker Verlag.
Werner, M., Gehrke, N. and Nüttgens, M. (2012). Business Process Mining and Recon-
struction for Financial Audits. In Proceedings of the 45th Hawaii International Con-
ference on System Sciences. 45th Hawaii International Conference on System Sci-
ences. Maui, pp. 5350–5359.
Werner, M., Gehrke, N. and Nüttgens, M. (2013). Towards Automated Analysis of Busi-
ness Processes for Financial Audits. In Proceedings of the 11th International Con-
ference on Wirtschaftsinformatik. Internationale Tagung Wirtschaftsinformatik.
Leipzig.
Werner, M. and Nüttgens, M. (2014). Improving Structure - Logical Sequencing of Process
Models. In Proceedings of the 47th Hawaii International Conference on System Sci-
ences. 47th Hawaii International Conference on System Sciences. Bis Island.
Weske, M. (2012). Business process management concepts, languages, architectures. Ber-
lin; New York: Springer.
Wilde, T. and Hess, T. (2007). Forschungsmethoden der Wirtschaftsinformatik. Wirt-
schaftsinformatik, 49(4), pp.280–287.
Wilde, T. and Hess, T. (2006). Methodenspektrum der Wirtschaftsinformatik: Überblick
und Portfoliobildung. Arbeitspapiere des Instituts für Wirtschaftsinformatik und
Neue Medien, LMU München, 2.
Winter, R. (2008). Design science research in Europe. European Journal of Information
Systems, 17(5), pp.470–475.
Wollnik, M. (1998). Ein Referenzmodell des Informationsmanagements. Information Man-
agement, 3(3).
Yin, R.K. (2008). Case Study Research: Design and Methods. Fourth Edition. Sage Publica-
tions.

95
APPENDIX A: PUBLICATIONS

10.2 Process Mining

Number 2
Title Process Mining
Appendix 10.2
Primary Related Chapters 4.1
Type Journal Paper
Journal wisu - das wirtschaftsstudium
Reference (Gehrke and Werner, 2013)
31)
Acceptance Rate
VHB JQ 2.1 Ranking E (2.86)
WKWI Ranking -
ERA 2010 -
-
CORE 2013
Review Procedure Blinded
Number of Reviews -
1. Nick Gehrke
Authors
2. Michael Werner
Dissertation Points 0.67
Authorship
Overall 95 %
Design 95 %
Realization 95 %
Writing 95 %
Status Published
Part of other Dissertations No
Link [Link]

31
Invited for publication

96
APPENDIX A: PUBLICATIONS

Process Mining

PROF. DR. NICK GEHRKE / DIPL.-WIRT.-INF. MICHAEL WERNER


Elmshorn

Abstract: The increasing integration of information systems for


the operation of business processes provides the basis for inno-
vative data analysis approaches. Information systems support or
even automate the execution of business transactions in modern
companies. Business intelligence aims to support and improve
decision making processes by providing methods and tools for
analyzing data. Process mining builds the bridge between data
mining as a business intelligence approach and business process
management. Its primary objective is the discovery of process
models based on available event log data. The discovered pro-
cess models can be used for a variety of analysis purposes.

1 Introduction
New opportunities Companies use information systems to enhance the processing
for data analysis of their business transactions. Enterprise resource planning
(ERP) and workflow management systems (WFM) are the pre-
dominant information system types that are used to support and
automate the execution of business processes. Business pro-
cesses like procurement, operations, logistics, sales and human
resources can hardly be imagined without the integration of in-
formation systems that support and monitor relevant activities
in modern companies. The increasing integration of information
systems does not only provide the means to increase effective-
ness and efficiency. It also opens up new possibilities of data ac-
cess and analysis. When information systems are used for sup-
porting and automating the processing of business transactions
they generate data. This data can be used for improving business
decisions.
Business intelli- The application of techniques and tools for generating infor-
gence approaches mation from digital data is called business intelligence (BI).
Prominent BI approaches are online analytical processing
(OLAP) and data mining (Kemper et al. 2010 pp. 1-5). OLAP tools
allow analyzing multidimensional data using operators like roll-
up and drill-down, slice and dice or split and merge (Kemper et
al. 2010 pp. 99-106). Data mining is primarily used for discover-
ing patterns in large data sets (Kemper et al. 2010 p. 113).

97
APPENDIX A: PUBLICATIONS

Blessing and curse The availability of data as a new source of information is not only
a blessing, it can also become a curse. The phenomena of infor-
mation overflow (Krcmar 2010 pp. 54-57), data explosion (Van
der Aalst 2011 pp. 1-3) and big data (Chen et al. 2012) illustrate
several problems that arise from the availability of enormous
amounts of data. Humans are only able to handle a certain
amount of information in a given time frame. When more and
more data is available, how can it actually be used in a meaning-
ful manner without overstraining the human recipient?
Aim of process Data mining is the analysis of data for finding relationships and
mining patterns. The patterns are an abstraction of the analyzed data.
Abstraction reduces complexity and makes information available
for the recipient. The aim of process mining is the extraction of
information about business processes (Van der Aalst 2011 p. 1).
Process mining encompasses “techniques, tools and methods to
discover, monitor and improve real processes (…) by extracting
knowledge from event logs (…)“ (Van der Aalst et al. 2012 p. 15).
The data that is generated during the execution of business pro-
cesses in information systems is used for reconstructing process
models. These models are useful for analyzing and optimizing
processes. Process mining is an innovative approach and a bridge
between data mining and business process management.
Process mining Process mining evolved in the context of analyzing software en-
research gineering processes by Cook/Wolf in the late 1990s (Cook/Wolf
1998). Agrawal/Gunopulos (Agrawal et al. 1998) and Herbst/Ka-
ragiannis (Herbst/Karagiannis 1998) introduced process mining
to the context of workflow management. Major contributions to
the field have been added during the last decade by van der Aalst
and others by developing mature mining algorithms and ad-
dressing a variety of topic related challenges (Van der Aalst
2011). This has led to a well-developed set of methods and tools
that are available for scientists and practitioners.
Question 1: Why is process mining important? Why can it be
seen as a bridge between data mining and busi-
ness process management?

98
APPENDIX A: PUBLICATIONS

2 Process Mining Basics


2.1 Process Models and Event Logs
Graphical represen- The aim of process mining is the construction of process models
tations of business based on available logging data. In the context of information
processes system science a model is an immaterial representation of its
real world counterpart used for a specific purpose (Becker/Pro-
bandt et al. 2012 pp. 1-3). Models can be used to reduce com-
plexity by representing characteristics of interest and by omit-
ting other characteristics. A process model is a graphical repre-
sentation of a business process that describes the dependencies
between activities that need to be executed collectively for real-
izing a specific business objective. It consists of a set of activity
models and constraints between them (Weske 2012 p. 7).
Different modeling Process models can be represented in different process model-
languages ing languages for example using the Business Process Model and
Notation (BPMN), Event Driven Process Chains (EPC) or Petri
Nets. Petri Nets represent the dominant modeling language in
the field of process mining (Tiwari et al. 2008). While the formal
expressiveness of the Petri Net language is strong it is less suita-
ble for addressees that are not familiar with its syntax and se-
mantic. BPMN provides more intuitive semantics that are easier
to understand for recipients that do not possess a theoretical
background in informatics. We therefore rely on BPMN models
for illustration in this article.
Figure 1 shows a business process model of a simple purchasing
process (we only use a subset of basic BPMN elements and es-
pecially do not include participants, data or artifacts for simplifi-
cation and easier understandability). It starts with the ordering
of goods. At some point of time the ordered goods get delivered.
After the goods have been received an invoice is issued by the
supplier that is finally paid by the company that ordered the
goods.

(A) Order (B) Receive (C) Receive (D) Pay


Goods Goods Invoice Invoice

Fig. 1: Ideal Purchasing Process


The illustrated process model was created manually. So we do
not know if the model actually reflects reality. There might be
occurrences for example that invoices are paid before the goods
and invoices have been delivered. For ordered services there
might even be no such step as a recorded delivery. The

99
APPENDIX A: PUBLICATIONS

question arises: How can we get reliable information about the


real execution of the business process?
Case ID Event ID Timestamp Activity
1 1000 01.01.2013 Order Goods
1001 10.01.2013 Receive Goods
1002 13.01.2013 Receive Invoice
1003 20.01.2013 Pay Invoice
2 1004 02.01.2013 Order Goods
1005 01.01.2013 Receive Goods
… … …
Fig. 2: Event Log Structure
The approach used in process mining for answering this question
bases on the exploitation of data stored in information systems
that is created during the processing of business transactions. An
information system stores data in log files or database tables
when processing transactions. In the case of issuing an order
data about the type and quantity of ordered goods, preferred
suppliers, time of ordering etc. gets recorded. The stored data
can be extracted from the information system and be made
available in so called event logs. They constitute the data basis
for process mining algorithms.
Cases, events and An event log is basically a table. It contains all recorded events
attributes that relate to executed business activities. Each event is mapped
to a case. A process model is an abstraction of the real world
execution of a business process. A single execution of a business
process is called process instance. They are reflected in the
event log as a set of events that are mapped to the same case.
The sequence of recorded events in a case is called trace. The
model that describes the execution of a single process instance
is called process instance model. A process model abstracts
from the single behavior of process instances and provides a
model that reflects the behavior of all instances that belong to
the same process. Cases and events are characterized by classi-
fiers and attributes. Classifiers ensure the distinctness of cases
and events by mapping unique names to each case and event.
Attributes store additional information that can be used for anal-
ysis purposes. An example of an event log is given in Figure 2.
Question 2: What is the difference between event log, case,
event, and trace? How do they relate to each
other?

100
APPENDIX A: PUBLICATIONS

2.2 Mining Procedure


Source information Figure 3 provides an overview of the different process mining ac-
systems tivities. Before being able to apply any process mining technique
it is necessary to have access to the data. It needs to be ex-
tracted from the relevant information systems. This step is far
from trivial. Depending on the type of source system the rele-
vant data can be distributed over different database tables. Data
entries might need to be composed in a meaningful manner for
the extraction. Another obstacle is the amount of data. Depend-
ing on the objective of the process mining up to millions of data
entries might need to be extracted which requires efficient ex-
traction methods. A further important aspect is confidentiality.
Extracted data might include personalized information and de-
pending on legal requirements anonymization or pseudonymiza-
tion might be necessary.

Data Data Filtering Mining and


Analysis
Extraction and Loading Reconstruction

Fig. 3: Mining Procedure


Data filtering Before the extracted event log can be used it needs to be filtered
and loaded into the process mining software. There are different
reasons why filtering is necessary. Information systems are not
free of errors. Data may be recorded that does not reflect real
activities. Errors can result from malfunctioning programs but
also from user disruption or hardware failures that leads to er-
roneous records in the event log. Another error source occurs
without incorrect processing. A specific process is normally ana-
lyzed for a certain time frame. When the data is extracted from
the source system process instances can get truncated that were
executed over the boundaries of the selected time frame. They
need to be deleted from the event log or extracted completely.
Otherwise they lead to erroneous results in the reconstructed
process models. Event logs commonly do not exclusively contain
data for a single process. Filtering is necessary to curtail the
event log in a way that it only contains events that belong to the
scrutinized process. Such a filtering needs to be conducted care-
fully because it can lead to truncated process instances as well.
A common criterion is the selection of activities that are known
to belong to the same process. Data filtering and loading is com-
monly supported by software tools and performed in a single
step. But it can also be done separately.

101
APPENDIX A: PUBLICATIONS

Process models Once the data is loaded into the process mining software the ac-
tual mining and reconstruction of the process model can take
place. The mining includes the discovery of relationships in the
event log whereas the reconstruction produces a process model
as a graphical representation. The mining and reconstruction are
commonly provided by the same software tool in a single step.
Analysis purposes When the process models are mined and reconstructed they can
be used for the intended purpose. We summarize this step with
the term analysis. A fundamental goal of process mining is the
discovery of formerly unknown processes. In this case the recon-
struction is the aim itself but not limited to it. The analysis might
aim at additional objectives like identifying opportunities for
process optimization, organizational aspects or conformance
and compliance analysis.

2.3 Mining Algorithms


The main component in process mining is the mining algorithm.
It determines how the process models are created. A broad va-
riety of mining algorithms does exist. The following three cate-
gories will be discussed in more detail:
— Deterministic mining algorithms
— Heuristic mining algorithms
— Genetic mining algorithms
Deterministic min- Determinism means that an algorithm only produces defined

the same input. A representative of this category is the y-Algo-


ing algorithms pro- and reproducible results. It always delivers the same result for
duce defined and
reproducible re- rithm (Van der Aalst et al. 2002). It was one of the first algo-
sults rithms that are able to deal with concurrency. It takes an event
log as input and calculates the ordering relation of the events
contained in the log.
Heuristic mining al- Heuristic mining also uses deterministic algorithms but they in-
gorithms take fre- corporate frequencies of events and traces for reconstructing a
quencies into ac- process model. A common problem in process mining is the fact
count that real processes are highly complex and their discovery leads
to complex models. This complexity can be reduced by disre-
garding infrequent paths in the models.
Genetic mining al- Genetic mining algorithms use an evolutionary approach that
gorithms mimic mimics the process of natural evolution. They are not determin-
natural evolution istic. Genetic mining algorithms follow four steps: initialization,
selection, reproduction and termination. The idea behind these
algorithms is to generate a random population of process mod-
els and to find a satisfactory solution by iteratively

102
APPENDIX A: PUBLICATIONS

selecting individuals and reproducing them by crossover and mu-


tation over different generations. The initial population of pro-
cess models is generated randomly and might have little in com-
mon with the event log. But due to the high number of models
in the population, selection and reproduction better fitting mod-
els are created in each generation.

Case ID Event ID Timestamp Activity Resource


1 1000 01.01.2013 (A) Order Goods Peter
1001 10.01.2013 (B) Receive Goods Michael
1002 13.01.2013 (C) Receive Invoice Frank
1003 20.01.2013 (D) Pay Invoice Tanja
2 1004 02.01.2013 (A) Order Goods Peter
1005 03.02.2013 (B) Receive Goods Michael
1006 05.02.2013 (C) Receive Invoice Frank
1007 06.02.2013 (D) Pay Invoice Tanja
3 1008 01.01.2013 (A) Order Goods Louise
1009 04.01.2013 (C) Receive Invoice Frank
1010 05.01.2013 (B) Receive Goods Michael
1011 10.01.2013 (D) Pay Invoice Tanja
4 1016 15.01.2013 (A) Order Goods Peter
1017 20.01.2013 (C) Receive Invoice Claire
1018 25.01.2013 (D) Pay Invoice Frank
5 1023 01.01.2013 (A) Order Goods Michael
1024 10.01.2013 (B) Receive Goods Michael
1025 13.01.2013 (C) Receive Invoice Michael
1026 20.01.2013 (D) Pay Invoice Michael
Fig. 4: Sample Event Log
The outcomes of mining differ depending on the used algorithm.
We use the event log displayed in Figure 4 to illustrate the mod-
els created by different mining algorithms and for getting an im-
pression how the mining works.
The event log contains a very limited number of events and
cases. But it is nevertheless suitable to demonstrate key aspects
that are relevant for process mining and the selection of mining
algorithms.

103
APPENDIX A: PUBLICATIONS

by applying the z-Algorithm to the sample event log. It was


Figure 5 shows a mined process model that was reconstructed

translated into a BPMN model for better comparability. Obvi-


ously this model is not the same as the model in Figure 1. The
reason for this is that the mined event log includes cases that
deviate from the ideal linear process execution that was as-
sumed for modeling in Figure 1. In case 3 the invoice is received
before the goods. Due to the fact that both possibilities are in-
cluded in the event log (goods received before the invoice in case
1, 2, 5 and invoice received before the ordered goods in case 3)
the mining algorithm assumes that these activities can be carried
out concurrently.

(B) Receive
Goods

(A) Order (C) Pay


Goods Invoice

(C) Receive
Invoice

Fig. 5: Mined Process Model Using the y-Algorithm (the origi-


nal model was generated using the ProM software)
When we look at the model in Figure 5 a little bit closer we can
see another important fact. Actually case 4 is not reflected in the
process model. No execution sequence in the model is able to
reproduce the trace of case 4. It only allows two possible execu-
tion sequences: ABCD and ACBD put not ACD. The end-gateway
after the “Order Goods“ activity requires that both following
branches are executed. Therefore it is not possible to have an
execution sequence without the execution of activity B. The
model has a poor fitness because it is not able to reflect all rele-
vant traces in the event log. The model shown in Figure 6 is able
to replay all traces. Due to the exclusive-gateways all three
traces ABCD, ACBD and ACD are possible. But now the problem
occurs that there are much more execution sequences possible
than reflected in the event log. In fact the process model allows
for an infinite set of sequences. Now loops are possible either
starting from B to C or from C to B leading to possible sequences
with infinite iterations of B and C or C and B. The sequence
ABCBCD would for example be possible although it is not in-
cluded as a trace in the event log. If a process model is too gen-
eral it is called under-fitting. It has a poor precision. A major chal-
lenge in process mining is finding an adequate solution between
fitness, precision, simplicity, and generalizability.

104
APPENDIX A: PUBLICATIONS

(B) Receive (D) Pay


Goods Invoice

(A) Order (C) Receive


Goods Invoice

Fig. 6: Under-fitting Process Model


Fuzzy miner algo- Various advanced mining algorithms do exist that can be used
rithm for different purposes. Figure 7 illustrates the mined model using
the heuristic fuzzy miner algorithm (Günther / Van der Aalst
2007). The model does not follow the BPMN notation, instead it
uses a dependency graph representation. It does not contain any
gateway operators but shows the dependencies between differ-
ent activities. The dependency graph illustrates for example that
A was followed three times by B and two times by C.
Question 3: Why do different process mining algorithms pro-
duce different process models?

3 (B) Receive
Goods 1 (C) Pay
4 Invoice
5
(A) Order 1
Goods
5 3
(C) Receive 4
Invoice
2 5

Fig. 7: Mined Process Model Using the Fuzzy Miner Algorithm


(the original model was generated using the Disco process
mining software)
It is important to identify which requirements need to be consid-
ered for achieving the intended objectives for each individual
process mining project. The appropriateness of an algorithm
should be evaluated depending on the area of application.

3 Application Areas
3.1 Process Discovery and Enhancement
New opportunities A major area of application for process mining is the discovery
for analysis and op- of formerly unknown process models for the purpose of analysis
timization or optimization (Van der Aalst et al. 2012 p. 13). Business pro-
cess reengineering and the implementation of ERP systems in or-
ganizations gained strong attention starting in the 1990s. Practi-
tioners have since primarily focused on designing and

105
APPENDIX A: PUBLICATIONS

implementing processes and getting them to work. With matur-


ing integration of information systems into the execution of busi-
ness processes and the evolution of new technical possibilities
the focus shifts to analysis and optimization.
Data stored in in- Actual executions of business processes can now be described
formation systems and be made explicit. The discovered processes can be analyzed
is more reliable for performance indicators like average processing time or costs
than manually col- for improving or reengineering the process. The major ad-
lected data vantage of process mining is the fact that it uses reliable data.
The data that is generated in the source systems is generally hard
to manipulate by the average system user. For traditional pro-
cess modeling necessary information is primarily gathered by in-
terviewing, workshops or similar manual techniques that require
the interaction of persons. This leaves room for interpretation
and the tendency that ideal models are created based on often
overly optimistic assumptions.
Detection, predic- Analysis and optimization is not limited to post-runtime inspec-
tion, recommenda- tions. Instead it can be used for operational support by detect-
tion ing traces being executed that do not follow the intended pro-
cess model. It can also be used for predicting the behavior of
traces under execution. An example for runtime analysis is the
prediction of the expected completion time by comparing the in-
stance under execution with similar already processed instances.
Another feature can be the provision of recommendations to the
user for selecting the next activities in the process. Process min-
ing can also be used to derive information for the design of busi-
ness processes before they are implemented.

3.2 Conformance Checking


A specific type of analysis in process mining is conformance
checking (Adriansyah et al. 2011). The assumption for being able
to conduct conformance checking is the existence of a process
model that represents the desired process. For this purpose it
does not matter how the model was generated either by tradi-
tional modeling or by process mining.
Event logs can be A given event log is then compared with the ideal model for iden-
replayed to identify tifying conform or deviant behavior. The process instances pre-
conform or deviant sented in the log as cases are replayed as simulations in the
behavior model. Cases that can be replayed are labeled conform and cases
that cannot be replayed deviant.
If we check the conformance of the event log in Figure 4 against
the process model in Figure 1 we observe that cases 1, 2 and 5
are conform with the model whereas cases 3 and 4 deviate. The
simple statement that cases conform or deviate is generally not

106
APPENDIX A: PUBLICATIONS

sufficient. The hypothetical case with the trace ABCDD meaning


that the invoice for an ordered and received good was paid twice
is probably a more significant deviation from the ideal process
then the incorrect sequence of activities of B and C in the trace
ACBD observable in case 3. In general local diagnostics can be
calculated that highlight the nodes in the model where devia-
tions took place and global conformance measures that quantify
the overall conformance of the model and the event log.
When conducting conformance checking it should be kept in
mind that not every deviation needs to be negative and should
therefore be eliminated. Major deviations from the ideal model
might also mean that the model itself does not reflect real world
circumstances and requirements.

3.3 Compliance Checking


Internal or external Compliance refers to the adherence of internal or external rules.
rules External rules primarily include laws and regulations but can also
reflect industry standards or other external requirements. Inter-
nal rules include management directives, policies and standards.
Compliance checking deals with investigating if relevant rules
are followed. It is especially important in the context of internal
or external audits.
New possibilities Process mining offers new and rigorous possibilities for compli-
ance checking (Ramezani et al. 2012). A major advantage is the
already mentioned reliability of used information. In the con-
text of compliance this has an even higher impact because indi-
viduals will normally not admit incompliant behavior in tradi-
tional information gathering techniques like interviews due to
likely negative consequences for themselves.
Violation discovery For illustrating the difference of conformance and compliance
checking we refer to a well know and common compliance rule
that is called the 4-Eyes-Principle. It means that at least two per-
sons should be involved in the execution of a business process in
order to prevent errors or fraud. Errors are more likely to be dis-
covered if a second person is involved and fraud is less likely
when conspiracy is needed among individuals.
Let us have a look at case 5 in Figure 4. We have observed in
section 3.1 that this case conforms with the ideal process model
in Figure 1. The trace does not show any deviations from the
ideal execution sequence in the model. However it is not com-
pliant to the 4-Eyes-Principle rule because all activities were ex-
ecuted by the same user. Compliance checking is a relatively
novel field of research in the context of process mining. It can

107
APPENDIX A: PUBLICATIONS

be distinguished between forward and backward compliance


checking (Ramezani et al. 2012). Forward compliance checking
can further be divided into design-time and runtime compliance
checking (Becker/Delfmann et al. 2012):
— Pre-runtime compliance checking is conducted when pro-
cesses are designed or redesigned and implemented. The
designed model is checked if relevant rules are violated.
— Runtime compliance checks if violations occur when a busi-
ness transaction is processed.
— Post-runtime compliance is applied over a certain period of
time, when the transactions have already taken place.
Approaches for pre-runtime compliance checking are available
and first approaches already do exist for post-runtime compli-
ance checking. But solutions for runtime compliance checking
still need to be developed before they will be available for prac-
titioners.

3.4 Organizational Mining


So far we have focused on the control-flow of process models by
inspecting the sequence of activities that are possible in a pro-
cess model. As illustrated in the example for compliance check-
ing other attributes in the event log provides rich opportunities
for investigation.
Organizational set- Organizational mining aims to analyze information that is rele-
tings and interac- vant from an organizational perspective. This includes the dis-
tions between re- covery of social networks, organizational structures and re-
sources source behavior (Song and Van der Aalst 2008). Metrics like im-
portance, distance, or centrality of individual resources (this can
be individuals as shown in the sample log in Figure 4 but also
information systems, machines or other resources) can be com-
puted.
Question 4: Explain the difference between conformance and
compliance checking.

108
APPENDIX A: PUBLICATIONS

4 Tool Support
Process mining tools are necessary for the application in prac-
tice. Figure 8 lists various available process mining tools.

Product Name Type Link


ARIS Process Performance
C [Link]
Manager
celonis business intelligence C [Link]

Disco C [Link]

Genet/Petrify O [Link]
Interstage Business Process
C [Link]
Manager
QPR ProcessAnalysizer C [Link]

ProM O [Link]

ProcessGold C [Link]

Rbminer/Dbminer O [Link]

ReflectOne C [Link]

ServiceMosaic O [Link]
Fig. 8: Process Mining Tools — C=commercial, O=open source
(adapted from Van der Aalst, p. 271)
Open source soft- A major academic and open-source tool is ProM. It provides
ware tool plug-ins for many different mining algorithms, as well as analysis,
conversion and export modules. Disco is a commercial applica-
tion that benefits from intuitive and easy usability. It also pro-
vides integrated functionality for filtering and loading of event
logs. It is therefore especially suited for novel process mining us-
ers. A non-commercial license is available for academic institu-
tions.
None of the tools support the extraction of event data from the
relevant source systems. This means that the data has to be ex-
tracted with specialized data extraction software or by using ex-
port functionalities of the source systems.

109
APPENDIX A: PUBLICATIONS

5 Challenges and Contemporary Research


Questions
5.1 Noise and Incompleteness
Noise refers to rare We mentioned that incorrect records in the event log can for ex-
and infrequent be- ample result from software malfunctioning, user disruptions,
havior hardware failures or truncation of process instances during the
data extraction. Erroneous records in the event log should be
distinguished from a phenomenon called noise. Noise refers to
correctly recorded but rare and infrequent behavior (Van der
Aalst 2011 p. 148). Noise leads to increased complexity in the
process model. Process mining approaches should therefore be
able to handle or filter out noise. But this requirement is debat-
able because the meaning and role of noise varies depending on
the objective of the conducted process mining project. While it
might be necessary to abstract from infrequent behavior for re-
ducing complexity in the reconstructed model, infrequent be-
havior and the detection of outliers are key aspects for conform-
ance or compliance checking. And how can noise actually be dis-
tinguished from erroneous records? It is therefore necessary to
consider the handling of noise for every project individually.
Incompleteness Behavior that is not recorded in the event log cannot be consid-
means that not all ered for mining a process model. On the other hand process
possible behavior is models typically allow for much more behavior than recorded in
recorded the event log. We have seen that the process model shown in
Figure 6 allows for an infinite number of sequences whereas the
event log only contains five cases with three different traces.
While noise refers to the problem of potentially too much data
recorded in the event log incompleteness refers to the problem
of having too little data. It is unrealistic to assume that an event
log includes all possible process executions. When conducting a
process mining project it should be ensured that sufficient event
log data is available to discover the main control-flow structure.

5.2 Competing Model Quality Criteria


Four main quality criteria can be used to specify different quality
aspects of reconstructed process models: fitness, simplicity, pre-
cision and generalization.
— Fitness addresses the ability of a model to replay all behav-
ior recorded in the event log.
— Simplicity means that the simplest model that can explain
the observed behavior should be preferred.

110
APPENDIX A: PUBLICATIONS

— Precision requires that the model does not allow additional


behavior that is very different from the behavior recorded
in the event log.
— Generalization means that a process model is not exclu-
sively restricted to display the eventually limited record of
observed behavior in the event log but that it provides an
abstraction and generalizes from individual process in-
stances.
These quality criteria compete with each other as shown in Fig-
ure 9. This means that it is normally not possible to perfectly
meet all criteria simultaneously. The model in Figure 5 for exam-
ple has a high precision because it does not allow for any behav-
ior that is not included in the log but a low fitness because it can-
not replay all the cases. The model in Figure 6 has a perfect fit-
ness because it is able to replay all cases in the event log. But it
has a poor precision because it allows for an infinite number of
execution sequences not present in the event log. An adequate
balance between the quality criteria should be achieved for
every process mining project depending on the intended out-
come and further use of the reconstructed process models.
„able to replay the log”
fitness simplicity

Process
Model

generalization precision
„not over-fitting the log” „not under-fitting the log”

Fig. 9: Quality Dimensions (adapted from Van der Aalst 2011,


p. 151)

5.3 Event Log Quality and Labeling


The quality of event logs is crucial for the quality of the mined
and reconstructed process models. Business process and work-
flow management systems provide the highest quality of event
logs (Van der Aalst et al. 2012). They primarily focus on support-
ing and automating the execution of business processes and
therefore most likely also store high quality event data that can
easily be used for process mining. Data from ERP systems in gen-
eral do not provide the same quality of event logs. The logging
of event data is more a byproduct than intended software func-
tionality.

111
APPENDIX A: PUBLICATIONS

The quality of an A fundamental assumption for contemporary process mining ap-


event log depends proaches is that events in an event log are already mapped to
on the source sys- cases. But it depends on the quality of the available event log if
tems ability to rec- this is indeed the case. ERP systems as one major source of
ord process rele- transactional data in organizations commonly do not store ex-
vant data plicit data that maps events to cases. An interesting probabilistic
approach for labeling event data is presented by (Fer-
reira/Gillblad 2009). A promising approach could also be the con-
sideration of the application domain context. Gehrke / Müller-
Wickop present a mining algorithm that is able to operate with
unlabeled events by exploiting characteristic structures in the
available data thereby mapping events to cases in the process of
mining (Gehrke/Müller-Wickop 2010).

5.4 Complex Process Models


Complexity reduc- The presented examples in Figure 3 and 5 display very simple
tion is an important process models. Real world processes are commonly much more
research area complex. Their graphical representation can lead to highly com-
plex and incomprehensible models as shown in Figure 10. Two
typical categories of complex process models are called lasagna
and spaghetti processes (Van der Aalst 2011 pp. 277-320) be-
cause of their intertwined appearance. The reduction of com-
plexity is a major challenge and subject to recent research
(Reichert 2012).

Fig. 10: Example of a Complex Process Model (Werner et al.


2012 p. 5)

5.5 Concept Drift


Processes change When processes are mined and reconstructed it is usually as-
over time sumed that they are stable over the time of observation. But this
might not be the case. A process might work differently over a
certain period of time. This is called concept drift.

112
APPENDIX A: PUBLICATIONS

Assuming stable business processes is therefore a simplistic view


and contemporary research introduces approaches how to deal
with concept drift in the context of process mining (Bose et al.
2011).
Question 5: Why is it necessary to ensure a balance of differ-
ent quality criteria for process mining?

References:
Van der Aalst, W.M.P.: Process Mining: Discovery, Conformance
and Enhancement of Business Processes, Berlin, Heidel-
berg 2011.
Van der Aalst, W.M.P./Andriansyah, A./Alves de Medeiros, A.
K./Arcieri, F./Baier, T./Blickle, T./Bose,
J.C./Van den Brand, P./Brandtjen, R./Buijs, J.: Process mining
manifesto. In: BPM 2011 Workshops Proceedings, pp.
169-194.
Van der Aalst, W.M.P./Weijters, A. J. M. M./Maruster, L.: Work-
flow Mining: Which Processes can be
Rediscovered. Presented at the Proc. Int'l Conf. Eng. and De-
ployment of Cooperative Information Systems (EDCIS),
2002 (Vol. 2480), pp. 45-63.
Adriansyah, A./Van Dongen, B.F./Van der Aalst, W.M.P.: Con-
formance checking using cost-based fitness analysis. In:
Enterprise Distributed Object Computing Conference
(EDOC), 2011, 15th IEEE International, pp. 55-64.
Agrawal, R./Gunopulos, D./Leymann, F.: Mining Process Models
from Workflow Logs. In: Proc. Sixth Int'l Conf. Extending
Database Technology, 1998, pp. 469-483.
Becker, J./Delfmann, P./Eggert, M./Schwittay, S.: Generalizabil-
ity and Applicability of Model-Based Business Process
Compliance-Checking Approaches — A State-of-the-Art
Analysis and Research Roadmap. In: BuR — Business Re-
search (5:2), 2012, pp. 221-247.
Becker, J./Probandt, W./Vering, O.: Grundsätze ordnungsmäßi-
ger Modellierung Konzeption und Praxisbeispiel für ein ef-
fizientes Prozessmanagement. Berlin, Heidelberg 2012.
Bose, R./Van der Aalst, W./Zliobaite, I./Pechenizkiy, M.: Han-
dling concept drift in process mining. In: Advanced Infor-
mation Systems Engineering, 2011, pp. 391-405.

113
APPENDIX A: PUBLICATIONS

Chen, H./Chiang, R.H.L./Storey, V.C.: Business Intelligence and


Analytics: From Big Data to Big Impact. In: MIS Quarterly
(36:4), 2012, pp. 1165-1188.
Cook, J. E./Wolf, A. L.: Discovering models of software pro-
cesses from event-based data. In: ACM Trans. Softw. Eng.
Methodol. (7:3), 1998, pp. 215-249.
Ferreira, D./Gillblad, D.: Discovering process models from unla-
belled event logs. In: Business Process Management,
2009, pp. 143-158.
fluxicon: Process Mining and Process Analysis. [Link]-
[Link].
Gehrke, N./Müller-Wickop, N.: Basic Principles of Financial Pro-
cess Mining A Journey through Financial Data in Account-
ing Information Systems. In: Proceedings of the 16th
Americas Conference on Information Systems, 2010. Pre-
sented at the Americas Conference on Information Sys-
tems, Lima, Peru.
Günther, C./Van der Aalst, W.: Fuzzy mining-adaptive process
simplification based on multi-perspective metrics. Busi-
ness Process Management, 2007, pp. 328-343.
Herbst, J./Karagiannis, D.: Integrating Machine Learning and
Workflow Management to Support Acquisition and Adap-
tation of Workflow Models. In: Proceedings 9th Interna-
tional Workshop on
Database and Expert Systems Applications 1998, pp. 745-752.
Kemper, H.-G./Mehanna, W./Baars, H.: Business intelligence —
Grundlagen und praktische Anwendungen? Eine Einfüh-
rung in die IT-basierte Managementunterstützung. Wies-
baden 2010.
Krcmar, H.: Informationsmanagement. Berlin 2010.
Process Mining Group: Process Mining. [Link]-
[Link].
Ramezani, E./Fahland, D./Van der Aalst, W. M. P.: Where Did I
Misbehave? Diagnostic Information in
Compliance Checking. In: Business Process Management Lec-
ture Notes in Computer Science, edited by A. Barros/A.
Gal/E. Kindler, 2012 (Vol. 7481) Springer, pp. 262-278.
Reichert, M.: Visualizing Large Business Process Models: Chal-
lenges, Techniques, Applications. In: 1st Int'l Workshop
on Theory and Applications of Process Visualization, Pre-
sented at the BPM 2012, Tallin.

114
APPENDIX A: PUBLICATIONS

Song, M./Van der Aalst, W.M.P.: Towards comprehensive sup-


port for organizational mining. Decision Support Systems
(46:1), 2008, pp. 300-317.
Tiwari, A./Turner, C.J./Majeed, B.: A review of business process
mining: state-of-the-art and future trends. Business Pro-
cess Management Journal (14:1), 2008, pp. 5-22.
Werner, M./Schultz, M./Müller-Wickop, N./Gehrke,
N./Nüttgens, M.: Tackling Complexity: Process Recon-
struction and Graph Transformation for Financial Audits.
In: Proceedings of 33rd International Conference on Infor-
mation Systems Orlando 2012.
Weske, M.: Business process management concepts, languages,
architectures. Berlin, New York 2012.

Question and Answers:


Question 1: Why is process mining important? Why can it be
seen as a bridge between data mining and busi-
ness process management?
Business intelligence as an important approach to use data
stored in information systems for improving decision making
processes and to overcome challenges like data explosion and
information overflow. Data mining is the analysis of data for find-
ing relationships and patterns. Process mining uses data mining
techniques in the context of business process management and
enables the application of innovative approaches for improving
the management of business processes.
Question 2: What is the difference between event log, case,
event, and trace? How do they relate to each
other?
An event log is a collection of events. A case is a record of events
that relate to a single executed process instance. An event is a
recorded execution of an activity. A trace is a recorded sequence
of events that belong to the same case. Events are mapped to
cases. Each case has a trace. Different cases can and commonly
do embody the same trace.
Question 3: Why do different process mining algorithms pro-
duce different process models?
The results of a process mining algorithm depend on its design.
Process mining algorithms use different approaches for mining.

115
APPENDIX A: PUBLICATIONS

They can be deterministic or non-deterministic. They can also


use a heuristic or genetic approach. The results differ according
to the used approach. Another source of variety is the modeling
language. Mining algorithms that use different modeling lan-
guages for representation will create different process models.
Question 4: Explain the difference between conformance and
compliance checking.
The aim of conformance checking is the identification of process
instances that do not conform to a given process model. Compli-
ance checking is the analysis if compliance rules are adhered to.
The adherence of compliance rules is commonly safeguarded by
controls. Process models and instances can be checked if con-
trols were effective or not.
Question 5: Why is it necessary to ensure a balance of differ-
ent quality criteria for process mining?
It is generally not possible to achieve a perfect match of oppos-
ing quality criteria in real world settings. The objectives and cir-
cumstances differ for each process mining project. They need to
be identified at the beginning of each project. The relevance of
each quality criterion should be evaluated in regard to these ob-
jectives and circumstances and they should be assessed accord-
ingly. A mining approach should then be chosen that leads to the
desired balance of relevant quality criteria.

116
APPENDIX A: PUBLICATIONS

10.3 Potentiale und Grenzen automatisierter Prozessprüfungen durch Pro-


zessrekonstruktionen

Number 3
Potentiale und Grenzen automatisierter Pro-
Title zessprüfungen durch Prozessrekonstruktio-
nen
Appendix 10.3
Primary Related Chapters 4.2, 5.1
Type Book Chapter
Book Forschung für die Wirtschaft
Reference (Werner and Gehrke, 2011)
32)
Acceptance Rate
VHB JQ 2.1 Ranking -
WKWI Ranking -
ERA 2010 -
CORE 2013 -
Review Procedure Blinded
Number of Reviews -
1. Michael Werner
Authors
2. Nick Gehrke
Dissertation Points 0.67
Authorship
Overall 95%
Design 95%
Realization 95%
Writing 95%
Status Published
Part of other Dissertations No
[Link]
Link logue/[Link]?lang=de&ID=8&ISBN=978-3-
8440-0684-1

32
Invited for Publication

117
APPENDIX A: PUBLICATIONS

Potentiale und Grenzen automatisierter Prozessprüfungen durch Prozessre-


konstruktionen

MICHAEL WERNER, NICK GEHRKE


NORDAKADEMIE – Hochschule der Wirtschaft, Elmshorn

Abstract: Die Jahresabschlussprüfung stellt einen komplexen und hoch spezia-


lisierten Prozess dar. Die Durchführung der Jahresabschlussprüfung wird vom
Gesetzgeber vorgesehen, um die Adressaten der Jahresabschlussprüfung vor
Fehlinformationen zu schützen. Im Zuge der Digitalisierung der Wirtschaft und
die damit kontinuierlich steigende Durchdringung der Geschäftstätigkeiten
durch Informationssysteme stellen die Jahresabschlussprüfer vor neue Heraus-
forderungen. Die zunehmende Automation der Geschäftsprozessverarbeitung
durch den Einsatz moderner Informationssysteme lässt traditionelle Prüfungs-
handlungen insbesondere zur Prüfung von Geschäftsprozessen und internen
Kontrollsystemen ineffektiv oder zumindest ineffizient werden. Auf Unterneh-
mensseite erfolgt die Verarbeitung von Geschäftsvorfällen zunehmend system-
basiert und automatisiert. Die Prüfungsprozeduren beim Jahresabschluss sind
hingegen geprägt durch manuelle Prüfungshandlungen. Damit ergibt sich eine
Diskrepanz zwischen einer automatisierten Transaktionsverarbeitung auf Sei-
ten der Unternehmen und manuellen Prüfungsprozeduren auf Seiten der Jah-
resabschlussprüfer. Eine Möglichkeit, dieser Diskrepanz zu begegnen, besteht
darin, Methoden der Prozessextraktion und Prozessrekonstruktion mit Metho-
den zur automatisierten Prüfung von im Informationssystem eingebetteten
Kontrollen zu kombinieren, um so adäquate systembasierte und automatisierte
Prüfungsprozeduren zu entwickeln. In diesem Artikel untersuchen wir die An-
wendungsmöglichkeiten sowie auch Grenzen dieses Ansatzes.

1 Einleitung
Die Prüfung der externen Finanzberichterstattung von Unternehmen in Form einer Jahres-
abschlussprüfung stellt in unserem Wirtschaftssystem eine wichtige Kontrolle dar, deren
Ziel es ist, die Adressaten des Jahresabschlusses vor Fehlinformationen zu schützen.33 Die
Jahresabschlussprüfung nimmt als Kontrollfunktion eine derart hervorgehobene Stellung
ein, dass vom Gesetzgeber die Zuständigkeit zur Durchführung der Jahresabschlussprüfung
einer speziellen Berufsgruppe zugewiesen ist, die eine entsprechende Qualifikation auf-
weist.34
Jahresabschlussprüfer, traditionell auf Themenschwerpunkte der Rechnungslegung ausge-
richtet, sehen sich mit der fortschreitenden Integration von Geschäftsprozessen und Infor-
mationssystemen neuen Herausforderungen gegenübergestellt. Gegenwärtige Prüfungs-
ansätze berücksichtigen bis zu einem gewissen Grad den Gedanken der Prozessorientie-

33
Die Pflicht zur Prüfung des Jahresabschlusses ergibt sich aus §316 HGB.
34
Nach §319 HGB erfolgt die Prüfung des Jahresabschlusses durch Wirtschaftsprüfer, Wirtschaftsprüfungs-
gesellschaften, vereidigte Buchprüfer oder Buchprüfungsgesellschaften.

118
APPENDIX A: PUBLICATIONS

rung und die Verflechtung von Geschäftsprozessen und Informationssystemen. Sowohl na-
tionale wie auch internationale Prüfungsstandards (vgl. IDW PS 261 [1] und ISA 315 [2])
verlangen die Anwendung risikoorientierter Prüfungsansätze. Bei diesen Ansätzen werden
zunächst die wesentlichen Risiken identifiziert, die zu fehlerhaften Darstellungen in der Fi-
nanzberichterstattung führen können. Ausgehend von dieser Risikoeinschätzung wird eva-
luiert, welche Kontrollen innerhalb eines Unternehmens vorhanden sind, um die vorhan-
denen Risiken zu minimieren. Bei dieser Betrachtung sind auch Kontrollen des internen
Kontrollsystems zu würdigen. Darüber hinaus ist die Berücksichtigung von Informationssys-
temen und deren Prüfung mittlerweile ebenfalls obligatorisch (vgl. IDW PS 330 [3] in Ver-
bindung mit IDW RS FAIT 1[4] bzw. ISA 315.81 [2]).
Eine grundlegende Problematik wird bei bisher angewendeten, risikoorientierten Prüfungs-
vorgehen jedoch nicht adäquat berücksichtigt. Mit zunehmender Integration von Ge-
schäftsprozessen in Informationssysteme auf Seiten der Unternehmen nimmt die Automa-
tisierung der Verarbeitung von Geschäftsvorfällen zu. Enterprise Resource Planing (ERP)
Systeme stellen die am weitesten verbreiteten Informationssysteme zur Unterstützung und
Automatisierung zur Transaktionsverarbeitung dar. Sie dienen jedoch nicht nur der Unter-
stützung und Automatisierung der Geschäftsprozessabwicklung, sondern bilden auch die
Basis für die interne und externe Finanzberichterstattung. Dies bedeutet, dass die Informa-
tionen, die während der Verarbeitung der Geschäftsvorfälle erzeugt und gespeichert wer-
den, zugleich die Datenbasis bilden für die Berichterstattung, die letztendlich eine Aggre-
gation der zu Grunde liegenden Transaktionsdaten darstellt.
In dem Maße, in dem die Integration der Geschäftsprozesse in ERP Systeme zunimmt, ge-
winnen diese an Bedeutung für die Finanzberichterstattung und somit auch für die Jahres-
abschlussprüfung, die letztendlich eine Aussage treffen soll, ob die dargestellten Informa-
tionen frei von wesentlichen Fehlern sind.
Der systembasierten und automatisierten Verarbeitung auf Unternehmensseite stehen
manuelle Prüfungsprozeduren auf Seiten der Jahresabschlussprüfer gegenüber. Die manu-
elle Durchführung von Prüfungshandlungen nimmt dabei einen großen Teil der Ressourcen
der Abschlussprüfer in Anspruch. Durch geeignete systembasierte und automatisierte Prü-
fungsprozeduren wäre es möglich, die notwendigen Prüfungshandlungen zur Prüfung inte-
grierter Geschäftsprozesse effektiver und effizienter zu gestalten, was letztendlich Ressour-
cen freisetzen würde für Prüfungshandlungen zu ungewöhnlichen, von Standardprozessen
abweichenden, und komplexen Geschäftsvorfällen, die inhärent ein höheres Risiko für Feh-
ler und Manipulation aufweisen.
In WERNER et al. [5] wird beschrieben, wie Methoden der Prozessextraktion (Process Mi-
ning) und der Prozessrekonstruktion in Verbindung mit Methoden zur automatisierten
Überprüfung von im Informationssystem eingebetteten Kontrollen (Application Controls
bzw. Anwendungskontrollen) kombiniert werden können, um systembasierte und automa-
tisierte Prüfungsprozeduren zu entwickeln und anzuwenden. Dieser Ansatz wird in diesem
Artikel aufgegriffen mit dem Ziel, darzustellen, welche Möglichkeiten sich insbesondere aus
der Anwendung der Prozessextraktion und Prozessrekonstruktion ergeben und wo sich
Grenzen bzw. Restriktionen bei deren Anwendung ergeben.
Im Abschnitt zwei dieses Artikels wird ein Überblick über den derzeitigen Stand der wissen-
schaftlichen Literatur zum behandelten Themenfeld erörtert. Abschnitt drei schließt mit

119
APPENDIX A: PUBLICATIONS

einer kurzen Erläuterung zum Prozess der Abschlussprüfung an, um darzustellen, in wel-
chen Teilprozessen mittels systembasierter Prüfungsprozeduren Effizienz- und Effektivi-
tätsgewinne erzielt werden können. In Abschnitt vier wird erläutert, wie systembasierte
und automatisierte Prüfungsprozeduren eingesetzt werden können. Diese Betrachtung ist
notwendig, um erschließen zu können, wie und auf welchen Ebenen sich neue Erkenntnisse
durch Prozessrekonstruktion erzielen lassen. In Verbindung mit den vorangegangenen Ab-
schnitten wird in Abschnitt fünf dargestellt, welche Grenzen sich beim Einsatz der entwi-
ckelten Prozeduren ergeben und welchen Restriktionen deren Einsatz unterliegen. Ab-
schnitt sechs schließt mit einer Zusammenfassung und einem Ausblick auf weitere For-
schungsinhalte. Die in diesem Artikel erläuterten Erkenntnisse basieren auf dem vom Bun-
desministerium für Bildung und Forschung finanzierten Forschungsprojekt „Virtual Ac-
counting Worlds“ [6].

2 Stand der Wissenschaft


Wesentliche Grundlage für die in diesem Artikel dargestellten Methoden sind Arbeiten zur
Prozessextraktion oder Process Mining, dessen Ursprung in den 1990er Jahren liegt. COOK
und WOLF [7], [8], [9] untersuchten Process Mining im Zusammenhang mit Software Engi-
neering. Sie beschreiben unterschiedliche Methoden zum Process Mining, jedoch ohne
Vorgehen zu präsentieren, die explizit eine Erstellung von Prozessmodellen erlauben.
AGRAVAL et al. [10] wenden Process Mining im Zusammenhang mit Workflow Systemen
an. HERBST und KARAGIANNIS adressieren Process Mining ebenfalls im Zusammenhang mit
Workflow Management unter Verwendung eines induktiven Ansatzes [11], [12], [13], [14],
[15], [16]. Weitere Betrachtungen zur Entwicklung von Mining Software wurden durchge-
führt von MAXEINER et al. [17] und SCHIMM [18], [19], [20], [21].
Wesentliche Beiträge zum Process Mining wurden von VAN DER AALST et al. [22], [23], [24],
[25], [26], [27], [28], [29], [30], [31], [32], [33], [34] veröffentlicht. Sie handeln vorwiegend
von rekonstruierten Prozessmodellen aus Event Logs, die mittels Petri-Netzen dargestellt
werden. Diese Arbeiten sind besonders relevant, da die entwickelten Methoden zum Pro-
cess Mining auf Basis von Event Logs adaptiert werden können für Extraktions- und Rekon-
struktionsmethoden zu Prozessen in Finanzbuchhaltungssystemen. Die Arbeiten von VAN
DER AALST et al. decken darüber hinaus relevante Themenbereiche zur Workflow Perfor-
mance, Nebenläufigkeit, Rauschen und Abweichungsanalysen im Zusammenhang mit Pro-
cess Mining ab. Obwohl diese Arbeiten hilfreiche Grundlagen bieten, sind einige Einschrän-
kungen zu berücksichtigen. Während bei den von VAN DER AALST et al. erarbeiteten Me-
thoden die Rekonstruktion von Prozessmodellen in Form von Petri-Netzen, die die Gesamt-
heit aller möglichen Prozesspfade eines Event Logs repräsentieren, im Mittelpunkt steht,
ist die Zielsetzung bei der Rekonstruktion von Finanzprozessen aus ERP Systemen unter-
schiedlich. Letztere beabsichtigt, bestimmte Geschäftsvorfälle für spezielle Fragestellungen
zu extrahieren und zu einem Prozessmodell zu aggregieren. Bei der Aggregation werden
insbesondere Zusammenhänge aus der Rechnungslegung herangezogen, indem die Ver-
knüpfung der zugrundeliegenden Belege über eine offene-Posten-Systematik genutzt wird.
Des Weiteren stellen ERP Systeme wesentlich detailliertere Informationen zur Verfügung
als herkömmliche Event Logs, die als Basis für spezielle Rekonstruktionen verwendet wer-
den können.

120
APPENDIX A: PUBLICATIONS

Generell lässt sich feststellen, dass die wissenschaftliche Literatur eine stark technische
Prägung aufweist. Die Verbindung zwischen Prozessflüssen, Prozessextraktion, Prozessre-
konstruktion und Prozessaggregation im Anwendungsfeld des Rechnungswesens ist in der
wissenschaftlichen Literatur wenig vertreten. Der Grund hierfür liegt wahrscheinlich in der
thematischen Distanz der klassischen Themenbereiche Rechnungswesen und Compliance
auf der einen sowie informationstechnisches Process Mining auf der anderen Seite.
In jüngerer Zeit befassen sich Arbeiten von GEHRKE et al. [35], [36], [37], [38], MÜLLER-
WICKOP et al. [39] sowie WERNER et al. [5] mit der Automatisierung von Prozessprüfungen
im Rahmen der Jahresabschlussprüfung. Hier werden Methoden entwickelt, die eine An-
wendung von Process Mining in ERP Systemen für Finanzprozesse ermöglichen. In WERNER
et al. [5] werden die angewendeten Methoden in einen Gesamtzusammenhang gesetzt.
Die vorliegende Arbeit greift insbesondere die dort vorgestellten Erkenntnisse auf und er-
weitert diese um Betrachtungen der Möglichkeiten und Grenzen die sich durch den Einsatz
von Prozessextraktion und -rekonstruktion im Rahmen der Jahresabschlussprüfung und
auch darüber hinaus ergeben.
Auch JANS et al. [40], [41] und ALLES et al. [42] behandeln den Zusammenhang zwischen
Process Mining und dessen Einsatz bei Accounting Information Systems und ERP Systemen.
JANS et al. betrachten Process Mining u.a. als Fortentwicklung von Data Mining Techniken
und Massendatenauswertungen für spezielle Fragestellungen hinsichtlich der Aufdeckung
krimineller Aktivitäten und der Prüfung von Geschäftsprozessen. Die Ansätze von GEHRKE
et al. [35], [36], [37], [38], MÜLLER-WICKOP et al. [39] und WERNER et al. [5] nehmen eine
unterschiedliche Betrachtung vor, indem die fachliche Perspektive im Sinne der Betrach-
tung der Zusammenhänge aus Rechnungslegungssicht als Basis für die zu entwickelnden
Extraktions-, Aggregations- und Auswertungsmethoden herangezogen wird. Dieser Ansatz
unterscheidet sich dahingehend von der technischen Erweiterung von Data Mining Metho-
den durch Methoden des Process Mining. Ziel des von GEHRKE et al., MÜLLER-WICKOP et
al. und WERNER et al. behandelten Forschungsansatzes ist es, automatisierte und system-
basierte Methoden für die Prüfung von integrierten Geschäftsprozessen zu entwickeln, die
in einem risikoorientierten Prüfungsansatz eine neue Bedeutung einnehmen und beste-
hende Methoden nicht nur erweitern. Trotz dieses Unterschieds werden insbesondere die
von JANS et al. [40] beschriebenen Einsatzfelder von Business Process Mining in diesem
Artikel aufgegriffen und auf Anwendbarkeit der speziell von GEHRKE et al. entwickelten
Methoden untersucht.

3 Grenzen manueller Prüfungsprozeduren bei hochintegrierten


Geschäftsprozessen
Bevor wir bewerten können, warum manuelle Prüfungsprozeduren bei Jahresabschluss-
prüfungen für Unternehmen, deren Geschäftsprozesse stark in ERP Systeme integriert sind,
ineffizient oder im Extremfall auch ineffektiv werden, ist ein Verständnis des Prozesses zur
Durchführung des Jahresabschlusses notwendig sowie ein Überblick der zum Einsatz kom-
menden Prüfungstechniken. Grundlegende Informationen im Hinblick auf die in diesem Ar-
tikel dargestellte Problematik sind bereits in WERNER et al. [5] aufgeführt, so dass an dieser
Stelle nur wesentliche Aspekte aufgegriffen werden.

121
APPENDIX A: PUBLICATIONS

3.1 Die Jahresabschlussprüfung


Unternehmen sind gesetzlich dazu verpflichtet, einen Jahresabschluss zu erstellen, der von
Jahresabschlussprüfern geprüft und testiert wird (§316 und §319 HGB). Die externen Ad-
ressaten des Jahresabschlusses umfassen insbesondere die Unternehmensleitung, Control-
ler, Aufsichtsrat oder Beirat, Finanzverwaltung, Kreditgeber, Gesellschafter und Anteilseig-
ner, Lieferanten, Arbeitnehmer, Kunden, externe Aufsichtsbehörden, Finanzanalysten,
Konkurrenz und die generelle Öffentlichkeit [43]. Der Jahresabschlussprüfer hat festzustel-
len, ob der Jahresabschluss ein getreues Bild der Vermögens-, Finanz- und Ertragslage des
Unternehmens wiedergibt und frei ist von wesentlichen Fehlern. Die Anforderungen für die
Prüfung des Jahresabschlusses sind in Prüfungsstandards definiert, die von nationalen
(IDW) und internationalen Institutionen (IAASB) erstellt werden, nach denen sich die Ab-
schlussprüfer zu richten haben. Gleiches gilt für Standards zur Rechnungslegung. Es bleibt
zu erwähnen, dass sowohl Rechnungslegungsstandards wie auch Prüfungsstandards von
Land zu Land variieren. Es ist jedoch zu beobachten, dass international eine Angleichung
der Standards erfolgt [44]. Eine Diskussion über die Konvergenz der Rechnungslegungs-
bzw. Prüfungsstandards ist nicht Gegenstand dieser Arbeit. Allerdings gelten die hier erar-
beiteten Ergebnisse nicht nur für den deutschen Raum sondern sind auch übertragbar auf
andere Länder, die westlich geprägten Wirtschaftssystemen folgen.
Die Jahresabschlussprüfung stellt einen Prüfungsprozess dar, der aus verschiedenen Teil-
prozessen besteht. Der erste Teilprozess umfasst im Allgemeinen die Informationssamm-
lung über das zu prüfende Unternehmen und die Identifikation und Bewertung potentieller
Fehlerrisiken für die Abschlusserstellung anhand der zur Verfügung stehenden Informatio-
nen. Auf die Analyse möglicher Fehlerrisiken folgt eine Informationserhebung im Unter-
nehmen vorhandener Geschäftsprozesse mit dem Ziel, interne Kontrollen in den Prozessen
zu identifizieren, die geeignet sind, relevante Risiken zu minimieren. Wenn diese Kontrollen
als geeignet beurteilt werden, schließt sich die Funktionsprüfung der Kontrollen an, in der
überprüft wird, ob die Kontrollen im jeweils zu betrachtenden Zeitraum effektiv waren, und
die Kontrollziele erreicht haben. Nachdem die Kontrollen getestet sind, wird evaluiert, wel-
che weiteren aussagebezogenen Prüfungshandlungen notwendig sind, um eine ausrei-
chende Prüfungssicherheit zu gewährleisten, so dass davon ausgegangen werden kann,
dass der geprüfte Abschluss mit hinreichender Sicherheit frei ist von wesentlichen Fehlern.
Der Prüfungsprozess endet mit der Erstellung des Prüfungsberichts und der Testierung. Der
Prozess der Jahresabschlussprüfung ist in Abbildung 3.1.1 dargestellt.
Informationssammlung und Bewertung der
Fehlerrisiken

Aufbauprüfung des internen Kontrollsystems

Funktionsprüfung des internen Kontrollsystems

Aussagebezogene Prüfungshandlungen

Berichterstattung und Testierung

Abbildung 3.1.1 Prüfungsprozess

122
APPENDIX A: PUBLICATIONS

Innerhalb der einzelnen Teilprozesse kommen verschiedene Prüfungsprozeduren zum Ein-


satz. Für die Aufbau- und Funktionsprüfung finden Informationserhebungen und Kontroll-
tests statt. Für die anschließenden aussagebezogenen Prüfungshandlungen werden analy-
tische und substantielle Prüfungshandlungen durchgeführt. Für die Durchführung der Prü-
fungshandlungen werden unterschiedliche Techniken verwendet wie Interviews, Beobach-
tung, Begutachtung von Unterlagen (Inspektion) und Wiederholung von Tätigkeiten (Re-
performance). Je höher der Verlässlichkeitsgrad der Information für eine bestimmte Prü-
fungshandlung angesetzt wird, desto investigativer muss die gewählte Prüfungstechnik
sein und desto höher ist der Prüfungsaufwand. Abbildung 3.1.2 veranschaulicht diesen Zu-
sammenhang.

Wiederholung
Verlässlichkeit

Inspektion

Beobachtung

Interview

Prüfungsaufwand

Abbildung 3.1.2 Prüfungstechniken


Zur Informationserhebung und Aufbauprüfung werden vorwiegend Interviews geführt, bei
denen Ansprechpartner mit ausreichender Sachkenntnis im Unternehmen befragt werden.
Bei der Beobachtung werden Mitarbeiter bei der Durchführung von Tätigkeiten aktiv beo-
bachtet, um daraus die gewünschten Informationen zu ziehen. Bei der Inspektion werden
vorhandene Ursprungsbelege gesichtet und evaluiert. Bei der Wiederholung werden Tätig-
keiten, die in der Regel durch Mitarbeiter des zu prüfenden Unternehmens als Bestandteil
normaler Tätigkeiten zur Geschäftsprozessabwicklung oder zu Kontrollzwecken durchge-
führt werden, vom Prüfer unabhängig wiederholt und damit nachvollzogen.

3.2 Manuelle Prüfungsprozeduren bei hochintegrierten Geschäftsprozessen


Geschäftsprozesse spielen bei der Jahresabschlussprüfung sowohl in der Aufbauprüfung als
auch in der Funktionsprüfung eine zentrale Rolle. Während der Aufbauprüfung werden vor
allem Interviews durchgeführt, um ein Verständnis der als wesentlich identifizierten Pro-
zesse zu erhalten. Die Durchführung der Interviews ist eine rein manuelle Tätigkeit und sehr
zeitintensiv. Während der Funktionsprüfung werden vor allem Beobachtungen am System,
Inspektionen von zur Verfügung gestellten Unterlagen zu Buchungsbelegen und Kontroll-
belegen durchgeführt, sowie Kontrollhandlungen nachvollzogen.
Zu Beginn der Prüfung identifiziert der Prüfer geeignete Interviewpartner, um Informatio-
nen über den Prozessablauf und implementierte Kontrollen zu erhalten. Nach Durchfüh-
rung der Interviews evaluiert er, ob Kontrollen vorhanden und angemessen sind, um die
Risiken, die sich aus dem Geschäftsprozess für die Finanzberichterstattung ergeben, zu mi-
nimieren. In einem zweiten Schritt sucht der Prüfer eine geeignete Grundgesamtheit, um

123
APPENDIX A: PUBLICATIONS

stichprobenartig eine Funktionsprüfung der Kontrollen durchzuführen. Für die Funktions-


prüfung werden bereitgestellte Unterlagen inspiziert oder direkt Prüfungen im System
durchgeführt, um die Funktionsfähigkeit eingebetteter Kontrollen (Application Controls
bzw. Anwendungskontrollen) zu verifizieren.
Bei manuellen Prozessen mit wenigen Geschäftsvorfällen und geringer Integration der In-
formationssysteme kann die manuelle Durchführung der notwendigen Prüfungshandlun-
gen eine angemessene Vorgehensweise darstellen. Bei Unternehmen, bei denen die Ge-
schäftsprozesse stark in die Informationssysteme integriert sind, ergeben sich eine Reihe
von Faktoren, die die Angemessenheit manueller Prüfungshandlungen in Frage stellen.
Mit zunehmender Integration und Automatisierung der Transaktionsverarbeitung wird der
Prozessablauf für die einzelnen beteiligten Personen intransparenter. Dies führt dazu, dass
die Ansprechpartner, die im Regelfall der Fachabteilung entstammen, häufig kein vollstän-
diges Verständnis der Verarbeitungsroutinen im System haben. Hinzu kommt, dass die An-
wendung von Stichprobenverfahren an Grenzen stößt, wenn die Grundgesamtheit der
Transaktionen sehr umfangreich wird. Bei Millionen von Transaktionen pro Geschäftsjahr
ist es fraglich, ob die Inspektion von Stichproben noch effektiv ist. Zumindest kann davon
ausgegangen werden, dass mit einer notwendigen Ausweitung der Stichproben aufgrund
der Zunahme der zu betrachtenden Geschäftsvorfälle die Effizienz der manuellen Prüfungs-
tätigkeiten abnimmt.

4 Anwendungsmöglichkeiten der Prozessrekonstruktion


In diesem Abschnitt werden die Anwendungsmöglichkeiten untersucht, die sich durch den
Einsatz von Methoden zur Prozessextraktion, -rekonstruktion und -aggregation ergeben
und wie diese mit Methoden zur automatischen Prüfung von Anwendungskontrollen kom-
biniert werden können. Um die Möglichkeiten der Anwendung zu untersuchen, ist es zu-
nächst notwendig, die grundlegenden Methoden sowie deren Zusammenwirken zu verste-
hen.

4.1 Methodenkombination
Methoden zur Prozessextraktion für Finanzprozesse erlauben, in ERP Systemen gespei-
cherte Daten zu extrahieren. Die extrahierten Daten können über entsprechende Algorith-
men zu Prozessinstanzen aggregiert und visualisiert werden (vgl. [36] und [37]). Grundlage
für die Extraktion und Rekonstruktion ist die Verknüpfung der rechnungslegungsrelevanten
Daten über offene-Posten-geführte Buchungen. Der Ausschnitt einer rekonstruierten und
visualisierten Prozessinstanz ist in Abbildung 4.1.1 dargestellt.

124
APPENDIX A: PUBLICATIONS

Abbildung 4.1.1 Prozessinstanz aus [5]


Die dargestellte Prozessrekonstruktion kann für beliebig viele und beliebig komplexe Pro-
zessinstanzen durchgeführt werden. Als Ergebnis der Rekonstruktion liegen je nach Aus-
wertungsziel eine bis alle in dem jeweiligen Betrachtungszeitraum im ERP System durchge-
führten Prozessinstanzen vor. Die extrahierten und rekonstruierten Instanzen müssen zu
Prozessmodellen aggregiert werden. Auf diese Weise werden eine höhere Abstraktionse-
bene und die Möglichkeit erreicht, die aggregierten Prozessmodelle mit den Ergebnissen
automatisierter Prüfung von Anwendungskontrollen zu vereinigen.
Steuerungs- und Kontrollmechanismen in ERP Systemen erlauben, die Durchführung von
Transaktionen im System zu beeinflussen und zu regulieren. Sie werden meist bei der Sys-
temeinführung entsprechend der Anforderungen des Nutzers konfiguriert. Anwendungs-
kontrollen steuern und überwachen die Durchführung aller Transaktionen, für die sie akti-
viert sind. Während bei manuell durchgeführten Prozessen manuelle Kontrollen dazu die-
nen, die vollständige und richtige Abarbeitung von Geschäftsvorfällen sicherzustellen, wird
diese Funktion bei integrierten Geschäftsprozessen von Anwendungskontrollen wahrge-
nommen. Die Herausforderung in der Praxis besteht darin, die in den ERP Systemen unter-
schiedlicher Softwareanbieter heterogen implementierten Kontrollen für die Zwecke der
automatisierten Prüfung auszuwerten. In [35] wird hierzu ein Lösungsansatz vorgestellt.
Mittels automatisierter Prozessrekonstruktion und der anschließenden Aggregation zu Pro-
zessmodellen in Kombination mit der Anwendung von Methoden zu automatisierten Prü-
fung von Anwendungskontrollen ist es möglich, Methoden zur automatisierten Prüfung
von Geschäftsprozessen zu entwickeln. Dieser Zusammenhang ist in Abbildung 4.1.2 gra-
phisch dargestellt.

125
APPENDIX A: PUBLICATIONS

Abbildung 4.1.2 Kombination Prozessrekonstruktion und automatisierte Prüfung von


Anwendungskontrollen

4.2 Möglichkeiten der Prozessrekonstruktion für die Jahresabschlussprüfung


Die im vorherigen Abschnitt präsentierten Methoden können nun für Zwecke der Automa-
tisierung der Prüfung von Geschäftsprozessen in Jahresabschlussprüfungen integriert wer-
den.
Zur Verdeutlichung der Integration der Methoden verwenden wir eine Analogie, die in [5]
vorgestellt wurde. Wir vergleichen hierbei den Jahresabschluss mit einem See, der mit
Wasser gefüllt ist. Das Wasser entspricht den rechnungslegungsrelevanten Daten in ERP
Systemen. Der See wird von Flüssen gespeist. Die Flüsse entsprechen den Geschäftsprozes-
sen, die das Wasser aufnehmen, bzw. die Transaktionen erzeugen, und diese in den See
bzw. die Finanzberichterstattung tragen. Die Jahresabschlussprüfung vergleichen wir mit
der Prüfung der Wasserqualität des Sees. Bei der Anwendung gegenwärtiger Prüfungsver-
fahren würde der Prüfer aus den Flüssen stichprobenartig Wasserproben entnehmen und
deren Qualität prüfen und diese Prüfung ggf. durch weitere Stichproben direkt im See er-
gänzen. Das Problem bei diesem Ansatz liegt darin, dass einerseits die Flüsse nur stichpro-
benartig zu bestimmten Zeitpunkten getestet werden und viel gravierender, dass der Prü-
fer gar nicht weiß, welche Flüsse, Parallelläufe und Verzweigungen tatsächlich existieren.
Hierfür kann er lediglich auf seine Erfahrungen aus ähnlichen Prüfungen aufbauen, die zu-
treffend sein können oder aber auch nicht.
Bei der Anwendung automatisierter und systembasierter Prüfungsmethoden werden diese
Informationen transparent. Der Prüfer erhält zunächst eine Karte über alle Zuflüsse, die
den See speisen (Prozessrekonstruktion). Des Weiteren erhält er Einsicht, durch welche

126
APPENDIX A: PUBLICATIONS

Kontrollstationen die Flüsse reguliert werden und ob diese funktionsfähig sind (automati-
sierte Prüfung von Anwendungskontrollen). Auf Basis dieser Kenntnisse kann der Prüfer
seine Prüfungshandlungen auf unkontrollierte oder ungewöhnliche Flussläufe konzentrie-
ren.
Eine konzeptionelle Darstellung unter der Verwendung der beschriebenen Analogie ist in
Abbildung 4.2 dargestellt. In Tabelle 4.2 sind Beispiele für die visualisierten Prozessflüsse
und Anwendungskontrollen enthalten.

Abbildung 4.2 Prozesskarte aus [5]

Tabelle 4.2 Beispielprozesse und Application Controls aus [5]


Die gewählte Analogie verdeutlicht das Potential des Einsatzes von Methoden der Prozess-
rekonstruktion in der Jahresabschlussprüfung. Bei zunehmender Automatisierung der
Transaktionsverarbeitung und Integration von Geschäftsprozessen in ERP Systeme stellt
die Anwendung automatisierter und systembasierter Kontrollen eine effektive Möglichkeit
dar, der Flut automatisierter Transaktionsabwicklungen mit automatischen Kontrollen zu
begegnen. Es ist davon auszugehen, dass deren Anwendung eine neue Qualität der Ergeb-
nisse der Prozessprüfung erreicht. Nicht mehr Erfahrung und manuelle Auswertung einzel-
ner Kontrollen bestimmen die Qualität einzelner Prozessprüfungen, sondern die Korrekt-
heit der Rekonstruktionsalgorithmen, die wissenschaftlich bewiesen werden können, so-
wie die Verfahren zur automatischen Prüfung von Anwendungskontrollen. Einschränkun-
gen zu letzteren werden in Abschnitt fünf erläutert. Allerdings lässt sich festhalten, dass
wesentliche Effektivitäts- und Effizienzgewinne insbesondere bei Prüfungen von hochinte-
grierten Geschäftsprozessen und hoher Transaktionsrate zu erwarten sind.

127
APPENDIX A: PUBLICATIONS

Einer der wesentlichen Vorteile des präsentierten Ansatzes ist die Systemunabhängigkeit
der Methoden. Dies ist darin begründet, dass die Rekonstruktion der Prozesse nicht unmit-
telbar auf Basis von Verknüpfungen der Daten auf Datenbankebene basiert sondern nur
mittelbar. Die Verknüpfung wird hergestellt über die offene-Posten-Buchhaltung rech-
nungslegungsrelevanter Buchungen. Die offene-Posten-Buchführung ist ein Erfordernis der
Rechnungslegung und muss somit von allen Informationssystemen angeboten werden, die
für Rechnungslegungszwecke eingesetzt werden. Eine Verknüpfung der Buchungen auf Da-
tenbankebene muss aus diesem Grund gewährleistet sein. Auch wenn die Umsetzung der
Verknüpfung technisch unterschiedlich ausfallen mag, kann davon ausgegangen werden,
dass sie stets für die Rekonstruktion genutzt werden kann, denn das jeweilige Informati-
onssystem muss auf solche Verknüpfungen zurückgreifen können, um eine offene-Posten-
Buchhaltung zu gewährleisten.

4.3 Unterstützung von substantiellen Prüfungshandlungen


Neben der Automatisierung der Prüfung von Geschäftsprozessen können Methoden zur
Prozessrekonstruktion eingesetzt werden, um dem Prüfer Hilfestellung zu substantiellen
Prüfungshandlungen zu geben. Über die Rekonstruktion und Visualisierung von Prozess-
flüssen wird offensichtlich, ob Verzweigungen von Prozessflüssen existieren, die auf Aus-
nahmen von Standardprozessen schließen lassen. Mittels der Auswertung nicht nur der
Anzahl von Transaktionen in einem Prozessfluss sondern auch über die Ermittlung der über
die Transaktionen abgewickelten Beträge können einzelne Transaktionen mit hohen Beträ-
gen und somit hoher Relevanz für den Jahresabschluss für dezidierte substantielle Prü-
fungshandlungen identifiziert werden.
Die Generierung von gezielten Stichproben z.B. auf Basis der Betragsgröße abgewickelter
Geschäftsvorfälle kann unterstützt werden. Diese Unterstützung ist insbesondere relevant
für Prozessflüsse, die nicht durch Anwendungskontrollen reguliert werden.

4.4 Analyse der Metadaten


JANS et al. [40] eröffnen einen interessanten Blickwinkel auf die Auswertung von Metada-
ten, die bei der Speicherung von Daten zu im System verarbeiteten Transaktionen erfolgen.
Metadaten sind hierbei als Daten zu verstehen, die nicht unmittelbar Daten zum bearbei-
teten Geschäftsvorfall betreffen, sondern die durch das System automatisch erzeugt und
gespeichert werden. Dies beinhaltet z.B. Daten zum Benutzer, der die Transaktion durch-
geführt hat, das Datum der Durchführung der Buchung im System (im Gegensatz zum vom
Benutzer eingegebenen Buchungsdatum) und Änderungshistorien zu Datenobjekten. Bei
manueller Bearbeitung eines Geschäftsvorfalls wären diese Informationen in der Regel
nicht verfügbar.
Die Auswertung der Metadaten erlaubt, Prüfungen zu unterstützen, die bei manueller Prü-
fung nicht möglich wären. Zum Beispiel betrifft dies Abweichungen zum zeitlichen Verlauf
von Prozessschritten und Manipulationen an Daten wie Preis oder Menge bestellter Wa-
ren.
Ein großer Vorteil bei der Auswertung dieser Daten liegt darin, dass die gespeicherten Da-
ten vor Manipulation geschützt sind, sofern Zugriffsrechte für Administratorberechtigun-
gen auf den relevanten Zugriffsebenen entsprechend eingeschränkt sind.

128
APPENDIX A: PUBLICATIONS

4.5 Funktionstrennung
Unter Funktionstrennung wird die Trennung kritischer Berechtigungen in ERP Systemen
verstanden. Durch die Trennung von Funktionen soll erreicht werden, dass mehrere kriti-
sche Transaktionen nicht durch ein und denselben Benutzer, z. B. für wirtschaftskriminelle
Tätigkeiten, unbemerkt verwendet werden können. Über die Funktionstrennung soll er-
reicht werden, dass ein Vier-Augen-Prinzip zu kritischen Geschäftsvorfällen eingehalten
wird.
Für die Überprüfung von Funktionstrennung und zur Entdeckung von Funktionstrennungs-
konflikten existieren anwendungsspezifische Softwarelösungen. Mittlerweile werden ent-
sprechende Funktionalitäten auch durch entsprechende Module verbreiteter Softwarean-
bieter angeboten. Diese Anwendungen erlauben jedoch lediglich zu analysieren, ob Funk-
tionstrennungskonflikte vorliegen, ob also Berechtigungen für die Durchführung kritischer
Transaktionen bei einzelnen Benutzern vorliegen. Sie erlauben aber in der Regel keine Aus-
wertung, ob diese Berechtigungen tatsächlich von dem jeweiligen Benutzer auch eingesetzt
wurden. Diese Prüfung muss manuell durchgeführt werden. Durch die Anwendung von Me-
thoden zu Prozessrekonstruktion und Auswertung bestimmter Daten wie Benutzername
und Transaktionscode wäre eine automatisierte Auswertung zu Funktionstrennungskon-
flikten bei tatsächlich eingetretenen Geschäftsvorfällen möglich.

4.6 Leistungsanalyse
Neben dem Einsatz für Prüfungszwecke könnte eine Anwendung zur Messung der Leistung
von in ERP Systemen integrierten Geschäftsprozessen ermöglichen. HUFGARD [45] bietet
eine umfangreiche Analyse über KPIs für SAP Systeme. Diese kann als Basis für die Entwick-
lung prozessbezogener und ERP-System-unabhängiger KPIs und Metriken für eine automa-
tisierte Auswertung verwendet werden. Die notwendigen Daten könnten über die Erwei-
terung der Extraktionsalgorithmen ausgewertet werden. Darüber hinaus ergibt sich die
Auswertungsmöglichkeit spezifischer KPIs wie z.B. die durchschnittliche Durchlaufzeit eines
nebenläufigen Prozessflusses.

4.7 Abweichungsanalyse und Prozessoptimierung


Die Erstellung von Prozesskarten erlaubt, Abweichungen von Soll-Prozessen zu identifizie-
ren. Sofern Soll-Prozesse definiert sind, können anhand visueller Auswertungen Abwei-
chungen erkannt werden. Die Analyse von Abweichungen erlaubt, deren Ursachen zu eva-
luieren und ggf. Änderungen zur Optimierung am Prozessablauf vorzunehmen.

5 Restriktionen
In den vorherigen Abschnitten wurden Potential und Anwendungsmöglichkeiten der Me-
thoden zur Rekonstruktion von Prozessen diskutiert. Dieser Abschnitt thematisiert ausge-
wählte Restriktionen zu deren Einsatz und zeigt wesentliche Anwendungsgrenzen auf.

5.1 Extraktionskomponente
Die Rekonstruktion der Prozessinstanzen erfolgt auf Basis der Verknüpfung der Buchungen,
die sich aus der offene-Posten-Buchführung ergibt. Diese Verknüpfung ist fachlicher Natur
und nicht systemspezifisch. Die Umsetzung der Verknüpfung erfolgt jedoch auf technischer

129
APPENDIX A: PUBLICATIONS

Ebene. Um die Rekonstruktion von Prozessinstanzen mit einem Softwareartefakt durchfüh-


ren zu können, ist die Extraktion der relevanten Daten samt derer Verknüpfungen aus den
jeweiligen zugrundeliegenden ERP Systemen notwendig.
Welche Daten für die Extraktion relevant sind, ist für jedes ERP System individuell zu spezi-
fizieren. Aufgrund der Heterogenität der ERP Systeme und deren technischer Umsetzung
ergibt sich die Restriktion, dass es keine generische Softwarekomponente für die Datenex-
traktion geben kann, zumindest solange keine Standards für eine relevante Datenbereit-
stellung etabliert sind. Über die Modularisierung und Entkopplung der Extraktionskompo-
nente und der Rekonstruktionskomponente in einem Softwareartefakt lässt sich die Uni-
versalität der Rekonstruktionskomponente bewahren. Bei einer softwaretechnischen Um-
setzung wird die Erstellung von ERP spezifischen Extraktionskomponenten allerdings nicht
zu vermeiden sein.
Die gleiche Problematik ergibt sich bei der automatisierten Auswertung von Anwendungs-
kontrollen. Auch hier wird die Entwicklung systemspezifischer Extraktionsmodule nicht ver-
mieden werden können, da die Speicherung von Systemeinstellungen zu Anwendungskon-
trollen systemspezifisch erfolgt.

5.2 Systemversionen
Eine ähnliche Problematik ergibt sich aus der Versionierung der ERP Systeme. Bei jeder
neuen Version eines ERP Systems ist prinzipiell zu prüfen, ob sich die zugrundeliegende
Struktur der Speicherung der relevanten Daten geändert hat. Sofern dies der Fall ist, müs-
sen die Extraktionskomponenten entsprechend angepasst werden.

5.3 Rechenkapazitäten
In ERP Systemen werden ggf. Informationen für Millionen von Transaktionen gespeichert.
Der Einsatz von Rekonstruktionsmethoden zielt darauf ab, gerade bei hochintegrierten Ge-
schäftsprozessen und hoher Transaktionsrate eingesetzt zu werden. Insofern ist mit großen
Datenvolumina zu rechnen, die ausgewertet werden müssen. Einerseits erfordert dies den
Einsatz leistungsfähiger Extraktionskomponenten. Andererseits werden die extrahierten
Daten ausgewertet, aggregiert und als Graphen visualisiert. Hierbei ergibt sich die Anfor-
derung, hinreichend leistungsfähige Algorithmen zu implementieren. Zum derzeitigen Ent-
wicklungsstand ist unklar, ob die verwendeten Algorithmen in der Lage sind, bei begrenzten
Rechnerkapazitäten hinreichend performante Berechnungen zu liefern. Des Weiteren ste-
hen Untersuchungen aus, ob beliebig komplexe Prozessinstanzen bei endlicher Rechenka-
pazität ausgewertet und repräsentiert werden können.

5.4 Datenschutzaspekte
Einen kritischen Aspekt stellt der Datenschutz dar. In den vorausgegangenen Abschnitten
wurden umfangreiche Einsatzmöglichkeiten der vorgestellten Methoden zu Auswertungs-
zwecken erläutert. Die Sammlung und Auswertung von Informationen bietet die Möglich-
keit zum Missbrauch. So könnten die erhaltenen Informationen z.B. unberechtigterweise
für Leistungsmessung und Beurteilung der Mitarbeiter genutzt werden. Der Einsatz der Me-
thoden ist auf Einsatzgebiete beschränkt, die im Einklang mit bestehenden Datenschutzbe-
stimmungen stehen.

130
APPENDIX A: PUBLICATIONS

Datenschutzaspekte sollten bereits bei der Konzeptionierung der Softwareartefakte in


Form von Anonymisierungs- oder Pseudonymisierungsverfahren berücksichtigt werden,
um das Potential von Missbrauch bereits in der Entwicklungsphase zu minimieren.

5.5 Sachlogische Verknüpfung


Eine wesentliche Anwendungsgrenze für die in diesem Artikel vorgestellten Methoden zur
Prozessrekonstruktion ergibt sich aus dem prinzipiell größten Vorteil der Methoden, der in
der Systemunabhängigkeit der Konstruktionsmethoden liegt. Dieser ist begründet in der
Nutzung der Verknüpfung von Buchungen durch die offene-Posten-Buchhaltung. Die Ver-
wendung dieser Verknüpfung bedingt allerdings auch, dass lediglich Geschäftsprozesse und
Teilprozesse ausgewertet werden können, die einer offenen-Posten-Buchhaltung folgen.
Für den Einsatz der Methoden für die Jahresabschlussprüfung stellt diese keine gravierende
Restriktion dar, da ein Großteil der rechnungslegungsrelevanten Geschäftsprozesse auf of-
fene-Posten-geführte Konten buchen. Hinsichtlich des Einsatzes der Methoden für Analyse-
oder Optimierungszwecke stellt dies jedoch eine Einsatzgrenze dar, da hier ggf. Kernwert-
schöpfungsprozesse von Bedeutung sind, bei deren Durchführung keine Buchungen auf of-
fene-Posten-geführten Konten stattfinden. Insofern stellt sich die Herausforderung, Ver-
knüpfungsmerkmale für Transaktionen bei Geschäftsprozessen zu identifizieren, für die
keine Buchung auf offene-Posten-geführte Konten stattfindet.

6 Zusammenfassung und Ausblick


In diesem Artikel wurden die Potentiale und Grenzen automatisierter Prozessprüfungen
durch Prozessrekonstruktionen untersucht und diskutiert. Das wesentliche Anwendungs-
szenario ergibt sich für den Einsatz von Methoden zu Prozessrekonstruktion im Rahmen
der Prüfung von Geschäftsprozessen bei Jahresabschlussprüfungen. In Kombination mit
Methoden zur automatisierten Prüfung von Anwendungskontrollen bieten diese die Mög-
lichkeit, die Diskrepanz zwischen einer automatisierten Transaktionsverarbeitung auf Sei-
ten der Unternehmen mit systembasierten und automatisierten Prüfungsmethoden auf
Seiten der Jahresabschlussprüfer zu begegnen. Insbesondere bei hochintegrierten Ge-
schäftsprozessen und hoher Transaktionsrate ist mit einer wesentlichen Effektivitäts- und
Effizienzsteigerung zu rechnen. Der Einsatz der Methoden erlaubt die Freisetzung und Nut-
zung von Ressourcen für die Prüfung ungewöhnlicher Geschäftsprozesse, die von Standard-
prozessen abweichen und ein inhärent höheres Risiko aufweisen. Neben dem Einsatz in der
Jahresabschlussprüfung können Methoden zur Prozessextraktion Anwendung finden für
Auswertungen auf Basis erzeugter Metadaten, die bei manueller Geschäftsprozessabwick-
lung außerhalb eines ERP Systems gar nicht zu Verfügung stehen. Es ergeben sich Einsatz-
felder für die Auswertung zu Funktionstrennungsverletzungen, Leistungsanalysen der Pro-
zesse sowie Abweichungs- und Optimierungsanalysen.
Neben vielfältigen Einsatzmöglichkeiten bestehen Restriktionen zum Einsatz der Methoden
durch die notwendige Systemspezifität der Extraktionskomponenten, die sich aus der He-
terogenität der ERP Systeme und unterschiedlicher Systemversionen ergeben. Ohne ent-
sprechende Testläufe kann derzeit schwer beurteilt werden, ob die entwickelten Algorith-
men leistungsfähig genug sind, um große Datenmengen auswerten und visualisieren zu

131
APPENDIX A: PUBLICATIONS

können. Weitere Einschränkungen zum Einsatz können sich aus Datenschutzaspekten er-
geben. Außerdem beschränken sich die Einsatzmöglichkeiten der Methoden bei jetzigem
Forschungsstand auf Geschäftsprozesse, bei deren Bearbeitung offene-Posten-geführte
Posten Konten involviert sind.
Zusammenfassend ist festzustellen, dass die Methoden zur Prozessrekonstruktion ein gro-
ßes Potential für die Erreichung wesentlicher Effektivitäts- und Effizienzgewinne bei der
Prüfung von in ERP Systeme integrierten Geschäftsprozessen aufweisen und darüber hin-
aus vielfältige weitere Einsatzmöglichkeiten bieten. Grenzen und Restriktionen für deren
Einsatz existieren und müssen durch die Beachtung entsprechender Anforderungen bei der
Umsetzung der Softwareartefakte berücksichtigt bzw. durch weitere Forschungstätigkeiten
bewältigt werden.35

Literaturverzeichnis
[1] IDW PS 261 Feststellung und Beurteilung von Fehlerrisiken und Reaktionen des Ab-
schlussprüfers auf die beurteilten Fehlerrisiken (Quelle: WPg 22/2006, S. 1433 ff.,
FN-IDW 11/2006, S. 710 ff., WPg Supplement 4/2009, S. 1 ff., FN-IDW 11/2009, S.
533 ff.) vom 09.09.2009
[2] ISA 315 Identifying and Assessing the Risks of Material Misstatement through Un-
derstanding in (Quelle: International Federation of Accountants (IFAC-I) (2008)
Handbook of international auditing, assurance, and ethics pronouncements 2008
Edition Part I [Link]
tio/[Link], abgerufen am 25.02.2010
[3] IDW PS 330 Abschlussprüfung bei Einsatz von Informationstechnologie (Quelle: WPg
21/2002, S. 1167 ff., FN-IDW 11/2002, S. 604 ff.) vom 24.09.2002
[4] IDW RS FAIT 1 Grundsätze ordnungsmäßiger Buchführung bei Einsatz von Informati-
onstechnologie (Quelle: WPg 21/2002, S. 1157 ff., FN-IDW 11/2002, S. 649 ff.) vom
24.09.2002
[5] M. Werner, M. Gehrke, N. Nüttgens: “Business Process Mining and Reconstruction
for Financial Audits”, in: Proceedings of the 45th Hawaii International Conference
on System Sciences (HICSS-45), Hawaii (2012 angenommen)
[6] Virtual Accounting Worlds (Quelle: [Link], abgerufen
am 07.12.2011
[7] J.E. Cook, and A.L. Wolf, “Discovering Models of Software Processes from Event-
Based Data,” ACM Trans. Software Eng. and Methodology, vol. 7, no. 3, pp. 215-249,
1998.
[8] J.E. Cook and A.L. Wolf, “Event-Based Detection of Concurrency,” Proc. Sixth Int’l
Symp. the Foundations of Software Eng. (FSE-6), pp. 35-45, 1998.

35
Die in dieser Veröffentlichung präsentierten Ergebnisse wurden im Forschungsprojekt Virtual Accounting
Worlds erarbeitet. Das Projekt wird vom Bundesministerium für Bildung und Forschung gefördert (För-
dernummer 01IS10041). Die Autoren sind verantwortlich für den Inhalt der Veröffentlichung.

132
APPENDIX A: PUBLICATIONS

[9] J.E. Cook, and A.L. Wolf, “Software Process Validation: Quantitatively Measuring the
Correspondence of a Process to a Model,” ACM Trans. Software Eng. and Methodol-
ogy, vol. 8, no. 2, pp. 147-176, 1999.
[10] R. Agrawal, D. Gunopulos, and F. Leymann, “Mining Process Models from Workflow
Logs,” Proc. Sixth Int’l Conf. Extending Database Technology, pp. 469-483, 1998.
[11] J. Herbst, “A Machine Learning Approach to Workflow Management,” Proc. 11th Eu-
ropean Conf. Machine Learning, pp. 183-194, 2000.
[12] J. Herbst, “Dealing with Concurrency in Workflow Induction,” Proc. European Con-
current Eng. Conf., U. Baake, R. Zobel, and M. Al- Akaidi, eds., 2000.
[13] J. Herbst, “Ein induktiver Ansatz zur Akquisition und Adaption von Workflow-Model-
len,” PhD thesis, Universität Ulm, Nov. 2001.
[14] J. Herbst, and D. Karagiannis, “An Inductive Approach to the Acquisition and Adap-
tation of Workflow Models,” Proc. Workshop Intelligent Workflow and Process
Management: The New Frontier for AI in Business, M. Ibrahim and B. Drabble, eds.,
pp. 52-57, Aug. 1999.
[15] J. Herbst, and D. Karagiannis, “Integrating Machine Learning and Workflow Manage-
ment to Support Acquisition and Adaptation of Workflow Models,” Proc. Ninth Int’l
Workshop Database and Expert Systems Applications, pp. 745-752, 1998.
[16] J. Herbst, and D. Karagiannis, “Integrating Machine Learning and Workflow Manage-
ment to Support Acquisition and Adaptation of Workflow Models,” Int’l J. Intelligent
Systems in Accounting, Finance, and Management, vol. 9, pp. 67-92, 2000.
[17] M.K. Maxeiner, K. Küspert, and F. Leymann, “Data Mining von Workflow-Protokol-
len zur teilautomatisierten Konstruktion von Prozessmodellen” Proc. Datenbanksys-
teme in Büro, Technik und Wissenschaft, pp. 75-84, 2001.
[18] G. Schimm, “Generic Linear Business Process Modeling,” Proc. ER 2000 Workshop
Conceptual Approaches for E-Business and The World Wide Web and Conceptual
Modeling, S.W. Liddle, H.C. Mayr, and B. Thalheim, eds., pp. 31-39, 2000.
[19] G. Schimm, “Process Mining Elektronischer Geschäftsprozesse,” Proc. Elektronische
Geschäftsprozesse, 2001.
[20] G. Schimm, “Process M6ning linearer Prozessmodelle - Ein Ansatz zur Automatisier-
ten Akquisition von Prozesswissen,” Proc. 1. Konferenz Professionelles Wissensman-
agement, 2001.
[21] G. Schimm, “Process Miner - A Tool for Mining Process Schemes from Event-Based
Data,” Proc. Eighth European Conf. Artificial Intelligence (JELIA), S. Flesca and G.
Ianni, eds., pp. 525-528, 2002.
[22] W. M. P. van der Aalst, „Business Alignment: Using Process Mining as a Tool for
Delta Analysis and Conformance Testing“, Requirements Engineering Journal, Bd.
10, Nr. 3, S. pp. 198-211, 2005.
[23] W. M. P. van der Aalst, Process Mining: Discovery, Conformance and Enhancement
of Business Processes, 1. Aufl. Springer Berlin Heidelberg, 2011.

133
APPENDIX A: PUBLICATIONS

[24] W. M. P. van der Aalst, A. J. M. M. Weijters, und L. Maruster, „Workflow Mining:


Which Processes can be Rediscovered“, presented at the Proc. Int’l Conf. Eng. and
Deployment of Cooperative Information Systems (EDCIS 2002), 2002, Bd. 2480, S.
45-63.
[25] W.M.P. van der Aalst, “The Application of Petri Nets to Workflow Management,”
The J. Circuits, Systems and Computers, vol. 8, no. I, pp. 21-66, 1998.
[26] W.M.P. van der Aalst, “Verification of Workflow Nets”, Application and Theory of
Petri Nets, P. Azema and G. Balbo, eds., pp. 407-426, Berlin: Springer-Verlag, 1997.
[27] W.M.P. van der Aalst, and B.F. van Dongen, “Discovering Workflow Performance
Models from Timed Logs,” Proc. Int’l Conf. Eng. and Deployment of Cooperative In-
formation Systems (EDCIS 2002), Y. Han, S. Tai, and D. Wikarski, eds., vol. 2480, pp.
45-63, 2002.
[28] W. M. P. van der Aalst, B. F. van Dongen, J. Herbst, L. Maruster, G. Schimm, and
A.J.M.M. Weijters, “Workflow Mining: a Survey of Issues and Approaches”, Beta:
Research School for Operations Management and Logistics, Working Paper 74,
2002.
[29] W.M.P. van der Aalst, A.J.M.M. Weijters, and L. Maruster, “Workflow Mining: Which
Processes can be Rediscovered?” BETA Working Paper Series, WP 74, Eindhoven
Univ. of Technology, Eindhoven, 2002.
[30] L. Maruster, W.M.P. van der Aalst, A.J.M.M. Weijters, A. van den Bosch, and W.
Daelemans, “Automated Discovery of Workflow Models from Hospital Data,” Proc.
13th Belgium-Netherlands Conf. Artificial Intelligence (BNAIC 2001), B. Kro¨se, M.
de Rijke, G. Schreiber, and M. van Someren, eds., pp. 183-190, 2001.
[31] L. Maruster, A.J.M.M. Weijters, W.M.P. van der Aalst, and A. van den Bosch, “Pro-
cess Mining: Discovering Direct Successors in Process Logs,” Proc. Fifth Int’l Conf.
Discovery Science (Discovery Science 2002), pp. 364-373, 2002.
[32] A.J.M.M. Weijters, and W.M.P. van der Aalst, “Process Mining: Discovering Work-
flow Models from Event-Based Data,” Proc. 13th Belgium-Netherlands Conf. Artifi-
cial Intelligence (BNAIC 2001), B. Kro¨ se, M. de Rijke, G. Schreiber, and M. van
Someren, eds., pp. 283-290, 2001.
[33] A.J.M.M. Weijters and W.M.P. van der Aalst, “Workflow Mining: Discovering Work-
flow Models from Event-Based Data,” Proc. ECAI Workshop Knowledge Discovery
and Spatial Data, C. Dousson, F. Höppner, and R. Quiniou, eds., pp. 78-84, 2002.
[34] A.J.M.M. Weijters, and W.M.P. van der Aalst, “Rediscovering Workflow Models
from Event-Based Data,” Proc. 11th Dutch-Belgian Conf. Machine Learning (Bene-
learn 2001), V. Hoste and G. de Pauw, eds., pp. 93-100, 2001.
[35] N. Gehrke, “The ERP AuditLab - A prototypical Framework for Evaluating Enterprise
Resource Planning System Assurance”, in: Proceedings of the 43th Hawaii Interna-
tional Conference on System Sciences (HICSS-43), Hawaii (2010).

134
APPENDIX A: PUBLICATIONS

[36] N. Gehrke, and N. Müller-Wickop, “Basic Principles of Financial Process Mining”,


Proceedings of the 16th Americas Conference on Information Systems, Lima, Peru,
2010.
[37] N. Gehrke, and N. Müller-Wickop, “Rekonstruktion von Geschäftsprozessen im Fi-
nanzwesen mit Financial Process Mining“, in: Lecture Notes in Informatics, Procee-
dings der Jahrestagung Informatik 2010, IT-supported Service Innovation and Ser-
vice Improvement, Leipzig, 2010.
[38] Gehrke, N., Wolf, P.: Towards Audit 2.0 - A Web 2.0 Community Platform for Audi-
tors, in: Proceedings of the 43th Hawaii International Conference on System Sci-
ences (HICSS-43), Hawaii (2010).
[39] Müller-Wickop, N., Schultz, M., Gehrke, N., Nüttgens, M.: Towards Automated Fi-
nancial Process Auditing: Aggregation and Visualization of Process Models, Proceed-
ings of the Enterprise Modelling and Information Systems Architectures (EMISA
2011), Germany, Hamburg, 2011
[40] M. Jans, M. Alles, und M. Vasarhelyi, „Process mining of event logs in auditing: op-
portunities and challenges“, Working paper. Hasselt University. Belgium, 2010.
[41] M. Jans, N. Lybaert, K. Vanhoof, und J. M. Van Der Werf, „Business process mining
for internal fraud risk reduction: Results of a case study“, 2008.
[42] M. Alles, M. Jans, und M. Vasarhelyi, „Process Mining: A New Research Methodol-
ogy for AIS“, in CAAA Annual Conference 2011, 2011.
[43] K. Küting, and M. Reuter, “Bilanzierung im Spannungsfeld unterschiedlicher Adres-
saten“, DSWR, Vol. 9, 2004.
[44] P. Leibfried, and P. Meixner, “Konvergenz der Rechnungslegung: Bestandsaufnahme
und Versuch einer Prognose“, Der Schweizer Treuhänder, Vol. 4, 2006.
[45] A. Hufgard, ROI von SAP-Lösungen verbessern, 1. Aufl. SAP PRESS, 2010.

135
APPENDIX A: PUBLICATIONS

10.4 Einsatzmöglichkeiten von Process Mining für die Analyse von Geschäfts-
prozessen im Rahmen der Jahresabschlussprüfung

Number 4
Einsatzmöglichkeiten von Process Mining für
Title die Analyse von Geschäftsprozessen im Rah-
men der Jahresabschlussprüfung
Appendix 10.4
Primary Related Chapters 4.2, 5.1
Type Book Chapter
Book Forschung für die Wirtschaft
Reference (Werner, 2012)
36)
Acceptance Rate
VHB JQ 2.1 Ranking -
WKWI Ranking -
ERA 2010 -
CORE 2013 -
Review Procedure Blinded
Number of Reviews -
Authors Michael Werner
Dissertation Points 1.00
Authorship
Overall 100%
Design 100%
Realization 100%
Writing 100%
Status Published
Part of other Dissertations No
[Link]
Link
tions/6274-forschung-fur-die-wirtschaft-2012

36
Invited for Publication

136
APPENDIX A: PUBLICATIONS

Einsatzmöglichkeiten von Process Mining für die Analyse von


Geschäftsprozessen im Rahmen der Jahresabschlussprüfung

MICHAEL WERNER
NORDAKADEMIE – Hochschule der Wirtschaft, Elmshorn

Abstract: Die Automatisierung der Durchführung von Geschäftsprozessen


durch den Einsatz von Informationssystemen erlaubt es Unternehmen, ihre
Prozesse effizienter und effektiver zu gestalten. Mit der Zunahme der Automa-
tisierung steigt die Verfügbarkeit von Daten, die im Zuge der Bearbeitung durch
Informationssysteme gespeichert werden. Business Intelligence Tools bieten
Funktionalitäten, um aus den immer stärker anwachsenden Datenbeständen
Informationen effizient zu extrahieren und in geeigneter Weise zur Verfügung
zu stellen. Während der Einsatz von Business Intelligence Software in den ver-
gangenen Jahren mit der gleichzeitigen Verfügbarkeit digital auswertbarer Da-
ten in vielen Industriezweigen zugenommen hat, ist dies für die Prüfungsbran-
che nicht oder nur in einem geringen Ausmaß zu beobachten gewesen. Wirt-
schaftsprüfer werden mit der Prüfung und Testierung der Jahresabschlüsse von
Unternehmen beauftragt. Ein wesentlicher Bestandteil der Jahresabschlussprü-
fung ist die Prüfung der Geschäftsprozesse, die für die Finanzberichterstattung
relevant sind. Für diese Prüfung werden durch den Wirtschaftsprüfer vorwie-
gend manuelle Tätigkeiten durchgeführt. Process Mining Methoden bieten
Möglichkeiten, aus vorhandenen Event Logs Prozessmodelle zu rekonstruieren.
Der Einsatz dieser Methoden eignet sich auch für die Prüfung von Geschäfts-
prozessen, indem Prozessmodelle automatisch rekonstruiert und analysiert
werden können. In diesem Beitrag wird veranschaulicht, wie spezielle Process
Mining Methoden auf die Anwendungsdomäne der Jahresabschlussprüfung an-
gewendet werden können.

1. Einleitung
Der Zweck von Business Intelligence37 besteht darin, umfangreiche Daten zu analysieren
und somit Informationen zugänglich zu machen, die andernfalls nicht zur Verfügung stehen
würden. Die Bedeutung von Business Intelligence hat in den vergangenen Jahren mit dem
Anwachsen auswertbarer, digital vorliegender Datenbestände stark zugenommen und Ein-
gang in Softwarewerkzeuge bedeutender Anbieter gefunden.38 Je stärker die Unterstützung
und Automatisierung der Verarbeitung von Geschäftsvorfällen in Informationssystemen
voranschreitet, desto wichtiger ist es, die verfügbaren Datenstände für Analysen auswer-
ten zu können. Primäres Anwendungsfeld von Business Intelligence Techniken ist die Un-
terstützung von Entscheidungsprozessen [5], [6]. Sie eignen sich allerdings auch zur Analyse
von Geschäftsprozessen für die Zwecke der Jahresabschlussprüfung.

37
Business Intelligence umfasst „Anwendungen und Techniken, die sich darauf konzentrieren, Daten aus
verschiedenen Quellen zu sammeln, zu speichern, zu analysieren und den Zugriff auf sie zu ermöglichen,
um den Benutzern zu helfen, bessere Entscheidungen zu treffen.“[1] S. 503.
38
Gängige Softwarelösungen werden zum Beispiel von IBM [2], Oracle [3] und SAP [4] angeboten.

137
APPENDIX A: PUBLICATIONS

Die Jahresabschlussprüfung ist ein Kontrollinstrument, das dazu dient, die Adressaten des
Jahresabschlusses vor Falschinformationen zu schützen. Sie stellt ein wichtiges Regulativ
im Rahmen der Wirtschaftsordnung westlicher Volkswirtschaften dar. Aufgrund ihrer Be-
deutung werden an die Durchführung der Prüfung und die Qualifikation der Prüfer hohe
Anforderungen gestellt, die durch den Gesetzgeber kodifiziert sind.39 Des Weiteren werden
die Anforderungen an die Durchführung der Prüfung in Prüfungsstandards konkretisiert.40
Eine bedeutende Tätigkeit während der Abschlussprüfung ist die Prüfung von Geschäfts-
prozessen. Den Buchungseinträgen auf den Konten der Bilanz und Gewinn- und Verlust-
rechnung liegen Geschäftsvorfälle zu Grunde, die im Unternehmen abgewickelt wurden.
Die Prüfung der Geschäftsprozesse soll sicherstellen, dass nur vollständige und richtige Ein-
träge auf den entsprechenden Buchhaltungskonten gebucht werden, die auch auf tatsäch-
lich stattgefundenen Geschäftsvorfällen basieren. Die Prüfung von Geschäftsprozessen
wird vorgenommen aufgrund der Annahme, dass wohlkontrollierte Geschäftsprozesse zur
vollständigen und richtigen buchhalterischen Erfassung der bearbeiteten Geschäftsvorfälle
führen. Von besonderer Bedeutung für Prüfung von Geschäftsprozessen sind die Internati-
onal Standards on Auditing 315 [9] und 330 [10] sowie auf nationaler Ebene der IDW Prü-
fungsstandard 261 [11]. In diesen wird festgelegt, wie die Prüfung von Geschäftsprozessen,
internen Kontrollen und relevanten Informationssystemen zu berücksichtigen ist.
Bei der Durchführung der Prüfung von Geschäftsprozessen stehen die Wirtschaftsprüfer
vor der Herausforderung, die Zuverlässigkeit der Kontrollmechanismen zu den betrachte-
ten Geschäftsprozessen zu würdigen. Dazu ist es notwendig, ein der Realität entsprechen-
des Verständnis der zugrundeliegenden Geschäftsprozesse sowie der involvierten internen
Kontrollen zu erhalten und letztere auf deren Angemessenheit und operative Funktionsfä-
higkeit zu prüfen. Dies erfolgt zumeist anhand manueller Prüfungstätigkeiten. Diese umfas-
sen die Durchführung von Interviews oder die Inspektion vorhandener Dokumente auf
Stichprobenbasis. Solche Prüfungshandlungen sind im Umfeld stark integrierter und auto-
matisierter Verarbeitung durch Informationssysteme kritisch zu hinterfragen. Werner und
Gehrke [12] zeigen, dass mit steigender Automatisierung der Geschäftsprozessabwicklung
und der Anzahl der zu verarbeiteten Geschäftsvorfälle manuelle Prüfungstätigkeiten inef-
fizient und ggf. ineffektiv werden.
Eine Lösung für dieses Problem stellt der Einsatz von Business Intelligence Methoden für
die Zwecke der Jahresabschlussprüfung dar. Ganzheitliche konzeptionelle Überlegungen
für den Einsatz von Business Intelligence Methoden zur Unterstützung und Teil-automati-
sierung der Geschäftsprozessprüfung finden sich in [13] und [14]. In diesem Beitrag wird
auf einen konkreten Teilaspekt des konzeptionellen Entwurfs von Werner et al. [13] einge-
gangen, indem illustriert wird, wie Process Mining Methoden als Analysemethoden für die
Aufbauprüfung von Geschäftsprozessen eingesetzt werden können.

39
§319 HGB regelt die Zuständigkeit der Prüfung des Jahresabschlusses und legt fest, dass diese ausschließ-
lich durch Wirtschaftsprüfer, Wirtschaftsprüfungsgesellschaften, vereidigte Buchprüfer oder Buchprü-
fungsgesellschaften durchgeführt werden darf.
40
Auf nationaler Ebene übernimmt das Institut der Wirtschaftsprüfer in Deutschland e.V. (IDW) [7] als Ver-
einigung von Wirtschaftsprüfern und Wirtschaftsprüfungsgesellschaften die fachliche Entwicklung der
Regeln zu deren Berufsausübung und veröffentlicht in diesem Zusammenhang anzuwendende Prüfungs-
standards. Auf internationaler Ebene übernimmt diese Aufgabe das International Auditing and Assurance
Standards Board (IAASB)[8] mit der Veröffentlichung der International Standards on Auditing (ISA).

138
APPENDIX A: PUBLICATIONS

Ziel der Aufbauprüfung als Bestandteil der Geschäftsprozessprüfung ist die Informations-
gewinnung über die für die Finanzberichterstattung relevanten Prozesse und Kontrollen.
Process Mining Methoden können eingesetzt werden, um durchgeführte Geschäfts-vor-
fälle als Modelle von Prozessinstanzen zu visualisieren und somit Informationen über die
Ausgestaltung der beteiligten Prozesse zu erhalten. Bei der Anwendung dieser Methoden
ist die Berücksichtigung von Anforderungen von großer Bedeutung, die sich aus der An-
wendungsdomäne ergeben.
Das Vorgehen sowie die Ergebnisse, die sich durch die Anwendung ausgewählter Metho-
den erzielen lassen, werden in diesem Beitrag erläutert. Die hier vorgestellten Methoden
können dabei zum derzeitigen Forschungstand nur als Meilenstein betrachtet werden auf
dem Weg zur Entwicklung ganzheitlicher automatisierter Analyse- und Prüfungsmethoden
für die Geschäftsprozessprüfung. Insofern liegt ein weiterer Schwerpunkt auf der Betrach-
tung, welche Hinweise sich für weitere Entwicklungsmöglichkeiten aus den dargestellten
Ergebnissen ableiten lassen.
In den folgenden Abschnitten zwei und drei werden der derzeitige Stand der Wissenschaft
zu den relevanten Themengebieten sowie der gewählte Forschungsansatz und Methoden
vorgestellt, die zur Erlangung der dargestellten Erkenntnisse herangezogen wurden. Um
die Bedeutung des Einsatzes der Process Mining Techniken für die Zwecke der Aufbauprü-
fung von Geschäftsprozessen zu veranschaulichen, bietet Abschnitt vier einen Überblick
über die Bedeutung der Analyse von Geschäftsprozessen für die Jahres-abschlussprüfung.
Für die Modellierung von Modellen ist es notwendig, eine angemessene Modellierungs-
sprache und Repräsentationsform zu wählen. Die hier dargestellten Modelle werden als
Petri-Netze modelliert. Hintergrund für diese Wahl sowie formale Aspekte werden in Ab-
schnitt fünf erörtert. In Abschnitt sechs werden schließlich bisher erzielte Auswertungser-
gebnisse dargestellt. Der Beitrag endet mit einer Zusammenfassung und Diskussion der
dargestellten Ergebnisse.

2. Stand der Wissenschaft


Ausgangspunkt der in diesem Beitrag dargelegten wissenschaftlichen Untersuchung sind
die Forschungsarbeiten aus dem Gebiet Process Mining. Erste Ansätze hierzu entstanden
zum Ende der 1990er Jahre durch Cook und Wolf [15–17] ursprünglich im Bereich des Soft-
ware Engineering. Seit Beginn der 2000er Jahre hat sich die Forschung auf diesem Gebiet
deutlich intensiviert, was zu einer Vielzahl von Veröffentlichungen geführt hat. Einen recht
aktuellen Überblick hierzu bietet Tiwari [18]. Ein bedeutender Vertreter dieser Disziplin ist
van der Aalst. In [19] bietet er einen umfassenden Überblick über Methoden und wesent-
liche Grundkenntnisse zum Process Mining.
In den letzten Jahren ist zu beobachten, dass sich wissenschaftliche Arbeiten vermehrt mit
einer Ausweitung des Anwendungskontextes befassen. Dies gilt auch für das An-wendungs-
gebiet von Process Mining Techniken für Prüfungszwecke. Hierbei ist zu unterscheiden zwi-
schen Arbeiten, die sich mit Conformance und solchen, die sich mit Compliance beschäfti-
gen. Erstgenannte Untersuchungen haben zum Ziel, Abweichungen von rekonstruierten
Geschäftsvorfällen anhand eines Soll-Modells zu identifizieren. Hierzu werden verfügbare
Event Logs gegen Sollmodelle verglichen und Abweichungen identifiziert. Dieser Ansatz

139
APPENDIX A: PUBLICATIONS

lässt sich unter anderem finden in [20]. Er ist interessant, für die in diesem Beitrag darge-
stellten Vorgehensweisen aber weitestgehend irrelevant. Die Anwendung von Confor-
mance Checking Methoden setzt das Vorhandensein eines Sollmodells voraus. Beim Einsatz
der in diesem Beitrag dargestellten Methoden geht es aber gerade darum, Prozessmodelle
anhand verfügbarer Daten erst zu entdecken und nicht rekonstruierte Modelle gegen ein
bereits existierendes Sollmodell zu vergleichen.
Anders verhält es sich mit Forschungsarbeiten, die sich mit Compliance-Aspekten beschäf-
tigen. Unter Compliance wird allgemein die Einhaltung von internen oder externen Vorga-
ben verstanden. Bei der Jahresabschlussprüfung geht es darum, festzustellen, ob externen
Vorgaben in Form von Gesetzen und Normen im Zuge der Rechnungslegung entsprochen
wurde. Insofern stellt die Prüfung der Einhaltung dieser Vorgaben einen Teilbereich der
Gesamtthematik dar, die unter dem Begriff Compliance zusammengefasst wird. Interes-
sante Ansätze für die Überprüfung der Einhaltung von Compliance-Regeln werden in [21–
24] gezeigt. Ihnen ist allerdings gemein, dass sie nicht speziell auf das Anwendungsfeld der
Jahresabschlussprüfung ausgerichtet sind. In diesem Anwendungsfeld ist insbesondere die
Berücksichtigung der Werteflüsse innerhalb der Prozesse notwendig. Diese werden bei den
zuvor genannten Arbeiten nicht explizit berücksichtigt.
Speziell für den Anwendungsbereich entwickelte Methoden präsentieren hingegen Gehrke
und Müller-Wickop [25]. Sie stellen einen Mining Algorithmus vor, der in der Lage ist, Mo-
delle von Prozessinstanzen zu rekonstruieren und dabei für die Prüfung relevante Informa-
tionen zu modellieren. Auf diese Methoden wird im Folgenden zurückgegriffen.

3. Methodik
Die Forschungstätigkeiten, die dieser Arbeit zugrunde liegen, folgen einem Design Science
Ansatz [26–28]. Bei dieser gestaltungsorientierten Herangehensweise liegt der Fokus auf
der Entwicklung von Artefakten. Der Erkenntnisgewinn ergibt sich als Ergebnis des Gestal-
tungsprozesses und der Evaluation der geschaffenen Artefakte. Primäre Artefakte bilden in
dieser Arbeit die Methoden, die zur Extraktion relevanter Daten sowie zur Rekonstruktion
von Modellen und deren Visualisierung entwickelt wurden.
Im Gegensatz zu behavioristischen Ansätzen wird gestaltungsorientierten Vorgehens-wei-
sen häufig eine fehlende Rigorosität vorgeworfen. Riege et al. [29] weisen in diesem Zu-
sammenhang auf die Notwendigkeit der Evaluation der Erkenntnisse sowohl gegen die
identifizierte Forschungslücke wie auch gegen die Realität hin, um eine ausreichende Aus-
sagekraft der Erkenntnisse nachweisen zu können. Um diesen Überlegungen Rechnung zu
tragen, wurden für die vorliegende Forschungsarbeit verschiedene Evaluationsmethoden
verwendet.
Abbildung 1 stellt die verschiedenen Phasen des Forschungsprozesses in Anlehnung an [28]
dar. Für die Analysephase wurden Experteninterviews durchgeführt, um generelle Anfor-
derungen für den Einsatz von Process Mining Methoden zu identifizieren. Primäre Quelle
für die Identifikation von Anforderungen bildeten die Recherche relevanter Literatur und
insbesondere die Sichtung von Standards und Normen zur Rechnungslegung und Jahresab-
schlussprüfung.

140
APPENDIX A: PUBLICATIONS

- Experteninterviews - Test anhand von Test-


- Literaturrecherche und Echtdaten
- Datenanalyse - Simulation

Analyse Entwurf Evaluation Diffusion


- Method - Presentation and
Engineering Publikation
- Prototypkonstruktion

Abbildung 1 Phasen des Forschungsprozesses


Für die technische Ebene der Entwicklung wurden vorhandene Datenbestände ausgewer-
tet. Die Entwicklung der Methoden erfolgte in Form eines Method Engineering [30], indem
bereits bestehende Methoden kombiniert wurden, wenn hiermit das gewünschte Ergebnis
erzielt werden konnte. Wenn notwendig wurden zudem neuartige Methoden entwickelt,
die prototypisch in einem Softwareartefakt implementiert wurden. Zur Evaluation wurden
diese anhand von Test- und Echtdaten getestet. Die Prozessmodelle, die mit der Software
erstellt werden können, wurden wiederum durch die Simulation der Erreichbarkeit rele-
vanter Schaltzustände auf ihre Korrektheit hin überprüft.
Bei dieser Vorgehensweise wurde der Forschungsprozess nicht einmalig durchlaufen son-
dern in mehrfachen Iterationen, in denen immer wieder Analyse, Entwurf, Evaluation und
Diffusion, letzteres in Form von Publikationen und Vorträgen, aufeinander aufbauten.

4. Geschäftsprozessprüfung in der Jahresabschlussprüfung


Um die Einsatzmöglichkeiten von Process Mining Methoden für Zwecke der Jahresab-
schlussprüfung bewerten zu können, ist es notwendig, die Bedeutung der Geschäftspro-
zessprüfung innerhalb der Jahresabschlussprüfung zu beleuchten.
Anhand der Abbildung 2 lässt sich veranschaulichen, warum das Prozessverständnis für die
Prüfung wichtig ist. Es stellt die typischen Phasen einer Jahresabschlussprüfung dar, die
nach der Auftragsannahme im Allgemeinen mit der Erhebung relevanter Informationen
und einer Beurteilung wesentlicher Risiken beginnt. Ziel der Prüfung ist es insgesamt, we-
sentliche Risiken für Fehler in der Finanzberichterstattung zu identifizieren und diese durch
gezielte Prüfungshandlungen zu adressieren. Ein wesentliches Risiko besteht in der Regel
darin, dass getätigte Buchungen nicht vollständig und richtig sind. Um dieses Risiko zu ad-
ressieren, ist es sinnvoll, die Prozesse zu prüfen, die zu den jeweiligen Buchungen geführt
haben. Hierbei wird ausgenutzt, dass Kontrollaktivitäten41 in den Prozessabläufen die kor-
rekte Verarbeitung und somit auch Buchung sicherstellen. Wenn diese Kontrollen einwand-
frei funktionieren, kann davon ausgegangen werden, dass auch die zugehörigen Buchungen
vollständig und fehlerfrei sind.

41
Ein häufig angeführtes Beispiel für eine solche Kontrolle ist der sogenannte Three-Way-Match. Bevor eine
Zahlung veranlasst werden kann, wird hierbei überprüft, ob Menge und Wert bei Bestellung, Warenliefe-
rung und eingegangener Rechnung übereinstimmen. Wenn dies nicht der Fall ist, wird die Zahlung ver-
hindert.

141
APPENDIX A: PUBLICATIONS

Aufbauprüfung
Validierung von Aussagebzogen
Informationssammlung der Prozesse Bericht-
internen e Prüfungs-
und Risikobeurteilung und internen erstattung
Kontrollen handlungen
Kontrollen

Abbildung 2 Prozessphasen der Jahresabschlussprüfung


Um auf die Prüfung der internen Kontrollen zurückgreifen zu können, muss der Prüfer zu-
nächst wissen, wie die Prozesse aussehen, um beurteilen zu können, welche Kontrollen
denn für eine Prüfung in Frage kommen. Dieser Zusammenhang wird aus Abbildung 2 deut-
lich. In der sogenannten Aufbauprüfung verschafft sich der Prüfer einen Überblick über die
relevanten Prozesse und internen Kontrollen. Erst wenn dies erfolgt ist, kann die Überprü-
fung der identifizierten Kontrollen erfolgen, wenn feststeht, welche Kontrollen überhaupt
relevant sind.
In Abbildung 3 ist das einfache Modell eines Bestellprozesses dargestellt. Die Rechtecke
symbolisieren Aktivitäten, die in einem Informationssystem ausgeführt werden. Wenn dies
erfolgt, erzeugen sie Buchungen auf bestimmten Konten der Finanzbuchhaltung. Diese sind
rechts und links der Aktivitäten dargestellt. Gemäß den Regeln der doppelten Buchführung
erfolgt für jede Buchung eine Soll- und eine Habenbuchung, so dass der Saldo der Buchung
null ist. Wenn in einem Unternehmen Waren bestellt werden, hat dies zunächst keine Bu-
chung auf Konten der Bilanz oder der Gewinn- und Verlustrechnung (GuV) zur Folge. Dies
geschieht erst mit dem Eingang der bestellten Ware. Dieser Vorgang führt zu einer Erhö-
hung der Rohstoffbestände und zu einer Gegenbuchung auf dem Wareneingangs-/Rech-
nungseingangskonto (WeRe-Konto). Erst wenn die Rechnung des Zulieferers dem Unter-
nehmen zugeht, entsteht eine Verbindlichkeit, die auf dem Kreditorenkonto für den ent-
sprechenden Lieferanten gebucht wird.
Die Gegenbuchung findet wiederum auf dem WeRe-Konto statt und gleicht die zugehörige
Haben-Buchung aus dem Wareneingang aus. Als nächste Aktivität findet die Zahlung statt,
mit der die Verbindlichkeit an den Kreditor beglichen und in diesem Fall das Bankkonto
belastet wird.

Warenbestellung

Rohstoffe Wareneingang / Rechnungseingang

Soll Haben Soll Haben


Wareneingang
78.419,65 78.419,65

Wareneingang / Rechnungseingang Kreditorenkonto

Soll Haben Soll Haben


Rechnungseingang
78.419,65 78.419,65

Kreditorenkonto Bankkonto

Soll Haben Soll Haben


Zahlung
78.419,65 78.419,65

Abbildung 3 Einfaches Prozessmodell

142
APPENDIX A: PUBLICATIONS

Bei dem dargestellten Beispiel handelt es sich um einen sehr einfachen Standardprozess.
Es illustriert aber, worauf es bei der Prüfung von Geschäftsprozessen in der Jahresab-
schlussprüfung ankommt. Für den Prüfer ist es wichtig zu verstehen, wie die Prozesse des
Unternehmens mit den Buchungen auf den Finanzbuchhaltungskonten zusammen-hängen.
In diesem Beispiel ist der Zusammenhang klar ersichtlich. Es ist zudem nach-vollziehbar,
welche Werte durch den hier dargestellten Geschäftsvorfall auf die jeweiligen Konten ge-
flossen sind.
Es stellt sich nun die Frage, wie der Prüfer Informationen über die Prozessabläufe und die
Zusammenhänge mit den Buchhaltungskonten für die Aufbauprüfung erhält und welche
Vorteile mit dem Einsatz von Process Mining Methoden in dieser Prüfungsphase erzielt
werden können. Die herkömmliche Vorgehensweise besteht in der Durchführung von In-
terviews mit Ansprechpartnern beim zu prüfenden Mandanten. Auf Basis der so erhaltenen
Informationen werden meist einfache Prozessdarstellungen in Form von Flussdiagrammen
wie in Abbildung 3 erstellt, die gegebenenfalls durch textuelle Beschreibungen erweitert
werden. Bestenfalls werden diese durch Prozessschreibungen, die beim Mandanten ver-
fügbar sind, ergänzt.
Werner und Gehrke [12], [13] erläutern, dass dieses Vorgehen ineffizient und eventuell in-
effektiv wird, wenn die Integration von Geschäftsprozessen und Informationssystemen zu-
nimmt. Mit der Zunahme der Integration erhöht sich in der Regel die Komplexität der Pro-
zesse. Des Weiteren können Aktivitäten vollständig automatisiert ablaufen, die für die ope-
rativ beteiligten Personen nicht mehr sichtbar sind, so dass sie auch keine verlässliche Aus-
kunft über den tatsächlichen Prozessablauf geben können. Bestenfalls kann mit dieser ma-
nuellen Erhebungsweise ein Sollmodell erstellt werden. Ob dieses mit der Realität überein-
stimmt, lässt sich nicht beurteilen. Erfahrungen aus der Praxis zeigen, dass dies in der Regel
nicht der Fall ist.
Vor diesem Hintergrund können Process Mining Techniken eingesetzt werden, um au-to-
matisiert Kenntnisse über den Zusammenhang von Prozessen und Buchungen auf Finanz-
buchhaltungskonten zu erhalten. Der Einsatz dieser Techniken erlaubt es, anhand der tat-
sächlich stattgefundenen Geschäftsvorfälle und deren zugehörigen Daten Prozessmodelle
zu erstellen. Damit besteht keine Abhängigkeit mehr zwischen der Verlässlichkeit der er-
halten Informationen und der Darstellung der zu prüfenden Informationen. Des Weiteren
erlaubt die automatisierte Erstellung eine wesentlich detaillierte Modellierung und bietet
die Möglichkeit für umfassende Auswertungen.

5. Modellierung und Petri-Netz Darstellung


Die im vorherigen Abschnitt genannten Einsatzmöglichkeiten sollen nun an einem Bei-spiel
illustriert werden. Für die Erstellung von Modellen spielt die Zweckmäßigkeit der gewähl-
ten Modellierungssprache eine wesentliche Rolle. Für die in dieser Arbeit dargestellten Pro-
zessmodelle wurde die Modellierung in Form von gefärbten Petri-Netzen gewählt. Petri-

143
APPENDIX A: PUBLICATIONS

Netze stellen eine formal robuste und mathematisch ausdrucksstarke Modellierungsspra-


che dar [31], [32]. Zudem kann durch die Modellierung von Petri-Netzen auf bereits um-
fangreiche Arbeiten zurückgegriffen werden.42
Abbildung 4 zeigt das Modell einer Prozessinstanz. Es basiert auf Daten, die aus einem SAP
Testsystem extrahiert und für die Rekonstruktion verwendet wurden. Für die Extraktion
und Rekonstruktion ist die Verwendung eines geeigneten Mining-Algorithmus notwendig.
Es gibt unterschiedliche Verfahren wie etwa den Alpha-Algorithmus, heuristische Verfah-
ren und genetische Algorithmen. Grundsätzlich sind diese Algorithmen geeignet, um Pro-
zessmodelle in verschiedenen Anwendungsdomänen zu rekonstruieren.
Bei der Rekonstruktion von Prozessmodellen für den Zweck der Jahresabschlussprüfung
kommt es jedoch darauf an, die Wertflüsse dazustellen und das Buchungsverhalten nach-
vollziehen zu können. Dies muss durch den einzusetzenden Mining-Algorithmus berück-
sichtigt werden. Zudem operieren die zuvor genannten Algorithmen auf Event Logs, die
unterschiedliche Geschäftsvorfälle (Cases) enthalten. Die rekonstruierten Prozessmodelle
bilden somit eine Abstraktion der im Event Log abgebildeten Prozessinstanzen. Hierbei
kann es zu Unter- wie auch Überdeckungen in der Prozessdarstellung kommen. D.h. es wer-
den unter Umständen Prozessabläufe modelliert, die gar nicht stattgefunden haben, oder
es werden ggf. existente Prozessabläufe nicht abgebildet. Aus Sicht der Jahresabschluss-
prüfung sind lediglich solche Informationen wünschenswert, die den tatsächlichen Verlauf
exakt widergeben. Der von Gehrke und Müller-Wickop [25], [33] entwickelte Algorithmus
erfüllt die zuvor genannten Anforderungen und wurde aus diesem Grund für die vorlie-
gende Arbeit verwendet und erweitert.
Der ursprüngliche Algorithmus stellt primär die Funktionalität dar, die Einträge im Event
Log einzelnen Prozessinstanzen (oder auf Log Ebene Cases) zuzuordnen und die Prozessin-
stanzen in einer mehr oder weniger informellen Darstellungsweise zu visualisieren. Für ers-
teres wird die Struktur von offene-Posten-geführten Buchungen verwendet [13]. Der Mi-
ning-Algorithmus wurde dahingehend erweitert, dass er nun in der Lage ist, die rekonstru-
ierten Prozessmodelle als Petri-Netze und somit in einer formalen Darstellungsform abzu-
bilden.

42
Tiwari et al. [18] zeigen, dass die Mehrheit der von ihnen untersuchten Artikel zum Thema Process Mining
Petri-Netze als Modellierungssprache einsetzten.

144
APPENDIX A: PUBLICATIONS

[5000004938] [100010291] [5100004733]

[5000004938] [100010291] [5100004733]

5000004938 191100 100010291 191100 5100004733 160000 160000 154000

MB01 FB1S MR1M


[562.42] [562.42] [562.42] [562.42] [646.78]

User 1 User 2 User 1 [646.78]


[2000000410]
[646.78] [2.53] 113101
[562.42] [0.00] [84.36]
790000 230051 154000 [2000000410]

2000000410
[5586.91]

F110
[16.87]
276000
User 2
[1900003421]
[5112.92]
[1900003421]
[5112.92] [153.39]
1900003421 160000 160000 276000

FB01
[5112.92]

User 3

[5112.92]
476100

Abbildung 4 Beispiel einer rekonstruierten Prozessinstanz


Ein Petri-Netz besteht grundsätzlich aus Transitionen und Stellen, die über Kanten mit-ei-
nander verbunden sind. Dies wird aus Abbildung 4 deutlich. Die Transitionen (Recht-ecke)
repräsentieren die Transaktionen, die im zugrundeliegenden ERP System ausgeführt wur-
den. Wenn Transaktionen in einem ERP System ausgeführt werden, erzeugen diese Bu-
chungseinträge. Diese Buchungseinträge werden im Petri-Netz als gefärbte Plätze (Kreise)
dargestellt. Die Färbung der Plätze ergibt sich aus der Kontonummer des Kontos, auf dem
die Buchung stattgefunden hat sowie der Information, ob es sich um eine Soll- oder Haben-
buchung gehandelt hat. Eine gestrichelte Kante mit Pfeil zeigt an, dass eine Transaktion die
entsprechende Buchung erzeugt hat. Gestrichelte Kanten ohne Pfeil zeigen, dass eine Bu-
chung durch eine weitere Transaktion ausgeglichen wurde. Bei diesen Kanten handelt es
sich um Testkanten. Bei einer normalen Kante von einer Transition zu einer Stelle wird beim
Schalten der Transition eine Markierung auf der verbundenen Stelle erzeugt. Die Beschrif-
tung der jeweiligen Kante legt dabei fest, welchen Wert (bzw. Farbe) diese Markierung er-
hält. Dabei muss die Farbe der Markierung mit der Farbe der Transition Übereinstimmen,
damit die Transition schalten kann. Bei einer Testkante wird hingegen überprüft, ob eine
Markierung mit einer der Beschriftung der Testkante entsprechendem Wert vorliegt. Diese
wird beim Schalten der zugehörigen Transition allerdings nicht konsumiert. Dieses Verhal-
ten spiegelt den buchhalterischen Vorgang der Auszifferung von offene-Posten-geführten
Buchungen wider. Hierbei wird der Posten zwar ausgeglichen, die ursprüngliche Buchung
bleibt aber erhalten und wird nicht etwa wieder gelöscht. Des Weiteren sind in dem Modell
zusätzliche Plätze enthalten. Diese sind mit Startmarkierungen belegt, die der Buchungs-
nummer der zugehörigen Transaktion entsprechen, und sorgen dafür, dass das rekonstru-
ierte Modell ein schaltfähiges Petri-Netz darstellt.
Das hier illustrierte Beispiel stellt somit das Modell einer Prozessinstanz in Form eines
schaltfähigen Netzes dar, das die Durchführung der tatsächlich stattgefundenen Transakti-

145
APPENDIX A: PUBLICATIONS

onen und Buchungen imitiert. MB01 ist eine SAP Transaktion, mit der Wareneingänge er-
fasst werden. Mit MR1M werden Eingangsrechnungen gebucht. Über F110 wird der Zahl-
lauf durchgeführt und mit FB1S werden offene Posten ausgeglichen. Mit der Transaktion
FB01 können allgemeine Buchungen vorgenommen werden. Das Modell zeigt somit eine
Instanz eines Einkaufsprozesses in dessen Zuge zuerst eine Ware eingegangen ist, dann die
Eingangsrechnung verarbeitet und hiermit der offene Posten durch den Wareneingang aus-
geglichen wurde. Nach Eingang der Rechnung wurde diese über einen Zahllauf beglichen.
Des Weiteren wurde mit dem Zahllauf eine weitere Verbindlichkeit ausgeglichen, deren
Gegenbuchung auf ein Aufwandskonto für EDV-Material gebucht wurde. Die Bedeutung
der einzelnen Transaktionscodes und die Kontenbezeichnungen sind in Tabelle 1 und 2 auf-
geführt.
Tabelle 1 Transaktionscodes
Transaktion Bedeutung
MB01 Wareneingang zur Bestellung buchen
FB1S Ausgleichen Sachkonto
MR1M Eingangsrechnung erfassen
F110 Parameter für maschinelle Zahlung
FB01 Beleg buchen

Tabelle 2 Kontenbezeichnungen
Kontonummer Kontobezeichnung
790000 Unfertige Erzeugnisse
191100 WE/RE-Verrechnung -Eigenfertigung-
230051 Erfolg Euro Umstellung / Beleg Differenz
154000 Eingangssteuer
160000 Kreditoren-Verbindlichkeiten Inland
113101 Deutsche Bank (Ausgangs-Schecks)
276000 Skonto-Ertrag
476100 EDV-Material

Zu Evaluationszwecken wurden rekonstruierte Modelle mit Hilfe der Software Renew [34]
untersucht. Mit dieser Software ist es möglich, die Ausführung von Petri-Netzen zu simu-
lieren. Die Simulation hat gezeigt, dass durch den verwendeten Algorithmus vollständig er-
reichbare und damit korrekt modellierte Netze erzeugt werden.

6. Bisherige Auswertungsergebnisse
Im vorausgegangenen Abschnitt wurden das Vorgehen zur Erzeugung und die verwendete
Repräsentationsform anhand eines einfachen Beispiels für eine Prozessinstanz eines Ein-
kaufsprozesses erläutert. Mittels der vorgestellten Methodik ist es möglich, Prozessinstan-
zen zu rekonstruieren. Für die Aufbauprüfung im Zuge der Jahresabschlussprüfung können
diese Verfahren eingesetzt werden, um den tatsächlichen Ablauf von Geschäftsvorfällen zu

146
APPENDIX A: PUBLICATIONS

visualisieren. Der Prüfer wird damit in die Lage versetzt, auf Basis der in den zugrundelie-
genden Systemen gespeicherten Daten Informationen über den Zusammenhang zwischen
der Ausführung von Prozessen und den Buchhaltungskonten zu erhalten, ohne auf manu-
elle Informationsgewinnung zurückgreifen zu müssen. Dies ist ohne den Einsatz der vorge-
stellten Process Mining Methoden nicht möglich.
Abbildung 5 zeigt ebenfalls eine Prozessinstanz für einen Einkaufsprozess. Hierbei ist zu
erkennen, dass das dargestellte Modell bereits erheblich komplexer ist. Im vorherigen Bei-
spiel beinhaltete die Instanz fünf Transaktionen. Die Prozessinstanz in Abbildung 5 hinge-
gen enthält 46 Transaktionen. Dabei ist zu erkennen, dass die Teilprozesse für Warenein-
gang, Rechnungseingang und Ausgleich der offenen Posten (MB01-FB1S-MR1M) mit den
Teilprozessen in Abbildung 4 vergleichbar sind. Beim ersten Beispiel wird durch den Zahl-
lauf ein Einkaufsteilprozess beendet, im zweiten Beispiel acht.
310000

310000
310000

[92855.72]
[15323.42]
[21735.02]

310000
5000005245 [5000005245]
[5000005245]

MB01
[8876.03]

User 2
[15323.42]
[8876.03] 191100 310000
191100 [5000005128]
[92855.72] [21735.02]

191100 310000
191100
[100010739] [5000005128]
[24337.49]
[100010740]
[8876.03]
[100010739] [100010486] [22087.81]
[21735.02] [100010740] [15323.42] 191100 5000005128
310000 [100010742] [100010741]
100010739 [92855.72] 100010740 MB01
310000 [100010742] [100010741] [100010486] [24337.49]
[100010829] 310000
FB1S FB1S [17097.60]
[24337.49] User 4
100010742 100010741
191100
User1 100010486
User1
[5000005280] [100010829] FB1S [22087.81]
[15890.95] FB1S
FB1S 191100 [17097.60]
[22292.33] [100010485]
[5000005280] User1
[22292.33] User1 [21735.02]
[22292.33] 100010829 191100
310000 [8876.03] User1
5000005280
191100 [100010485]
[100010828] [92855.72]
FB1S [22087.81]
MB01 191100 [15323.42]
[9162.35] 191100 191100
[100010828] [24337.49] 100010485
[9162.35] User1 191100
User 2
100010828 [17097.60]
191100 FB1S
310000 [20472.13]
[9162.35] FB1S [22292.33] [21735.02]
[15890.95]
[8876.03] [92855.72] User1
[20472.13] 100010484
191100 191100 [15323.42] [100010484]
[100010830] User1
[22087.81]
191100 5100004917 191100 FB1S [100010484]
[9162.35]
191100 154000
[100010830]
[15890.95] MR1M [5100004866] [24337.49] User1
[5100004917]
[22206.43] [5100004917] 191100 [17097.60]
100010830
User 2 [5100004866] [22087.81]
[22292.33]
[20472.13] FB1S
[9162.35] [0.01] 5100004866 [100010743]
191100 230051
[15890.95] [17097.60]
User1 230051 [100010743] 310000
100010831 MR1M
5100004961 154000 [160996.61] 154000
[15890.95]
[100010831] FB1S [0.01] [10163.66] 100010743
191100 User 3 191100
[100010831] MR1M
[10850.84]
FB1S 310000
User1 160000
[20472.13] [20472.13] User 2 [27200.73]
[73686.57] [27200.73]
User1
[5100004961]
160000 230051 191100 [27200.73] [27200.73]
[0.00] [78668.60] [27814.28]
[5100004961]
160000 154000 5000005246
230051 191100
[100010832]
154000 MB01
[100010832] [160996.61] 310000
100010832 [2000000539] 100010744 [25360.08]
154000 [0.00] [27200.73] [25360.08]
[73686.57] User 2
FB1S [628.30]
[78668.60] [2000000539] 113101 FB1S [25360.08]
191100
191100 [12860.01]
310000 5100004929
191100 2000000539 [5000005246]
[27977.89]User1 [27977.89] [703669.77] 160000
[25360.08] User1 [100010744] [27814.28]

[13393.39] 160000 MR1M [25360.08] [100010744] 191100


F110
[100010834] [27977.89] [5000005246]
[93235.10] [93235.10]
[97102.10] User 2
5100004962 User 1
310000 [29552.67] [100010834] [27814.28]
191100
[27977.89] 160000
100010834 191100
[97102.10] [27814.28]
MR1M [708224.95] [3926.88] [5100004929] 100010745
191100 230051 [83757.38] 154000
5000005281 FB1S 276000
[26178.14] [29552.67] [29552.67] 160000
User 2 [68081.79] [27814.28] FB1S
[0.01] [5100004929]
MB01 [29552.67] User1 [52696.81]
[29552.67]
User1 [100010745]
[26178.14] [5100004962] 160000
User 2 [100010745]
310000 [27977.89] 191100 [11552.74]
[100010833] [83757.38]
[26178.14] 160000
191100 [5100004962] 154000 191100
[5000005281] [100010833] 5100004997

100010833 [26178.14]
[5100004997] MR1M 100010933
[68081.79] [5100004997] [24092.07]
[5000005281] [26178.14]
FB1S 230051 [24092.07] [100010933]
[9390.59] User 2 FB1S
191100 [100010933]
230051 [22905.88]
User1 191100 [52696.81]
5100004996 User1
154000
[0.01] [25206.69]
[8017.06] MR1M [22905.88] [24092.07]
[0.00] 100010934 191100
[8017.06]
100010929 5100004865 [7268.53] 191100
User 2
191100 [17547.54] FB1S
FB1S [5100004996] 191100
[19505.79] MR1M
[100010932] User1 191100 310000
[100010929] 191100 [5100004996] [22905.88]
[100010929] [5100004865] [12429.51]
User1 [5100004865] User 3
[13620.82]
[100010932] [17547.54] [24092.07]
[25206.69] [100010934]
[6012.79]
191100 [22905.88]
100010932 [100010930] [13053.28] [22905.88]
191100 [13932.70]
[8017.06] 100010935 [100010934]
[19505.79] 191100 5000005337
FB1S [100010930]
191100
FB1S 310000
100010930 [12429.51] MB01
191100 User1 [100010480]
[6012.79] 191100
User1 [25206.69] [25206.69]User 2 [25206.69]
FB1S [100010480]
[17547.54] [100010482] [13053.28]
[13620.82] [100010483]
191100 100010480
User1 [100010482] [100010935] [5000005337]
[100010483] [24092.07]
[100010931] [13932.70] 100010482 FB1S
[19505.79] 100010483 310000
[8017.06] 191100
[100010931] [100010481] [100010935]
[5000005337]
[5000005336] 100010931 FB1S User1
[17547.54] [100010481] FB1S
[5000005336] 100010481 User1
FB1S
[19505.79] User1
310000 5000005336 [6012.79]
User1 FB1S
191100 [13053.28] 191100
MB01 [13620.82]
[17547.54] User1 191100 [12429.51]
User 2 [13620.82]
191100
[13932.70]
[19505.79] [8017.06]
310000 191100
[13620.82] 310000 [6012.79]
[13053.28]
310000 [12429.51]

[13932.70]
5000005127
310000
MB01

[5000005127] [12429.51]
[5000005127] User 4

[13932.70]
[6012.79] 310000
310000 [13053.28]

310000

Abbildung 5 Umfangreiche Prozessinstanz eines Einkaufsprozesses


Für die Bewertung der Anwendbarkeit der dargestellten Process Mining Methode wurden
neben den dargestellten Prozessinstanzen umfangreiche Event Logs ausgewertet. Hierbei
handelte es sich um Testdaten aus dem SAP IDES System [35] sowie um Echtdaten aus pro-
duktiven Systemen von Unternehmen. Hiervon ist eins in der Einzelhandelsbranche tätig
sowie das zweite in der verarbeitenden Industrie. Tabelle 3 gibt einen Überblick über den
Umfang der verwendeten Daten.

147
APPENDIX A: PUBLICATIONS

Tabelle 3 Übersicht Event Logs


#1 #2 #3
Event Logs
SAP IDES Einzelhandel Verarbeitende Industrie
Buchhaltungsbelege 115,060 92,487 1,764,773

Belegsegmente 419,106 222,901 7,395,434


Anzahl rekonstruierter Pro-
81,171 40,130 1,035,805
zessinstanzen

Die rekonstruierte Instanz mit der höchsten Komplexität umfasste ca. 470.000 Transaktio-
nen. Die in Abbildung 6 dargestellte Instanz beinhaltet 2.966 Transaktionen. Obwohl die
bisher entwickelten Methoden sicherlich nutzbringend für die Aufbauprüfung von Prozes-
sen eingesetzt werden können, zeigt dieses Beispiel deutlich deren Grenzen. Denn anhand
der bloßen Visualisierung einer solch komplexen Instanz können keine für die Prüfung ver-
wertbaren Informationen gewonnen werden. Hierfür bedarf es weiterer Betrachtungen.

Abbildung 6 Komplexe Prozessinstanz


Erste Ansatzmöglichkeiten ergeben sich aus der Analyse der Verteilungsfunktionen der An-
zahl der Instanzen über deren Komplexität43. Diese zeigen, dass es nur extrem wenige In-
stanzen mit einer extrem hohen Anzahl von Transaktionen gibt, wohingegen die über-wie-
gende Mehrzahl der Instanzen aus sehr wenigen Transaktionen bestehen. Ebenso ist zu
beobachten, dass auch in sehr großen Prozessinstanzen nur relative wenige Konten ver-
wendet werden.
Diese Beobachtungen bieten Ansatzpunkte für mögliche Verfahren zur Komplexitätsredu-
zierung der rekonstruierten Modelle. Dabei ist zu berücksichtigen, dass die Vereinfachung
von rekonstruierten Modellen weder im Forschungsfeld der Graphentheorie noch im Ge-
biet der Forschung zur Modellierung von Geschäftsprozessen ein neuartiges Problem dar-

43
Gemessen als Anzahl der Knoten und Kanten der jeweiligen Graphen.

148
APPENDIX A: PUBLICATIONS

stellt. Algorithmen aus der Graphentheorie befassen sich mit der Identifizierung isomor-
pher Teilgraphen [36], [37]. Im Forschungsbereich zum Thema Process Mining wird für
diese Problematik der anschauliche Begriff „Spaghetti-Prozesse“ verwendet [38] und mit-
tels Abstraktionsmethoden adressiert [39]. Wie diese Methoden für Process Mining einge-
setzt werden können, bleibt zukünftiger Forschung überlassen.

7. Zusammenfassung und Diskussion


Durch die Automatisierung der Verarbeitung von Geschäftsvorfällen in Unternehmen wer-
den Daten in den beteiligten Informationssystemen gespeichert. Der Einsatz von Business
Intelligence Software ermöglicht die Auswertung dieser Daten. Wirtschaftsprüfer stehen
im Zuge der Jahresabschlussprüfung vor der Herausforderung, komplexe und zunehmend
in Informationssysteme integrierte Geschäftsprozesse prüfen zu müssen. Traditionelle ma-
nuelle Prüfungsmethoden stoßen hierbei an ihre Grenzen. Dies führt zu einem Ungleichge-
wicht zwischen automatisierter Transaktionsverarbeitung auf der Seite der zu prüfenden
Unternehmen und manuellen Prüfungshandlungen auf der Seite der Jahresabschlussprü-
fer.
Zielgerichtete Prozess Mining Methoden bieten eine vielversprechende Möglichkeit, dieses
Ungleichgewicht zu reduzieren. In diesem Beitrag wurde gezeigt, wie Process Mining Me-
thoden für die Aufbauprüfung von Prozessen eingesetzt werden können, um anhand re-
konstruierter Modelle von Prozessinstanzen Informationen über den Zusammenhang der
untersuchten Geschäftsprozesse und der Buchhaltungskonten zu erlangen.
Neben den erzielten Ergebnissen zeigen die Auswertungen von Prozessmodellen beste-
hende Grenzen zu deren praktischen Einsatzmöglichkeiten. Sobald es sich um komplexe
Prozessinstanzen handelt, bietet deren Modellierung und Visualisierung mit den vorgestell-
ten Mitteln wenig Mehrwert. Hier gilt es Methoden zu finden, welche für die Komplexitäts-
reduktion verwendet werden können. Ein weiterer Aspekt betrifft die Zusammenfassung
gleichartiger Prozessinstanzen zu Prozessmodellen. Obwohl bestehende Algorithmen die
Erstellung von Prozessmodellen erlauben, ist zu prüfen, inwiefern diese den Anforderun-
gen der Anwendungsdomäne genügen und ob diese im Sinne eines Method Engineering
zweckmäßig verwendet werden können. 44

Literaturverzeichnis
[1] K. C. Laudon, J. P. Laudon, und D. Schoder, Wirtschaftsinformatik : eine Einführung.
München [u.a.]: Pearson Studium, 2006.
[2] IBM, „IBM Cognos Business Analytics und Performance Management Software -
Deutschland“. [Online]. Available: [Link]
nos/. [Accessed: 17-Sep-2012].

44
Die in dieser Veröfentlichung präsentierten Ergebnisse wurden im Forschungsprojekt Virtual Accounting
Worlds erarbeitet. Das Projekt wird vom Bundesministerium für Bildung und Forschung gefördert (För-
dernummer 01IS10041). Die Autoren sind verantwortlich für den Inhalt der Veröffentlichung.

149
APPENDIX A: PUBLICATIONS

[3] Oracle, „Oracle Business Intelligence Enterprise Edition“. [Online]. Available:


[Link]
view/[Link]. [Accessed: 17-Sep-2012].
[4] SAP, „SAP Deutschland - Komponenten & Werkzeuge von SAP NetWeaver: SAP
NetWeaver Business Intelligence“. [Online]. Available: [Link]
many/plattform/netweaver/components/businessintelligence/[Link]. [Ac-
cessed: 17-Sep-2012].
[5] E. Turban, J. E. Aronson, T.-P. Liang, und R. Sharda, Decision support and business
intelligence systems. Upper Saddle River, N.J.; London: Pearson Education Interna-
tional, 2007.
[6] C. Vercellis, Business intelligence. Chichester: Wiley, 2009.
[7] IDW, „IDW Aktuell - Willkommen im Institut der Wirtschaftsprüfer“. [Online]. Avail-
able: [Link] [Accessed: 17-Sep-2012].
[8] IAASB, „IAASB | International Accounting | Auditing Standards | Quality Assurance -
IFAC“. [Online]. Available: [Link] [Accessed: 17-Sep-
2012].
[9] International Federation of Accountants, „ISA 315 (Revised), Identifying and As-
sessing the Risks of Material Misstatement through Understanding the Entity and Its
Environment“. 2012.
[10] International Federation of Accountants, „ISA 330 The Auditor’s Responses to As-
sessed Risks“, in Handbook of international quality control, auditing, review, other
assurance, and related services pronouncements., Bd. 1, New York, NY: Interna-
tional Federation of Accountants, 2010.
[11] IDW, IDW PS 261 Feststellung und Beurteilung von Fehlerrisiken und Reaktionen
des Abschlussprüfers auf die beurteilten Fehlerrisiken. 2009.
[12] M. Werner und N. Gehrke, „Potentiale und Grenzen automatisierter Prozessprüfun-
gen durch Prozessrekonstruktionen“, in Forschung für die Wirtschaft, Shaker Verlag,
2011.
[13] M. Werner, N. Gehrke, und M. Nüttgens, „Business Process Mining and Reconstruc-
tion for Financial Audits“, in Hawaii International Conference on System Sciences,
2012, S. 5350–5359.
[14] W. van der Aalst, K. van Hee, J. M. van der Werf, A. Kumar, und M. Verdonk, „Con-
ceptual model for online auditing“, Decision Support Systems, Bd. 50, Nr. 3, S. 636–
647, Feb. 2011.
[15] J. E. Cook und A. L. Wolf, „Software process validation: quantitatively measuring the
correspondence of a process to a model“, ACM Trans. Softw. Eng. Methodol., Bd. 8,
Nr. 2, S. 147–176, 1999.
[16] J. E. Cook und A. L. Wolf, „Discovering models of software processes from event-
based data“, ACM Transactions on Software Engineering and Methodology, Bd. 7,
Nr. 3, S. 215–249, Juli 1998.

150
APPENDIX A: PUBLICATIONS

[17] J. E. Cook und A. L. Wolf, „Event-based detection of concurrency“, SIGSOFT Softw.


Eng. Notes, Bd. 23, Nr. 6, S. 35–45, 1998.
[18] A. Tiwari, C. Turner, und B. Majeed, „A review of business process mining: state-of-
the-art and future trends“, Business Process Management Journal, Bd. 14, Nr. 1, S.
5–22, 2008.
[19] W. M. P. van der Aalst, Process Mining: Discovery, Conformance and Enhancement
of Business Processes, 1st Edition. Springer Berlin Heidelberg, 2011.
[20] M. de Leoni, F. M. Maggi, und W. M. P. van der Aalst, „Aligning Event Logs and De-
clarative Process Models for Conformance Checking“, in Business Process Manage-
ment, 2012, Bd. 7481, S. 82–97.
[21] J. M. E. M. van der Werf, H. M. W. Verbeek, und W. M. P. van der Aalst, „Context-
Aware Compliance Checking“, in Business Process Management, 2012, Bd. 7481, S.
98–113.
[22] E. Ramezani, D. Fahland, und W. M. P. van der Aalst, „Where Did I Misbehave? Diag-
nostic Information in Compliance Checking“, in Business Process Management,
2012, Bd. 7481, S. 262–278.
[23] S. Banescu, M. Petković, und N. Zannone, „Measuring Privacy Compliance Using Fit-
ness Metrics“, in Business Process Management, 2012, Bd. 7481, S. 114–119.
[24] R. Accorsi und A. Lehmann, „Automatic Information Flow Analysis of Business Pro-
cess Models“, in Business Process Management, 2012, Bd. 7481, S. 172–187.
[25] N. Gehrke und N. Müller-Wickop, „Basic Principles of Financial Process Mining A
Journey through Financial Data in Accounting Information Systems“, in Proceedings
of the 16th Americas Conference on Information Systems, Lima, Peru, 2010.
[26] A. R. Hevner, S. T. March, J. Park, und S. Ram, „Design science in information sys-
tems research“, Mis Quarterly, S. 75–105, 2004.
[27] S. T. March und G. F. Smith, „Design and natural science research on information
technology“, Decision support systems, Bd. 15, Nr. 4, S. 251–266, 1995.
[28] H. Österle, J. Becker, U. Frank, T. Hess, D. Karagiannis, H. Krcmar, P. Loos, P. Mer-
tens, A. Oberweis, und E. J. Sinz, „Memorandum zur gestaltungsorientierten Wirt-
schaftsinformatik“, Schmalenbachs Zeitschrift für betriebswirtschaftliche Forschung,
Bd. 62, Nr. 9, S. 662–672, 2010.
[29] C. Riege, J. Saat, und T. Bucher, „Systematisierung von Evaluationsmethoden in der
gestaltungsorientierten Wirtschaftsinformatik“, Wissenschaftstheorie und gestal-
tungsorientierte Wirtschaftsinformatik, S. 69–86, 2009.
[30] S. Brinkkemper, „Method engineering: engineering of information systems develop-
ment methods and tools“, Information and Software Technology, Bd. 38, Nr. 4, S.
275–280, 1996.
[31] W. M. P. van der Aalst und C. Stahl, Modeling business processes : a petri net-ori-
ented approach. Cambridge, Mass.: MIT Press, 2011.

151
APPENDIX A: PUBLICATIONS

[32] R. Valk, „Lecture Notes: Formale Grundlagen der Informatik II (FGI 2) Modellierung
& Analyse paralleler und verteilter Systeme“. Universität Hamburg, 2008.
[33] N. Gehrke und N. Müller-Wickop, „Rekonstruktion von Geschäftsprozessen im Fi-
nanzwesen mit Financial Process Mining“, in Lecture Notes in Informatics, Procee-
dings der Jahrestagung Informatik, Leipzig, 2010.
[34] University of Hamburg, „Renew - The Reference Net Workshop“, 2012. [Online].
Available: [Link] [Accessed: 03-Mai-2012].
[35] SAP, „SAP-UCC“, 2012. [Online]. Available: [Link] [Accessed:
03-Mai-2012].
[36] J. Huan, W. Wang, und J. Prins, „Efficient mining of frequent subgraphs in the pres-
ence of isomorphism“, in Data Mining, 2003. ICDM 2003. Third IEEE International
Conference on, 2003, S. 549–552.
[37] X. Yan und J. Han, „gSpan: graph-based substructure pattern mining“, 2002, S. 721–
724.
[38] C. Günther und W. van der Aalst, „Fuzzy mining–adaptive process simplification
based on multi-perspective metrics“, Business Process Management, S. 328–343,
2007.
[39] M. Reichert, „Visualizing Large Business Process Models: Challenges, Techniques,
Applications“, in 1st Int’l Workshop on Theory and Applications of Process Visualiza-
tion, Tallin, 2012.

152
APPENDIX A: PUBLICATIONS

10.5 Business Process Mining and Reconstruction for Financial Audits

Number 5
Business Process Mining and Reconstruction
Title
for Financial Audits
Appendix 10.5
Primary Related Chapters 5.1
Type Conference Paper
45th Hawaii International Conference
Conference
on System Sciences (HICSS 2012)
Reference (Werner et al., 2012a)
Acceptance Rate 54 %
VHB JQ 2.1 Ranking C (6.44)
WKWI Ranking B
ERA 2010 A
CORE 2013 A
Review Procedure Double blinded
Number of Reviews 6
1. Michael Werner
Authors 2. Nick Gehrke
3. Markus Nüttgens
Dissertation Points 0.50
Authorship
Overall 90%
Design 90%
Realization 90%
Writing 90%
Status Published
Part of other Dissertations No
[Link]
Link [Link]/stamp/[Link]?tp=&ar-
number=6149542&tag=1

153
APPENDIX A: PUBLICATIONS

Business Process Mining and Reconstruction for Financial Audits

MICHAEL WERNER NICK GEHRKE MARKUS NÜTTGENS


Nordakademie, University of Nordakademie, University of University of Hamburg, Ger-
Applied Sciences, Germany Applied Sciences, Germany many
[Link]@ [Link]@ [Link]@
[Link] [Link] [Link]

Abstract: In modern companies business processes and information systems


are highly integrated and transactions are executed system based and auto-
mated. The data generated in the course of processing transactions commonly
provides the basis for internal and external financial reporting. The financial
statements are subject to audits due to regulatory requirements. Contempo-
rary audit approaches take into account internal control frameworks over rele-
vant business processes and underlying information systems, but they lack ad-
equate audit procedures needed to handle voluminous data flows when busi-
ness processes are highly integrated and automated. We face a discrepancy be-
tween an integrated and automated transaction processing on the one side and
manual audit procedures on the other. Financial audits would be more effective
and efficient if an audit approach with system based and automated proce-
dures would be applied. This article describes how business process mining and
reconstruction of mined processes can be used to overcome this discrepancy.

1. Introduction
The execution of business processes in companies is regularly based on information sys-
tems. The integration between business process and information systems ranges from sup-
port for manual executions to completely automated processing. Enterprise Resource Plan-
ning (ERP) systems represent the dominant type of information systems that are imple-
mented to support and automate transaction processing. Depending on the industry and
types of business processes that are integrated into the ERP systems millions or even bil-
lions of transactions may be processed within a financial period.
ERP systems do not only support or automate the execution of transactions but they also
commonly provide the data basis for the internal and external financial reporting as well as
integrated functionality for preparing the financial statements including the balance sheet
and profit and loss statements. This means that the financial statements represent an ag-
gregation of the information stored in the ERP system that is made up by the processing of
myriad numbers of transactions.
Companies are required to prepare financial statements in order to inform addressees pri-
marily about the financial situation of the company. To protect addressees from misinfor-
mation the financial statements are subject to independent audits by financial auditors.
The requirements for the audit are specified in local or international laws, regulations and
standards.

154
APPENDIX A: PUBLICATIONS

International Standards on Auditing require the application of a risk based audit approach
that takes into account the internal control framework over relevant business processes
and underlying information systems (ISA 315). A risk based audit approach requires the
identification of relevant risks for material misstatements and the evaluation how internal
controls are able to mitigate existing risks. When business processes are integrated with
ERP systems application controls represent a significant type of internal controls that have
to be considered in the audit. Contemporary audit approaches consider internal controls
and application controls as a special type of internal controls that are embedded in the
integrated system. Although relevant business processes and internal controls are taken
into account contemporary audit procedures are generally not system based. The selection
and test of controls is done manually.
This situation leads to a discrepancy. Business processes are highly integrated with ERP
systems and transactions are processed automatedly. Auditors identify significant risks and
manually evaluate relevant internal controls that are in place to mitigate these risks. On
the company side we observe system based and automated processing and on the auditor
side a risk based approach with manual evaluation procedures. An approach with auto-
mated and system based audit procedures would lead to more efficient and effective au-
dits.
The need for automated audit procedures has been pointed out by major market partici-
pants [17], but an approach that includes system based and automated audit procedures
has not been developed yet due to the fact that adequate methods and software artifacts
have not been available.
Recent research by GEHRKE et al. [14], [15] has revealed how financially relevant infor-
mation can be extracted via financial business process mining from information systems.
The mined information can be used to reconstruct processes based on that information.
This paper presents how business process mining and reconstruction can be used to apply
a risk based audit approach with system based and automated audit procedures.
The article starts with an overview of related theoretical work in section 2. Section 3 pro-
vides a brief summary of contemporary audit approaches and their limitations in system
based and highly integrated environments. Section 4 presents the concepts of business
process mining and reconstruction. Section 5 discusses the relevance of application con-
trols for highly integrated and automated business processes. In section 6 we discuss how
the concepts of business process mining and reconstruction can be combined with auto-
mated application control testing for developing system based and automated audit pro-
cedures. Section 7 closes with a discussion how stakeholders benefit, which limitations ex-
ist and what further developments are needed.

2. Related Work
The idea of process mining evolved in the 1990s. COOK and WOLF [9], [10], [11] investi-
gated process mining in the context of software engineering. They describe different meth-
ods for process discovery. However, they do not provide an approach to generate explicit
process models. The idea to apply process mining in the context of workflow management
was first introduced by AGRAVAL et al. [8]. Further research was undertaken by MAXEINER
et al. [28] and SCHIMM [33], [34], [35], [36] who developed mining tools. HERBST and

155
APPENDIX A: PUBLICATIONS

KARAGIANNIS also address process mining in the context of workflow management using
an inductive approach [18], [19], [20], [21], [22], [23]. Substantial research has been pub-
lished by VAN DER AALST et al. [1], [2], [3], [4], [5], [6], [7], [26], [27], [38], [39], [40] that
deals with the mining and rediscovery of process models from event logs. This research is
especially relevant because the provided methods and algorithms allow the reconstruction
of petri nets which represent the process models mined from event logs. They also cover
considerations of workflow performance, concurrency, noise and conformance checking.
Although this research is valuable for the research subject of this article several limitations
have to be considered. The process mining research by VAN DER AALST et al. focuses on
event based logs and intends to reconstruct graphs that completely represent the pro-
cesses that produce these event logs. Mining of business processes in ERP systems for the
financial audit entails different environmental settings and intends to achieve different
aims. First, the stored data in ERP systems includes much more detailed information than
common event logs from workflow systems. They store accounting information about jour-
nal entries that provide more specific data usable for process mining. Second, for the pro-
posed mining we intend to rediscover single representative process instances that are ag-
gregated into process models, we do not intend to rediscover complete representations of
the mined data.
The discipline of process mining is characterized by technical research approaches. The
connection between process flows, process mining, process reconstruction and accounting
has not been extensively covered in scientific work so far. We assume that the connection
between informatics oriented process mining and the business management topics ac-
counting and compliance has not been the focus of interest so far due to the thematically
distance of the two disciplines.
An exception is the research work by GEHRKE et al. [13], [14], [15], [16] that has been de-
rived from the research project Virtual Accounting Worlds [37]. The developed methods
and concepts bridge the gap between accounting, compliance and process mining. They
represent an application of fundamental concepts from VAN DER AALST et al. for mining of
business processes that are relevant for financial accounting. Within this paper we relate
to these methods and concepts and include them in a wider consideration in order to
demonstrate how they can be applied for developing an audit approach that includes sys-
tem based and automated audit procedures.
Further relevant research covers topics like process data warehousing by EDER et al. [12]
and ZUR MÜHLEN et al. [29], [30], [31] which can be used for developing software artifacts
needed to apply the discussed approach in practice.

3. Risk based audit approach in the context of integrated and au-


tomated business processes
Before we can understand the shortcomings of the application of contemporary audit ap-
proaches in business environments with integrated and automated business processes we
need to clarify their relevant characteristics.
Companies are required to apply generally accepted accounting principles (GAAP) when
preparing their financial statements. This ensures that the published statements display

156
APPENDIX A: PUBLICATIONS

correct and comparable information to the addressees. External addressees are sharehold-
ers, creditors, tax and regulatory authorities, employees, clients, financial analysts, com-
petitors and the general public [24].
In order to protect addressees from misinformation the financial statements are audited
by financial auditors. The obligation to engage auditors for auditing the financial state-
ments is generally mandated by law. Auditors follow standards on auditing to ensure that
adequate audit procedures are applied. They assess if the audited statements give a fair
and true view of the financial situation of a company and if the statements are free of ma-
terial misstatements.
Standards on accounting and standards on auditing are issued by regulatory bodies such as
the International Accounting Standards Board (IASB) for the International Financial Report-
ing Standards (IFRS) or International Auditing and Assurance Standards Board (IAASB) for
the International Standards on Auditing (ISA). Laws, regulations and standards differ be-
tween countries. But in recent years we observe a convergence between internationally
significant accounting frameworks especially between the IFRS and US GAAP [25]. We do
not intend to focus on differences of the accounting and audit frameworks in this article.
For our purpose it is sufficient to point out that a risk based audit approach is mandated by
ISA (e.g. ISA 315) as well as local regulations or standards such as the Sarbanes-Oxley-Act
in the USA. For the remainder of this article we primarily refer to IFRS and ISA while pointing
out that the same considerations and conclusions provided in this article are applicable to
other accounting and audit frameworks.
ISA 315 requires the application of a risk based audit approach: “The objective of the audi-
tor is to identify and assess the risks of material misstatement, whether due to fraud or
error (…) through understanding the entity and its environment, including the entity’s in-
ternal control (…)” (ISA 315.3). The auditor has to identify and to evaluate the risks that
might lead to material misstatements. The auditor further needs to identify if internal con-
trols do exist that mitigate existing risks: “The auditor shall obtain an understanding of in-
ternal control relevant to the audit (…)” (ISA 315.12). The underlying axiom of the approach
is the assumption that well organized and controlled processes lead to correct financial
reporting.
In practice the audit takes place by identifying risks that are significant to the audit. A gen-
eral significant risk is that business transactions are not recorded completely or correctly.
Following a risk based approach it is not necessary to consider all business processes within
a company but only those where errors in the processing might lead to a material misstate-
ment in the financial statements.
Typical business processes relevant for financial accounting are purchase, sales, payroll,
production and logistics processes.
When the scope of the audit is determined and relevant processes identified the auditor
has to gain an understanding of the processes and the internal control over these pro-
cesses. The auditor has to evaluate if the controls are properly designed and operative to
achieve the desired control objectives. The procedure to understand business processes,
to evaluate and test internal controls is a manual and highly time-consuming activity. It
generally includes interviews with knowledgeable contact persons and manual reviews of
provided documentation.

157
APPENDIX A: PUBLICATIONS

We illustrate the procedure for the following example. For producing goods a company
creates purchase requisitions and orders individually and with paper based forms. The or-
ders have to be approved by signature by the purchase representative. The responsible
warehouse worker checks if the amount and quantity of the received goods equal the
amount and quantity of the purchase order when the goods are delivered. When the in-
voice for the delivered goods is received from the supplier a responsible person in the ac-
counting department checks if the billed amount and quantity equal the amount and quan-
tity of purchase order and the goods received.
For understanding the process and for evaluating the relevant internal controls an auditor
first performs interviews with the persons involved in the process. He evaluates if the con-
trols in place are adequate to control the process and to achieve the desired control objec-
tives. Based on the understanding of the process and the controls he performs tests to
evaluate if the controls are carried out continuously and effectively throughout the rele-
vant reporting period. The testing of the operating effectiveness requires the review of rel-
evant documents. In the mentioned example the auditor would draw a representative sam-
ple of purchase transactions and verify if check marks and signatures are available on the
provided documentation.
The example illustrates that the audit procedures for auditing business processes and in-
ternal controls are highly manual and time-consuming in nature.
The described procedure is practical for manually executed business processes but it is in-
sufficient when business processes are highly integrated with ERP systems and executed
automated. Under such conditions contact persons from the relevant business functions
generally lack sufficient knowledge about the integration and type of automation with the
underlying systems. A common observation is that provided information does not correctly
reflect the process implementation within the ERP systems. Second, with an increasing
number of executed transactions manual review of available evidence becomes increas-
ingly inefficient or even ineffective. A manual review of even hundreds of documents does
not provide sufficient audit comfort when millions of such transactions are executed within
the relevant period.
We illustrate the shortcomings of contemporary audit procedures in an integrated and au-
tomated business process environment with a second example.
A company has integrated its production and purchase processes in an ERP system. Pro-
duction orders automatically initiate purchase requisitions and purchase orders based on
item lists maintained in the system. The program routines initiate purchase orders only if
required items are not available in the warehouse. Purchase requisitions and orders are
approved automatically up to a certain amount. Only purchase orders exceeding that
amount are subject to a system based approval by the purchase department. The system
further blocks purchase orders randomly for manual but system based approval in order to
prevent manipulation. The warehouse clerk can only accept received goods if the quantity
and amount match the purchase order (two-way-match). Otherwise an exception handling
sub-process is initiated. The accounting department can only process incoming invoices if
the billed amount and quantity matches the amount and quantity of the purchase order
and the goods received (three-way-match). Otherwise an exception handling sub-process
is initiated.

158
APPENDIX A: PUBLICATIONS

The example demonstrates a highly integrated business process with automated execu-
tions and automated and systems based internal controls also referred to as application
controls. In the described environment performing interviews with contact persons from
the functional departments might not provide sufficient information because they may lack
the information how transactions are processed automatically, when no human interaction
occurs, and especially which application controls do exist. A common occurrence is that
contact persons think application controls are in place and effective which in fact is not the
case. A second dilemma becomes obvious when controls actually get tested. The automa-
tion of execution means that paper based evidence might not be available. In such a situa-
tion it is necessary to manually evaluate and to test relevant application controls. The eval-
uation and testing of application controls requires a specialized knowledge of the ERP sys-
tem in use. Furthermore, the review of control settings, commonly based on the customiz-
ing settings, requires extensive access rights and is a manual time-consuming work. Third,
even if these procedures are applied no information is available if the controls really cover
complete transaction flows or if controls are bypassed by concurrent transaction flows dif-
fering from the general transaction flows, by manual journal entries or manipulation.
Fourths, generally it is hard to test if the application controls were effective over the whole
relevant period, for example if specific application controls were disabled for a specific
timeframe.
Computer assisted audit techniques (CAAT) for supporting the testing of application con-
trols and business processes integrated into ERP systems do exist [13]. But they only sup-
port the manual execution of tests or provide functionality to analyze mass data for journal
entry testing. The described fundamental problems are not solved.
An audit approach is needed that counters the system based and automated processing by
applying system based and automated audit procedures.

4. Business process mining and process reconstruction


Business process mining and process reconstruction provides the concepts and methods
needed to implement system based and automated audit procedures.
VAN DER AALST et al. [1], [3], [4], [5], [6], [26], [27], [38], [39] focus on event logs in order
to mine and reconstruct workflows. We can rely on these concepts to mine and reconstruct
process models in ERP systems. In contrast to the systems and log files described by VAN
DER AALST et al. ERP systems provide much more detailed information about the processed
transactions. Every financially relevant business transaction executed in an ERP system is
recorded as a journal entry posted to an account in a main or sub ledger. Basically, entries
in the accounting of an ERP system are structured in a simple way [32]. Each entry consists
of an accounting document and at least two items posted as credits and debits.
Technically documents and items are stored as entries in data tables in the underlying da-
tabase of the ERP system. The stored data for each transaction contains information that
allows identifying relationships between the transactions. The transactions can be traced
back to the executed instance of a business process they belong to.
Mining of business processes that are relevant for financial accounting can be applied
where open item accounting is enabled, which is the case for most relevant processes. If

159
APPENDIX A: PUBLICATIONS

open item accounting is in use for a particular account, each item contains a flag that indi-
cates if the item has already been cleared or not. If an item has been cleared, it also con-
tains a reference to the entry / document which cleared the item.

Figure 1 General data structure of an open item accounting entry


Figure 1 shows the general data structure of an accounting entry in a database. One docu-
ment consists of two or more items posted on different accounts. Each item can be linked
to one (other) document (=item cleared) or does not refer to (another) document (=item
still open).
We illustrate the data structure with the example of an execution of a purchase process.
The receipt of an ordered material (transaction 1) is recorded as a debit posting on raw
materials and a credit posting on the goods received / invoices received account. Upon
receipt of the incoming invoice (transaction 2) the posting on the goods received / invoices
received account is cleared with a corresponding credit posting on the creditor account.
The document number of transaction 2 becomes the clearing document number for the
posting item from transaction 1. With the payment run (transaction 3) the creditor account
is cleared. The document number of transaction 3 becomes the clearing document number
for the posting item of transaction 2. The example illustrates how the execution of a busi-
ness process instance is recorded in the system and how the processing produces a chain
of journal entries within the system that is traceable.
If we abstract from this example we can conclude that transactions processed in an ERP
system leave a digital trace within the system. GEHRKE AND MÜLLER-WICKOP [14], [15]
present an algorithm that is able to mine these traces. The algorithm takes off with a start
document and iteratively mines corresponding documents by identifying the relevant
clearing document. If no further documents can be found the algorithm terminates. The
mined information can be used for graphically reconstructing and representing the mined
process instance.

160
APPENDIX A: PUBLICATIONS

Figure 2 Reconstruction of a mined process instance


Figure 2 shows a section of the reconstruction for a mined purchase process instance from
a SAP system.
Transactions are represented as rounded orange rectangles. Simple rectangles represent
items of financial entries involved in open item accounting. The rectangles with black bor-
ders represent cleared items of financial entries. Items illustrated as rectangles without
borders are not cleared and still open. Hexagons represent entries not involved in open
item accounting. The color of the rectangles and hexagons indicates if the item is a debit
or credit posting on a general ledger or a profit and loss account. An arrow from a business
activity to an item means that the business activity has produced the item as a part of the
complete entry. An arrow from an item to a business activity means that the item has been
cleared by the corresponding accounting document. Detailed information such as item
number, account and amount is displayed for each item.
The shown process instance starts with the transaction MB01 (post goods receipt for pur-
chase order). The open items are cleared by the transaction MR1M (enter incoming in-
voice). When entering an invoice via MR1M it is possible to explicitly reference correspond-
ing open items. In the mined instance MR1M was executed without such explicit refer-
ences. In this case the items are cleared by the ERP system automatically via the execution
of transaction FBS1 (clear G/L account). FBS1 does not represent a separate business trans-
action and therefore does not follow the structure displayed in Figure 1. The open item
posted by MR1M is cleared by transaction F110 (payment run).

161
APPENDIX A: PUBLICATIONS

5. Application controls
ERP systems provide control mechanisms in order to govern and control the processing
within the system. Control mechanisms that are inherently embedded in software are
called application controls. Application controls represent a type of internal controls [13].
Examples are automatic reconciliation procedures, prevention of entering duplicate trans-
actions, system forced approvals or system based two- and three-way-matches.
Application controls play a key role for auditing system based and automated processes.
They provide a means for overcoming the problem that manual testing of business trans-
actions becomes inefficient for integrated and automated processes. Instead of testing sin-
gle business transactions it is possible to test the design and effectiveness of application
controls that cover whole process flows independent from the number of business trans-
actions that are processed.
By relying on application controls provided by the system for the purpose of the financial
audit automated processing of transactions can be countered by automated control mech-
anisms.
Unfortunately the testing of application controls itself is a manual and time-consuming pro-
cedure. Application controls are generally configured and enabled during the implementa-
tion of the system by setting relevant customizing settings. These settings need to be eval-
uated. However, settings for application controls are stored within the ERP systems and
methods and software artifacts exist that allow to extract relevant settings and to test them
in an automated way [13].

6. System based and automated audit procedures


In the previous sections we illustrated how process instances can be automatically mined
and reconstructed. We point out that business process mining and reconstruction does not
merely represent another CAAT. Indeed it allows introducing a new audit procedure that is
adequate for the audit of companies with highly integrated and automated business pro-
cesses as shown by the following considerations.
When applying a contemporary audit approach the auditor decides which financial ac-
counts are in scope from a risk perspective and evaluates which business processes have
to be considered. This decision is based on professional judgment derived primarily from
experience. The scoping may be appropriate or not. It is a manual procedure and highly
dependent on the knowledge and experience of the auditor. Information from the under-
lying information systems is not considered or only to a marginal extent although the in-
formation systems indeed provide all necessary information for a precise scoping.
Business process mining and reconstruction provides the possibility to analyze how the
transactions flow throughout the system and which processes effect relevant accounts. The
methods to reconstruct and visualize a single execution of a business process were pre-
sented in section 4. In order to implement automated and system based audit procedures
it is necessary to aggregate mined process instances to process models that represent the

162
APPENDIX A: PUBLICATIONS

process flows within the system. The aggregation of mined process instances is possible
but adequate algorithms for automated aggregations are currently still under research.
We define a process flow as a collection of executed similar business activities. The process
flows can be made explicit and the auditor can virtually see how they interact with the
relevant accounts.
For illustrating the possibilities that business process mining offers, we use the following
analogy. We compare the financial statements of a company to a lake of water. The auditor
has to provide an opinion if the lake only contains water from specific sources with a de-
fined quality. This requirement is the analogy to the real life requirement that financial
statements present a fair and true view of the financial situation of the company. Rivers
feed our imaginary lake. The rivers represent transaction flows and the water transaction
data. Following a contemporary audit approach the auditor would manually take samples
from different places in the lake to verify the water quality (=substantive testing). Based on
experience and professional judgment he would also choose several rivers for inspection
and verify manually if control mechanisms are in place (=manual controls testing) that reg-
ulate the flow and quality of the water, but without knowing which rivers and concurrent
flows indeed exist and how much water they actually carry into the lake.
By applying business process mining and reconstruction the auditor first develops a map
with all relevant rivers that flow into the lake with information which control mechanisms
control the flow and quality of the water. The auditor gathers information about how much
water each river carries and which rivers or concurrent streams flow uncontrolled (=busi-
ness process mining and reconstruction). Based on this understanding the auditor can de-
cide precisely which rivers are significant and instead of taking random samples the auditor
can decide specifically which control mechanisms should be tested to cover relevant pro-
cess flows (=automated controls testing). Process flows with no application controls in
place can be identified for targeted samples.

Figure 3 Process flow map

163
APPENDIX A: PUBLICATIONS

Figure 3 illustrates on an aggregated level a map of process flows. It illustrates how differ-
ent process flows feed the financial statements. The diagram further shows how applica-
tion controls interact with the process flows and how they control them.
In the upper left corner of the diagram we illustrate how a typical automated purchase
process as already described in section 3 would be represented.
The process flow of purchase orders (process flow 1) is controlled by a system based ap-
proval (application control A), the combined process flow of purchase orders and goods
receipt (process flow 2) is controlled by an automated two-way-match (application control
B). When transaction flows 1 and 2 combine with the process flow of incoming invoices
(process flow 4) the combined flow is controlled by an automated three-way-match (appli-
cation control C). The diagram also shows that the concurrent process flow of receipt ser-
vices (process flow 3) is not controlled by the two-way-match (application control B). The
reason is that commonly no receipt data for the delivery of services is available that could
be subject of an automated control activity. A matching for delivered services commonly
takes place between the purchase order and the invoice via application control C.

7. Conclusion
In modern companies business processes and information systems are highly integrated
and transactions are executed system based and automatedly. In section 3 we have shown
that contemporary audit procedures for financial audits are not adequate in environments
where business processes are highly integrated and automated. It is ineffective and ineffi-
cient to audit automated business processes and internal controls with manual audit pro-
cedures.
Business process mining and reconstruction provides methods and procedures that base
directly on the information stored in the underlying ERP systems. These methods and pro-
cedures combined with methods and procedures usable for automated application control
testing can be applied to implement system based and automated audit procedures for
financial audits. Business process mining and reconstruction allows visualizing process
flows within an ERP system and how these process flows interact with application controls
embedded in the system. The automated and system based analysis and its graphical rep-
resentation enables auditors to handle the complexity of integrated and system based busi-
ness processes and internal controls.
Auditing firms have recognized the need to introduce automated audit procedures in order
to keep up with technological progress [17]. Via the application of system based and auto-
mated audit procedures as introduced in this article it is possible to meet this requirement.
It is expected that the introduction of system based and automated audit procedures will
lead to significant gains in effectiveness and efficiency of financial audits.
The conclusions presented in this paper base on the research work derived from the re-
search project Virtual Accounting Worlds (VAW) [37]. A major market participant of the
auditing industry participates in the project as an associated project partner. The prototype
for financial business process mining and reconstruction as well as the prototype for auto-

164
APPENDIX A: PUBLICATIONS

mated application control testing has been developed within the VAW project. It is in-
tended to develop a software artifact providing the functionality for system based and au-
tomated audit procedures as described in this article in further research.
Internal and external auditors are not the only stakeholder that would benefit from the
availability of methods and artifacts for automated auditing. The process mining, recon-
struction and visualization provides the basis for analyses, performance and optimization
consideration that are of interest for process owners and managers, risk management and
business management in general.
We have to point out that in order to implement system based and automated audit pro-
cedures further issues need to be researched. The prototypes referred to in this article [13],
[14] provide software artifacts that proof the correctness and applicability of the underly-
ing concepts and methods. Nevertheless no information is available how the artifacts will
behave and perform in real live environments. The described methods for process mining
allow the mining of single process instances. In order to analyze and to visualize recon-
structed process flows it is necessary to aggregate mined process instances to process mod-
els. Adequate algorithms for the automated aggregation of mined process instances are
still under research as well as methods for an automated visualization. Further attention
has to be paid to the selection of representative process instances if the complete mining
of all instances is not a viable option due to the amount of processed instances which will
be a common occurrence in real life settings.
The aspects discussed in this paper focus on methods for automating business process and
internal controls testing. The overall audit of financial statements is a complex and difficult
task carried out by qualified experts. Process and controls testing is only a part of a financial
audit. The possibilities for automating audit procedures are limited to the extent to how
the underlying transaction processing is automated. Due to the fact that companies act as
market participants in changing and volatile environments there will always be unique busi-
ness transactions such as mergers or acquisitions that need to be evaluated by manual and
substantive audit procedures. The aim of introducing system based and automated audit
procedures is to counter automated processing with adequate audit procedures and to set
free resources for more sophisticated audits of unique, uncontrolled or exceptional trans-
actions that generally comprise higher risks than standard transactions.
Although further research is needed for developing mature software artifacts that allow
the implementation of system based and automated audit procedures for financial audits
presented in this article we conclude that the basic methods have already been developed
and proofed valid.
Further research has to focus on issues such as process instance selection and automated
aggregation, visualization, complexity and viability of developed algorithms. These aspects
will be the focus of further research within the VAW project that has already been initiated.
Specific information concerning these topics and the developed software artifact will be
provided in subsequent publications.

165
APPENDIX A: PUBLICATIONS

8. References
[1] W.M.P. van der Aalst, “Business Alignment: Using Process Mining as a Tool for Delta
Analysis and Conformance Testing”, Requirements Engineering Journal, volume 10,
issue 3, pp. 198-211, 2005.
[2] W.M.P. van der Aalst, Process Mining: Discovery, Conformance and Enhancement of
Business Processes, 1st Edition, Springer Berlin Heidelberg, 2011.
[3] W.M.P. van der Aalst, “The Application of Petri Nets to Workflow Management,”
The J. Circuits, Systems and Computers, vol. 8, no. I, pp. 21-66, 1998.
[4] W.M.P. van der Aalst, “Verification of Workflow Nets”, Application and Theory of
Petri Nets, P. Azema and G. Balbo, eds., pp. 407-426, Berlin: Springer-Verlag, 1997.
[5] W.M.P. van der Aalst, and B.F. van Dongen, “Discovering Workflow Performance
Models from Timed Logs,” Proc. Int’l Conf. Eng. and Deployment of Cooperative In-
formation Systems (EDCIS 2002), Y. Han, S. Tai, and D. Wikarski, eds., vol. 2480, pp.
45-63, 2002.
[6] W. M. P. van der Aalst, B. F. van Dongen, J. Herbst, L. Maruster, G. Schimm, and
A.J.M.M. Weijters, “Workflow Mining: a Survey of Issues and Approaches”, Beta:
Research School for Operations Management and Logistics, Working Paper 74,
2002.
[7] W.M.P. van der Aalst, A.J.M.M. Weijters, and L. Maruster, “Workflow Mining: Which
Processes can be Rediscovered?” BETA Working Paper Series, WP 74, Eindhoven
Univ. of Technology, Eindhoven, 2002.
[8] R. Agrawal, D. Gunopulos, and F. Leymann, “Mining Process Models from Workflow
Logs,” Proc. Sixth Int’l Conf. Extending Database Technology, pp. 469-483, 1998.
[9] J.E. Cook, and A.L. Wolf, “Discovering Models of Software Processes from Event-
Based Data,” ACM Trans. Software Eng. and Methodology, vol. 7, no. 3, pp. 215-249,
1998.
[10] J.E. Cook and A.L. Wolf, “Event-Based Detection of Concurrency,” Proc. Sixth Int’l
Symp. the Foundations of Software Eng. (FSE-6), pp. 35-45, 1998.
[11] J.E. Cook, and A.L. Wolf, “Software Process Validation: Quantitatively Measuring the
Correspondence of a Process to a Model,” ACM Trans. Software Eng. and Methodol-
ogy, vol. 8, no. 2, pp. 147-176, 1999.
[12] J. Eder, and G.E. Olivotto, and W. Gruber, “A Data Warehouse for Workflow Logs,”
Proc. Int’l Conf. Eng. and Deployment of Cooperative Information Systems (EDCIS
2002), Y. Han, S. Tai, and D. Wikarski, eds., pp. 1-15, 2002.
[13] N. Gehrke, “The ERP AuditLab - A prototypical Framework for Evaluating Enterprise
Resource Planning System Assurance”, in: Proceedings of the 43th Hawaii Interna-
tional Conference on System Sciences (HICSS-43), Hawaii (2010).
[14] N. Gehrke, and N. Müller-Wickop, “Basic Principles of Financial Process Mining”,
Proceedings of the 16th Americas Conference on Information Systems, Lima, Peru,
2010.

166
APPENDIX A: PUBLICATIONS

[15] N. Gehrke, and N. Müller-Wickop, “Rekonstruktion von Geschäftsprozessen im Fi-


nanzwesen mit Financial Process Mining“, in: Lecture Notes in Informatics, Procee-
dings der Jahrestagung Informatik 2010, IT-supported Service Innovation and Ser-
vice Improvement, Leipzig, 2010.
[16] Gehrke, N., Wolf, P.: Towards Audit 2.0 - A Web 2.0 Community Platform for Audi-
tors, in: Proceedings of the 43th Hawaii International Conference on System Sci-
ences (HICSS-43), Hawaii (2010).
[17] K. Heese, H. Kreisel, “Püfung von Geschäftsprozessen“, Die Wirtschaftsprüfung, Vol.
18, pp. 907-919, 2010.
[18] J. Herbst, “A Machine Learning Approach to Workflow Management,” Proc. 11th Eu-
ropean Conf. Machine Learning, pp. 183-194, 2000.
[19] J. Herbst, “Dealing with Concurrency in Workflow Induction,” Proc. European Con-
current Eng. Conf., U. Baake, R. Zobel, and M. Al- Akaidi, eds., 2000.
[20] J. Herbst, “Ein induktiver Ansatz zur Akquisition und Adaption von Workflow-Model-
len,” PhD thesis, Universität Ulm, Nov. 2001.
[21] J. Herbst, and D. Karagiannis, “An Inductive Approach to the Acquisition and Adap-
tation of Workflow Models,” Proc. Workshop Intelligent Workflow and Process
Management: The New Frontier for AI in Business, M. Ibrahim and B. Drabble, eds.,
pp. 52-57, Aug. 1999.
[22] J. Herbst, and D. Karagiannis, “Integrating Machine Learning and Workflow Manage-
ment to Support Acquisition and Adaptation of Workflow Models,” Proc. Ninth Int’l
Workshop Database and Expert Systems Applications, pp. 745-752, 1998.
[23] J. Herbst, and D. Karagiannis, “Integrating Machine Learning and Workflow Manage-
ment to Support Acquisition and Adaptation of Workflow Models,” Int’l J. Intelligent
Systems in Accounting, Finance, and Management, vol. 9, pp. 67-92, 2000.
[24] K. Küting, and M. Reuter, “Bilanzierung im Spannungsfeld unterschiedlicher Adres-
saten“, DSWR, Vol. 9, 2004.
[25] P. Leibfried, and P. Meixner, “Konvergenz der Rechnungslegung: Bestandsaufnahme
und Versuch einer Prognose“, Der Schweizer Treuhänder, Vol. 4, 2006.
[26] L. Maruster, W.M.P. van der Aalst, A.J.M.M. Weijters, A. van den Bosch, and W.
Daelemans, “Automated Discovery of Workflow Models from Hospital Data,” Proc.
13th Belgium-Netherlands Conf. Artificial Intelligence (BNAIC 2001), B. Kro¨se, M.
de Rijke, G. Schreiber, and M. van Someren, eds., pp. 183-190, 2001.
[27] L. Maruster, A.J.M.M. Weijters, W.M.P. van der Aalst, and A. van den Bosch, “Pro-
cess Mining: Discovering Direct Successors in Process Logs,” Proc. Fifth Int’l Conf.
Discovery Science (Discovery Science 2002), pp. 364-373, 2002.
[28] M.K. Maxeiner, K. Küspert, and F. Leymann, “Data Mining von Workflow-Protokol-
len zur teilautomatisierten Konstruktion von Prozessmodellen” Proc. Datenbanksys-
teme in Büro, Technik und Wissenschaft, pp. 75-84, 2001.

167
APPENDIX A: PUBLICATIONS

[29] M. zur Mühlen, “Process-Driven Management Information Systems—Combining


Data Warehouses and Workflow Technology,” Proc. Int’l Conf. Electronic Commerce
Research (ICECR-4) B. Gavish, ed., pp. 550-566, 2001.
[30] M. zur Mühlen, “Workflow-Based Process Controlling-Or: What You Can Measure
You Can Control,” Workflow Handbook 2001, Workflow Management Coalition, L.
Fischer, ed., pp. 61-77, Lighthouse Point, Fla.: Future Strategies, 2001.
[31] M. zur Mühlen, and M. Rosemann, “Workflow-Based Process Monitoring and Con-
trolling - Technical and Organizational Issues,” Proc. 33rd Hawaii Int’l Conf. System
Science (HICSS-33), R. Sprague, ed., pp. 1-10, 2000.
[32] M. Romney, and M. Steinbart, Accounting Information Systems, International Ver-
sion, 11th Edition, Prentice Hall, pp. 56-69, 2008.
[33] G. Schimm, “Generic Linear Business Process Modeling,” Proc. ER 2000 Workshop
Conceptual Approaches for E-Business and The World Wide Web and Conceptual
Modeling, S.W. Liddle, H.C. Mayr, and B. Thalheim, eds., pp. 31-39, 2000.
[34] G. Schimm, “Process Mining Elektronischer Geschäftsprozesse,” Proc. Elektronische
Geschäftsprozesse, 2001.
[35] G. Schimm, “Process M6ning linearer Prozessmodelle - Ein Ansatz zur Automatisier-
ten Akquisition von Prozesswissen,” Proc. 1. Konferenz Professionelles Wissensman-
agement, 2001.
[36] G. Schimm, “Process Miner - A Tool for Mining Process Schemes from Event-Based
Data,” Proc. Eighth European Conf. Artificial Intelligence (JELIA), S. Flesca and G.
Ianni, eds., pp. 525-528, 2002.
[37] Virtual Accounting Worlds, [Link], retrieved June 14,
2011
[38] A.J.M.M. Weijters, and W.M.P. van der Aalst, “Process Mining: Discovering Work-
flow Models from Event-Based Data,” Proc. 13th Belgium-Netherlands Conf. Artifi-
cial Intelligence (BNAIC 2001), B. Kro¨ se, M. de Rijke, G. Schreiber, and M. van
Someren, eds., pp. 283-290, 2001.
[39] A.J.M.M. Weijters, and W.M.P. van der Aalst, “Rediscovering Workflow Models
from Event-Based Data,” Proc. 11th Dutch-Belgian Conf. Machine Learning (Bene-
learn 2001), V. Hoste and G. de Pauw, eds., pp. 93-100, 2001.
[40] A.J.M.M. Weijters and W.M.P. van der Aalst, “Workflow Mining: Discovering Work-
flow Models from Event-Based Data,” Proc. ECAI Workshop Knowledge Discovery
and Spatial Data, C. Dousson, F. Höppner, and R. Quiniou, eds., pp. 78-84, 2002.

168
APPENDIX A: PUBLICATIONS

10.6 Colored Petri Nets for Integrating the Data Perspective in Process Au-
dits

Number 6
Colored Petri Nets for Integrating the Data
Title
Perspective in Process Audits
Appendix 10.6
Primary Related Chapters 5.2
Type Conference Paper
32nd International Conference on
Conference
Conceptual Modeling (ER 2013)
Reference (Werner, 2013)
Acceptance Rate 32 %45)
VHB JQ 2.1 Ranking B45) (7.59)
WKWI Ranking B45)
ERA 2010 A45)
CORE 2013 A45)
Review Procedure Double Blinded
Number of Reviews 3
Authors Michael Werner
Dissertation Points 1.00
Authorship
Overall 100%
Design 100%
Realization 100%
Writing 100%
Status Published
Part of other Dissertations No
[Link]
Link
ter/10.1007%2F978-3-642-41924-9_31

45
Submitted as full paper (14 pages), accepted as short paper (8 pages), presented at the ER 2013 and pub-
lished in the conference proceedings, overall acceptance rate for all papers was 32 %

169
APPENDIX A: PUBLICATIONS

Colored Petri Nets for Integrating the


Data Perspective in Process Audits

MICHAEL WERNER
University of Hamburg, Germany
[Link]@[Link]

Abstract: The complexity of business processes and the data volume of pro-
cessed transactions increase with the ongoing integration of information sys-
tems. Process mining can be used as an innovative approach to derive infor-
mation about business processes by analyzing recorded data from the source
information systems. Although process mining offers novel opportunities to an-
alyze and inspect business processes it is rarely used for audit purposes. The
application of process mining has the potential to significantly improve process
audits if requirements from the application domain are considered adequately.
A common requirement for process audits is the integration of the data per-
spective. We introduce a specification of Colored Petri Nets that enables the
modeling of the data perspective for a specific application domain. Its applica-
tion demonstrates how information from the application domain can be used
to create process models that integrate the data perspective for the purpose of
process audits.
Keywords: Business Process Audits, Petri Nets, Business Process Modeling, Pro-
cess Mining, Business Intelligence, Business Process Intelligence

1 Introduction
The integration of information systems for supporting and automating the operation of
business processes in organizations opens up new ways for data analysis. Business intelli-
gence is an academic field that investigates how data can be used for analysis purposes. It
provides a rich set of analysis methods and tools that are well accepted and applied in a
variety of application domains, but it is rarely used for auditing purposes.
This article deals with the application of Colored Petri Net models that combine the control
flow and data perspective for process mining in the context of process audits. We refer to
the example application domain of financial audits for illustration purposes. The benefit of
this application domain is the fact that the event data which is necessary for the application
of process mining displays structural characteristics that are particularly suitable to be used
for the integration of a data perspective. These characteristics relate to the structure of
financial accounting and are independent from the used source system. The objective of
this article is to illustrate an approach for the integration of the data perspective into pro-
cess models in the context of process mining and process audits. This approach is not re-
stricted to the illustrated example application domain but can be applied in a variety of
application scenarios where information about the involved data is available and valuable
for process mining purposes.

170
APPENDIX A: PUBLICATIONS

2 Related Research
Process Mining is a research area that emerged in the late 1990s. Tiwari et al. provide a
good overview of the state-of-the-art in process mining until 2008 [1] and Van der Aalst [2]
provides a comprehensive summary of basic and advanced mining concepts that have been
researched during the last decade. Jans et al. investigate the application of process mining
for auditing purposes [3–5]. They provide an interesting case study [6] for compliance
checking by using the Fuzzy Miner [7] implemented in the ProM software framework [8].
The study shows how significant information can be derived from using process mining
methods for internal audits. The research results are derived from analyzing the control
flow of discovered process models and the interaction of users. We are not aware of any
implementations or case studies that consider the data perspective in the context of pro-
cess mining for process audits. One of the reasons may be the fact that the data perspective
in process mining has generally not been investigated extensively yet in the academic com-
munity [9, 10]. Exceptions are the research results published by Accorsi and Wonnemann
[11] and de Leoni and van der Aalst [10]. The research presented by Accorsi and Wonne-
mann is motivated by finding control mechanisms to identify information leaks in process
models. The authors introduce information flow nets (IFnets) as a meta-model based on
Colored Petri Nets that are able to model information flows. Instead of using tokens exclu-
sively for the modeling of the control flow colored tokens are used to represent data items
that are manipulated during the process execution. De Leoni and van der Aalst use a differ-
ent approach. Their intention is to incorporate a data perspective for analyzing why a cer-
tain path in a process model is taken for an individual case. The modeled data objects in-
fluence the course of routing. They introduce Petri Nets with data (DPN-nets) that base on
Petri Nets but that are extended by a set of data variables that are modeled as graphical
components in the DPN-nets.

3 Application Domain and Requirements


Financial information is published to inform stakeholders about the financial performance
of a company. The published information is prepared based on data that is recorded by
information systems in the course of transaction processing. Public accountants audit fi-
nancial statements for ensuring that the financial information is prepared according to rel-
evant rules and regulations. The understanding of business processes plays a significant
role in financial audits [12]. The rationale of considering business processes is the assump-
tion that well controlled business processes lead to complete and correct recording of en-
tries to the financial accounts. Auditors traditionally collect information about business
processes manually by performing interviews and inspecting available process documenta-
tion. These procedures are extremely time-consuming and error-prone [13]. Process min-
ing allows an effective and efficient reconstruction of reliable process models. Its applica-
tion would significantly improve the efficiency and effectiveness of financial process audits
[14].
Most mining algorithms focus on the reconstruction of control flows in process models that
determine the relation and sequence of process activities [9, 10]. Information about control
flows is important for auditors to understand the structure of a business process. But this

171
APPENDIX A: PUBLICATIONS

information alone is not sufficient from an audit perspective because an auditor addition-
ally needs to understand how the business processes relate to the entries in the financial
accounts [15, 16]. It is therefore necessary to receive information on how the execution of
activities in a business process relate to recorded financial entries. This can be achieved by
incorporating the data perspective. The relationship between transactions, journal entries
and financial accounts is illustrated in Figure 1.
Fig. 1. Accounting Structure Entity-Relationship-Model

Transaction 1 0...N Posting Document 1 2...N Journal Entry Item 0...N 1 Financial Account
creates contains posted on
TransactionCode DocumentNr DocumentNr AccountNr
UserName PositionNr AccountType
PostingDate AccountNr Balance
TransactionCode 0...N 0...1 Amount
is cleared
PostingText CreditOrDebit
ClearingDocNo

A second important concept in financial process audits is materiality [15, 16]. Auditors just
inspect those transactions that could have a material effect on the financial statements. To
be able to identify which business processes are material information is needed about the
amounts that were posted on the different accounts. It is therefore also necessary to model
the value of the posted journal entries in the produced process models.

4 Integrating the Data Perspective


The relevant data objects in financial audits are journal entries. They are created during the
execution of a business process but their values do not influence the course of routing.
They can be interpreted as passive information objects and we therefore refer to the ap-
proach used by Accorsi and Wonnemann [11] for integrating them into the process models.
Petri Net places and tokens are normally used in process mining to model the control flow.
The general approach for integrating the data perspective is to model data objects as col-
ored tokens that are stored in specific places.

be expressed as a tuple CPN = ( , , , Σ, , , , , ) [17]. Table 1 presents the formal def-


The integration is illustrated in the following specification. A Colored Petri Net can formally

inition of each tuple element, the used net components, and their meaning when applied
in the context of this paper. For ease of reference we refer to this type of nets as Financial
Petri Nets (FPN) for the remainder of this paper.

172
APPENDIX A: PUBLICATIONS

Table 1. Specification of Colored Petri Nets for Process Mining in Financial Audits
is a finite set of transitions
The transitions represent the activities that were executed in the process. They dis-
Transactioncode

play the name of the activity. Further information, for example the transaction code
Activity Name

name, can be added.


is a finite set of places
Places in the FPN represent financial accounts and control places.
Control Places Control places determine the control flow in a process model. For
every process model one source place is modeled that connects to the
start transactions. The sink place marks the termination of the pro-
cess. The control places between the start and end transition deter-
Source Sequence Sink
Place Place Place
mine the execution sequence of the process model. A control place
belongs to the set of control places CP.
Account Places
Account
debit side of a Account
credit side of The account places represent financial accounts
balance sheet balance sheet that are affected by the execution of activities in a
account account process. The symbol color indicates the meaning of
Account
credit side of a Account
debit side of a an account. Account places belong to the set of ac-
profit and loss profit and loss
count places AP.
∈ × ∪ × is a set of arcs also called flow relation
account account

Control arcs connect control places with transitions. They model the control flow
Control in the model.
Arc
Posting arcs illustrate the relationship between activities represented as transi-
Posting tions in the model and financial accounts that are modeled as account places.
Arc
Clearing arcs are used to model that an activity cleared an entry on the corre-

tical abbreviation for two arcs (p,t) and (t,p).


Clearing sponding account. Clearing arcs are double-headed arcs and are used as an syntac-
Arc
is a set of non-empty color sets
The color set contains the possible values that are posted or
colset VALUES = double
colset ACCOUNTS = string
cleared.
The color set contains all account numbers.
colset ACOUNTTYPE = boolean
The color set contains {1,0} indicating if the represented ac-
count is a balance sheet or a profit and loss account.
colset CREDorDEB = boolean
The color contains {1,0} indicating if the account place is a
representation of the debit or credit side of an account.
colset EXECUTIONS = int
The color contains {1,…,n} indicating how often a path was

ACCOUNTPLACES is the color set as a product of VALUES *


chosen in the FPN.

colset ACCOUNTPLACES ACCOUNTS * ACCOUNTTYPE * CREDorDEB


* is a finite set of typed variables such that Type[,] ∈ for all variables , ∈ *.

and = {}.
In FPN models arc inscriptions are modeled as constants. Variables are therefore not necessary

173
APPENDIX A: PUBLICATIONS

1: → is a color set function that assigns a color set to each place.


678 9 : ;< 4 ∈
(4) 5
= 7 68: ;< 4 ∈
The color set function in FPN assigns dif-
ferent color sets to places depending if
they belong to the group of control or
account places:
>: → ?@ A* is a guard function that assigns a guard to each transition B such that
CDE [>(B)] = FGGHEIJ.
= K is the set of expressions that is provided by the used inscription language. FPN do not ex-
plicitly include guards because they do not model dynamic behavior of transitions that depends

for FPN is therefore defined as (L) = LMNO <PM QRR L ∈ .


on specific input but illustrate the processing of already executed processes. The guard function

?: → ?@ A* is an arc expression function that assigns an arc expression to each arc I such
that CDE [?(I)] = 1(D)ST , where p is the place connected to the arc I.
The arc expressions in a FPN are constants. The arc expression function assigns to each posting
and clearing arc a set of constants that denote the posted or cleared value, the account num-
ber, account type and an indicator if it is a credit or debit posting. For each control flow arc the
number of execution times is assigned indicating how often this path was chosen in the process

{VQR ∈ 97 :, QWW ∈ 678 :, QWWLX4 ∈ 678 Y , WPMZ ∈ K [PM[ \}


model.

];Lℎ X4O [ (Q)] = (4)_` = 678 9 : ;< 4 ∈


(Q) U
{Oa ∈ = 7 68: ];Lℎ X4O [ (Q)] = (4)_` = = 7 68: ;< 4 ∈

b: → ?@ A∅ is an initialization function that assigns an initialization expression to each


place D such that CDE [b(D)] = 1(D)ST

ef Oa ∈ = 7 68: ];Lℎ X4O [ (4)] = (4)_` = = 7 68: ;< 4 = gPNMWO


The initialization function of a FPN assigns initialization expressions to each place as follows:

(4) d
∅_` PLℎOM];gO
Only the source place is initialized in a FPN. The initialization expression for 4 = gPNMWO gener-
ates e tokens in the initial marking hi (4), one for each connected start transition. The inscrip-
tion of each token is a member of the set = 7 68:.

Figure 2 illustrates a FPN model for a purchasing process. The model includes:
= {h\01, h K6, }110}
= {Source, S1, S2, Sink, 100_D, 200_D, 200_C, 300_D, 300_C, 400_D}
Transitions:

⊆ = {Source, S1, S2, Sink}


Places:

⊆ = {100_D, 200_D, 200_C, 300_D, 300_C, 400_D}.


VALUES = {50,000}, ACCOUNTS {100, 200, 300, 400}, ACOUNTTYPE =
{0,1}, CREDorDEB = {0,1}; EXECUTIONS = {1}; ACCOUNTPLACES =
Color sets:

VALUES * ACCOUNTS * ACCOUNTTYPE * CREDorDEB.


Initialization: (4) = {1f 1 <PM 4 = gPNMWO QeZ ∅_` <PM 4 ≠ gPNMWO

The model shows that the processing of received goods created journal entries with the
amount of 50,000 on the raw materials and the goods receipt / invoices receipt (GR/IR)

174
APPENDIX A: PUBLICATIONS

account. The receipt of the corresponding invoice for the purchased goods was processed
with activity MIRO which led to journal entries on the GR/IR account and a creditor account.
It also cleared the open debit item on the GR/IR account that was posted by MB01. The
received invoice was finally paid by executing the activity F110 which posted a clearing item
on the creditor account and a debit entry on the bank account.
100_D 200_C 200_D 400_D
Raw Materials Good Receipt / Invoices Receipt Good Receipt / Invoices Receipt Bank Account

[50,000] [50,000]
[50,000] [50,000] [50,000]

MB01 MIRO F110

[1] Post Goods [1] [1] Enter Incoming [1] [1] Automatic [1]
[1]
Receipt for PO Invoice Payment
Source S1 S2 Sink

[50,000]

[50,000] [50,000]

300_C 300_D
Creditor Account Creditor Account

Fig. 2. Simple FPN Example of a Purchase Process


The example in Figure 2 demonstrates how the used Petri Net specification can be used to
model the control flow and the data perspective simultaneously in a single model. The tran-
sitions in the model create colored tokens when they fire. They store information on the
values that are posted on the connected account places. The illustrated model does not
only mimic the execution behavior of the involved activities but also the creation of journal
entries on the financial accounts. The model shows the execution sequence and the value
flows that are produced.

5 Implementation and Experimental Evaluation


We applied the described FPN specification on real world data to evaluate if the theoretical
constructs can actually be used in real world settings. The data base for the evaluation in-
cluded about one million cases of process executions from a company operating in the
manufacturing industry. The raw data was extracted from a SAP ERP system and checked
for first and second order data defects [18]. We used an adjusted implementation of the
Financial Process Mining (FPM) algorithm [19] that is able to produce FPN models. The min-
ing was limited to 100,000 process instances that affected a specific raw materials account.
The mining resulted in 113 process variants. They were analyzed in the evaluation phase
by observation using the yEd Graph Editor [20] for verifying if the process models presented
the desired information adequately. Selected models were further tested using the Renew
software [21] for evaluating if correct FPN were created by simulating the execution of the
models. The evaluation demonstrated that FPN can be created correctly by using an
adapted FPM algorithm. The produced FPN are able to adequately model the control and
data flow based on the used real world data.
The same modeling procedure was used in a different scenario to analyze technical cus-
tomer service processes which is not described in this paper due to place restrictions.

175
APPENDIX A: PUBLICATIONS

6 Summary and Conclusion


Process Mining is an innovative approach for analyzing business processes but it is rarely
used in the context of process audits. The successful application of process mining requires
the consideration of domain specific requirements. A common requirement is the incorpo-
ration of a data perspective which has not been addressed extensively yet in the academic
community. We have introduced a specification of Colored Petri Nets that allows the mod-
eling of the control flow and data perspective simultaneously. We referred to the applica-
tion domain of financial audits as a representative example to demonstrate how the data
perspective can be included by referring to relevant application domain requirements.
The evaluation of mined process models shows the suitability of the presented specifica-
tion in real world settings. The evaluation included the data from a SAP system of a single
company. It can therefore not be concluded that the results also hold true for other com-
panies or ERP systems. But the presented specification bases on the structure of financial
accounting and is therefore independent from any proprietary ERP software implementa-
tion or industry. Evaluation results from further current research indicate that the pre-
sented results are also applicable in other settings.46

7 References
1. Tiwari, A., Turner, C.J., Majeed, B.: A review of business process mining: state-of-
the-art and future trends. Business Process Management Journal. 14, 5–22 (2008).
2. Van der Aalst, W.M.P.: Process Mining: Discovery, Conformance and Enhancement
of Business Processes. Springer, Berlin Heidelberg (2011).
3. Jans, M., Alles, M., Vasarhelyi, M.: Process mining of event logs in auditing: opportu-
nities and challenges. Working paper. Hasselt University. Belgium (2010).
4. Jans, M.J.: Process Mining in Auditing: From Current Limitations to Future Chal-
lenges. In: Daniel, F., Barkaoui, K., and Dustdar, S. (eds.) Business Process Manage-
ment Workshops. pp. 394–397. Springer, Berlin, Heidelberg (2012).
5. Jans, M., van der Werf, J.M., Lybaert, N., Vanhoof, K.: A business process mining ap-
plication for internal transaction fraud mitigation. Expert Systems with Applications.
38, 13351–13359 (2011).
6. Jans, M., Alles, M., Vasarhelyi, M.: Process Mining of Event Logs in Internal Auditing:
A Case Study. 2nd International Symposium on Accounting Information Systems
(2011).
7. Günther, C., van der Aalst, W.: Fuzzy mining–adaptive process simplification based
on multi-perspective metrics. Business Process Management. 328–343 (2007).
8. Process Mining Group: ProM, [Link]

46
The research results presented in this paper were developed in the research project EMOTEC sponsored
by the German Federal Ministry of Education and Research (grant number 01FL10023). The authors are
responsible for the content of this publication.

176
APPENDIX A: PUBLICATIONS

9. Stocker, T.: Data flow-oriented process mining to support security audits. Service-
Oriented Computing-ICSOC 2011 Workshops. pp. 171–176 (2012).
10. De Leoni, M., van der Aalst, W.M.: Data-Aware Process Mining: Discovering Deci-
sions in Processes Using Alignments. (2013).
11. Accorsi, R., Wonnemann, C.: InDico: Information Flow Analysis of Business Pro-
cesses for Confidentiality Requirements. Security and Trust Management. pp. 194–
209. Springer (2011).
12. International Federation of Accountants: ISA 315 (Revised), Identifying and As-
sessing the Risks of Material Misstatement through Understanding the Entity and Its
Environment, (2012).
13. Werner, M.: Einsatzmöglichkeiten von Process Mining für die Analyse von Ge-
schäftsprozessen im Rahmen der Jahresabschlussprüfung. In: Plate, G. (ed.) For-
schung für die Wirtschaft. pp. 199–214. Cuvillier Verlag, Göttingen (2012).
14. Werner, M., Gehrke, N., Nüttgens, M.: Business Process Mining and Reconstruction
for Financial Audits. Hawaii International Conference on System Sciences. pp. 5350–
5359. , Maui (2012).
15. Schultz, M., Müller-Wickop, N., Nüttgens, M.: Key Information Requirements for
Process Audits – an Expert [Link]. Proceedings of the 5th International
Workshop on Enterprise Modelling and Information Systems Architectures. , Vienna
(2012).
16. Müller-Wickop, N., Schultz, M., Peris, M.: Towards Key Concepts for Process Audits -
A Multi-Method Research Approach. Proceedings of the 10th International Confer-
ence on Enterprise Systems, Accounting and Logistics. , Utrecht (2013).
17. Jensen, K., Kristensen, L.M.: Coloured petri nets. Springer (2009).
18. Kemper, H.-G., Mehanna, W., Baars, H.: Business intelligence - Grundlagen und
praktische Anwendungen : eine Einführung in die IT-basierte Managementunter-
stützung. Vieweg + Teubner, Wiesbaden (2010).
19. Gehrke, N., Müller-Wickop, N.: Basic Principles of Financial Process Mining A Jour-
ney through Financial Data in Accounting Information Systems. Proceedings of the
16th Americas Conference on Information Systems. , Lima, Peru (2010).
20. yWorks GmbH: yEd - Graph Editor, [Link]
ucts_yed_about.html.
21. University of Hamburg: Renew - The Reference Net Workshop, [Link]
[Link]/.

177
APPENDIX A: PUBLICATIONS

10.7 Tackling Complexity: Process Reconstruction and Graph Transfor-


mation for Financial Audits

Number 7
Tackling Complexity: Process Reconstruction
Title and Graph Transformation for Financial Au-
dits
Appendix 10.7
Primary Related Chapters 5.3
Type Conference Paper
33rd International Conference on
Conference
Information Systems (ICIS 2012)
Reference (Werner et al., 2012b)
Acceptance Rate 29 % 47)
VHB JQ 2.1 Ranking A47) (8.48)
WKWI Ranking A47)
ERA 2010 A47)
CORE 2013 A*47)
Review Procedure Double Blinded
Number of Reviews 5
1. Michael Werner
2. Martin Schultz
Authors 3. Niels Müller-Wickop
4. Nick Gehrke
5. Markus Nüttgens
Dissertation Points 0.33
Authorship
Overall 80%
Design 80%
Realization 80%
Writing 80%
Status Published
Part of other Dissertations No
[Link]
Link
ticle=1230&context=icis2012

47
Submitted an accepted as Research-in-Progress paper (12 pages), presented at the ICIS 2014 and pub-
lished in the conference proceedings

178
APPENDIX A: PUBLICATIONS

Tackling Complexity: Process Reconstruction and Graph Transformation for Fi-


nancial Audits
Research-in-Progress

MICHAEL WERNER MARTIN SCHULTZ


University of Hamburg University of Hamburg
Chair for Information Systems Chair for Information Systems
Max-Brauer-Allee 60 Max-Brauer-Allee 60
D-22765 Hamburg D-22765 Hamburg
[Link]@[Link] [Link]@[Link]

NIELS MÜLLER-WICKOP NICK GEHRKE


University of Hamburg Nordakademie
Chair for Information Systems Chair for Information Systems
Max-Brauer-Allee 60 Köllner Chaussee 11
D-22765 Hamburg D-25337 Elmshorn
[Link]-wickop@[Link] [Link]@[Link]

MARKUS NÜTTGENS
University of Hamburg
Chair for Information Systems
Max-Brauer-Allee 60
D-22765 Hamburg
[Link]@[Link]

Abstract: A key objective of implementing business intelligence tools and meth-


ods is to analyze voluminous data and to derive information that would other-
wise not be available. Although the overall significance of business intelligence
has increased with the general growth of processed and available data it is al-
most absent in the auditing industry. Public accountants face the challenge to
provide an opinion on financial statements that are based on the data produced
by the automated processing of countless business transactions in ERP systems.
Methods for mining and reconstructing financially relevant process instances
can be used as a data analysis tool in the specific context of auditing. In this
article we introduce and evaluate an algorithm that effectively reduces the
complexity of mined process instances. The presented methods provide a part
of the foundation for implementing automated analysis and audit procedures
that can assist auditors to perform more efficient and effective audits.
Keywords: business process modeling, data mining, process mining, business
intelligence, data analysis, enterprise resource planning systems, financial au-
dits

179
APPENDIX A: PUBLICATIONS

1 Introduction
Enterprise resource planning (ERP) systems are key components for supporting and auto-
mating the processing of business transactions in modern companies. While ERP systems
are primarily used for supporting and automating business processes they commonly also
provide the functionality to prepare the financial statements that companies are required
to publish. The provided information plays a critical role in the economic system. It enables
stakeholders to acquire information about the financial situation of the entity they are in-
terested in. Due to their informative significance governments and regulatory institutions
have issued laws and regulations that intend to safeguard the correctness of published fi-
nancial statements. They have entrusted public accountants to carry out audits for ensuring
that accounting standards are adhered to and that the published information is free of ma-
terial misstatements. But when auditors perform their audits they encounter a significant
problem. While on the one hand business transactions are processed automatically in ERP
systems auditors on the other hand apply mainly manual audit procedures to achieve their
audit comfort. With increasing integration of ERP systems and rising numbers of processed
transactions manual audit procedures become inefficient and ineffective.
This situation is quite astonishing as the precondition for implementing automated audit
procedures is given in every case where the mere number of processed transactions makes
manual audit procedures inefficient or even ineffective. Whenever this is the case systems
must be involved that support or automate the processing. Otherwise the number of trans-
actions could still be effectively audited with manual audit procedures. While processing
ERP systems produce data that is stored in the systems’ databases. The stored data includes
the journal entries and other logging information that can be used to reconstruct relation-
ships between the stored data entries and their corresponding originally executed transac-
tions. The content of the stored financially relevant data in ERP systems is generally suitable
for automated process mining and analysis purposes. It is clearly structured and needs to
comply with accounting requirements like completeness and accuracy. It is possible to cap-
italize on the available structures by using purposeful mining and analysis methods.
Gehrke and Müller-Wickop (2010a, 2010b) present an algorithm for mining business pro-
cesses from journal entries stored in ERP databases. The approach is similar to mining pro-
cess models from event logs (van der Aalst 2011) but applied to financially relevant busi-
ness processes. Werner et al. (2012) show how these process mining methods can be com-
bined with automated testing of controls that are embedded in ERP systems. The combi-
nation generally allows implementing system based and automated audit procedures that
are efficient and effective to audit highly integrated and automated financially relevant
business processes. This way the imbalance between automated processing on the compa-
nies’ side and manual audit procedures on the auditors’ side can be overcome.
The application of these methods can be seen as a business intelligence tool that enables
especially auditors to receive information from available mass data that they otherwise do
not have access to. Their application would enable the auditor to efficiently gather infor-
mation for understanding the relationship between the business processes and the finan-
cial statements for the audited entity. They would also allow the identification of unusual
business transactions that normally inherit higher risk than standard transactions (Werner
and Gehrke 2011).

180
APPENDIX A: PUBLICATIONS

In this paper we focus on the results and issues that arise from analyzing mined process
instances. We therefore address a specific sub-problem on the path towards developing
automated analysis and audit methods for financial audits. Analyzing mined processes from
real life data shows that process instances do exist that contain up to tens of thousands
executed transactions leading to very complex graphs consisting of hundreds of thousands
elements. The complexity of these graphs does not allow sophisticated interpretation for
the purpose of auditing. It is necessary to find mechanisms to reduce their complexity.
In this article we illustrate how the complexity of mined process instances can be reduced
by using a graph transformation system. We start with an overview of related work and
chosen research methodology. We continue with an illustration of how mined process in-
stances can be represented as Petri nets by using an illustrative example of a mined process
instance. Based on this representation we introduce an aggregation algorithm that oper-
ates as a graph transformation system by aggregating net elements within a process in-
stance. We graphically show the aggregation results for the used example and provide eval-
uation results that were derived from applying the aggregation algorithm to test and real
life data. The paper closes with a summary and conclusion of the presented methods and
derived results.
We focus our attention on financial audits in order to stay in reasonable limits but we like
to point out that the application of the discussed process mining, reconstruction and graph
transformation methods is not restricted to financial audits but is also relevant for perfor-
mance and optimization considerations (Werner and Gehrke 2011).

2 Related Work and Research Methodology


The concept of process mining forms the foundation of the research presented in this pa-
per. Process mining first evolved in the 1990s for investigating process mining in the con-
text of software engineering (Cook and Wolf 1998a, 1998b, 1999). The idea of applying
process mining to workflow logs was first introduced by Agrawal et al. (1998). Maxeiner et
al. (2001) and Schimm (2000; 2001a; 2001b; 2002) developed mining tools whereas Herbst
and Karagiannis (Herbst 2000a; 2000b; 2003; Herbst and Karagiannis 1998; 1999) ad-
dressed process mining in the context of workflow management. Substantial research work
exists for mining and rediscovering process models from event logs (van der Aalst 1997,
1998, 2005, 2011; van der Aalst et al. 2002a, 2002b; van der Aalst and Dongen 2002; van
der Aalst and Weijters 2002; Maruster et al. 2001a, 2001b; Weijters and van der Aalst
2001). Their work covers a variety of aspects like delta analysis, conformance, concurrency,
workflow verification and performance. Gehrke and Müller-Wickop (2010a, 2010b) build a
bridge between process mining and financial accounting by applying process mining to re-
construct financially relevant business processes.
A traditional application area of business intelligence is the support for decision making
processes (Turban et al. 2007; Vercellis 2009). Anandarajan et al. (2004) illustrate business
intelligence techniques from an accounting and finance perspective but we are not aware
of publications that address business intelligence tools from an explicit auditing perspec-
tive. For developing an aggregation method we refer to the theory of graph grammars and
transformation (Heckel 2006; Rozenberg 1997).

181
APPENDIX A: PUBLICATIONS

The research presented in this paper follows a design science approach (March and Smith
1995; Österle et al. 2010; Hevner et al. 2004). For the purpose of developing adequate
complexity reduction methods we implemented a software artifact. The artifact allows re-
constructing and observing process instances. Based on the observation derived from ex-
tensive real life data we engineered formal methods (Brinkkemper 1996) and finally evalu-
ated them against test and real life data.

3 Automated Reconstruction and Petri Net Representation


The mined processes presented by Gehrke and Müller-Wickop (2010a, 2010b) use a more
or less informal presentation that focusses on illustrating essential process components.
Müller-Wickop et al. (2011) introduce a Business Process Model and Notation (BPMN)
based representation which primarily aims to integrate the process and financial perspec-
tive by including graph elements that capture financially relevant aspects of modeled pro-
cess instances. In order to be able to develop comprehensive methods for complexity re-
duction we need a formally robust presentation. We chose a Petri net representation due
to the fact that Petri nets constitute a formally sound and mathematically powerful mod-
eling language (Valk 2008). Modeling mined process instances as colored Petri nets makes
it possible to apply already existing research for event log mining (van der Aalst 2011).
Figure 1 shows a reconstructed instance of a purchasing process that was mined from an
ERP system by using the mining algorithm developed by (Gehrke and Müller-Wickop
2010b). This means that a purchasing business transaction was executed and recorded in
the examined ERP system. The relevant data was extracted and the process instance was
reconstructed. The illustrated example represents an executable colored Petri net that
mimics the behavior of the originally processed instance. A colored Petri net can formally
be expressed as a tuple N = (P, T, F, C, cd, W, m0). The specific meaning of each tuple ele-
ment for the constructed instance is described in Table 1 and can be illustrated by referring
to Figure 1. The transitions of the Petri net shown in Figure 1 represent executed transac-
tions that created journal entries in the ERP system. MB01 are transactions that recorded
the receipt of ordered material, MR1M processed the received invoices and F110 executed
the payment run. FB1S represent automatic clearing transactions. The set of transition T
for this Petri net contains all the transitions that represent the executed transactions
{MB01_5000004383, MB01_5000004384, …}. Journal entry items are modeled as places
with inscriptions identifying the account the corresponding item was posted to. The differ-
ent colors illustrate whether the item was a debit or credit posting. The created-relation-
ships between items and transactions are illustrated as dotted arrows. The implemented
mining algorithm exploits the open-item-accounting structure of journal entries. Journal
entries with enabled open-item-accounting are either cleared or not cleared. When a pro-
cess has been successfully terminated all open items are cleared. Otherwise the process
has not been completely executed. A clearing document exists for each cleared item. By
following this connection between journal entries the whole process instance can be re-
constructed. The cleared-relationships between transactions and items are illustrated as
dotted arcs. The inscriptions on the arcs represent the amounts that were posted or cleared
on the accounts the items were placed on. The modeling of start places marked with tokens
representing the document number that was created by the corresponding transaction en-
sures that each transition in the Petri net model can actually be executed. The flow relation

182
APPENDIX A: PUBLICATIONS

F for this Petri net contains all arcs between the transitions and places. The set of places
includes all illustrated places (start places and places for journal entry items). The set of
colors includes the set of account numbers {310000, 191100, …}, indicators for credit or
debit posting {credit, debit} and document numbers {5000004383, 5000004384, …}. The
function cd maps the colors to the individual places. The function W maps the arc inscrip-
tions that represent the booking values {13907.17, 10880.29, …} to the arcs for relations
between transitions and places representing journal entry items or the document numbers
to the arcs between start places and transitions.

Figure 1. Simple Process Instance A

Table 1. Formal Petri Net Representation

Tuple elements Specific meaning in the representation context


P Set of places Start places and places representing journal entry items
T Set of transitions Transactions executed in the ERP system
F Flow relation Arcs between the nodes indicating their type of relationship
C Set of colors Set of place characteristics, posting values, document numbers
cd Color domain mapping Mapping of characteristics to places
W Arc inscription Mapping of posting values or document numbers to arcs
Tokens marking start places with the corresponding document
m0 Initial marking
number

Example A contains twelve transitions. The instance is clearly interpretable by observation.


It changes for more complex process instances. Example B illustrated in Figure 2 shows a
mined process instance like example A and consists of the same Petri net elements as de-
scribed in Table 1 but it contains 358 transitions. It is still a small instance compared to the
most complex instances in our sample database that included up to 27,177 transitions. Ta-
ble 2 provides an overview of the characteristics of both examples. When looking at Figure

183
APPENDIX A: PUBLICATIONS

2 it is obvious that the interpretation and analysis of reconstructed instances by simple


observation becomes impossible when they include more than a few transactions.

Table 2. Net Characteristics for


Process Instance Examples
Example Instance A B
Number of transitions 12 358
Number of places 43 1804
Number of arcs 60 2549
Sum of net elements 115 4711

Figure 2. Complex Process Instance B

4 Graph Transformation for Complexity Reduction


The complexity of a graph can be defined differently, depending on the relevant perspec-

sets G = (V,E) where V is the set of vertices and E is the set of edges with E ⊆ V2 (Diestel
tive und purpose (Neel and Orrison 2006). A graph is generally defined as a pair of disjoint

2010). For the purpose of this paper we consider the complexity of the graph simply as a
function of its number of edges and vertices. The mined process instances are modeled as
Petri nets. We therefore use the term net elements for the sum of vertices and edges that
determine the complexity of the Petri net under review. A promising approach for reducing
complexity of voluminous process instances is to aggregate similar net elements and
thereby reducing their total number. A key requirement for any type of transformation is
that financially relevant information remains unchanged compared to the originally recon-
structed process. In the context of auditing it is crucial to be able to trace any transaction
from the point of origin to the final ledger posting. The path that allows tracing a transac-
tion through an information systems is called audit trail (Romney and Steinbart 2008 p.
687). In the context of process mining for audit purposes this means that the analyzed data
has to mirror exactly the transactions that were actually executed and that no existing
paths in the graph may be deleted or new ones be added. Müller-Wickop et al. (2011) point
out that the behavior of the process instances may not be changed by any aggregation
method. If an aggregation algorithm changed the behavior, this fundamental requirement
would be violated. The behavior of a Petri net can be expressed as the set of possible firing
sequences. For designing an aggregation method it is necessary to ensure that the set of
firing sequences remains unchanged.
Requirement I: The set of firing sequences has to stay constant
Different arc types in the Petri net mean that transactions interact differently with journal
entry items. If the arc types between transitions and places were altered the resulting
graph would no longer reflect the original relationship between these elements. The arc

184
APPENDIX A: PUBLICATIONS

types and therefore the type of interaction between places and transitions may not be al-
tered.
Requirement II: Different arc types may not be merged
The methods of financial process mining by Gehrke and Müller-Wickop (2010b) have been
developed to exploit the structure of accounting data and to capture financially relevant
information. This information primarily concerns the value flow within the processes. Ana-
lyzing this information actually allows to identify which amounts have been posted on the
different accounts by a process instance and to evaluate how significant they are from a
materiality perspective.
Requirement III: The arc inscriptions representing the value of posted journal
entries have to be preserved
We reconstructed and analyzed processes instances from different data sets originating
from two companies that operate in the retail and manufacturing industries and from the
SAP IDES test system (SAP 2012). Figure 3 provides an overview of the frequency distribu-
tion of approximately 40,000 mined process instances originating from a corporation in the
retail industry with a logarithmic scaling on the x- and y-axes. The mean value of the distri-
bution is 2.30, the standard deviation is 3.37, the median value is 2 and the maximum value
is 596. The frequency distributions for the other data sets show the same pattern with
slightly different values.

100000 Median
10000 Mean Maximum
Number of
Instances

1000
100
10
1
1 10 100 596.00 1000
2.00 2.30 Number of Accounts

Figure 3. Frequency Distribution of Instances over Accounts


Figure 3 shows that the number of different accounts in process instances is relatively small
compared to the instance size. This observation is comprehensible when considering how
business processes generally affect accounts. The execution of a certain business process
normally produces journal entries only on a very limited subset of the overall available ac-
counts. These are the accounts related to that business process. The execution of a stand-
ard procurement process would most likely lead to journal entries on the expense accounts
but not on the sales accounts. Although this observation cannot be generalized without
further research it is reasonable to assume that it is true also for other industries not cov-
ered by the analyzed data sets.
In the Petri net models journal entry items are represented as places carrying an inscription
denoting the account they were posted to. Due to the fact that only few accounts are used
within the same process instance it is reasonable to assume that the aggregation of places
that represent journal entry items on the same accounts might significantly reduce the
overall number of net elements. When designing an algorithm for aggregating places re-
quirements I to III have to be considered. The following cases listed in Table 3 illustrate the

185
APPENDIX A: PUBLICATIONS

different constellations that can occur when two places representing journal entry items
are to be aggregated. Places are only considered for aggregation if they carry the same
account number and credit or debit flag. The case description contains the possible con-
stellations according to the net definition used in this paper and shows the cases for arcs
directed from transitions to places (incoming arcs). These constellations also have to be
considered for arcs directed from places to transitions (outgoing arcs). Designing an algo-
rithm based on these cases ensures that the set of firing sequences is not changed because
the relationship between the transactions is preserved as well as the different arc types
connecting the places and transitions. By maintaining inscriptions from merged arcs we
ensure that the financially relevant information on posting values does not get lost.

Table 3. Case Distinction for Place Aggregation


Input Description Output
P1 and P2 are connected to the
same transition T1. The arc types of
both arcs are equal. The arcs (T1,P1)
Case 1
and (T1,P2) can be aggregated. The
inscriptions [b] is added to [a] result-
ing in [a];[b]. Place P2 is deleted.
P1 and P2 are connected to the
same transition T1. But the arc types
are different. The arcs cannot be ag-
Case 2
gregated. Arc (T1,P2) has to be redi-
rected to P1 resulting in (T1,P1).
Place P2 is deleted.
P1 and P2 are connected to different
transitions T1 and T2. The arc types
of both arcs are equal. The arcs
(T1,P1) and (T2,P2) cannot be aggre-
gated in order to preserve the infor-
Case 3
mation that the represented items
were actually created (or cleared) by
different transactions. Arc (T2,P2)
has to be redirected to P1 resulting
in (T2,P1). Place P2 is deleted.
P1 and P2 are connected to different
transitions T1 and T2. The arc types
of both arcs are different. The arcs
(T1,P1) and (T2,P2) cannot be aggre-
Case 4 gated. The arc types are different
and they originate from different
transitions. Arc (T2,P2) has to be re-
directed to P1 resulting in (T2,P1).
Place P2 is deleted.

The following algorithm in Listing 1 implements the aggregation of places according to the
description included in Table 3.

186
APPENDIX A: PUBLICATIONS

Listing 1. Aggregation Algorithm


PItem ⊆ P
PItemAgg= ∅
set of all places representing journal entry items in the net
initially empty set for aggregated places
FArcs set of all arcs in the net

While PItem≠ ∅
Aggregate Places

Take pi ∈ PItem
Select all pj ∈ PItem with cd(pi)=cd(pj)
Merge arcs for each pi and pj
Add pi to PItemAgg
Remove pi and pj from PItem
Set PItem=PItemAgg

arcs FIncArcsI ⊆ FArcs for pi


Merge Arcs

arcs FIncArcsJ ⊆ FArcs for pj


Get incoming

ai(tm,pi) ∈ FIncArcsI and arc aj(tn,pj) ∈ FIncArcsJ


Get incoming
For each arc
If the arc type of ai = arc type of aj
And if tm = tn then add W(aj) to W(ai)and /Case 1
remove aj from FArcs
Else set aj(tm,pj) /Case 2,3 and 4
Get outgoing arcs FOutArcsI ⊆ FArcs for pi
Get outgoing arcs FOutArcsJ ⊆ FArcs for pj
For each arc ai(pi,tm) ∈ FOutArcsI and arc aj(pj,tn) ∈ FOutArcsJ
If the arc type of ai = arc type of aj
And if tm = tn then add W(aj) to W(ai)and /Case 1
remove aj from FArcs
Else set aj(pj,tm) /Case 2,3 and 4

The algorithm represents a graph transformation production that iteratively substitutes


sub-graphs consisting of two nodes and an arbitrary number of arcs into sub-graphs con-
sisting of one node and the same or smaller set of arcs.

5 Evaluation
The application of the aggregation algorithm on example instance A leads to the graph il-
lustrated in Figure 4. The number of net elements was reduced from 115 to 79. We used
the software Renew (University of Hamburg 2012) for verification purposes. Renew allows
to simulate the execution of Petri nets. The testing of samples showed that the algorithm
works properly and generates fully reachable Petri nets. We applied the algorithm to two
different data sets. Data set 1 originates from the SAP IDES database. The database is avail-
able for universities participating in the SAP University Alliance Program (SAP 2012). The
database contains over 115,000 journal entries from approximately 81,000 process in-
stances. We chose the SAP IDES database to enable interested readers to reproduce our
results. Data set 2 originates from a corporation operating in the retail industry. The used
database contains approximately 90,000 journal entries constituting about 40,000 process
instances for a period of one year. Table 4 illustrates the characteristics of the distribution
of instance frequency over the number of net elements before and after applying the ag-
gregation algorithm. The mean value for the number of net elements per instance is re-
duced by 23.2 % from 15.31 to 11.76 for the SAP IDES data set and by 23.4 % from 21.18 to
16.22 in data set 2. We achieve an average complexity reduction of approximately a quar-

187
APPENDIX A: PUBLICATIONS

ter. The complexity reduction is more effective for complex processes with many net ele-
ments. We therefore analyzed the effectiveness of the algorithm by selecting a subset from
the available data set 2 containing process instances with 100 or more net elements. The
mean value of the original subset was 929.48 net elements per instance. The mean value
of the aggregated instances was reduced by 44% to 520.

Figure 4. Aggregated Process Instance A


Table 4. Aggregation Results
Data Set 1 SAP IDES Data Set 2 Retail Company
original aggregated original aggregated
Mean value of net elements per instance 15.31 11.76 21.18 16.22
Median value of net elements per instance 9.00 7.00 9.00 9.00
Maximum of net elements per instance 32,519 18,602 275,870 146,620
Standard deviation of number of net elements 127.83 73.33 1,380.57 743.81

6 Conclusion
Public auditors face the challenge of auditing financial statements that base on the data
generated by automated transaction processing in ERP systems. While business intelli-
gence techniques have introduced new ways of analyzing data in the general corporate
context such techniques are missing in the auditing industry. The mining and reconstruc-
tion of financially relevant processes provides the basis for analyzing data from ERP systems
for the purpose of financial audits. We have presented an aggregation algorithm in this
paper that operates on mined process instances modeled as Petri nets that can be applied
to any process instance mined with the used mining algorithm. It aggregates places within
the instances and thereby reduces the number of net elements in the graph. The results
are less complex models. The evaluation of the algorithm on the basis of test and real life
data shows that significant complexity reductions can be achieved. The reduction effect is
higher for process instances encompassing many net elements. With the application of the
presented mining methods and representation form it is possible to derive and visualize
information about executed processes and their effect on the financial accounts that the
public accountant has to audit. They provide a means to overcome the imbalance between
mainly manual audit procedures on the auditor side and the automated and system based
processing on the company side. The usage of these methods could make financial audits

188
APPENDIX A: PUBLICATIONS

more efficient and set free resources for the evaluation of unusual transactions that com-
monly constitute higher risks than standard transactions.
When using the methods several limitations should be taken into account. The availability
of necessary data is the first. The data needs to be extracted from the ERP system, trans-
formed into a data format that can be processed by the mining algorithm and loaded into
a database where it can be accessed for mining purposes. This procedure is called ETL (ex-
tract, transform, load) process and requires efficient extraction tools that are able to han-
dle voluminous data. When analyzing the data confidentiality and privacy aspects need to
be considered. The use of pseudonymization procedures in the ETL process might address
this restriction. A second limitation derives from the current scope of application for the
mining algorithm. The algorithm for mining financially relevant processes is only applicable
for transactions that affect open-item-operated accounts. A purchase requisition sub-pro-
cess for example does not directly affect open-item-operated accounts but might be rele-
vant for financial audits. Further development for incorporating financially relevant pro-
cesses and sub-processes that do not affect open-item-operated accounts might signifi-
cantly enlarge the scope of application.
The presented methods provide a new approach of analyzing financial data in ERP systems.
The presented algorithm achieves significant complexity reductions. But although the num-
ber of net elements for the most complex process instance on our sampled data was re-
duced by almost half from 275,870 to 146,620 the aggregated instance is still too complex
for meaningful interpretation and evaluation. First observations show that similar struc-
tures among sub-graphs within instances might be common. Frequent sub-graph mining is
a well-studied data mining problem (Huan et al. 2003; Yan and Han 2002) The development
of aggregation procedures that merge sub-graphs within an instance by considering exist-
ing sub-graph mining algorithms might be a promising approach for further complexity re-
duction. Another obstacle to efficient evaluation is the huge amount of process instances
that need to be considered. The analysis of mined process instances leads to the assump-
tion that a great amount of process instances are very similar in structure. Further research
is needed to investigate if efficient procedures that create clusters or categories across dif-
ferent process instances can be developed.
We have shown how a process mining approach can be used as a data analysis tool espe-
cially in the context of financial audits and how the complexity of mined instances can be
reduced effectively. Further research is needed to answer open questions and to overcome
existing limitations. But the methods discussed in this paper provide a meaningful mile-
stone on the way for designing system based and automated analysis and audit procedures
that cannot only be used in the context of auditing but also in the wider business context
for example for performance measurement or optimization purposes.

7 Acknowledgement
The research results presented in this paper were developed in the research project Virtual
Accounting Worlds. The project is sponsored by the German Federal Ministry of Education
and Research. The authors are responsible for the content of this publication (grant num-
ber 01IS10041).

189
APPENDIX A: PUBLICATIONS

8 References
van der Aalst, W. M. P. 1997. “Verification of Workflow Nets,” In Application and Theory
of Petri Nets, P. Azema and G. Balbo (eds.), Berlin: Springer-Verlag, pp. 407–426.
van der Aalst, W. M. P. 1998. “The Application of Petri Nets to Workflow Management,”
Journal of Circuits, Systems and Computers (8:I), pp. 21–66.
van der Aalst, W. M. P. 2005. “Business Alignment: Using Process Mining as a Tool for
Delta Analysis and Conformance Testing,” Requirements Engineering Journal (10:3),
pp. 198–211.
van der Aalst, W. M. P. 2011. Process Mining: Discovery, Conformance and Enhancement
of Business Processes, (1st ed, )Springer Berlin Heidelberg.
van der Aalst, W. M. P., and Dongen, B. F. 2002. “Discovering Workflow Performance
Models from Timed Logs,” In Engineering and Deployment of Cooperative Infor-
mation Systems, Y. Han, S. Tai, and D. Wikarski (eds.), (Vol. 2480)Berlin, Heidelberg:
Springer Berlin Heidelberg, pp. 45–63.
van der Aalst, W. M. P., van Dongen, B. F., Herbst, J., Maruster, L., Schimm, G., and Weij-
ters, A. J. M. M. 2002. “Workflow Mining: a Survey of Issues and Approaches,” Beta:
Research School for Operations Management and Logistics (Working Paper 74).
van der Aalst, W. M. P., and Weijters, A. 2002. “Rediscovering Workflow Models from
Event-Based Data,” In Proceedings of the Third International NAISO Symposium on
Engineering of Intelligent Systems (EIS 2002)Presented at the Third International
NAISO Symposium on Engineering of Intelligent Systems (EIS 2002).
van der Aalst, W. M. P., Weijters, A. J. M. M., and Maruster, L. 2002. “Workflow Mining:
Which Processes can be Rediscovered?,” BETA Working Paper Series (Working Pa-
per 74).
Agrawal, R., Gunopulos, D., and Leymann, F. 1998. “Mining Process Models from Work-
flow Logs,” In Proc. Sixth Int’l Conf. Extending Database Technology, pp. 469–483.
Anandarajan, M., Anandarajan, A., and Srinivasan, C. A. 2004. Business intelligence tech-
niques : a perspective from accounting and finance, Berlin, Germany; New York:
Springer-Verlag.
Brinkkemper, S. 1996. “Method engineering: engineering of information systems develop-
ment methods and tools,” Information and Software Technology (38:4), pp. 275–
280.
Cook, J. E., and Wolf, A. L. 1998a. “Discovering models of software processes from event-
based data,” ACM Trans. Softw. Eng. Methodol. (7:3), pp. 215–249.
Cook, J. E., and Wolf, A. L. 1998b. “Event-based detection of concurrency,” SIGSOFT
Softw. Eng. Notes (23:6), pp. 35–45.
Cook, J. E., and Wolf, A. L. 1999. “Software process validation: quantitatively measuring
the correspondence of a process to a model,” ACM Trans. Softw. Eng. Methodol.
(8:2), pp. 147–176.
Diestel, R. 2010. Graph theory, (4th ed, )Heidelberg; New York: Springer.

190
APPENDIX A: PUBLICATIONS

Gehrke, N., and Müller-Wickop, N. 2010a. “Rekonstruktion von Geschäftsprozessen im Fi-


nanzwesen mit Financial Process Mining,” In Lecture Notes in Informatics, Procee-
dings der Jahrestagung Informatik Presented at the Jahrestagung Informatik 2010,
IT-supported Service Innovation and Service Improvement, Leipzig.
Gehrke, N., and Müller-Wickop, N. 2010b. “Basic Principles of Financial Process Mining A
Journey through Financial Data in Accounting Information Systems,” In Proceedings
of the 16th Americas Conference on Information SystemsPresented at the Americas
Conference on Information Systems, Lima, Peri.
Heckel, R. 2006. “Graph Transformation in a Nutshell,” Electronic Notes in Theoretical
Computer Science (148:1), pp. 187–198.
Herbst, J. 2000a. “A Machine Learning Approach to Workflow Management,” In Proceed-
ings ofMachine Learning: 11th European Conference Machine Learning, R. López de
Mántaras and E. Plaza (eds.), (Vol. 1810)Berlin, Heidelberg: Springer Berlin Heidel-
berg, pp. 183–194.
Herbst, J. 2000b. “Dealing with concurrency in workflow induction,” In European Concur-
rent Engineering Conference. SCS Europe.
Herbst, J. 2003. Ein induktiver Ansatz zur Akquisition und Adaption von Workflow-Model-
len, Tenea Verlag Ltd.
Herbst, J., and Karagiannis, D. 1998. “Integrating Machine Learning and Workflow Man-
agement to Support Acquisition and Adaptation of Workflow Models,” In Proceed-
ings 9th International Workshop on Database and Expert Systems Applications
(DEXA’98), pp. 745–752.
Herbst, J., and Karagiannis, D. 1999. “An inductive approach to the acquisition and adap-
tation of workflow models,” In Proceedings of the IJCAI (Vol. 99), pp. 52–57.
Hevner, A. R., March, S. T., Park, J., and Ram, S. 2004. “Design science in information sys-
tems research,” Mis Quarterly (28:1), pp. 75–105.
Huan, J., Wang, W., and Prins, J. 2003. “Efficient mining of frequent subgraphs in the pres-
ence of isomorphism,” In Data Mining, 2003. ICDM 2003. Third IEEE International
Conference on, pp. 549–552.
March, S. T., and Smith, G. F. 1995. “Design and natural science research on information
technology,” Decis. Support Syst. (15:4), pp. 251–266.
Maruster, L., Van Der Aalst, W. M. P., Weijters, A., van den Bosch, A., and Daelemans, W.
2001. “Automated discovery of workflow models from hospital data,” In Proceed-
ings of the 13th Belgium-Netherlands Conference on Artificial Intelligence (BNAIC
2001), pp. 183–190.
Maruster, L., Weijters, A., van der Aalst, W., and van den Bosch, A. 2001. “Process mining:
Discovering direct successors in process logs,” In Discovery Science, pp. 32–36.
Maxeiner, M. K., Küspert, K., and Leymann, F. 2001. “Data Mining von Workflow-Protokol-
len zur teilautomatisierten Konstruktion von Prozessmodellen,” In Datenbanksys-
teme in Büro, Technik und Wissenschaft (BTW), 9. GI-Fachtagung, Springer-Verlag,
pp. 75–84.

191
APPENDIX A: PUBLICATIONS

Müller-Wickop, N., Schultz, M., Gehrke, N., and Nüttgens, M. 2011. “Towards Automated
Financial Process Auditing: Aggregation and Visualization of Process Models,” In
Proceedings of the Enterprise Modelling and Information Systems Architectures
Presented at the EMISA 2011, Germany.
Neel, D. L., and Orrison, M. E. 2006. “The linear complexity of a graph,” the electronic
journal of combinatorics (13:R9), pp. 1.
Österle, H., Becker, J., Frank, U., Hess, T., Karagiannis, D., Krcmar, H., Loos, P., Mertens, P.,
Oberweis, A., and Sinz, E. J. 2010. “Memorandum on design-oriented information
systems research,” European Journal of Information Systems (20:1), pp. 7–10.
Romney, M. B., and Steinbart, P. J. 2008. Accounting Information Systems, (11th Revised
edition (REV), )Prentice Hall.
Rozenberg, G. 1997. Handbook of Graph Grammars and Computing by Graph Transfor-
mation, (illustrated ed, , Vols. 1-3, Vol. Volume 1 Foundations)Singapore: World Sci-
entific Pub Co.
SAP. 2012. “SAP-UCC,” available at [Link] accessed on 29th May 2012.
Schimm, G. 2000. “Generic Linear Business Process Modeling,” In Conceptual Modeling
for E-Business and the Web, S. W. Liddle, H. C. Mayr, and B. Thalheim (eds.), (Vol.
1921)Berlin, Heidelberg: Springer Berlin Heidelberg, pp. 31–39.
Schimm, G. 2001a. “Process Mining Elektronischer Geschäftsprozesse,” In Proceedings
Elektronische Geschäftsprozesse.
Schimm, G. 2001b. “Process Mining linearer Prozessmodelle - Ein Ansatz zur Automati-
sierten Akquisition von Prozesswissen,” In Proceedings 1. Konferenz Professionelles
Wissensmanagement.
Schimm, G. 2002. “Process Miner — A Tool for Mining Process Schemes from Event-Based
Data,” In Logics in Artificial Intelligence, S. Flesca, S. Greco, G. Ianni, and N. Leone
(eds.), (Vol. 2424)Berlin, Heidelberg: Springer Berlin Heidelberg, pp. 525–528.
Turban, E., Aronson, J. E., Liang, T.-P., and Sharda, R. 2007. Decision support and business
intelligence systems, Upper Saddle River, N.J.; London: Pearson Education Interna-
tional.
University of Hamburg. 2012. “Renew - The Reference Net Workshop,” available at
[Link] accessed on 29th May 2012.
Valk, R. 2008. “Lecture Notes: Formale Grundlagen der Informatik II (FGI 2) Modellierung
& Analyse paralleler und verteilter Systeme”, Universität Hamburg.
Vercellis, C. 2009. Business intelligence, Chichester: Wiley.
Weijters, A., and Van der Aalst, W. M. P. 2001. “Process mining: discovering workflow
models from event-based data,” In Proceedings of the 13th Belgium-Netherlands
Conference on Artificial Intelligence (BNAIC 2001), pp. 283–290.
Werner, M., and Gehrke, N. 2011. “Potentiale und Grenzen automatisierter Prozessprü-
fungen durch Prozessrekonstruktionen,” In Forschung für die Wirtschaft Shaker Ver-
lag.

192
APPENDIX A: PUBLICATIONS

Werner, M., Gehrke, N., and Nüttgens, M. 2012. “Business Process Mining and Recon-
struction for Financial Audits,” In Proceedings of the 45th Hawaii International Con-
ference on System Sciences Presented at the 45th Hawaii International Conference
on System Sciences, pp. 5350–5359.
Yan, X., and Han, J. 2002. “gSpan: graph-based substructure pattern mining,” IEEE Com-
put. Soc, pp. 721–724.

193
APPENDIX A: PUBLICATIONS

10.8 Improving Structure: Logical Sequencing of Mined Process Models

Number 8
Improving Structure: Logical Sequencing of
Title
Process Models
Appendix 10.8
Primary Related Chapters 5.4
Type Conference Paper
47th Hawaii International Conference
Conference
on System Sciences (HICSS 2014)
Reference (Werner and Nüttgens, 2014)
Acceptance Rate 56 %
VHB JQ 2.1 Ranking C (6.44)
WKWI Ranking B
ERA 2010 A
CORE 2013 A
Review Procedure Double Blinded
Number of Reviews 4
1. Michael Werner
Authors
2. Markus Nüttgens
Dissertation Points 0.67
Authorship
Overall 95%
Design 95%
Realization 95%
Writing 95%
Status Published
Part of other Dissertations No
[Link]
Link
ings/hicss/2014/2504/00/[Link]

194
APPENDIX A: PUBLICATIONS

Improving Structure: Logical Sequencing of Mined Process Models

MICHAEL WERNER MARKUS NÜTTGENS


University of Hamburg, Germany University of Hamburg, Germany
[Link]@[Link] [Link]@[Link]

Abstract: The increasing availability of digital data offers new opportunities for
analyzing business processes. Process aware information systems like Enter-
prise Resource Planning systems store data in the course of transaction pro-
cessing. This data can be exploited by using process mining techniques. Process
mining algorithms produce process models by analyzing recorded event logs. A
fundamental challenge in process mining is the creation of purpose-oriented
and useful process models. Process mining algorithms commonly refer to the
temporal ordering of events for determining the control flow in reconstructed
process models. We show how the logical sequence of events can be used in-
stead of the temporal for reconstructing the control flow in mined process
models. The exploitation of the logical structure of available event log records
opens up new ways to receive purpose-oriented, less complex, and more in-
formative process models.

tools to deal with big data. Process mining is


1 Introduction a BI approach that uses recorded event log
Data explosion [1] and information over- data to provide information about business
flow [2] are well known phenomena that ac- processes. It can be seen as a bridge be-
company the increasing integration of infor- tween BI and business process manage-
mation technology into society and busi- ment [6]. Process mining algorithms pro-
ness. The availability of huge amounts of duce process models. They reconstruct
digital data provides extensive opportuni- models by analyzing the available source
ties for novel data analyses. But the han- event logs. A fundamental challenge in pro-
dling and examination of voluminous data cess mining is the creation of process mod-
sets also creates new challenges that need els that are useful for the intended purpose.
to be addressed thoroughly to provide well- Mined process models are often too com-
researched solutions that can be applied in plex for simple interpretation and analysis.
practice. The analysis of voluminous data Extremely complex process models are re-
has recently gained increased attention in ferred to as lasagna and spaghetti processes
the academic community under the term in the process mining community due to
Big Data [3]. Its implications are not limited their graphical characteristics of layered or
to the field of information systems or com- intertwined arcs in the respective process
puter science but are also relevant for other models [2]. Complexity reduction of process
scientific disciplines like biology [4] or phys- models has been addressed by scholars fol-
ics [5]. Business intelligence (BI) is a re- lowing different approaches. Advanced
search domain that provides methods and mining algorithms like the Fuzzy Miner [7]

195
APPENDIX A: PUBLICATIONS

provide functionality to abstract from infre- audits due to their significance for the well-
quent events and execution paths and are functioning of economic markets. An im-
able to deliver less complex process models. portant step in financial audits is the audit-
But infrequent behavior might indeed be ing of business processes. It is assumed that
relevant for compliance [8] and conform- well-controlled business processes most
ance checking purposes [9] and their omis- likely lead to correct recording of financially
sion in the process model might lead to pro- relevant transactions in the financial ac-
cess models that are of little use in these counts. It is more efficient to audit the over-
contexts. Other scholars suggest abstrac- all business process structure than inspect-
tion methods like aggregation and reduc- ing individual business transactions. A criti-
tion for reducing complexity in process cal requirement for the application of pro-
models [10], [11]. Complexity is just one cri- cess mining in financial audits is the reliabil-
terion that is relevant for evaluating the ity of produced process models and the
quality of a process model. The quality and preservation of the audit trail [15]. The au-
value of a process model has to be inter- dit trail describes the path in an information
preted in the context of the intended pur- system that allows following an entry on the
pose and cannot be generally defined [12]. financial accounts back to its point of origin
A process model is purpose-oriented if it ful- and vice versa. Auditors use process models
fills the requirements that can be derived to gain an understanding of the audited
from the aimed goal or objective of a pro- business processes and to identify incorrect
cess mining attempt. transaction processing. These process mod-
els are commonly created using traditional
We suggest an alternative approach. Our
audit procedures like interviews and inspec-
aim is to provide process models that are
tions of available documents which are
less complex and that provide information
highly time-consuming and error-prone.
on the logical structure of business pro-
Process mining can be used to produce reli-
cesses by exploiting the logical structure of
able process models very efficiently. But er-
recorded events in an event log. This ap-
rors or omissions in the models that can be
proach creates more informative models
a result of the mining process might lead to
according to relevant requirements of the
misinterpretations in the audit. Mining algo-
application domain compared to traditional
rithms that produce unfitting and imprecise
approaches that use the temporal structure
[16] process models are therefore generally
of events. Common process mining algo-
not suitable for the application in financial
rithms like deterministic [2], heuristic [13]
audits.
and genetic [14] algorithms use the tem-
poral ordering of recorded events for recon- We illustrate in this article how the logical
structing the control flow in a process structure of recorded events can be used for
model. We illustrate in this article how in- reconstructing the control flow in mined
formation from the application domain can process models. Entries in financial ac-
be used to enable the reconstruction of pro- counts exhibit a specific structure. Enter-
cess models based on the logical ordering. prise Resource Planning (ERP) systems pro-
duce entries on financial accounts for each
The selected application domain refers to
financially relevant transaction that is pro-
the auditing industry. Companies publish fi-
cessed in the system [17]. The relationships
nancial information for informing stake-
between journal entries can be exploited to
holders about the financial performance.
mine process models [18], [19].
These financial statements are subject to

196
APPENDIX A: PUBLICATIONS

The research presented in this paper follows Four dominant application domains for pro-
a design science research approach [20]– cess mining are process discovery, enhance-
[22]. The reason for choosing such an ap- ment, conformance [29] and compliance
proach is the proximity of the research checking [8]. Process discovery, conform-
question to the practical challenge of creat- ance and compliance checking are areas
ing useful process models and the objective that are especially relevant for audit pur-
to deliver design artifacts that are valuable poses. Process mining has already been suc-
for the application domain. We have col- cessfully applied in the context of internal
lected extensive real world data from differ- audits [30]–[33] and financial audits [34].
ent companies for experiments and test But we are not aware of research ap-
purposes to develop a mining algorithm proaches that investigate the logical struc-
that is able to reconstruct the logical order- ture of events for the reconstruction of the
ing of events. control flow in the context of process min-
ing. Approaches for organizational mining
The following section includes a brief review
and social network analysis [35] focus on
of related work to provide an overview of
the resources that interact in a business
the positioning of the presented work in the
process and exploit data referring to the re-
context of already published research. Sec-
lationship between process participants
tion three provides an example of a simple
and activities to create models that illus-
business process instance that is used for il-
trate organizational structures and social
lustrating the differences between tem-
networks. Organizational mining also uses
poral and logical ordering of events in sec-
information from the event log other than
tion four and five. Section six describes the
the temporal ordering of events for dis-cov-
evaluation efforts that were integrated into
ering models but with the objective to dis-
the research work to demonstrate the rigor
cover the interaction of process partici-
of achieved results. The article closes with a
pants. It therefore differs from the ap-
brief summary and outlook in the last sec-
proach used in this paper that aims to re-
tion.
construct the control flow based on the log-
2 Related Work ical structuring of events.
We use the Fuzzy Miner algorithm [7] for il-
Process mining is a research area that lustrating the traditional reconstruction of
emerged in the context of software engi- the control flow that bases on the temporal
neering [23], [24]. It was first applied to sequence of recorded events. The Fuzzy
workflow management in the late 1990s Miner is an advanced heuristic mining algo-
[25]–[27] and has matured significantly in rithm that is able to create less complex pro-
recent years. cess models by abstracting from infrequent
It would be out of the scope of this paper to behavior and events. It has been imple-
provide an all-embracing overview on pro- mented in academic [36] and commercial
cess mining. The Process Mining Manifesto software tools [37].
[28] provides a comprehensive overview of The reconstruction of control flows based
contemporary challenges in process mining, on the logical sequence of events is demon-
and an overview of basic and advanced con- strated by using the Financial Process Min-
cepts on process mining can be found in [2]. ing (FPM) algorithm that was developed
specifically for the context of financial au-

197
APPENDIX A: PUBLICATIONS

dits [18]. Financial accounts and journal en- Table 1 Example event log
tries are key concepts for process audits
[38]. Each execution of a financially relevant Event ID Timestamp Activity
activity creates an accounting journal entry Post Received
0050155443 2010/01/02
that is recorded in an ERP system. Goods
Post Received
0050155250 2010/02/08
Goods
Post Received In-
0015975223 2010/02/17
Raw Materials
Goods Received /
Invoices Received
Trade Payables voice
10,000 cleared 10,000 10,000 cleared 10,000 open Post Received In-
0015975224 2010/02/18
voice
Post Received In-
0015975221 2010/02/19
Receive Goods Receive Invoice voice
0095348327 2010/02/20 Clear Postings
Figure 1 Accounting Structure
0095348517 2010/02/21 Clear Postings
A journal entry consists of at least two jour- Post Received
nal entry items, one on the debit and one on 0050157332 2010/08/16
Goods
the credit side of a financial account. Open Post Received In-
0015980342 2010/09/03
items on an account are cleared by items voice
belonging to other journal entries. These 0012490379 2010/09/04 Payment
are created by the execution of activities
that belong to the same process instance. 0007904673 2010/09/05 Post with Clearing
This structure is illustrated in Figure 1. It
0095359370 2010/09/07 Clear Postings
shows the financial accounts with the corre-
sponding journal entry items and the activi-
ties that created them. The analysis of the Process mining algorithms usually produce
relationships between activities, journal en- process models. A business process is a set
tries and journal entry items allows the re- of connected activities that in combination
construction of the logical control flow realize a specific business goal [39]. A busi-
which will be discussed in detail in the fol- ness process model is an abstraction of a
lowing sections. business process and consists of a set of ac-
tivity models and execution constraints be-
3 Process Example tween them [40]. A single execution of a
business process is called process instance.
We start with an example of a simple pro- Process models represent the behavior of a
cess instance for demonstrating the differ- set of process instances that belong to the
ences between temporal and logical order- same business process. A model represent-
ing. Table 1 provides the event log for this ing a single process instance is called pro-
example that was derived from a company cess instance model. Process models and
operating in the manufacturing industry. process instance models generally include
every activity only once. The represented
activity models in process instance and pro-
cess models are already abstractions of a set
of executed activities. A process instance
graph [41] resides on the lowest level of ab-
straction. Each event is represented as a sin-

198
APPENDIX A: PUBLICATIONS

gle activity in the model. The differences be- types of places (circles). The yellow and blue
tween the abstraction levels of process colored places represent financial accounts.
models, instance models and instance The dotted arrows leading from a transition
graphs are important for the interpretation to an account place denote that the corre-
of the used example and mining algorithm sponding activity posted a journal entry
outputs in the subsequent sections. We will item on the connected account. The values
only refer to the instance graph and process of the posted items are included as inscrip-
instance level in this article for ease of illus- tions for each arrow. The dotted edges (test
tration. But it is important to note that the arcs) without arrows denote that an entry
same results are also valid for the process item was cleared on the respective account
model level. by the connected activity. The value of the
0001400100 0004000070 0001400100 0004000070 0001400100 0004000070
cleared item is also displayed as an inscrip-
tion for the corresponding edge. Each tran-
sition is further connected to a control place
[846.60] [68.08] [8918.00] [493.94] [4233.00] [296.11]

1 0050157332 1 0050155250 1 0050155443

Post
Received Goods
Post
Received Goods
Post
Received Goods that carries a token indicating how often the
2010/02/16 2010/11/02 2010/11/02

[914.68] [9411.94] [4529.11] activity was executed in the process in-


0002810200 0002810200
0002810200
stance.
[914.68] [9411.94] [4529.11]

1 0095359370

Clear Postings
1 0095348327

Clear Postings
1 0095348517

Clear Postings
The explicit modeling of these places cre-
2010/09/03 2010/02/19 2010/02/19 ates life CPNs that adequately represent the
[4529.11]

observed behavior of the process instance.


[914.68] 0004000070 [9411.94]

0002810200
0002810200 0002810200

[914.68] [1158.69] [9411.94] [4529.11] It ensures that each activity in the net can
1 0015980342

Post
1 0015975224

Post
1 0015975223

Post
1 0015975221

Post
only fire once as recorded in the event log.
Received Invoice Received Invoice Received Invoice Received Invoice

The model represents an instance of a pur-


2010/09/03 2010/02/19 2010/02/19 2010/02/19

[3.69] [1158.69] [1038.68] [8373.26] [4586.87] [57.76]


[910.99]

chase process. It shows that three different


0004000070 0002811000 0004000070 0002811000 0002811000 0004000070

0002811000

[910.99] [1158.69] [8373.26] [4586.87] events occurred for the recording of re-
1 0012490379
ceived goods. The receipt of goods recorded
Payment
[15029.81]
0002811000 by the activity Post Received Goods with the
2010/09/03

[15029.81]
event ID 0050157332 for example led to
0001900113

journal entry item postings on the accounts


1
[15029.81]

0007904673
0001400100, 0004000070 and
Post
with Clearing [15062.42]
0001900111 0002810200. An invoice was received for
0005004040
2010/09/03

0005035200
each obtained good. An additional invoice
processed by the activity with event ID
[15.14] [17.47]

0001900113
[15029.81] 0015975224 was received with no corre-
sponding recording of received goods. The
Figure 2 Example process instance open items on the related accounts were
Table 1 includes the event log of a single cleared by using the dedicated activity Clear
process instance. Figure 2 shows an in- Postings. All received invoices were subse-
stance graph of the used example. The pic- quently cleared by the same payment. In-
ture shows a Colored Petri Net (CPN) [42]. It termediate postings were finally shifted to
provides information on the activities that the final accounts by using the activity Post
were executed and involved financial ac- with Clearing.
counts. The transitions (rectangles) repre- The diagram provides a detailed overview of
sent activities that were executed in the the structure of the process instance and il-
process. The CPN includes two different lustrates information on financial accounts

199
APPENDIX A: PUBLICATIONS

as well as clearing and posting relationships


between account places and transitions
that are relevant to reconstruct the logical
ordering of events.

4 Temporal Sequence
If we use the Fuzzy Miner to discover a
model based on the event log in Table 1 we
obtain a model as displayed in Figure 3. The
mining algorithm produces an instance
model. It includes every activity only once
(contrary to the instance graph in Figure 2)
and therefore provides a higher level of ab-
straction than the model represented in Fig-
ure 2. The shown figure is not a process
model because it only illustrates the behav-
ior of a single process instance.
The mined model shows a different struc-
ture than expected when compared to the
graph in Figure 2. Figure 2 actually shows
four different branches that represent sub- Figure 3 Discovered Fuzzy Miner process
processes of received goods and invoices instance model
that were all paid by the same payment run.
We can also observe a back loop from the
The invoice in each branch was received af-
Clear Postings activity to Post Received
ter the receipt of goods was recorded. We
Goods. The reason for the difference be-
would therefore expect a process model
comes obvious when comparing Figure 4. It
that shows the corresponding sequence of
shows the temporal dependencies between
Post Received Goods → Post Received In-
the different events based on the recorded
voice → Clear Postings → Payment → Post
timestamps.
with Clearing. The model in Figure 3 instead
shows several short loops indicating that The Fuzzy Miner algorithm reconstructs the
the execution of an activity was followed by control flow according to the temporal se-
the execution of the same activity. quence of events. The event Post Received
Goods with the event ID 0050155443 was
the first event that occurred at the
2010/01/02.

200
APPENDIX A: PUBLICATIONS

that are purpose-oriented and easy to inter-


pret.
0050157332 0050155250 0050155443
1
Post
Received Goods
Post
Received Goods
Post
Received Goods We satisfy this requirement by modeling
1
2010/08/16 2010/02/08 2010/01/02
the logical control flow. It can be recon-
0095359370 0095348327 0095348517 structed by analyzing which journal entry
1
Clear Postings 1 1 Clear Postings
1
Clear Postings item was cleared by another activity. An
2010/09/07 2010/02/20
1
2010/02/21
open item can only be cleared if it has been
0015980342 0015975224 0015975223 0015975221 posted before the clearing item. It is there-
Post
Received Invoice
Post
Received Invoice
1 Post
Received Invoice
Post
Received Invoice fore possible to derive the causal relation-
2010/09/03 2010/02/18

1
2010/02/17 2010/02/19
ship between events by analyzing which ac-
0012490379
1 tivities cleared items created by other activ-
1
Payment
1
ities.
2010/09/04

1 0001900113
1 1
0007904673
[15029.81] [15029.81]
1 Post 0012490379 0007904673
0002811000 0005035200
with Clearing
Post
2010/09/05 Payment
[15029.81] with Clearing [17.47]

2010/09/03 2010/09/03
0005004040
0001900111
Figure 4 Temporal sequence [15062.42]
[15.14]

This event was followed by Post Received 0001900113

Goods with the event ID 0050155250 at the [15029.81]

2010/02/08. The algorithm interprets this


temporal dependency for reconstructing Figure 5 Logical dependency
the control flow from event 0050155443 → Figure 5 provides an illustration of this ap-
0050155250. This approach produces a pro- proach. The activity Payment posted an
cess model that adequately illustrates the item with the value of 15,029.81 on the ac-
temporal ordering of events. But it depends count 0001900113. This item was cleared
on the application scenario if this infor- by the activity Post with Clearing. The logical
mation is indeed useful for the intended sequence for these activities can therefore
purpose. be derived as Payment → Post with Clear-
ing.
5 Logical Sequence
This procedure can be applied to all activi-
5.1 Defining Logical Dependencies ties in the event log for determining the
complete logical control flow of the process
In financial audits it is important to get an instance. It is further necessary to identify
understanding of the control flow of pro- end and start nodes for creating a complete
cesses and the relation of business pro- process model. Start nodes can be deter-
cesses to the financial statements [38], [43]. mined by identifying activities that do not
The auditor needs to get an understanding have any incoming control arcs and end
what the different processes do and how nodes do not have any outgoing control
they interact with the financial accounts. arcs.
The first step in a process audit is to under-
stand the process. Information on the logi- 5.2 Clearing Deadlocks
cal structure is in this context more useful
A problem arises with activities that do not
than the temporal sequence of events. The
post any items but only clear items posted
key requirement is to create process models
by other activities. This constellation, which

201
APPENDIX A: PUBLICATIONS

we call a clearing dead-lock, is illustrated in


Activity B Activity C Activity B Activity C
Figure 6.
Clearing Clearing
0002810200 0002810200 Activity X Activity X

1 1 1
[914.68] [914.68] [914.68] [914.68] Activity A Activity A

0050157332 0095359370 0015980342

Post Post a) Deadlock Constellation b) Deadlock Resolution


Clear Postings
Received Goods Received Invoice

2010/02/16 2010/09/03 2010/09/03

[846.60]
[68.08] [910.99]
[3.69] Figure 7 Deadlock resolution
Figure 7a) provides an illustration of a typi-
0001400100 0004000070 0002811000 0004000070
cal constellation. Items from activities A and
Figure 6 Clearing deadlock B were cleared by clearing activity X. One or
more additional items from activity B were
Due to the fact that the Clear Postings activ- also cleared by activity C. Because activity A
ity did not create any other posted item the has no further outgoing control arcs it is rea-
logical control flow ends at this activity and sonable to assume that it is logically ordered
it would be defined as an end node. This is a before activity B and that the process con-
problem. The clearing activity is actually not tinued with activity C after A, B and X took
an end node because the process does not place. The related sub-graph can be substi-
end before the activity Post with Clearing. tuted by an amended sub-graph as illus-
The results from our evaluations show that trated in Figure 7b).
this constellation is very common at least The logical sequence can be rearranged un-
for data extracted from SAP systems. It can- ambiguously if only one of the posting activ-
not be neglected because the resulting ities has posted an item that was cleared by
models would be of little value for the user. another activity. It is not possible to define
The additional end nodes would confuse the a definite logical sequence locally by observ-
process model and imply invalid infor- ing the direct neighbor activities if more
mation about the process structure. than one of the posting activities has posted
This outcome can be prevented by using items that were cleared by other activities.
graph transformations [44]. The idea is to Such constellations require non-local
identify clearing deadlocks and to apply searches in the graph. The evaluation re-
graph transformations for removing them. sults showed that such constellations rarely
Clearing activities first get identified by se- occur and can be neglected for applications
lecting all activities that did not post any in practice.
items but only cleared items from other ac- Although the presented transformation im-
tivities. Clearing activities commonly clear proves the usefulness of the produced mod-
items from two or more other activities. At els it violates an important requirement
least one of these activities must have from the application domain that relates to
posted an additional item that was cleared the preservation of the audit trail and pos-
by another activity different from the clear- tulates that mining algorithms should not
ing activity for a deadlock to appear. Other- alter the original data or provide infor-
wise the clearing activity would represent a mation that is not reflected in the recorded
valid end node and such a constellation event log [11]. The transformation of the
would not be considered a clearing dead- modeled logical sequence is indeed an alter-
lock. ation of the original sequence. But negative
impacts should be minimal. The presented

202
APPENDIX A: PUBLICATIONS

approach is intended to be applied for the clearing deadlocks and (3) defining start and
discovery of process models to provide end nodes were implemented in a software
models that allow auditors to easily under- prototype.
stand relevant processes. The amendment
Figure 8 shows the effect of the logical se-
of the logical control flow for resolving
quencing. The positioning of the activities
clearing deadlocks is not critical for this pur- has been kept constant compared to Figure
pose but it might not be suitable for detect-
2 and 4. The first, third and fourth branch
ing conformance or compliance violations
show the logical sequence Post Received
that require completely unchanged process Goods → Post Received Invoice → Clear
models.
Postings. The second branch only consists of
the sequence Post Received Invoice. All
5.3 Logically Sequenced Process
Models branches follow the subsequent sequence
Payment → Post with Clearing. The illus-
The original FPM algorithm [18] is only able trated instance therefore contains the logi-
to reconstruct simple graphical representa- cal sequences:
tions of process instances. We used an
amended algorithm that produces CPN and α: Post Received Goods → Post Received
is able to create models on process instance Invoice → Clear PosCngs → Payment
and process level [11]. → Post with Clearing
β: Post Received Invoice → Payment →
Post with Clearing
0050157332 0050155250 0050155443

Post Post Post Figure 9 shows the model for the example
Received Goods Received Goods Received Goods

2010/08/16 2010/02/08 2010/01/02


instance produced by the FPM algorithm. It
only exhibits the same logical sequences α
0095359370 0095348327 0095348517

Clear Postings Clear Postings Clear Postings


and β that also characterize the control flow
1 1 1

2010/09/07 2010/02/20 2010/02/21


in the model displayed in Figure 8.
1 1 1
1 1
0015980342 0015975224 0015975223 0015975221

Post Post Post Post Post Post


Clear Postings Payment
Post
Received Invoice Received Invoice Received Invoice Received Invoice 3 Received Goods 3 Received Invoice 3 3 1 with Clearing 1

2010/09/03 2010/02/18 2010/02/17 2010/02/19

1
0012490379 1 Figure 9 Discovered FPM process instance
1
model48
1
Payment

2010/09/04

1 The provided model in Figure 9 that was


0007904673

Post
mined using the logical sequence of events
with Clearing
is less complex and provides more useful in-
2010/09/05
formation for the purpose of understanding
the structure of a process than the model
derived from using the traditional approach
Figure 8 Logical sequence of determining the control flow by analyzing
The formerly explained procedures for (1) the temporal sequence that is illustrated in
defining causal dependencies, (2) removing Figure 3.

48
Process models produced by the FPN usually in- Figure 2. We have omitted this information and
clude information on the data perspective by only present a simplified dependency graph that
modeling account places and their connections can be easily compared with the output of the
to the activities similar to the graph presented in Fuzzy Miner illustrated in Figure 3.

203
APPENDIX A: PUBLICATIONS

6 Evaluation ing the same activity label. This constella-


tion leads to a short loop and deadlock in
Evaluation is a significant part of design sci- the model. But it only occurred in very com-
ence research [22], [45]. We have chosen an plex process models and very rarely. Re-
experimental setup for evaluating the min- search for addressing this problem is cur-
ing results for the described mining ap- rently in progress. A solution could be the
proach. A laboratory experiment is a suita- prevention of aggregating activities if this
ble evaluation method for artificial ex post would lead to a short loop and to allow sub-
evaluations [46] and was therefore chosen sequent activities in the model that carry
for the presented research work. We ex- the same activity label.
tracted event logs from productive SAP sys-
The aim of using the logical structure of
tems from three different companies oper-
event logs for process mining is the provi-
ating in the manufacturing, media and retail
sion of less complex and more informative
industries.
process models according to the require-
An overview of the used data sets is pro- ments of the application domain. Several
vided in Table 2. The event logs served as in- metrics exist for measuring the complexity
put for the Fuzzy Miner and FPM algorithm. of process models [16]. The metric struc-
We evaluated the proper functioning of the tural appropriateness measures the com-
implemented FPM and checked if the min- plexity of a process model by counting the
ing results fitted the expected outcome. The number of included tasks [50]. Structural
adjusted FPM should create sound [47] precision and structural recall measure the
CPNs. Produced CPNs were tested for amount of causality relations that a mined
soundness by using the academic software model has in common with a reference
Renew [48]. The graph editor yEd [49] pro- model. The metrics duplicates precision and
vides powerful automatic layout functional- duplicates recall measure how many dupli-
ity and was used to graphically represent cate tasks a mined model has in common
mined models. The produced models were with a reference model [14]. The metrics
inspected and compared with the models structural precision, structural recall, dupli-
produced by the Fuzzy Miner. The FPM cates precision, and duplicates recall are not
models generally showed a lower complex- suitable for measuring the complexity in our
ity due to fewer loops in the models com- experimental setup because they require a
pared to the Fuzzy Miner models and com- reference model for comparison. The metric
plied with the expected mining outcomes. structural appropriateness only considers
Table 2 Overview of evaluation data sets the number of modeled activities. Using the
logical structure of event logs for process
Data Process Process mining has no effect on the number of rep-
Industry
Set Instances Models
1 Manufacturing 1,035,805 841
resented activities because only the se-
2 Media 18,975 516 quence of activities is different. We there-
3 Retail 40,634 307 fore used the number of arcs as a criterion
for measuring the complexity of a process
The evaluation revealed that unexpected, model. The vast majority of mined models
unsound process models were created un- from our data sets represent trivial process
der certain circumstances. This occurred instances that consist of only one activity
when an activity created a posted item that [34]. Trivial instance models do not differ in
was cleared by a subsequent activity carry- complexity if the temporal or logical struc-
ture is used because the sequence is always

204
APPENDIX A: PUBLICATIONS

the same if just one activity is involved. We It remains to discuss if the resulting models
therefore focused on nontrivial models that are more informative compared to those
consist of more than four activities assum- using a temporal ordering. A key require-
ing that the complexity reduction is higher ment for process audits in the context of fi-
for more complex models. A quantitative nancial audits is the understanding of the
evaluation for measuring the complexity re- structure of a process and its interaction
duction was carried out by using a specifi- with the financial accounts. Figure 3 shows
cally designed software artifact. It produced the temporal sequence of recorded activi-
instance models using the FPM algorithm ties. The model is only of limited use be-
with the traditional temporal structuring cause the intertwined temporal sequence
and the alternative logical structuring as of activities that belong to different sub-
presented in this paper. Using the same al- processes as indicated in Figure 4 leads to
gorithm ensured comparability of the re- loops that make it difficult to identify the
sults. main sequence of activities in the process.
Table 3 Mean values of the number of arcs This is different for the model shown in Fig-
ure 9. It clearly shows the logical structure
included in mined instance models
and facilitates the interpretation of the pro-
Nontrivial Mod- cess. It will be evaluated in future research
Data Complex Models
els work if this argument based assessment is
Set Tem- Tem-
poral
Logical %
poral
Logical % also supported by experts from the applica-
1 9.67 9.55 1 24.90 16.33 34 tion domain.
2 12.25 11.26 8 23.17 20.19 13
3 12.23 10.70 12 32.90 21.70 34
7 Conclusion and Outlook

Table 3 summarizes the evaluation results The value of a process model depends on its
for the used data sets.49 It shows the mean ability to satisfy the requirements that are
values of the number of arcs that are in- relevant to the application domain. A funda-
cluded in the mined instance models. The mental challenge in process mining is to cre-
table is split into two parts. The left columns ate process models that are useful for the
show the mean values as a representative user. Mined models are often too complex
for the model complexity for nontrivial for the intended purpose. A common ap-
models that contain five or more activities. proach to produce more suitable process
The overall complexity reduction is modest models is the application of complexity re-
ranging from 1 to 12 percent. The right col- duction methods. This includes the develop-
umns show the results for complex instance ment of mining algorithms like the Fuzzy
models that contain eight or more activities. Miner that are able to abstract from infre-
The reduction is significantly higher for quent behavior and events to provide more
more complex models ranging from 13 to 34 abstract and therefore less complex mod-
percent. The complexity reduction for the els. Other approaches use abstraction tech-
most complex models was 51 percent for niques like aggregation or reduction meth-
data set 1, 36 percent for data set 2 and 55 ods for reducing complexity.
percent for data set 3.

49
A random sample consisting of 100,000 instances
was used for data set 1 due to computational
constraints.

205
APPENDIX A: PUBLICATIONS

The research presented in this paper follows the data sets are extensive they only in-
a different approach. The aim is to use infor- cluded event data from SAP systems. It can
mation from the application domain to dis- therefore not be concluded that the results
cover the control flow in process models by are also valid for other ERP systems. But due
exploiting the logical structure of events ra- to the fact that the chosen methods exploit
ther than the temporal structure that is the generic structure of accounting entries
used in traditional approaches. The exploi- which need to be supported by all infor-
tation of the logical structure of events pro- mation systems used for accounting it is
vides the opportunity to mine process mod- very likely that they are also applicable for
els that fit the requirements from the appli- other systems. Our evaluation proofed that
cation domain better than traditional ap- the designed methods work correctly. But
proaches and that produce less complex we did not gain information if they will also
models. We have illustrated how infor- be accepted and useful in real world organ-
mation on the structure of event data can izational settings. Additional research like
be used by referring to the application do- field experiments will be conducted in fu-
main of financial audits. Information on the ture research to address this aspect.50
structure of accounting data can be used to
reconstruct the logical control flow in pro- 8 References
cess models that are useful for financial au- [1] H. Krcmar, Informationsmanage-
dits. This application domain should only be ment. Berlin; Heidelberg: Springer,
seen as an example of how such additional 2010.
information from the event data can be
used to improve the produced process mod- [2] W. M. P. van der Aalst, Process Min-
els. A prerequisite for using domain specific ing: Discovery, Conformance and En-
information is its availability. The domain hancement of Business Processes, 1st
knowledge has to be explicit and formal- Edition. Berlin Heidelberg: Springer,
ized. Ontologies are known as a suitable 2011.
representation form for such purposes. The [3] H. Chen, R. H. L. Chiang, and V. C.
development of an ontology for formalizing Storey, “Business Intelligence and An-
key concepts in process audits in the con- alytics: From Big Data to Big Impact,”
text of financial audits will be addressed in MIS Q., vol. 36, no. 4, pp. 1165–1188,
future research. It cannot be guaranteed Dec. 2012.
that similar results can also be achieved for
[4] D. Howe, M. Costanzo, P. Fey, T. Go-
different application scenarios. But our re-
jobori, L. Han-nick, W. Hide, D. P. Hill,
search shows a new route that can be cho-
R. Kania, M. Schaeffer, S. St Pierre, S.
sen for innovative improvements of mined
Twigger, O. White, and S. Yon Rhee,
process models.
“Big data: The future of biocuration,”
The presented methods have been evalu- Nature, vol. 455, no. 7209, pp. 47–50,
ated by using extensive data for testing the Sep. 2008.
implemented methods and for inspecting
the achieved outcomes. The data was de-
rived from different companies. Although
50
The research results presented in this paper were Ministry of Education and Research. The authors
developed in the research projects VAW (grant are responsible for the content of this publica-
number 01IS10041) and EMOTEC (grant number tion.
01FL10023) sponsored by the German Federal

206
APPENDIX A: PUBLICATIONS

[5] C. Lynch, “Big data: How do your data Model., vol. 11, no. 4, pp. 557–569,
grow?,” Nature, vol. 455, no. 7209, 2012.
pp. 28–29, Sep. 2008. [13] A. Weijters, W. M. P. van der Aalst,
[6] W. M. P. van der Aalst, “Using Pro- and A. K. A. de Medeiros, “Process
cess Mining to Bridge the Gap Be- mining with the heuristics miner-al-
tween BI and BPM,” Computer, vol. gorithm,” Tech. Univ. Eindh. Tech
44, no. 12, pp. 77–80, 2011. Rep WP, vol. 166, 2006.
[7] C. Günther and W. van der Aalst, [14] A. K. A. de Medeiros, “Genetic Pro-
“Fuzzy mining–adaptive process sim- cess Mining,” Eindhoven University of
plification based on multi-perspective Technology, Eindhoven, 2006.
metrics,” Bus. Process Manag., pp. [15] M. B. Romney and P. J. Steinbart, Ac-
328–343, 2007. counting Information Systems, 11th
[8] J. Becker, P. Delfmann, M. Eggert, Revised edition (REV). Prentice Hall,
and S. Schwittay, “Generalizability 2008.
and Applicability of Model- Based [16] A. Rozinat, A. K. A. de Medeiros, C.
Business Process Compliance-Check- W. Günther, A. Weijters, and W. M.
ing Approaches – A State-of-the-Art van der Aalst, “The Need for a Pro-
Analysis and Research Roadmap,” cess Mining Evaluation Framework in
BuR - Bus. Res., vol. 5, no. 2, pp. 221– Research and Practice,” in Business
247, Nov. 2012. Process Management Workshops,
[9] W. M. van der Aalst and A. K. A. de 2008, pp. 84–89.
Medeiros, “Pro-cess mining and secu- [17] M. Werner, N. Gehrke, and M.
rity: Detecting anomalous process ex- Nüttgens, “Business Process Mining
ecutions and checking process con- and Reconstruction for Financial Au-
formance,” Electron. Notes Theor. dits,” in Hawaii International Confer-
Comput. Sci., vol. 121, pp. 3–21, ence on System Sciences, Maui,
2005. 2012, pp. 5350–5359.
[10] M. Reichert, “Visualizing Large Busi- [18] N. Gehrke and N. Müller-Wickop,
ness Process Models: Challenges, “Basic Principles of Financial Process
Techniques, Applications,” in 1st Int’l Mining A Journey through Financial
Workshop on Theory and Applica- Data in Accounting Information Sys-
tions of Process Visualization, Tallin, tems,” in Proceedings of the 16th
2012. Americas Conference on Information
[11] M. Werner, M. Schultz, N. Müller- Systems, Lima, Peru, 2010.
Wickop, N. Gehrke, and M. Nüttgens, [19] N. Gehrke and N. Müller-Wickop,
“Tackling Complexity: Process Recon- “Rekonstruktion von Geschäftspro-
struction and Graph Transformation zessen im Finanzwesen mit Financial
for Financial Audits,” in Proceedings Process Mining,” in Lecture Notes in
of 33rd International Conference on Informatics, Proceedings der Jahres-
Information Systems, Orlando, 2012. tagung Informatik, Leipzig, 2010.
[12] W. M. P. van der Aalst, “What makes
a good process model?,” Softw. Syst.

207
APPENDIX A: PUBLICATIONS

[20] A. R. Hevner, S. T. March, J. Park, and Models,” in Proceedings 9th Interna-


S. Ram, “De-sign Science in Infor- tional Workshop on Database and Ex-
mation Systems Research,” MIS Q., pert Systems Applications (DEXA’98),
vol. 28, no. 1, pp. 75–105, Mar. 2004. 1998, pp. 745–752.
[21] S. T. March and G. F. Smith, “Design [28] W. M. P. van der Aalst, A. Andrian-
and natural science research on in- syah, A. K. de Medeiros, F. Arcieri, T.
formation technology,” Decis Support Baier, T. Blickle, J. C. Bose, P. van den
Syst, vol. 15, no. 4, pp. 251–266, Brand, R. Brandtjen, and J. Buijs,
1995. “Process Mining Manifesto,” in BPM
2011 Workshops Proceedings, 2012,
[22] H. Österle, J. Becker, U. Frank, T.
pp. 169–194.
Hess, D. Karagiannis, H. Krcmar, P.
Loos, P. Mertens, A. Oberweis, and E. [29] W. M. P. van der Aalst, “Process Min-
J. Sinz, “Memorandum on design-ori- ing: Overview and Opportunities,”
ented information systems research,” ACM Trans. Manag. Inf. Syst., vol. 99,
Eur. J. Inf. Syst., vol. 20, no. 1, pp. 7– no. 99, pp. 1–16, Feb. 2012.
10, 2010. [30] M. Jans, J. M. van der Werf, N. Lyba-
[23] J. E. Cook and A. L. Wolf, “Discovering ert, and K. Vanhoof, “A business pro-
models of software processes from cess mining application for internal
event-based data,” ACM Trans Softw transaction fraud mitigation,” Expert
Eng Methodol, vol. 7, no. 3, pp. 215– Syst. Appl., vol. 38, no. 10, pp.
249, 1998. 13351–13359, 2011.
[24] J. E. Cook and A. L. Wolf, “Software [31] M. Jans, N. Lybaert, K. Vanhoof, and
process validation: quantitatively J. M. Van Der Werf, “Business Pro-
measuring the correspondence of a cess Mining for Internal Fraud Risk
process to a model,” ACM Trans Reduction: Results of a Case Study,”
Softw Eng Methodol, vol. 8, no. 2, pp. 2008.
147–176, 1999. [32] M. Jans, M. Alles, and M. Vasarhelyi,
[25] R. Agrawal, D. Gunopulos, and F. Ley- “Process Mining of Event Logs in In-
mann, “Mining Process Models from ternal Auditing: A Case Study,” in 2nd
Workflow Logs,” in Proc. Sixth Int’l International Symposium on Ac-
Conf. Extending Database Technol- counting Information Systems, 2011.
ogy, 1998, pp. 469–483. [33] M. Jans, M. Alles, and M. Vasarhelyi,
[26] Herbst and Karagiannis, “Integrating “Process mining of event logs in au-
Machine Learning and Workflow diting: opportunities and challenges,”
Management to Support Acquisition Working paper. Hasselt University.
and Adaptation of Workflow Mod- Belgium, 2010.
els,” Int’l J Intell. Syst. Account. Fi- [34] M. Werner, N. Gehrke, and M.
nance Manag., vol. 9, no. 2, pp. 67– Nüttgens, “Towards Automated Anal-
92, Jun. 2000. ysis of Business Processes for Finan-
[27] J. Herbst and D. Karagiannis, “Inte- cial Audits,” presented at the
grating Machine Learning and Work- Wirtschaftsinformatik, Leipzig, 2013.
flow Management to Support Acqui- [35] M. Song and W. M. P. van der Aalst,
sition and Adaptation of Workflow “Towards comprehensive support for

208
APPENDIX A: PUBLICATIONS

organizational mining,” Decis. Sup- Transformation, Illustrated edition.,


port Syst., vol. 46, no. 1, pp. 300–317, vol. Volume 1 Foundations, 3 vols.
2008. Singapore: World Scientific Pub Co,
1997.
[36] Process Mining Group, “ProM,” 2013.
[Online]. Available: [Link] [45] S. Gregor and A. R. Hevner, “Position-
[Link]/prom/start. [Ac- ing and Present-ing Design Science
cessed: 08-Apr-2013]. Research for Maximum Impact,” MIS
Q., vol. 37, no. 2, pp. 337–355, 2013.
[37] fluxicon, “Process Mining and Process
Analysis - Fluxicon,” 2013. [Online]. [46] J. Venable, J. Pries-Heje, and R. Bas-
Available: [Link]. [Ac- kerville, “A comprehensive frame-
cessed: 01-Feb-2013]. work for evaluation in design science
re-search,” Des. Sci. Res. Inf. Syst.
[38] N. Müller-Wickop, M. Schultz, and M.
Adv. Theory Pr., pp. 423–438, 2012.
Peris, “To-wards Key Concepts for
Process Audits – A Multi-Method Re- [47] W. M. P. van der Aalst and C. Stahl,
search Approach,” in Proceedings of Modeling business processes : a petri
the 10th International Conference on net-oriented approach. Cambridge,
Enterprise Systems, Accounting and Mass.: MIT Press, 2011.
Logistics, Utrecht, 2013. [48] University of Hamburg, “Renew - The
[39] M. Reichert and B. Weber, Enabling Reference Net Workshop,” 2013.
flexibility in process-aware infor- [Online]. Available: [Link]
mation systems challenges, methods, [Link]/. [Accessed: 03-May-2012].
technologies. Berlin; New York: [49] yWorks GmbH, “yEd - Graph Editor,”
Springer, 2012. 2013. [Online]. Available:
[40] M. Weske, Business process manage- [Link]
ment concepts, languages, architec- ucts_yed_about.html. [Accessed: 12-
tures. Berlin; New York: Springer, Aug-2012].
2012. [50] A. Rozinat and W. M. P. van der Aalst,
[41] B. F. van Dongen and W. M. P. van “Conformance Checking of Processes
der Aalst, “Multi-phase process min- Based on Monitoring Real Behavior,”
ing: Aggregating instance graphs into Inf. Syst., vol. 33, no. 1, pp. 64–95,
EPCs and Petri nets,” in PNCWB 2005 Mar. 2008.
workshop, 2005, pp. 35–58.
[42] K. Jensen and L. M. Kristensen, Col-
oured petri nets. Springer, 2009.
[43] M. Schultz, N. Müller-Wickop, and M.
Nüttgens, “Key Information Require-
ments for Process Audits – an Expert
[Link],” in Proceedings of
the 5th International Workshop on
Enterprise Modelling and Information
Systems Architectures, Vienna, 2012.
[44] G. Rozenberg, Handbook of Graph
Grammars and Computing by Graph

209
APPENDIX A: PUBLICATIONS

10.9 Towards Automated Analysis of Business Processes for Financial Audits

Number 9
Towards Automated Analysis of Business Pro-
Title
cesses for Financial Audits
Appendix 10.9
Primary Related Chapters 6
Type Conference Paper
11th International Conference on
Conference
Wirtschaftsinformatik (WI 2013)
Reference (Werner et al., 2013)
Acceptance Rate 26%51)
VHB JQ 2.1 Ranking C (6.73)
WKWI Ranking A
ERA 2010 C
CORE 2013 C
Review Procedure Double Blinded
Number of Reviews 5
1. Michael Werner
Authors 2. Nick Gehrke
3. Markus Nüttgens
Dissertation Points 0.50
Authorship
Overall 90%
Design 90%
Realization 90%
Writing 90%
Status Published
Part of other Dissertations No
[Link]
Link ings/WI2013%20-%20Track%203%20-
%[Link]

51
Acceptance rate for the relevant track 21.6 %

210
APPENDIX A: PUBLICATIONS

Towards Automated Analysis of Business Processes for Financial Audits

MICHAEL WERNER1, NICK GEHRKE2, AND MARKUS NÜTTGENS1


1
University of Hamburg, Germany
{[Link], [Link]}@[Link]

2
NORDAKADEMIE, Elmshorn, Germany
[Link]@[Link]

Abstract. Financial audits play a significant role in the economy by safeguarding


the correctness of published financial information. Public auditors face the
challenge to audit financial statements that are created by increasingly inte-
grated and complex information systems. This paper addresses a specific prob-
lem in the auditing process. A major challenge in this process is the analysis and
audit of business processes that produce financial entries. We illustrate results
from applying business process mining techniques to extensive test and real life
data and discuss gained insights from the application for the development of
automated business process analysis methods in the context of financial audits.
Keywords: Process Mining, Financial Audits, Business Process Analysis

1 Introduction
Financial audits play a significant role in modern economies. Companies publish financial
statements in order to inform relevant stakeholders. For preventing misinformation of the
addressees financial statements are subject to audits that are mandated by law and speci-
fied in regulatory requirements. The audits are carried out by public auditors who follow
specific audit approaches for planning and executing their audits. Audit standards require
that auditors consider and test relevant business processes during the audit [1]. The re-
quirement derives from the assumption that well controlled transaction processing will
lead to valid entries on the balance sheet and profit and loss statements. When business
transactions are carried out in a correct and controlled manner they will most likely lead to
complete and accurate journal entries.
With increasing integration of the execution of business processes in information systems
and the accompanied progress in automation of transaction processing it becomes more
and more challenging to audit these processes. Contemporary audit approaches take into
account the relevance of business processes, supporting information systems and internal
control frameworks, but they basically rely on manual audit procedures to analyze and test
them. The manual procedures primarily include interviews for obtaining information and
manual test activities for evaluating relevant controls. With increasing integration of infor-
mation systems for supporting and automating transaction processing audit activities like
interviews and manual audit activities become inefficient or even ineffective due to the
increasing complexity and the mere volume of processed transactions [2].

211
APPENDIX A: PUBLICATIONS

An alternative would be the application of automated analysis and audit procedures as a


business intelligence tool that supports the auditor in the auditing process. [3] conceptually
illustrate how process mining methods can be combined with automated application con-
trol testing methods for designing automated audit methods. A requisite for such a devel-
opment are methods that allow an automated analysis of business processes. The analysis
results can then be used for automated testing purposes.
When information systems are used to support or automate the transaction processing
they also provide information that can be used for an automated analysis. By using process
mining techniques [4] and specific mining algorithms for financially relevant business pro-
cesses [5] executed process instances can be mined, reconstructed and analyzed.
In this paper we focus on the aspect of automated analysis. We apply an existing mining
algorithm for financially relevant business processes to test data and real life data. The aim
of this research is to evaluate which insights can be derived by analyzing the application of
the implemented algorithm. We statistically analyze the mined business processes in-
stances that are reconstructed from the available data to identify which further research
and improvement is needed on the path towards automated analysis methods.
We start with an illustration of related work in section two, followed by a brief description
of the applied research methodology in section three. The mined process instances are
represented as Petri nets. The used representation, the chosen mining method and the
experimental setup are explained in section four. Section five provides the results from
analyzing the process instances that were mined from the test and real life data. A discus-
sion of the gained results and an illustration of identified limitations followed by a brief
summary and conclusion close the paper.

2 Related Work
Of particular interest for the research laid out in this paper are publications from the field
of process mining. Research on process mining started in the late 1990s by [6] and has
gained extensive attention in the last decades. Significant research work has been pub-
lished by van der Aalst et al. leading to a comprehensive basic publication on process min-
ing that covers major aspects of the research domain [4].
From a financial accounting and auditing perspective requirements are outlined in relevant
audit standards. The major international standard setting body is the International Auditing
and Assurance Standards Board (IAASB) which publishes the International Standards on
Auditing (ISA). The ISA 315 (Revised) “Identifying and Assessing the Risks of Material Mis-
statement through Understanding the Entity and Its Environment“ outlines the require-
ment to consider business processes and related internal controls in order to assess the
risk for material misstatement (ISA 315.18) [1].
The role of information systems for accounting is well researched but few authors address
the role of information systems in the context of auditing. [7] describes techniques to audit
enterprise resource planning (ERP) systems, but the exploitation of information that is
available in information systems for the purpose of automated analyses is a relatively novel
field of research as illustrated by [8].

212
APPENDIX A: PUBLICATIONS

Specific research on process mining for auditing purposes has gained increased attention
over the last two to three years. [9] offers an overview of current limitations and future
challenges of process mining in the context of audits whereas [10] illustrates opportunities
of online auditing. [11–13] focus on fraud and outline possibilities of process mining for
fraud detection and auditing thereby highlighting the potential of process mining as a new
toolkit for internal audits. [5, 14] developed a mining algorithm that is able to exploit the
structure of financial journal entries for the purpose of process mining in the context of
financial audits. [15] further introduce automated audit methods for testing application
controls in ERP systems. [2, 3] finally conceptually describe how process mining techniques
for financially relevant business processes can be combined with methods for automated
control testing.
For the purpose of the research work of this paper we implemented the mining algorithm
introduced by [5, 14]. Their mining technique includes the extraction of financially relevant
information of journal entry values that are relevant for the purpose of auditing.
An alternative approach is used by [16]. They provide an interesting case study about the
examination of mined instances of a procurement process. Their approach differs from the
research presented in this paper as they actually perform a deviation analysis of the mined
process instances with a manually evaluated ideal process. They base their analysis on a
predefined set of process instances. The aim of the research work illustrated in this paper
does not focus on providing a case study for auditing mined business processes but intends
to reveal general possibilities and limitations to discover and analyze business processes
from event log data without further knowledge of the underlying processes in the context
of financial audits. As illustrated in [3] the ultimate aim is to develop methods that show
which processes are mirrored in the available logs and how they affect the financial state-
ments.
For analyzing mined process instances these need to be modeled in a purposeful modeling
language. [17] suggest using a BPMN representation for mined process instances in the
context of financial audits. Although BPMN process models might be easier to interpret for
end users we have chosen Petri nets as a modeling language for the research presented in
this paper. On the one hand a broad variety of the aforementioned research work from the
field of process mining relies on Petri nets as the choice of modeling language [18]. And
referring to Petri nets opens up the opportunity to incorporate these already existing re-
search results and techniques for the purpose of mining and analyzing. On the other hand
Petri nets have a mathematical foundation and offer a formal graphical notation. These
characteristics allow the development of sophisticated analysis methods. They therefore
constitute the preferred modeling language for the research outlined in this paper. In the
context of this paper we primarily refer to the publications of [19] and [20] concerning the
theoretical foundation for the application of Petri nets.

3 Research Methodology
The research presented in this paper follows a design science approach [21–23]. A common
critic in the academic arena refers to the perceived lack of rigor concerning design science
oriented research. In order to address this aspect we have obtained extensive test and real

213
APPENDIX A: PUBLICATIONS

life data for testing the designed artifacts. Actually the key aspect of this paper is to illus-
trate the results of evaluating already designed methods against this data. The illustrated
work follows a research process as suggested by [23] consisting of the phases analysis, de-
sign, evaluation and diffusion. The requirements for an adequate representation and mod-
eling of mined processes were investigated by considering specific, already existing litera-
ture [17] and by analyzing available test and real life data. The used mining methods were
engineered [24] by assembling parts of already available methods and by developing new
concepts where no adequate solutions were available yet.52 The analysis results and re-
quirements for further development constitute the primary outputs produced in the re-
search process that is laid out in this paper. The engineered methods were implemented in
a software prototype for evaluation purposes. We rigorously tested the software artifact
with test and real life data in order to validate it against the relevant research questions
addressed by the research work [25]. The content of this article discusses the results and
insights that have been generated by applying the designed methods to this voluminous
data.

4 Representation and Experimental Setup


[5] introduced a simple, deterministic and unsupervised mining algorithm that is suitable
for extracting data from information systems and for reconstructing executed process in-
stances. When using process mining in the context of financial audits it is necessary to mine
information that is relevant from an audit perspective and to ensure that the received in-
formation precisely reflects the executed transactions. Financial transactions in ERP sys-
tems create journal entries when they are executed. The chosen algorithm exploits the
open-item-accounting structure of journal entries that can be used to link transactions to
a process instance.53 Journal entries consist of an accounting document and at least two
entry items that are posted as credits and debits. When open-item-accounting is enabled
each cleared item has a reference to the accounting document that cleared it. The algo-
rithm starts with an arbitrary journal entry and reconstructs the links between journal en-
tries that cleared each other. It matches the events in the event log to cases that represent
process instances.54
The original mining algorithm produced directed graphs representing the mined process
instances. We extended the mining algorithm with a function for mapping the mined cases
to Petri nets and implemented it in a software artifact. The software prototype was written
in Java using the Java NetBeans IDE [26]. It provides functionality to export Petri net models
in different data formats for visual representation. The open source software Renew [27]
was used for verifying that the software artifact reconstructs reachable and therefore cor-
rect Petri nets. The yEd Graph Editor [28] was used for graphical representation and auto-
matic layout of mined process instances.
Figure 1 displays a colored Petri net (CPN) model of a reconstructed process instance. The
example shows an instance of a purchasing process. Executed transactions are modeled as
52
The engineering of the applied methods is not part of this paper. Details are available in [5].
53
The open-item-accounting is a fundamental concept of the double-entry bookkeeping which needs to be
supported by every information system used for double-entry bookkeeping.
54
Compared to other mining algorithms like the α-Algorithm [4] the implemented algorithm does not rely
on the temporal ordering of events but on their logical structure.

214
APPENDIX A: PUBLICATIONS

rectangles (Petri net transitions). The journal entry items produced by executing the trans-
actions are modeled as circles (Petri net places). The places are colored by the account
number according to the account the item was posted to. Two different types of connec-
tions are possible between transactions and journal entry items. A dotted arrow (Petri net
arc) means that a transaction has posted the connected journal entry item. A dotted line
(Petri net test arc) illustrates that a journal entry item was cleared by the connected trans-
action. Arc inscriptions play a significant role. They denote the values that are associated
to the connection between transactions and journal entry items. Each transition is accom-
panied by a start place containing a token colored with the original document number of
the journal entry. They are connected to the corresponding transaction with a simple arrow
(Petri net arc). This actually leads to an enabled CPN that mimics the behavior of the origi-
nally executed process instance.
The example in Figure 1 shows that a transaction for receiving goods (MB01) was processed
by user 2. It led to entries on the raw material account (310000) and on the goods received
/ invoices received account (191100) with the amount of 17,874.76. The invoice for these
received goods were processed (MR1M) and a payment run (F110) executed that cleared
the items posted by the MR1M transactions. The FB1S transaction was executed for clear-
ing the items that were created by MB01 and MR1M.
160000

[5000004968] [100010317] [5100004749] [2000000441]


[20555.98]
[5000004968] [100010317] [5100004749] [2000000441]
191100 191100 160000
5000004968 100010317 5100004749 2000000441

MB01 FB1S MR1M F110


[17874.76] [17874.76] [17874.76] [17874.76] [20555.98] [20555.98]
User 2 User 1 User 2 User 1

[17874.76] [2681.21] [0.00] [0.01] [19939.30] [536.24] [80.44]

310000 154000 230051 230051 113101 276000 154000

Fig. 1. Example of a Reconstructed Purchase Process Instance


For applying the mining algorithm the necessary data of the executed process instances
was extracted from the available ERP systems. The setup of the experiment including con-
ducted activities, involved software modules and input and output for each activity is illus-
trated in Figure 2. The relevant data was extracted by using a configurable extraction mod-
ule. It retrieves the event log from the ERP system by extracting data from relevant data-
base tables. The usage of a separate module provides the benefit that only the extraction
component needs to be adjusted when data is extracted from different ERP systems. The
mining algorithm operates independently from the underlying data structure of the indi-
vidual ERP systems. The extraction module loads the extracted data into an event log da-
tabase that can be accessed by the mining module. The mining module matches events in
the log to cases, reconstructs executed instances and provides functionalities for analyzing
them. It also produces output files in different formats (Extensible Graph Modeling Lan-
guage (XGML) and Petri Net Markup Language (PNML)) that can be imported into subse-
quent software for verification (Renew) and graphical presentation (yEd Graph Editor) pur-
poses. The reconstructed instances are modeled as Petri nets and stored as separate data
objects in the mining module.

215
APPENDIX A: PUBLICATIONS

Input ERP Data Event Log Process Instances Process Instances

Activity Data Extraction Mining Verification Representation

Output Event Log Process Instances Petri Nets Petri Nets

Software
Extraction Module Mining Module Renew yEd
Components

Fig. 2. Experimental Setup


We used three different data sets for analysis. The first set was extracted from the SAP IDES
test system. The test system is available for universities participating in the SAP University
Alliance Program [29]. Postings in ERP systems are stored as data entries with information
relating to the whole posting (journal entry) and the single entries that were posted to
different accounts (journal entry items). The data set contained 115,060 journal entries and
419,106 journal entry items. 81,171 process instances could be reconstructed by executing
the implemented mining algorithm. The data set included all transaction data that was
available in the test system covering a period of 17 years.
The second data set was extracted from a SAP system of a retail company. The set included
the data of all executed transactions but only for a time period of one year. The volume of
92,487 journal entries and 222,901 related journal entry items that can be traced back to
40,130 process instances over just one year illustrates the high amount of transactions that
are processed in real life environments.
This observation becomes even more evident for the third data set. It originated from a
SAP system of a manufacturing company in the health sector. It contains 1,764,773 journal
entries and 7,395,434 journal entry items. 1,035,805 process instances could be recon-
structed using the mining algorithm. Table 1 provides an overview of the different data
sets.
Table 1. Overview of Data Sets
#1 #2 #3
Data Set
SAP IDES Retail Manufacturing
Number of journal entries 115,060 92,487 1,764,773
Number of journal entry
419,106 222,901 7,395,434
items
Number of process in-
81,171 40,130 1,035,805
stances
Covered period 17 years 1 year 1 year

5 Mining Results Analysis


The mining and reconstruction of process instances from the available data sets provide
the basis for analyzing the created Petri net models. The aim of this analysis is the identifi-
cation of patterns that might help in further improvement of the mining algorithm, the
gaining of insights concerning which further research is needed for developing automated
analysis methods that can be applied in real life scenarios, and what kind of limitations for
developing such methods might exist.

216
APPENDIX A: PUBLICATIONS

The following sub-sections illustrate results from statistical analyses of the mined process
instances for all three used data sets. Due to place restrictions we limit the presentation of
results to those aspects that we consider relevant for the aforementioned aim.

5.1 Distribution of Net Size


Figures 3, 4 and 5 show the distribution of the number of process instances over the num-
ber of net elements with logarithmic scaling on the x- and y-axes for the different data
sets.55 Net elements include transitions, places and arcs. The number of net elements gives
an impression of the size and complexity of a process instance.
The charts illustrate that the distribution of the number of net elements over the number
of instances exhibit the same pattern for all data sets. Only very few instances consist of
very many net elements. The vast majority of instances consist of relatively few net ele-
ments. Table 2 provides an overview of specific characteristic values of the distributions.
The shown distributions by themselves do not allow drawing conclusions concerning the
design of automated business process analysis procedures. But the observation that the
majority of instances actually consist of relatively few elements and that this is the case for
all evaluated data sets constitutes a useful insight in combination with the analysis results
highlighted in the following sub-sections.

Fig. 3. Data Set #1 Distribution of Number of Net Elements over the Number of Instances

55
The diagrams do not show values for instances with less than three net elements. Each net consists at
least of one transaction represented by a transition in the model. Each transition is accompanied by a
start place that is connected to the transition and that enables the transition to fire. Therefore the mini-
mum number of net elements per net is three.

217
APPENDIX A: PUBLICATIONS

Fig. 4. Data Set #2 Distribution of Number of Net Elements over the Number of Instances

Fig. 5. Data Set #3 Distribution of Number of Net Elements over the Number of Instances

Table 2. Overview of Net Size Distribution Characteristics


#1 #2 #3
Data Set
SAP IDES Retail Manufacturing
Mean value of net ele-
15.31 21.28 20.30
ments per instance
Median value of net ele-
9 9 7
ments per instance
Maximum of net ele-
32,519 275,870 4,769,379
ments per instance
Standard deviation of
127.83 1,380.41 4,688.30
number of net elements

5.2 Distribution of Transaction Code Combinations


Figures 6, 7 and 8 show the distribution of transaction code combinations over the number
of instances. The y-axes follow a logarithmic scaling. Each number on the x-axis represents
a transaction code combination. This means for example that the transaction code combi-
nation number 27 (FB1S, MB01, F110, FB05, MIRO) was executed in 1,069 process instances
in data set three. These instances did not include any other transaction codes.

218
APPENDIX A: PUBLICATIONS

The distributions for all data sets show the same pattern. The majority of instances only
contain very few different transaction code combinations. Taking into account the results
from analyzing the distribution of net sizes from the previous section it is reasonable to
assume that the majority of instances are very limited in size and reveal the same transac-
tion code combinations.
A clustering of instances that exhibit the same size and the same transaction code combi-
nation could be a starting point for automatically analyzing a large amount of the mined
instances. Each cluster could be reviewed for the value that it contributes to the financial
statements and if the cluster constitutes a material process flow that needs to be further
evaluated from a materiality perspective. Such an analysis would provide useful infor-
mation to the auditor about how the processes in a company actually affect the financial
statements.
The clustering of isomorphic graphs into clusters would also enable to search for applica-
tion controls that affect the cluster under review. In combination with automated applica-
tion control testing the isomorphic process instances in a cluster could be automatically
audited.

Fig. 6. Data Set #1 Distribution of the Number of Instances for Different Transaction
Code Combinations

Fig. 7. Data Set #2 Distribution of the Number of Instances for Different Transaction
Code Combinations

Fig. 8. Data Set #3 Distribution of the Number of Instances for Different Transaction
Code Combinations

219
APPENDIX A: PUBLICATIONS

5.3 Distribution of Accounts


We suggested in the previous section that clustering could be a promising solution for an-
alyzing similar and small instances. This leads to the question how large and complex in-
stances should be handled. The instance in Figure 1 contains four transitions and consists
of 38 net elements. An instance of this size and complexity can be evaluated by simple
observation. This is not the case anymore for more complex instances. Figure 9 shows a
process instance containing 3,057 transitions and consisting of 15,319 net elements. These
kinds of instances cannot be evaluated without further consideration.
Figures 3 to 5 show that the number of net
elements extending 100 net elements is rel-
atively small in all the data sets. Only 691 in-
stances in data set one consist of 100 or
more net elements, 428 in data set two and
7,470 in data set three. Although the num-
ber of complex instances is relatively low
they might represent a material amount of
transactions that affect the financial state-
ments and can therefore not be neglected.
For analyzing and evaluating these in-
stances it is necessary to reduce their com-
plexity. The analysis of the accounts that are
actually used within an instance provides a
starting point for investigating complexity Fig.9. Complex Process Instance
reduction possibilities. The distributions of
the accounts over the number of process instances in Figures 10 to 12 show on how many
accounts journal entry items were posted to in a single process instance. For example 222
instances in data set two used seven accounts to post items.56
The distributions for all three data sets again display the same pattern. Table 3 provides an
overview of characteristic values of these distributions. The maximum number of used ac-
counts is relatively low compared to the maximum net sizes. For interpreting the illustrated
distributions and their characteristic values it is necessary to consider that only one in-
stance in data set two uses the maximum number of 596 accounts and one in data set three
the maximum of 399 accounts. All other instances do not use more than 63 accounts in
data set two and no more accounts than 45 in data set three.
The observation of the distributions reveals that journal entry items are posted to relatively
few accounts. This is reasonable because a specific process generally only uses a subset of
the available set of accounts. The execution of a purchase process for example would likely
lead to journal entries on expense accounts but not on sales accounts.
Based on the observation that the number of used accounts is relatively small it might be
useful to aggregate journal entry items that were posted to the same accounts and thereby
reducing the net size and complexity. Items are modeled as places in the Petri net and col-
ored by the account number reflecting the financial account the item was posted on. The

56
A logarithmic scaling is again used for the x- and y-axes.

220
APPENDIX A: PUBLICATIONS

places carrying the same color (account number) could be folded leading to Petri net mod-
els that contain significantly less net elements and therefore a reduced complexity. Further
research would actually be needed for verifying if a sufficient complexity level can be
reached by this approach.
Table 3. Overview of Account Distribution Characteristics

#1 #2 #3
Data Set
SAP IDES Retail Manufacturing
Mean value 2.95 2.30 2.33
Median value 2 2 2
Maximum value 36 596 399
Standard deviation 1.61 3.37 0.92

Fig. 10. Data Set #1 Distribution of the Number of Instances over Number of Accounts

Fig. 11. Data Set #2 Distribution of the Number of Instances over Number of Accounts

Fig. 12. Data Set #3 Distribution of the Number of Instances over Number of Accounts

221
APPENDIX A: PUBLICATIONS

6 Discussion and Limitations


The previous sections lay out insights that can be derived from analyzing results from the
application of a mining algorithm. When considering these results it is necessary to keep in
mind that they only provide starting points for further research and that no designed meth-
ods and their implementations do exist yet that could prove or disprove the postulated
assumptions.
A further restriction is the number of used data sets. [16] relied on an event log represent-
ing 26,185 procurement process instances from a single company in the financial sector.
We therefore consider the size of the data sets as being sufficient and we also did not limit
our analysis to a specific type of business process but we actually only cover two industries
– retail and manufacturing. We do not know if the identified characteristics also hold true
for other industries.
As a result of the analysis of used accounts we proposed the folding of places in a Petri net
carrying the same account number for complexity reduction purposes. From an audit per-
spective it is crucial that the received information exactly mirrors the transaction that really
took place and therefore to maintain the audit trail [30]. Applying a folding approach would
actually lead to a graph transformation of the mined process model and it would need to
be ensured that the behavior of the net remains stable [17].
A further limitation relates to the execution of the programmed software and the time that
is needed to calculate the mined instances. Dedicated research on the performance of the
implemented software prototype is currently outstanding but our experience from analyz-
ing the data sets presented in this paper lets us assume that the complete reconstruction
of all process instances might be unrealistic in a real life scenario and that a sampling ap-
proach based on a materiality perspective might be needed.

7 Summary and Conclusion


Research on process mining has advanced and matured significantly over the past two dec-
ades and is now increasingly applied to specific application domains. The focus of the re-
search presented in this paper lies on the domain of financial audits. The audit of business
processes is a mandatory step in the auditing process that becomes increasingly challeng-
ing with the ongoing integration of information systems and automation of transaction
processing. The application of process mining for supporting the auditor comprises a prom-
ising alternative to counter the growing complexity and amount of processed transactions
that need to be evaluated during an audit.
We implemented a mining algorithm in a software artifact and evaluated it against volumi-
nous test and real life data. Based on the results derived from the analysis of reconstructed
process instances we gained insights that can be used for developing automated process
analysis methods in the context of financial audits.
The majority of process instances is small and consists of only few net elements. The clus-
tering of instances that exhibit the same size and the same transaction code combination
could be a starting point for automatically analyzing a large amount of the mined instances.
Each cluster could be reviewed for the value that it contributes to the financial statements.

222
APPENDIX A: PUBLICATIONS

Such an analysis would provide useful information to the auditor about how the processes
in a company actually affect the financial statements.
Process instances containing more than a few processed transactions cannot be evaluated
manually by simple observation. Although the overall number of complex process instances
is very limited they cannot be neglected from an audit perspective. A single instance may
already contain extremely high volumes of transactions that might be material. A starting
point for reducing complexity could be the consideration of used accounts. Even for large
and complex process instances the number of used accounts in a single process instance is
relatively small. The folding of equally colored places in process instances could lead to
significant complexity reduction especially for large process instances.
The research presented in this paper can be seen as a step towards the development of
automated business process analysis methods. The accounting scandals of major compa-
nies over the last years illustrates that the audit industry is currently lacking adequate so-
lutions for safeguarding the correctness of published financial statements. The usage of
automated analysis and audit methods constitutes a necessary requirement to overcome
the existing imbalance between automated processing on the companies’ side and manual
audit procedures on the auditors’ side. The introduction of automated audit procedures
has the potential to leverage this imbalance.
Limitations concerning the applicability of the implemented algorithm exist especially in
regard to processing time and the amount of process instances that can be mined and an-
alyzed. At this point of time it also remains unclear if the identified starting points can ac-
tually be transferred successfully into the design of automated business process analysis
methods. But first results from consecutive research that bases on the results presented in
this paper that will be published in forthcoming articles provide a positive indication.57

References
1. International Federation of Accountants: ISA 315 (Revised), Identifying and As-
sessing the Risks of Material Misstatement through Understanding the Entity and Its
Environment, (2012).
2. Werner, M., Gehrke, N.: Potentiale und Grenzen automatisierter Prozessprüfungen
durch Prozessrekonstruktionen. Forschung für die Wirtschaft. Shaker Verlag (2011).
3. Werner, M., Gehrke, N., Nüttgens, M.: Business Process Mining and Reconstruction
for Financial Audits. Hawaii International Conference on System Sciences. pp. 5350–
5359. Maui (2012).
4. Van der Aalst, W.M.P.: Process Mining: Discovery, Conformance and Enhancement
of Business Processes. Springer, Berlin, Heidelberg (2011).
5. Gehrke, N., Müller-Wickop, N.: Basic Principles of Financial Process Mining A Jour-
ney through Financial Data in Accounting Information Systems. Proceedings of the
16th Amer-icas Conference on Information Systems. Lima, Peru (2010).

57
The presented results were developed in the research project Virtual Accounting Worlds. The project is
sponsored by the German Federal Ministry of Education and Research (grant number 01IS10041). The
authors are responsible for the content of this publication.

223
APPENDIX A: PUBLICATIONS

6. Agrawal, R., Gunopulos, D., Leymann, F.: Mining Process Models from Workflow
Logs. Proc. Sixth Int’l Conf. Extending Database Technology. pp. 469–483 (1998).
7. Musaji, Y.F.: Integrated Auditing of ERP Systems. John Wiley & Sons (2002).
8. Alles, M., Jans, M., Vasarhelyi, M.: Process Mining: A New Research Methodology
for AIS. CAAA Annual Conference 2011 (2011).
9. Jans, M.J.: Process Mining in Auditing: From Current Limitations to Future Chal-
lenges. In: Daniel, F., Barkaoui, K., and Dustdar, S. (edr.) Business Process Manage-
ment Workshops. pp. 394–397. Springer, Berlin, Heidelberg (2012).
10. Aalst, W.M.P. van, Hee, K.M. van, Werf, J.M. van, Verdonk, M.: Auditing 2.0: Using
Process Mining to Support Tomorrow’s Auditor. Computer. 43, pp. 90–93 (2010).
11. Jans, M., Lybaert, N., Vanhoof, K., Van Der Werf, J.M.: Business process mining for
in-ternal fraud risk reduction: Results of a case study. (2008).
12. Jans, M., Alles, M., Vasarhelyi, M.: Process mining of event logs in auditing: oppor-
tunities and challenges. Working paper. Hasselt University, Belgium (2010).
13. Jans, M., Van der Werf, J.M., Lybaert, N., Vanhoof, K.: A business process mining
application for internal transaction fraud mitigation. Expert Systems with Applica-
tions. 38, pp. 13351–13359 (2011).
14. Gehrke, N., Müller-Wickop, N.: Rekonstruktion von Geschäftsprozessen im Finanz-
wesen mit Financial Process Mining. Lecture Notes in Informatics, Proceedings der
Jahrestagung Informatik. Leipzig (2010).
15. Gehrke, N.: The ERP Auditlab-A Prototypical Framework for Evaluating Enterprise
Re-source Planning System Assurance. 43rd Hawaii International Conference on Sys-
tem Sciences (HICSS) (2010).
16. Jans, M., Alles, M., Vasarhelyi, M.: Process Mining of Event Logs in Internal Audit-
ing: A Case Study. 2nd International Symposium on Accounting Information Systems
(2011).
17. Müller-Wickop, N., Schultz, M., Gehrke, N., Nüttgens, M.: Towards Automated Fi-
nancial Process Auditing: Aggregation and Visualization of Process Models. Proceed-
ings of the Enterprise Modelling and Information Systems Architectures. Germany
(2011).
18. Tiwari, A., Turner, C., Majeed, B.: A review of business process mining: state-of-the-
art and future trends. Business Process Management Journal. 14,pp. 5–22 (2008).
19. Van der Aalst, W.M.P., Stahl, C.: Modeling business processes : a petri net-oriented
approach. MIT Press, Cambridge, Mass. (2011).
20. Valk, R.: Lecture Notes: Formale Grundlagen der Informatik II (FGI 2) Modellierung
& Analyse paralleler und verteilter Systeme. University of Hamburg (2008).
21. Hevner, A.R., March, S.T., Park, J., Ram, S.: Design science in information systems
re-search. Mis [Link]. 75–105 (2004).
22. March, S.T., Smith, G.F.: Design and natural science research on information tech-
nology. Decision support systems. 15,pp. 251–266 (1995).

224
APPENDIX A: PUBLICATIONS

23. Österle, H., Becker, J., Frank, U., Hess, T., Karagiannis, D., Krcmar, H., Loos, P.,
Mertens, P., Oberweis, A., Sinz, E.J.: Memorandum on design-oriented information
systems research. European Journal of Information Systems. 20, pp.7–10 (2010).
24. Brinkkemper, S.: Method engineering: engineering of information systems develop-
ment methods and tools. Information and Software Technology. 38, pp.275–280
(1996).
25. Riege, C., Saat, J., Bucher, T.: Systematisierung von Evaluationsmethoden in der ge-
staltungsorientierten Wirtschaftsinformatik. Wissenschaftstheorie und gestaltungs-
orientierte Wirtschaftsinformatik. pp.69–86 (2009).
26. Oracle: Welcome to NetBeans, [Link]
27. University of Hamburg: Renew - The Reference Net Workshop, [Link]
[Link]/.
28. yWorks GmbH: yEd - Graph Editor, [Link]
ucts_yed_about.html.
29. SAP: SAP-UCC, [Link]
30. Romney, M.B., Steinbart, P.J.: Accounting Information Systems. Prentice Hall
(2008).

225
APPENDIX A: PUBLICATIONS

10.10 Multilevel Process Mining for Financial Audits

Number 10
Multilevel Process Mining for
Title
Financial Audits58)
Appendix 5.5, 6
Primary Related Chapters 10.10
Type Journal Paper
Journal IEEE Transactions on Services Computing
Reference (Werner and Gehrke, 2015)
59)
Acceptance Rate
VHB JQ 2.1 Ranking -
WKWI Ranking A
ERA 2010 -
CORE 2013 -
Review Procedure Blinded
Number of Reviews 360)
1. Michael Werner
Authors
2. Nick Gehrke
Dissertation Points 0.67
Authorship
Overall 95%
Design 95%
Realization 95%
Writing 95%
Status Published61)
Part of other Dissertations No
[Link]
Link
[Link]?reload=true&arnumber=7277120

58
The following article is the first revised version of the originally submitted manuscript. It reflects the status
at the time of the submission of this dissertation. The final version was published in December 2015 and
partly differs from the version presented in this thesis.
59
The journal does not communicate acceptance rates for special or regular issues.
60
The number of reviews refers to the first review round, three more reviews were received during the sec-
ond review round.
61
Paper was accepted for publication for the IEEE Transactions on Services Computing special issue on Pro-
cesses Meet Big Data and published in December 2015, initial submission in August 2013, passed the first
review round in December 2013 and second review round in July 2015.

226
APPENDIX A: PUBLICATIONS

Multilevel Process Mining for Financial Audits

MICHAEL WERNER, NICK GEHRKE


Abstract: The relevance of business intelligence increases with the growing
amount of recorded data. The research on business intelligence has led to a
mature set of methods and tools that are used in many application areas, but
they are almost absent in the auditing industry. The audit of business processes
is a critical part in financial audits. Public accountants face the challenge of be-
ing obligated to audit increasingly complex business processes that process
huge amounts of transaction data. Process mining can be used as a business
intelligence approach in the context of process audits to exploit this data. Key
requirements for the application of process mining in financial audits are the
reliability of the mining results, the integration of a data flow perspective and
the ability to inspect data from the point of origin to the final output on the
financial accounts. We introduce a process mining algorithm that integrates the
control flow and data flow perspective. It operates on different abstraction lev-
els, creates precise and fitting process models, accepts specific unlabeled event
logs as input, and considers data relationships for inferring the control flow.
Index Terms: Business intelligence (BI), Financial Audits, Business Process Intel-
ligence, Process Mining, Data Mining, Data Analysis, Business Process Model-
ing, ERP Systems, Design Science

Standard on Auditing (ISA) 315 (Revised)


1 Introduction mandates the consideration of business
Financial audits are important for the processes and related information systems
smooth functioning of economic markets. for financial audits: “The auditor shall ob-
They are a control mechanism to prevent tain an understanding of the information
the publication of false financial infor- system, including the related business pro-
mation. The reliability of published financial cesses, relevant to financial reporting (…)”
statements is crucial for stakeholders to di- [1, p. 7]. It is much more efficient to investi-
rect their decisions. Governments have is- gate the structure and embedded controls
sued laws and regulations to ensure that in business processes than to inspect single
companies prepare their financial state- transactions. The task of auditing business
ments truthfully and fairly. Public account- processes is getting more and more chal-
ants act as referees ensuring the adherence lenging with the increasing integration of in-
to laws, regulations and accounting stand- formation systems for the automation of
ards by auditing financial statements. A sig- transaction processing and the growing
nificant part of a financial audit is the audit- amount of produced data. Traditional audit
ing of business processes. The rationale of procedures like interviews and inspections
auditing business processes is the assump- of selected documents become inefficient
tion that well-controlled business processes or even ineffective in such audit environ-
will lead to complete and correct postings ments [2]. Interview partners may no longer
on the financial accounts. The International have overall information about a business
process if part of it is operated automated

227
APPENDIX A: PUBLICATIONS

and nontransparent in the information sys- technologies, systems, practices, methodol-


tem without any human interaction. Fur- ogies, and applications that analyze critical
thermore it is questionable if the inspection business data to help an enterprise better
of relatively few samples is a sufficient audit understand its business and market and
procedure when millions of transactions are make timely business decisions.” [4, p.
processed. 1166].
The current situation in auditing engage- BI provides mature methods and tools that
ments leads to an imbalance. Companies on are used in many application areas. But they
the one hand use highly integrated infor- are almost absent in the auditing industry.
mation systems to support and automate Software tools that assist the auditor are
the operation of business transactions lead- called computer assisted audit tools (CAAT).
ing to huge amounts of processed data. The Braun and Davis state that CAAT “include
auditors on the other hand apply traditional any use of technology to assist in the com-
and mainly manual audit procedures that pletion of an audit” [9, p. 726]. This is a very
are not appropriate anymore in highly auto- broad definition and it would mean that
mated environments. The results are ineffi- simple project management and documen-
cient or even ineffective audits. tation tools could also be considered as
CAAT. Gehrke describes a more specific ap-
Problems arising from the increase of pro-
proach by illustrating how information on
cessed and available data are not idiosyn-
controls that are embedded in information
cratic to the field of financial audits. Data
systems can be tested using a prototypical
explosion is a phenomenon that accompa-
software tool for the purpose of financial
nies the expanding availability of digital
audits [10]. But the presented approach
data storage capacity [3]. The challenges
only uses a very small part of the data that
and prospects that arise from the handling
is stored in information systems and merely
and analysis of huge data amounts are cur-
focusses on the control perspective. The
rently being discussed and intensively inves-
richest source of information is the rec-
tigated under the umbrella of the term Big
orded transaction data. This source remains
Data [4]. Big Data does not only occur in the
mostly untouched by existing CAAT. Trans-
field of information systems or computer
action data stored in Enterprise Resource
science but also in very diverse scientific dis-
Planning (ERP) systems exhibit characteris-
ciplines like biology [5] and physics [6]. Hu-
tics of Big Data due to its volume and the
man recipients are only able to handle a cer-
velocity of data accumulation [4].
tain amount of information before infor-
mation overload hinders any additional in- A promising solution to exploit the available
formation reception [7, pp. 53–57]. transaction data in financial audits is the use
of process mining. Process mining deals
Business intelligence (BI) provides solutions
with the discovery, monitoring and en-
to handle Big Data and to prevent infor-
hancement of business processes by ex-
mation overload. It is a scientific domain
tracting information from event logs [11].
that traditionally researches how digital
The usage of process mining would enable
data can be analyzed and presented to en-
an auditor to receive reliable information
hance decision making processes [8]. Chen
for the audited business processes very effi-
et al. discuss the evolution of business intel-
ciently and effectively. Van der Aalst pro-
ligence and analytics (BI&A) and state that it
vided a conceptual model for the integra-
is commonly “referred to as the techniques,
tion of process mining into auditing [12] and

228
APPENDIX A: PUBLICATIONS

scholars have applied process mining in the specific data relationships in this data can
context of internal audits [13]. But process be exploited for process mining purposes
mining tools still have not been widely ac- [19]. The mining algorithm presented in this
cepted and applied in the auditing industry. paper combines and visualizes the control
flow and the data flow perspective. It ac-
A main reason is the specific characteristic
cepts unlabeled event log data from ERP
of the application domain. Process mining
systems as input, produces perfectly fitting
algorithms can only be applied usefully if
and precise process models, and uses data
they fit the requirements of the application
dependencies to determine the control
domain. The research area of process min-
flow. It is especially suitable as a special pur-
ing has matured during the last decade with
pose mining algorithm for financial audits.
the development of powerful general pur-
But the presented approach for combining
pose mining algorithms such as heuristic
the control flow and the data flow perspec-
[14], fuzzy [15] or genetic [16] mining algo-
tive in a single model might also be valuable
rithms. However, a significant aspect has
for other application domains that exhibit
not been investigated intensively yet. The
similar requirements.
vast majority of mining algorithms focusses
on the control flow perspective. Other per-
spectives like the data flow perspective are
2 Research Methodology and
neglected [17]. But the data flow is very im- Structure
portant for financial audits. If a process is
The research presented in this paper follows
not compliant the auditor needs to assess
a design science research approach (DSR).
the impact on the financial accounts. This
This approach was chosen because of the
requires the integration of financially rele-
proximity of the investigated research ques-
vant information. Process mining is com-
tions to the practical problems and the in-
monly used to condense information by ab-
tention to develop artifacts that have a high
straction. Process models abstract from the
contribution and relevance for the applica-
observed behavior of single process execu-
tion domain. Österle et al. suggest following
tions. It is generally necessary to weigh
a research process that consists of the four
competing quality criteria against each
phases analysis, design, evaluation and dif-
other when mining process models. Simple
fusion [20]. The structure of this article fol-
process models are normally preferred to
lows these phases. The analysis phase is
complex ones even if this means that these
represented in the subsequent sections that
simple models are not as precise and fitting.
illustrate the state-of-the-art of related sci-
Auditors have to rely on the correctness of
entific work and discuss the requirements
mined models to identify and to assess com-
for the development of the presented min-
pliance violations. They therefore require
ing algorithm. The design of the mining al-
perfectly fitting and precise process models.
gorithm as the core contribution of the pa-
The auditor also has to be able to inspect in-
per is presented in section five. Scholars like
dividual process executions to find out
Hevner et al. stress the importance of re-
which business transactions caused a viola-
search rigor in DSR [21]. We have therefore
tion. Common mining algorithms use la-
included an evaluation section that illus-
beled event logs. They contain ordered
trates the results that have been achieved
events that are mapped to cases [18]. Finan-
by applying the presented algorithms to ex-
cially relevant transaction data stored in
tensive real world data. The paper closes
ERP systems cannot be used to create such
an event log without prior preparation. But

229
APPENDIX A: PUBLICATIONS

with a discussion of the research results, rel- 3 Related Work


evant limitations, outlook to future re-
search and a brief summary. The research presented in this paper deals
with the question as to how process mining
We used different research methods with
can be used as a BI approach for financial
varying paradigmatic orientation during the
audits. Process mining is a research domain
research process to reduce the risk of para-
that has matured over the past decades. It
digmatic bias and to investigate the re-
would go beyond the scope of this paper to
search problem from different angles [22].
provide a complete overview of process
The results presented in this paper are part
mining research. Instead we will refer to lit-
of larger research effort that has been car-
erature that provides a good overview. Ti-
ried out by several researchers over the re-
wari et al. provide a survey of the state-of-
cent years. The development of the pre-
the-art and future trends in process mining
sented algorithm bases on the application
until the year 2008 [30]. Basic and advanced
domain requirements that were identified
process mining concepts have been com-
in empirical studies using expert interviews
prehensively summarized by van der Aalst
and surveys [23]. The main research meth-
[3]. He provides an extensive collection of
ods that were used to develop the pre-
main research results that have been
sented algorithm were method engineering
achieved in recent years and presents an
[24] and prototyping [25]. Method engi-
overview of contemporary opportunities
neering is an approach for assembling novel
and challenges [31].
methods based on already existing or newly
developed method fragments. Research Compliance checking is a process mining ap-
outcomes from prior research have been in- plication area that is of particular interest to
corporated as input for the presented re- the research at hand. Compliance can be
search to develop a novel algorithm. Rele- defined as the adherence to internal or ex-
vant research results that have been consid- ternal rules. The main objective of financial
ered in the course of method engineering audits is to ensure that companies adhere
include solutions for the calculation of in- to accounting standards, laws and regula-
stance graphs, projections, aggregation tions. Debreceny and Gray suggest the ap-
graphs [26], and causal matrices [16] as well plication of data mining methods for analyz-
as previously published research results ing journal entries and provide an extensive
from the authors concerning the integration case study [32]. Becker et al. have re-
of the data perspective into process models searched the applicability of model-based
[27], case matching [19], complexity reduc- business process compliance-checking ap-
tion of mined models [28], and data de- proaches and developed a classification
pendent control flow inference [29]. The de- framework [33]. They provide a literature
signed algorithm was instantiated in a soft- review and distinguish between forward-
ware prototype to enable laboratory simu- and backward-compliance checking ap-
lation experiments. Data sets from industry proaches. Process mining for financial au-
partners were used as input for the proto- dits is a backward-compliance checking ap-
type to evaluate the proper functioning of proach because its objective is the disclo-
the implemented algorithm and to analyze sure of compliance violations after they
the mining results. have occurred and been recorded in the
event log. The application of process mining
as a compliance checking approach has al-
ready been addressed by different scholars.

230
APPENDIX A: PUBLICATIONS

Alles et al. propose the application of pro- colored tokens in Colored Petri Nets (CPN)
cess mining in accounting information sys- [39]. We follow a similar approach and use
tems [34]. Jans et al. highlight opportunities CPN to include the data perspective by
and challenges for using process mining as modeling data objects as colored tokens
an audit tool [35] and provide interesting [27]. This allows us to present the control
case studies [13], [36]. They focus on the flow and data flow perspective in a single
control flow and organizational perspective. model.
An important difference between internal
A fundamental challenge for process mining
and external audits is the relevance of the is the balancing between competing quality
data perspective. Müller-Wickop et al. con-
criteria [11]. Process mining is generally
ducted empirical research which shows that
used to reduce complexity by visual repre-
internal and external auditors do not neces-
sentation and abstraction. A process model
sarily share the same perspective on the im-
represents a set of process executions
portance of different application domain
which are called process instances.62 Multi-
constructs [23], [37]. We will discuss the re- ple executions of a business process com-
quirements needed to use process mining in
monly do not occur in exactly the same
financial audits in greater detail in the sub-
manner. Variety in the execution leads to
sequent section. But it is important to men-
differing process instances. Every organiza-
tion that the inclusion of the data perspec-
tion needs flexibility to adapt business activ-
tive is a key requirement which has not
ities to changing customer demands and
been considered by prior research. The
market influences. A certain degree of devi-
lion’s share of process mining research
ation is therefore neither surprising nor
deals with the discovery of the control flow
damaging. But it is an obstacle to process
whereas the integration of the data per-
mining because the objective of using pro-
spective in process mining has generally not
cess mining is to discover process models
been investigated extensively in the aca-
that describe the real business processes in
demic community yet, apart from very few
the best possible way. Variance in the
scientific publications. This observation is
course of execution means that it can get
supported by Stocker [38] and de Leoni and
impossible to create a model that unambig-
van der Aalst [17]. De Leoni and van der
uously describes the represented process.
Aalst use the data flow perspective to dis-
This leads to the phenomenon of miss-fit-
cover rules that explain why instances of the
ting process models. Models are either un-
same process follow different execution
der- or over-fitting. A model is under-fitting
paths. They introduce variables as net com-
if it allows execution paths in the process
ponents for the extension of Petri Nets
model that are not represented in the event
(DPN-nets). Trcka et al. follow a different re-
log and over-fitting if they do not allow for
search question to discover data flow errors
any additional behavior that is not included
but apply a similar approach by using ex-
in the event log. Rozinat et al. provide a
tended workflow nets (WFD-nets). Accorsi
framework for the evaluation of process
and Wonnemann choose a different ap-
mining algorithms. They identify four qual-
proach to identify information leaks in pro-
ity criteria for the evaluation: fitness, preci-
cess models. They include data objects as

62
The term “process instance” and “case” are used ness processes whereas “cases” represent a rec-
ambiguously among scholars. We refer to “pro- ord of a process instance in an event log. A case
cess instances” as real world executions of busi- is therefore a purposeful abstraction of a process
instance represented as a data record.

231
APPENDIX A: PUBLICATIONS

sion, generalization and structure [40]. Fit- 4 Requirements


ness indicates if a model is able to represent
all cases in the event log. Precision is the This section describes the requirements for
complementary criterion. It indicates if a using process mining algorithms in financial
process model does not allow additional be- audits. A business process consists of activi-
havior that was not observed in the log. ties. Information systems that support or
Generalization addresses the capability of a automate the execution of activities create
model to express more behavior than rec- journal entries that are posted to the rele-
orded in the log. It is generally desirable to vant financial accounts. Fig. 1 illustrates this
create a process model that shows an ade- relationship and provides an example of a
quate degree of generalization. Structure simple purchase process that consists of
refers to the graphical representation of a four activities.
business process and depends on the graph- Goods Received /
ical components of the target language. Raw Materials Invoices Received
Trade Payables Bank Account

10,000
10,000 10,000 10,000 10,000 10,000
Other scholars use the closely related crite-
rion of simplicity instead of structure [11].
The following section will show that con- Order Goods Receive Goods Receive Invoices Pay Invoices

trary to many other application domains fi-


nancial audits require perfectly fitting and Fig. 1. Simple purchase process
as precise process models as possible. De Three of the four illustrated activities create
Medeiros et al. suggest the clustering of journal entries on different accounts. Mül-
cases that exhibit similar traces in the event ler-Wickop et al. conducted empirical inves-
log.63 Process models are then generated for tigations using expert interviews and sur-
each cluster [41]. This approach prevents veys to identify key concepts and infor-
over-generalization and is very useful be- mation requirements for process audits
cause it does not require using a specific [23], [44]. They show that the process flow
mining algorithm. A similar trace clustering is a central concept and highly relevant from
approach is used by Song et al. [42]. Van der an audit perspective in practice. To audit a
Aalst et al. provide a powerful mining algo- process it is necessary to understand how a
rithm for balancing between over- and un- process is structured and which control ac-
der-fitting by using the theory of regions to tivities are included that safeguard the cor-
create Petri Nets from transition systems rect processing of transactions. This re-
[43]. The approaches by de Medeiros et al., quirement can be satisfied by modeling the
Song et al. and van der Aalst et al. are very control flow perspective. Two further im-
valuable but they do not consider the data portant concepts for external auditors are
perspective and the idiosyncratic structure financial statements and materiality64.
of journal entries that form the basis for the Knowledge about the interaction among ac-
event log in financial audits. tivities alone is not sufficient. For the audi-

63 64
A trace is the recorded sequence of executed ac- Materiality is defined in ISA 320: “Misstatements,
tivities in a process instance. Every case has a including omissions, are considered to be mate-
specific trace but different cases can exhibit iden- rial if they, individually or in the aggregate, could
tical traces. reasonably be expected to influence the eco-
nomic decisions of users taken on the basis of
the financial statements.” [45].

232
APPENDIX A: PUBLICATIONS

tor it is necessary to understand how the ac- false negative audit results and unnecessary
tivities relate to the financial accounts as il- investigations by the auditor. The quality
lustrated in Fig. 1. criteria identified by Rozinat et al. [40] are
useful to express the requirement of accu-
Only those business transactions are in-
rate process models in terms that are appli-
spected in a financial audit that can have a
cable for the process mining research do-
material effect on the financial statements.
main. Simplicity of a process model is a pre-
It is therefore necessary to receive infor-
ferred characteristic but it is not a key re-
mation on the value flow that is created by
quirement in financial audits. Auditors cur-
the audited business process to decide if it
rently spent weeks trying to understand a
needs to be audited from a materiality per-
business process with the use of traditional
spective or if it can be neglected.
audit procedures and by reviewing hun-
Another critical requirement in financial au- dreds of documents. It is therefore accepta-
dits is the preservation of the audit trail. The ble if process models are complex and in ex-
audit trail is a fundamental concept in finan- treme cases only com comprehensible to
cial accounting. It is a path in an information experts. Nevertheless a mining algorithm
system that allows tracing a transaction should be able to deliver process models as
from the point of origin to the final output. simple as possible. Generalization should be
It is used to verify the accuracy and validity minimized and precision maximized to pre-
of journal entries [46]. Translating this re- vent false negative compliance testing re-
quirement into the context of process min- sults. Process models should be perfectly
ing implies that a mining algorithm may not fitting to possibly represent all recorded be-
alter the original data during the mining havior and to prevent that incompliant be-
process. A suitable mining algorithm must havior that actually occurred remains unde-
further be able to present the unchanged tected and therefore reduce the audit effec-
source data to the auditor for investigation tiveness. Different metrics can be used to
purposes. But on the other hand the mining measure fitness (completeness [47], PFcom-
algorithm should also be able to present in- plete [16], fitness (f) [48], parsing measure
formation at an adequate abstraction level (PM) and continuous parsing measure
to provide an overview of the control and (CPM) [14]. The metrics completeness and
data flow as discussed before. If process PM measure calculate the percentage of
mining is used in financial audits the gener- traces in the log that can be replayed by the
ated models are used to discover incompli- model. The other three metrics consider
ant behavior. The provided process models both traces and tasks in a model. A process
should therefore be as precise and fitting as model has a perfect fitness if the metrics
possible. If the produced process models have a value of one. The precision dimen-
are over-fitting certain behavior recorded in sion can also be measured by using different
the event log is not represented in the pro- metrics (soundness [47], behavioral appro-
cess model. The auditor would therefore as- priateness [48], and behavioral precision
sume that no incompliant behavior has oc- [16]). A perfect precision is reached if the
curred. But in reality it is just not repre- relevant metrics also take on the value of
sented in the miss-fitting process model one.
which would eventually lead to false posi-
tive audit results. If the process model is too
general, process behavior is illustrated that
actually did not occur. This would lead to

233
APPENDIX A: PUBLICATIONS

5 Multilevel Process Mining The algorithm is called Multilevel-Process-


Mining algorithm (MLPM) due to its main
5.1 Mining Algorithm feature of being able to produce process
The previous section discussed the require- models at different abstraction levels. The
ments for using process mining in financial MLPM first matches event data to cases,
audits. These requirements are partially then creates instance graphs, calculates
conflictive. Listing 1 shows an abstract and process instance models and aggregates in-
simplified version of the mining algorithm stance models to process models. It is di-
that was developed based on the identified vided into four main sections. Each section
requirements. A fundamental aspect of the is discussed in detail in the following subsec-
presented solution is to satisfy different re- tions.
quirements at different abstraction levels.

Listing 1 Multilevel Process Mining Algorithm


46. Mine Cases (Section 1)
47. D set of all posting document numbers
48. J set of all journal entry item numbers

Di = ∅ initially empty set of document numbers belonging to case i ∈ ID


49. ID set of all case IDs

Ji = ∅ initially empty set of journal entry item numbers for case i ∈ ID


50.

52. While D ≠ ∅
51.

Remove d ∈ D from D and insert d into Di


Insert all j ∈ J posted by d into Ji and remove j from J
53.

Insert all d ∈ D that cleared j ∈ Ji into Di and remove d from D


54.

Repeat 54. and 55. for all d ∈ Di and j ∈ Ji


55.
56.

IG = ∅
57. Reconstruct Instance Graphs (Section 2)
58. initially empty set of instances graphs

Ni ≠ ∅, Ei⊆ Ni× Ni is the set of arcs, Li the set of task labels


59. IGi instance graph (Ni,Ei,Li,li) for case i with the set of nodes

and li: Ni→ Li is a labeling function mapping nodes onto Li


60.

62. For all i ∈ ID


61.

Create n ∈ Ni for each d ∈ Di with li(n)= transaction code of d


Create e(nj,nk) ∈ Ei for each dj and dk ∈ Di if dk cleared an item
63.

j ∈ Ji that was posted by dj


64.
65.
66. Insert IGi into IG

IM = ∅
67. Reconstruct Instance Models (Section 3)

instance model (Ti,Pi,Ai,Σi,Vi,Ci,Gi,Ei,Ii)for case i ∈ ID


68. initially empty set of instances models

70. For all I i ∈ ID


69. IMi

For each e(nj,nk)∈ Ei create p ∈ Pi, a(tj,p), a(p,tk)


71. Set Ti = Ni

For each d ∈ Di
72.

Create p ∈ Pi for each j ∈ Ji that was posted by d


73.

Create a(t,p) ∈ Ai for each j ∈ Ji that was posted by d and


74.

a(t,p), a(p,t) ∈ Ai for each j ∈ Ji that was cleared by d


75.

Aggregate all places pk and pj ∈ Pi if Ci(pk) = Ci(pj)


76.

Aggregate all transitions tk and tj ∈ Ti if li(tk) = li(tj)


77.
78.
79. Insert IMi into IM

234
APPENDIX A: PUBLICATIONS

PM = ∅
80. Mine Process Models (Section 4)

82. Compute causal matrix CM(IMi) for all i ∈ ID


81. initially empty set of process models

83. While IM ≠ ∅

For each IMk ∈ IM


84. Remove IMj from IM and insert IMj into PM
85.

Merge IMj and IMk with Tjk=Tj∪Tk, Pjk=Pj∪Pk, Ajk=Aj∪Ak, Σjk=Σj∪ Σk


86. If CM(IMj) = CM(IMk)

Aggregate all places pl and pm ∈ Pjk if Cjk(pl) = Cjk(pm)


87.

Aggregate all tl and tm ∈ Tjk if ljk(tl) = ljk(tm)


88.
89.
90. Remove IMk from IM

5.2 Case Mining and Event Log Struc- Posting Document 1 2...N Journal Entry Item
contains
ture DocumentNr DocumentNr
UserName PositionNr
PostingDate AccountNr
The mining algorithm uses recorded trans- TransactionCode 0...1
is cleared
0...N Amount
PostingText CreditOrDebit
action data from ERP systems as the data ClearingDocNo

source for the mining. This kind of data only


has a medium maturity level from a process Fig. 2. Entity-relationship-model for ac-
mining perspective because ERP systems counting data structure
like SAP or JD Edwards do not use a system-
atic approach to link and store process rele- It is possible to exploit this data chain for
vant data [11]. The event log data in such each process instance. The procedure is il-
systems is stored in different database ta- lustrated in line 1 to 11 in Listing 1. The al-
bles and needs to be composed in a mean- gorithm maps event log entries to cases. It
ingful manner before it can be used as input starts with an arbitrary document number
for process mining algorithms. Data related and mines all related journal entry items by
to financially relevant business transactions using the contains relationship illustrated in
exhibits specific characteristics that can be Fig. 2. It then searches for all document
used for process mining purposes [27]. The numbers that have cleared these items by
execution of financially relevant transac- using the is cleared relationship. All posted
tions in an ERP system creates specific data items for every clearing document are then
records. Every execution of an activity cre- searched in further iterations. The loop ter-
ates a posting document. Each document minates when all documents and items that
contains at least two journal entry items. belong to the same process instance have
This connection is illustrated by the contains been found. The loop itself is repeated until

is an event log ⋃•••‘{[• , Ž• } with [• =


relationship in the ER model presented in all case IDs have been mined. The outcome

{[PWN’OeL8M“‘ ,
Fig. 2. Journal entries that follow an open-

… , [PWN’OeL8M“• } ∀ 4PgL;e— ZPWN’OeLg


item-accounting principle are linked to each
•mn˜o
Z‘ … Z• ™ššš› ; . The event log also includes
other if they belong to the same process in-
stance. If a process instance is terminated
each open journal entry item has been the data attributes associated to each post-
cleared by another posting document. This ing document and journal entry item that
dependency is modeled in Fig. 2 via the is are listed as entity attributes in Fig. 2. In
cleared relationship. contrast to traditional event logs the events
belonging to a single case do not follow a
strict linear order. The causal de-

235
APPENDIX A: PUBLICATIONS

pendencies between the events have to be a) Process Instance Graphs b) Process Instance Models c) Process Model

determined in a separate step. A


1
1
B
1
1 1 B'
1
1

1 D A' D' 2 B' 2


1 2 1 2
1
A 1 C 1 C' 1 A' 3 D'
1 1 2

5.3 Instance Graphs and Abstraction 1 1 1 1


1 C' 1 1

Levels A 1 B 1 D1 A' 1 B' 1 D'


1

The application of process mining for finan-


Fig. 3. Abstraction levels in the context of
cial audits requires algorithms that enable
process mining66
the investigation of a financially relevant
transaction from its point of origin to the Fig. 3 illustrates how the different abstrac-
posting on the financial accounts. On the tion levels in the context of business process
other hand an auditor should be able to re- models relate to each other. Process in-
ceive an overview to get an all-encompass- stance graphs reside on the lowest level of
ing understanding of the process structure abstraction. They are graphical representa-
and its effect on the financial accounts. Both tions of the source data from executed and
requirements can be met by using different recorded process instances. Fig. 3a shows
abstraction levels. Scholars commonly dis- two instance graphs. Each rectangle repre-
tinguish between four levels of horizontal sents the execution of an activity and the
abstraction in business process manage- arcs between two activities denote the
ment [49]. The instance level represents causal relationships between executed ac-
tangible entities that are involved in busi- tivities. The numbers indicate how often an
ness processes which include executed ac- activity was executed and how often the
tivities, resources, concrete data values etc. path from one activity to another was cho-
A set of similar business processes are rep- sen. The process instance model is an ab-
resented as business process models on the straction of the instance graph (Fig. 3b). Ex-
model level. A model is expressed using ecutions of identical activities are aggre-
constructs of a meta-model. These can gated into activity models but an instance
themselves again be defined at the meta- model still represents just a single execution
meta-model level. Process mining algo- of a business process. A process model is an
rithms commonly operate on the model abstraction of a set of similar process in-
level. They generate process models that stances (Fig. 3c). The process model is im-
are abstractions of the individual process in- portant for the auditor to get an overview
stances they represent.65 about the structure of the process, its rela-
tionships to the financial accounts and to as-
sess the overall materiality of a business
process. Process instance models and pro-
cess instance graphs are useful for following
the audit trail and for inspecting individual

65
An exception is the multi-phase process mining have already been made and all relationships in
approach that was developed to generate event- Fig. 3a have the semantic of an AND split or join.
driven-process-chains [26]. It operates on the in- The models presented in Fig. 3b and 3c represent
stance and model level. A similar approach is abstractions of the lower level models. The rela-
used for the design of the MLPM. tionships on these level can generally represent
66
The models represent simple directed and labeled AND, OR and XOR splits and joins. An algorithm
graphs for illustration purposes without logical to transform these models into EPC or Petri Nets
operators. It is not necessary to model choice at is described in [26].
the instance graph level because all decisions

236
APPENDIX A: PUBLICATIONS

denoted as \ → • because of 6ž ∩ ¡ ≠
∅ with LO’ž ∈ 6ž ∧ LO’ž ∈ ¡ . The re-
process instances with respect to involved
activities, users, data values etc.
sults of these operations are instance
The second part of the mining algorithm
graphs in form of directed graphs equiva-
ranging from line 12 to 21 creates instance
lent to the graph shown in Fig 2a.
graphs. The algorithm first creates a node
for every document number labeled with
5.4 Instance Models and Colored Pe-
the transaction code that was used to cre- tri Nets
ate the posting document in the ERP system
(line 18). Different nodes in the instance The third section of the mining algorithm re-
graph can carry the same label (compare constructs process instance models as CPN.
Fig. 2a) if they were created by the same Petri nets are the predominant modeling
transaction of the ERP system. The algo- language in the process mining research do-
rithm then infers the causal dependency be- main [30]. They are suitable for the model-
tween activities (line 19 and 20). Traditional ing of business processes and offer a formal
mining algorithms rely on the time stamp of as well as graphical notation that can even
events to infer the control flow. This is not be understood by non-experts [52]. They
suitable for the event log created by the provide a sound mathematical foundation
MLPM algorithm. The same posting docu- for the simulation and verification of Petri
ment can clear journal entry items belong- Net models [53]. The majority of process
ing to different other posting documents mining algorithms rely on low-level Petri
leading to parallelism in the event log. Using Nets. An exception is the approach used by
the temporal ordering of events can lead to Accorsi and Wonnemann [39]. They use CPN
intertwining in parallel branches resulting in and model data objects as colored tokens.
models that are of little use to the user.67 We use a similar approach to generate pro-
The algorithm uses instead a data depend- cess models that model the control flow and
ent approach for determining the control data flow perspective simultaneously in a
flow. Sun and Zhao introduced an approach single model. A Colored Petri Net is formally

( , , , Σ, , , , , ) [53, p. 87], with:


to derive the control flow by modeling the expressed by the tuple CPN =

activity V can be expressed as œ“• ( “• , 6• ) ,


data flow [51]. The data dependencies of an

where • is the input for V, and 6• the out-


1) is a finite set of transitions

put. Z denotes the type of data depend- ∈ × ∪ × is a set of directed


2) is a finite set of places
3)
ency. The algorithm analyses in lines 19 and arcs

* is a finite set of typed variables such


20 how activities relate to each other based 4) is a set of non-empty color sets

ity has cleared LO’ž posted by activity that Type[,] ∈ for all variables , ∈ *
on their data relationships. It checks if activ- 5)

\. If the condition is true a control arc from 1: → is a color set function that as-
\ to is inserted. can only have occurred
6)

if \ took place before. Otherwise there >: → ?@ A* is a guard function that


signs a color set to each place

would have been no LO’ž that could have assigns a guard to each transition B such
7)

that CDE [>(B)] = FGGHEIJ


mandatory dependency between and \
been cleared by . This is equivalent to a

67
A discussion of this aspect is beyond the scope of with undesired side-effects like the duplication of
this paper but it is illustrated in detail in [29]. The events in the log, and it is not able to deal with
event logs can be converted into linear event the data perspective.
logs [50]. But this transformation is accompanied

237
APPENDIX A: PUBLICATIONS

8) ?: → ?@ A* is an arc expression place 4 if the transition posted a journal en-

to each arc I such that CDE [?(I)] =


function that assigns an arc expression try item on the represented financial ac-

1(D)ST , where D is the place con- Two arcs Q(L, 4) ∧ Q(4, L) are inserted if the
count. They are referred to as posting arcs.

nected to the arc I.


9) b: → ?@ A∅ is an initialization func-
transition has cleared an item on the re-
spective account. These arcs have the se-

sion to each place D such that


tion that assigns an initialization expres- mantic of testing arcs because they do not

CDE [b(D)] = 1(D)ST


consume the tokens on the connected
places. Double-headed arcs are used for vis-

two arcs Q(L, 4) ∧ Q(4, L). They are called


ualization as a syntactical abbreviation for
We integrate the data perspective by mod-

eled as constants ( = {}). The arc expres-


eling colored places in the CPN that repre- clearing arcs. The arc inscriptions are mod-
sent financial accounts [27]. An instance
graph produced by section Reconstruct In- sion function assigns to each posting and
stance Graphs is first transformed into a clearing arc a set of constants that denote
CPN. The nodes of the instance graph be- the posted or cleared value, the account
come the transitions of the CPN (line 26).68 type, account number and an indicator
We distinguish between two types of whether it is a credit or debit posting. Each
places. Control places model the control control arc is assigned a number that indi-
flow and account places model the data cates how often this represented control
flow. Each arc from the instance graph is path was chosen in the instance.69 The
transformed into a combination of a control source place is initialized by the initializa-

instance model. A source place 4¢£¡¤˜¥ is in- for 4¢£¡¤˜¥ generates e tokens in the initial
place and two connecting control arcs in the tion function . The initialization expression

serted with arcs Q¦ §4¢£¡¤˜¥ , L¦ ¨ ∀ L¦ ∈ ∧ marking hi (4), one for each connected
• L¦ = ∅ and a sink place 4¢••ª is with arcs
Qª (Lª , 4¢••ª ) ∀ Lª ∈ ∧ Lª • = ∅. The data
start transition.
After having integrated the data perspec-
perspective is integrated and visualized by tive by creating a CPN it is necessary to ag-
creating account places for each journal en- gregate model components to create in-
try item in the event log (line 29). The color stance models. The first step is aggregating

places 4ª and 4¦ can be aggregated if they


set function assigns different color sets to places representing financial accounts. Two
places depending on whether they belong

The set of color sets Σ includes the color sets if they carry the same color (4ª ) = (4¦ )
to the group of control or account places.69 represent the same account. This is the case

for all possible journal entry values, account (line 32). The relationships between the
numbers, account types, credit or debit in- places and connected transitions need to be
dicators and execution numbers. maintained in respect to the type and in-
Account places are connected to related scription of each arc.70 The second step ag-

Q(L, 4) is inserted from transition L to the


transitions (line 30 and 31). A simple arc gregates the transitions (line 33). The used
method for aggregating transitions is based
on the algorithm described by van Dongen

is therefore defined as (L) = true for all L ∈ .


68 70
Guards are not needed and the guard function A complete description of this procedure is be-
yond the scope of this paper. It is described in
69
A detailed description of the color set function [28].
and arc expression function is provided in [27].

238
APPENDIX A: PUBLICATIONS

gregated if they carry the same label R(L¦ ) =


and van der Aalst [26]. Transitions are ag- gregation procedure is repeated for all in-

R(Lª ). The result of the third section of the


stance models until the set of instance mod-
els is empty (line 38).
mining algorithm is a set of instance process Aggregating process instance models that
models represented as CPN. exhibit identical control flow patterns does
not mean that the resulting process model
5.5 Process Models and Clustering
is identical to the source process instance
The final section of the mining algorithm models. The data values in each process in-
produces process models (lines 35 to 45). A stance model are unique and are preserved
key requirement for the application of pro- during the aggregation process. A process
cess mining in the context of financial audits model therefore shows the aggregated data
is the creation of perfectly fitting and pre- values and data flow of all represented pro-
cise process models. Some researchers have cess instances. The final outcome of the
developed methods to create precise and mining algorithm is a set of process models
fitting process models by clustering traces in where each model represents process in-
the event log [41],[42]. Their approaches stances with the same control flow pattern.
are not directly applicable because they do Each process model provides the aggregate
not take into account the data perspective view of the data flow and relationship of ac-
and because they are not suitable for the tivities with the financial accounts.
given event log structure. But clustering is
an approach that is also used for the crea- 6 Mining Results and Evalua-
tion of process models in the final section of tion
the presented mining algorithm. But instead
of clustering traces we cluster and aggre- Many scholars stress the importance of
gate process instance models. We use evaluation for rigorous DSR [20], [21], [55].
causal matrices to identify isomorph pro- Riege et al. provide a categorization of eval-
cess instances. Causal matrices are key com- uation methods and point out that it is not
ponents for heuristic [14] and genetic min- sufficient in DSR to evaluate research arti-
ing algorithms [54, p.6]. A causal matrix de- facts against the epistemological objective
scribes the causal relation between activi- but also against the design objective [56].

in column ® and row ¯ designates that a


ties in a model. An entry in a causal matrix Venable et al. provide a framework for the
evaluation in DSR that can be used to sup-

® and ¯. The causal matrices of two instance


causal relation exists between the activities port the selection of adequate evaluation
methods [57]. The aim of the conducted
models are identical if they show exactly the evaluation is to test whether the created ar-
same behavior. The algorithm computes the tifact is able to produce the estimated re-
causal matrices for all instance models (line sults. An artificial evaluation is appropriate
37) that were created by the previous oper- for the research at hand because it is not
ations. It then searches for identical causal necessary to observe the behavior in organ-
matrices (line 41) and aggregates two mod- izational contexts at this stage. We have
els if they have the same causal matrix (lines chosen an artificial laboratory experiment
42 to 44). The aggregation procedures are as one of the evaluation methods suggested
identical to the procedures used at the in- for this setting by Venable et al. [57]. Such
stance model level (line 32 and 33). The ag- an evaluation requires the existence of in-
stantiated artifacts that can be used for the
experiment. The different sections of the

239
APPENDIX A: PUBLICATIONS

mining algorithms were therefore imple- dustry The MLPM was used to mine all pro-
mented in a software artifact in iterative cy- cess instance graphs, process instance mod-
cles. The experiment itself was divided into els and process models for the three data
the phases data extraction, mining and re- sets. The models were inspected on a sam-
sults analysis. We used a separate extrac- ple basis by observation and comparison
tion module for extracting relevant data with the original event log data. The soft-
from ERP systems that can be adjusted to ware yEd – Graph Editor [58] provides pow-
different source systems. We checked the erful automatic layout functionality and is
data for first and second order defects [8, free for use for noncommercial purposes. It
pp. 29–32]. uses the GXML and GraphML formats as in-
put and was used in the experiment to
Three data sets were used for the evalua-
graphically represent the process models.
tion. Their characteristics are listed in Table
The models were further tested for sound-
1. We extracted data from the productive
ness (proper completion, option to com-
SAP systems of three companies operating
plete, absence of dead transitions, and safe-
in the retail, manufacturing and media in-
ness for all places except account places) by
using the CPN simulation tool Renew [59].
TABLE 1 Evaluation Data Sets
Journal Entry Process Process
Set Industry Journal Entries
Items Instances Models
1 Manufacturing 1,764,773 7,395,434 1,035,805 841
2 Media 156,604 559,506 18,975 516
3 Retail 92,487 222,901 40,634 307
0005035200
[5,208.68]

0004002000
0002811000
[117.28] [35.55] 0001900101
[36,937.40]
0004000070
[35.55]
[308.86] [7,426.98] 0005004900 [648,155.17]
[35.55] 0001900103 0003160100
[36,922.40] [36,922.40] [4,429.73]
[12] [12]
0004000070
[49,864.85] [9,364.44] 0005004900 0001900311
0001900313
(B) Enter
Incoming Invoice
[35.55]
0001421100
0002811000 [366,410.06] [8]
[365,748.82] (E) Post with
[365,748.82] Clearing [36,922.40] 0001900103
(D) Payment
[593,846.02] [648,190.72] [648,190.72] [17] [17]

0004000010 0004000110
0001900111
[83] [1,078.42]
[83] [475.10]
[8] (A) Post Goods [245,171.63]
[8] [96] [96] [83]
Receipt [643,402.01] (C) Clear 0001900113
Account [83]
0001900113
0002810200
0002810200
[245,008.85] [245,008.85]
[245,008.85]
[645,022.01] [645,022.01] 0001900313
[643,402.01] 0005004040
[25] [25] [365,748.82]
[60.07]

Fig. 4. Example of a mined process model71

Fig. 471 shows an example of a mined pro- a CPN as specified in section 5.4. The transi-
cess model from data set 1. It is modeled as tions illustrated as rectangles represent the
activities. The mined model represents a

71
The arc inscriptions for posting and clearing arcs is the case for the inscriptions of the connected
only display the assigned constant for the posted account places that only show the account num-
or cleared value. The inscriptions for the account ber.
type, account number and credit or debit indica-
tor are omitted for better readability. The same

240
APPENDIX A: PUBLICATIONS

purchasing process that was executed by 8 measures can be calculated by replaying


identical process instances. It contains the cases from the event log and by counting if
activities Post Goods Receipt, Enter Incom- tokens are missing in the CPN in order to ex-
ing Invoice, Clear Account, Payment and ecute the simulation or remain uncon-
Post with Clearing that were executed in a sumed in the model after the simulation is
sequential order. Control places are ren- terminated. These matrices are not directly
dered with a bold black border. Simple ar- applicable to the used CPN. Colored data to-
rows between transitions and control kens remain in the model by definition even
places model the control flow and sequence after all transitions have fired. The used
of activities. The inscriptions for these arcs events in the event log do not follow a strict
show how often the path was chosen in the linear order for each case. It is therefore un-
represented instances. Account places have clear which sequence of events should be
a solid or dashed border. The type and color used for replaying a case. The traces can be
of the border line indicate if the repre- transformed into strict linear sequences
sented account is a balance sheet or profit [50] but these sequences would not fit to
and loss account and if it represents the the mined model anymore. The possible
debit or credit side.72 Account places and transformation into linearized traces would
transitions are connected by dotted arcs. A introduce choice at the instance level that
simple dotted arrow is a posting arc. Dou- did not occur in reality. The model illus-

The paths 4‘ = \ → [ and 4° = \ →


ble-headed dotted arcs are clearing arcs. trated in Fig. 4 does not show any choice.

→ [ are parallel and not optional paths.


Posting and clearing arcs visualize the data

→\→[→
flow in the process model. The Post Goods
Receipt activity for example creates a token The trace for example,
with the value of 593,842.00 representing a which would be an output of the transfor-
journal entry item posting on the raw mate- mation, cannot be replayed by the model
rials account 0001421100. Another token is from Fig. 4.73
Calculating token based measures like (<)
created on the account 0002810200. This is
and (Qž ) for the produced process model
cleared by the subsequent activity of Clear
Account but without consuming the token
would require many artificial adjustments
on the respective account. The modeled
which are not reflected by the actual source
CPN mimics the behavior of the represented
data. We alternatively use metrics that are
process. Transitions create colored tokens
suitable to take into account possible paral-
on the account places representing the
lelism at the instance level and that directly
posted journal entries. The control places
compare the execution paths in the differ-
define the control flow in the model.
control flow paths from 4¢£¡¤˜¥ → 4¢••ª
ent models. We use the percentage of the
The model is perfectly fitting and precise be-
cause it can replay all process instances and that are present in the instance graphs to

the fitness denoted as <lmno . And we com-


does not allow any additional behavior. Fit- those in the process model for measuring

different metrics such as fitness (<) and be-


ness and precision can be measured using

havioral appropriateness (Qž ) [48]. These


pute the percentage of the control flow

72
Solid line = balance sheet account, dashed line = by using the procedures described in [28]. But
profit and loss account, black line = debit side of this would result in a much more complex mod-
an account, gray line = credit side of a account. els and a negative effect on the model precision.
73
The model itself can be transformed in such a
way that it is able to replay all linearized traces

241
APPENDIX A: PUBLICATIONS

paths in the process model that are not pre- log. It is an acceptable result compared to

the precision denoted as 4lmno with:


sent in the instance graphs for measuring outcomes of artificial simulations from
other scholars that use comparable metrics
[40]. But it has to be validated in further re-

|{4 | 4 ∈ ∧∈ ±² }|
search if this precision is high enough for the
<lmno =
l_
| ±² |
application in real world scenarios.
The models produced by the MLPM from

| l_ | − |{4 | 4 ∈ l_ ∧ ∉ ±² }|
data set 1 to 3 were analyzed using descrip-

4lmno =
tive statistics. Fig. 5 to 8 show selected re-
| l_ | sults for data sets 1 and 2. Fig. 5 and 7 show
the frequency distributions of the mined

= {4‘ , … , 4• } is the set of distinct con-


process models depending on their model
l_
trol flow paths in the process model h and
complexity measured as the number of in-

±² = {4‘ , … , 4• } is the set of all distinct


cluded transitions. They show that the pro-
cess models are distributed comparably to a

that belong to h.
control flow paths from all instance graphs normal distribution. Fig. 6 and 8 present
scatter diagrams for data set 1 and 2. They
illustrate the distribution of the number of
Van Dongen and van der Aalst proof that
represented instance models in a process
the used aggregation procedures are path
model depending on the model size. The di-
preserving [26]. This means that all control
agrams show that the vast majority of pro-
cess model with <lmno = 1. We calculated
flow paths are also represented in the pro-
cess instances actually belong to very sim-
4lmno for the mined models from data set 1
ple process models that contain only a few
transitions. This observation confirms pre-
to 3. The average value for this measure was
liminary results from prior research work
0.81.74 This means that 19 % of the paths in
[60].
the process models may actually represent
behavior that was not recorded in the event
20

1.0e+06

100000
Number of Represented Instances
15

10000
Frequency

1000
10

100
5

10

1
0

0 5 10 15 0 5 10 15
Number of Transitions Model Complexity

Fig. 5. Frequency distribution for data set 175 Fig. 6. Scatter diagram for data set 176

74 76
Process models including loops were excluded The dependent variables in Fig. 6 and 8 use a log-
from the calculation arithmic scaling.
75
Process models including loops were excluded
from the calculation

242
APPENDIX A: PUBLICATIONS

20
1.0e+06

100000

Number of Represented Instances


15

10000
Frequency

1000
10

100
5

10

1
0

0 5 10 15 20 25 0 5 10 15 20 25
Number of Transitions Model Complexity

Fig. 7. Frequency distribution for data set 2 Fig. 8. Scatter diagram for data set 276

Table 2 Descriptive Statistics


Parameter Data Set 1 Data Set 2 Data Set 3
Mean value of transitions per process model 4.61 4.95 3.98
Median value of transitions per process model 5 5 4
Standard deviation of transitions per process model 2.10 2.43 2.24
Maximum value of transitions per process model 15 25 21
Minimum value of transitions per process model 1 1 1

The diagrams for the third data set are not other to promote the progress of infor-
included due to space restrictions. They fol- mation systems research [55]. DSR has the
low similar patterns. Some relevant descrip- potential to create artifacts that are of prac-
tive statistical values are listed in Table 2 for tical relevance and prescriptive nature.
all data sets. These artifacts can in return be the subject
of descriptive science. Gregor and Hevner
7 Discussion provide a useful framework for the catego-
rization of knowledge contribution by DSR
The main contribution of this paper is the in- [61]. They categorize research work into the
troduction of a process mining algorithm four domains routine design, improvement,
that is able to discover the control flow and exaptation and invention. The presented re-
the data flow perspective in process models search work can be assigned to the exapta-
by simultaneously providing perfectly fitting tion quadrant because the main objective is
and precise models at different abstraction to provide a solution for a new application
levels. The innovation of the presented arti- area by partly using already existing
fact is achieved by combining already exist- knowledge. But it also affects the improve-
ing and newly developed methods that lead ment quadrant by introducing a new
to a novel solution for a new application method to model the control flow and data
area. It is often questioned if research re- simultaneously in mined process models.
sults derived from DSR can be equally im- Gregor and Hevner further differentiate be-
portant as results provided by other re- tween three levels of contribution types
search approaches. March and Smith em- that range from abstract, complete and ma-
phasize that both, design and natural sci- ture knowledge on the highest level to more
ence, have to coexist and benefit from each specific, limited and less mature knowledge

243
APPENDIX A: PUBLICATIONS

on the lowest level. The research results when account places are neglected. The re-
presented in this paper are mainly located maining models would then represent
on the second level providing constructs sound workflow nets but without repre-
and methods for the mining of process senting the data flow perspective. The data
models and on the first level presenting an sets were all extracted from SAP systems. It
instantiated software artifact. The results can therefore not be concluded that the re-
that can be achieved by analyzing the min- search results also hold true for other data
ing outcomes can also be input for the third sources. But an important advantage of the
and highest knowledge contribution level. used mining algorithm is its independence
The distributions of the number of instances from the implemented data structures of a
over the number of transitions in Fig. 6 and particular ERP system because it bases on
8 for example can lead to the assumption the general structure of accounting entries.
that their distribution curves are very simi- Some process models showed loops that oc-
lar. But the descriptive statistics in Table 2 cur when a transaction has cleared a journal
shows that the mean values for the number item that was posted by the same transac-
of transitions in the process model differ tion or by a transaction located in the sub-
quite significantly from each other with 4.61 sequent execution path. This constellation
for data set 1, 4.95 for set 2 and 3.98 for set leads to a deadlock in the process model
3. It could hypothesized that the complexity which is not critical for the interpretation of
of the mined process models relates to the the model but generally not desired for the
maturity of the mined business processes modeling of correct process models. A solu-
following the assumption that a mature pro- tion could be the prevention of aggregating
cess is more integrated into information transitions carrying the same label if this
systems than a less mature. This research would result in a loop.
question surely needs further investigation, The mining algorithm produces precise and
but it highlights how the presented results fitting process models at the cost of lacking
can be the starting point for further theoret- generalization. It is therefore not applicable
ical research. for scenarios with highly variable business
The presented mining algorithm is able to processes. In the worst case scenario all
discover process models in accordance with process instances show a different behav-
the identified requirements to a large ex- ior. The mining algorithm would then not be
tent. The mined models are not absolutely able to aggregate any instance models and
precise. It needs to be validated in further the set of process instances models would
research if the achieved level of precision is be identical to the set of process models re-
sufficient in practice. Several other limita- sulting in no or little information gain. The
tions need to be taken into account. The data presented in Table 1 shows that this
process models do not represent sound risk is not acute for the given application
workflow nets according to commonly used area. Business processes are usually stand-
definitions [3, p. 39]. This handicap is not ardized to a certain degree when they are
too severe because the objective of process supported by ERP systems. The data shows
mining for financial audits is the adequate that the number of process models ranges
modeling of the control and data flow per- from 307 for the smallest data set to 841
spective with precise and fitting process models for the largest. This may still seem
models. Formally well-structured process to be a big number. 63 process models in
models are of minor interest. But if it should data set 1 only consist of one transition.
be necessary soundness could be achieved

244
APPENDIX A: PUBLICATIONS

These represent trivial processes. The activ- business processes is a significant part in the
ities in these processes were mostly carried financial audit. Process mining can be ap-
out by using a single general purpose trans- plied as a BI approach to support and auto-
action. They are of little interest from a pro- mate the audit of business processes. The
cess perspective but highly important from selection of a process mining algorithm
an audit perspective because this category should be founded on the analysis of rele-
of process models represent 96% of process vant application domain requirements. For
instances in data set 1, 70% in set 2 and 51% the case of financial audits it is crucial that a
in set 3. It is clear that further analytical pro- mining algorithm is able to model both the
cedures are necessary to address this cate- control flow and data flow. The algorithm
gory. A starting point could be the clustering should preserve the audit trail and produce
of process models that use the same ac- perfectly fitting and as precise process mod-
counts. The process models that reflect the els as possible. We have designed and eval-
major business processes are those that uated a multilevel process mining algorithm
contain many transitions and represent a that meets these requirements to a large
high number of process instances. Data set extent. It introduces novel constructs and
1 contains 135 process models consisting of methods for the mining of process models
5 transitions. But just two of them already and an instantiated software artifact. The
represent 62% of the instances of this cate- results derived by exposing the designed ar-
gory. It can therefore be assumed that the tifact to extensive real life data can be the
majority of instances for more complex pro- starting point for future theory building.
cess models only represent very infrequent Data sets from the SAP systems of three dif-
behavior and can be tested traditionally by ferent companies operating in diverse in-
inspecting individual journal entries. The dustries were used for the evaluation of the
process models that represent many pro- designed artifact. It cannot be concluded
cess instances and create a high value flow that the results hold true for other ERP sys-
are interesting from a materiality perspec- tems and industries but the implemented
tive and can be audited by including the mining algorithm exploits the structure of
testing of embedded application controls accounting entries that is system-independ-
[62]. ent and should therefore be generally appli-
cable. The extension to other ERP systems
8 Conclusion will be covered in future research.
The amount of available data increases with Public accountants face the challenge to au-
the integration of information systems for dit increasingly complex and integrated
the support and automation of business ac- business processes that process huge
tivities. BI is a research domain that pro- amount of data. The presented mining algo-
vides mature methods and tools that can be rithm exploits large data sets that are cre-
used to exploit and handle the growing ated during the operation of business pro-
amount of data. While it is commonly used cesses. It provides a suitable solution for an-
in many application scenarios it is almost alyzing business processes in financial au-
absent in the auditing industry. Traditional dits but it can also be applied in application
audit procedures are not efficient and effec- contexts that exhibit similar requirements.
tive in audit environments with highly inte-
A basic limitation of the presented algo-
grated information systems and an increas-
rithm is the lack of generalization. This is a
ing amount of processed data. The audit of
desired characteristic for financial audits

245
APPENDIX A: PUBLICATIONS

but it is not appropriate in settings with [7] H. Krcmar, Informationsmanage-


highly variable business processes. Other ment. Berlin; Heidelberg: Springer,
mining approaches using a two-step ap- 2010.
proach [43] or fuzzy mining [15] might be [8] H.-G. Kemper, W. Mehanna, and H.
more appropriate in these settings. Some Baars, Business intelligence - Grund-
restrictions still exist in certain process con- lagen und praktische Anwendungen :
stellations that create deadlocks in the pro- eine Einführung in die IT-basierte Ma-
cess models. This phenomenon and ade- nagementunterstützung. Wiesbaden:
quate solutions still need to be researched Vieweg + Teubner, 2010.
in the future.
[9] R. L. Braun and H. E. Davis, “Com-
References puter-assisted audit tools and tech-
[1] International Federation of Account- niques: Analysis and perspectives,”
ants, “ISA 315 (Revised), Identifying Managerial Auditing Journal, vol. 18,
and Assessing the Risks of Material no. 9, pp. 725–731, 2003.
Misstatement through Understand- [10] N. Gehrke, “The ERP Auditlab - A Pro-
ing the Entity and Its Environment.” totypical Framework for Evaluating
2012. Enterprise Resource Planning System
[2] M. Werner and N. Gehrke, “Potenti- Assurance,” in Proceedings of the
ale und Grenzen automatisierter Pro- 43th Hawaii International Conference
zessprüfungen durch Prozessrekon- on System Sciences, Kauai, 2010, pp.
struktionen,” in Forschung für die 1–9.
Wirtschaft, G. Plate, Ed. Aachen: [11] W. M. P. van der Aalst, A. Andrian-
Shaker Verlag, 2011. syah, A. K. de Medeiros, F. Arcieri, T.
[3] W. M. P. van der Aalst, Process Min- Baier, T. Blickle, J. C. Bose, P. van den
ing: Discovery, Conformance and En- Brand, R. Brandtjen, and J. Buijs “Pro-
hancement of Business Processes, 1st cess Mining Manifesto,” in BPM 2011
Edition. Berlin Heidelberg: Springer, Workshops Proceedings, 2012, pp.
2011. 169–194.

[4] H. Chen, R. H. L. Chiang, and V. C. [12] W. M. P. van der Aalst, K. M. van Hee,
Storey, “Business Intelligence and An- J. M. van Werf, and M. Verdonk, “Au-
alytics: From Big Data to Big Impact,” diting 2.0: Using Process Mining to
MIS Quarterly, vol. 36, no. 4, pp. Support Tomorrow’s Auditor,” Com-
1165–1188, Dec. 2012. puter, vol. 43, no. 3, pp. 90–93, Mar.
2010.
[5] D. Howe, M. Costanzo, P. Fey, T. Go-
jobori, L. Han-nick, W. Hide, D. P. Hill, [13] M. Jans, M. Alles, and M. Vasarhelyi,
R. Kania, M. Schaeffer, S. St Pierre, S. “Process Mining of Event Logs in In-
Twigger, O. White, and S. Yon Rhee, ternal Auditing: A Case Study,” in 2nd
“Big data: The future of biocuration,” International Symposium on Ac-
Nature, vol. 455, no. 7209, pp. 47–50, count-ing Information Systems, 2011.
Sep. 2008. [14] A. Weijters, W. M. P. van der Aalst,
[6] C. Lynch, “Big data: How do your data and A. K. A. de Medeiros, “Process
grow?,” Nature, vol. 455, no. 7209, mining with the heuristics miner-al-
pp. 28–29, Sep. 2008. gorithm,” Technische Universiteit

246
APPENDIX A: PUBLICATIONS

Eindhoven, Tech. Rep. WP, vol. 166, Process Audits – A Multi-Method Re-
2006. search Approach,” in Proceedings of
the 10th International Conference on
[15] C. Günther and W. van der Aalst,
Enterprise Systems, Accounting and
“Fuzzy mining–adaptive process sim-
Logistics, Utrecht, 2013.
plification based on multi-perspective
metrics,” Business Process Manage- [24] S. Brinkkemper, “Method engineer-
ment, pp. 328–343, 2007. ing: engineering of information sys-
tems development methods and
[16] A. K. A. de Medeiros, “Genetic Pro-
tools,” Information and Software
cess Mining,” Eindhoven University of
Technology, vol. 38, no. 4, pp. 275–
Technology, Eindhoven, 2006.
280, 1996.
[17] M. de Leoni and W. M. van der Aalst,
[25] T. Wilde and T. Hess, “Forschungsme-
“Data-Aware Process Mining: Discov-
thoden der Wirtschaftsinformatik,”
ering Decisions in Processes Using
Wirtschaftsinformatik, vol. 49, no. 4,
Alignments,” 2013.
pp. 280–287, 2007.
[18] D. Ferreira and D. Gillblad, “Discover-
[26] B. F. van Dongen and W. M. P. van
ing Process Models from Unlabelled
der Aalst, “Multi-phase process min-
Event Logs,” Business Process Man-
ing: Aggregating instance graphs into
agement, pp. 143–158, 2009.
EPCs and Petri nets,” in PNCWB 2005
[19] N. Gehrke and N. Müller-Wickop, workshop, 2005, pp. 35–58.
“Basic Principles of Financial Process
[27] M. Werner, “Colored Petri Nets for
Mining A Journey through Financial
Integrating the Data Perspective in
Data in Accounting Information Sys-
Process Audits,” in Proceedings of
tems,” in Proceedings of the 16th
32nd International Conference on
Americas Conference on Information
Conceptual Modeling (ER 2013),
Systems, Lima, Peru, 2010.
Hong Kong, China, 2013, pp. 387–
[20] H. Österle, J. Becker, U. Frank, T. 394.
Hess, D. Karagiannis, H. Krcmar, P.
[28] M. Werner, M. Schultz, N. Müller-
Loos, P. Mertens, A. Oberweis, and E.
Wickop, N. Gehrke, and M. Nüttgens,
J. Sinz, “Memorandum on design-ori-
“Tackling Complexity: Process Recon-
ented information systems research,”
struction and Graph Transformation
European Journal of Information Sys-
for Financial Audits (Research in Pro-
tems, vol. 20, no. 1, pp. 7–10, 2010.
gress),” in Proceedings of 33rd Inter-
[21] A. R. Hevner, S. T. March, J. Park, and national Conference on Information
S. Ram, “Design science in infor- Systems, Orlando, 2012.
mation systems research,” MIS Quar-
[29] M. Werner and M. Nüttgens, “Im-
terly, pp. 75–105, 2004.
proving Structure - Logical Sequenc-
[22] J. Mingers, “Combining IS research ing of Process Models,” in Proceed-
methods: to-wards a pluralist meth- ings of the 47th Hawaii International
odology,” Information systems re- Conference on System Sciences, Big
search, vol. 12, no. 3, pp. 240–259, Island, 2014.
2001.
[30] A. Tiwari, C. J. Turner, and B. Majeed,
[23] N. Müller-Wickop, M. Schultz, and M. “A review of business process mining:
Peris, “To-wards Key Concepts for

247
APPENDIX A: PUBLICATIONS

state-of-the-art and future trends,” European Conference on Information


Business Process Management Jour- Systems, Utrecht, 2013.
nal, vol. 14, no. 1, pp. 5–22, 2008. [38] T. Stocker, “Data flow-oriented pro-
[31] W. M. P. van der Aalst, “Process Min- cess mining to support security au-
ing: Overview and Opportunities,” dits,” in Service-Oriented Computing-
ACM Transactions on Management ICSOC 2011 Workshops, 2012, pp.
Information Systems, vol. 99, no. 99, 171–176.
pp. 1–16, Feb. 2012. [39] R. Accorsi and C. Wonnemann, “In-
[32] R. S. Debreceny and G. L. Gray, “Data Dico: Information Flow Analysis of
mining journal entries for fraud de- Business Processes for Confidentiality
tection: An exploratory study,” Inter- Requirements,” in Security and Trust
national Journal of Accounting Infor- Management, Springer, 2011, pp.
mation Systems, vol. 11, no. 3, pp. 194–209.
157–181, Sep. 2010. [40] A. Rozinat, A. K. A. de Medeiros, C.
[33] J. Becker, P. Delfmann, M. Eggert, W. Günther, A. Weijters, and W. M.
and S. Schwittay, “Generalizability van der Aalst, “The Need for a Pro-
and Applicability of Model- Based cess Mining Evaluation Framework in
Business Process Compliance-Check- Research and Practice,” in Business
ing Approaches – A State-of-the-Art Process Management Workshops,
Analysis and Research Roadmap,” 2008, pp. 84–89.
BuR - Business Research, vol. 5, no. 2, [41] A. K. A. de Medeiros, A. Guzzo, G.
pp. 221–247, Nov. 2012. Greco, W. M. P. van Der Aalst, A.
[34] M. Alles, M. Jans, and M. Vasarhelyi, Weijters, B. F. Van Dongen, and D.
“Process Mining: A New Research Saccà, “Process Mining Based on
Methodology for AIS,” in CAAA An- Clustering: A Quest for Precision,” in
nual Conference 2011, 2011. Business Process Management Work-
shops, 2008, pp. 17–29.
[35] M. J. Jans, “Process Mining in Audit-
ing: From Current Limitations to Fu- [42] M. Song, C. W. Günther, and W. M.
ture Challenges,” in Business Process Van der Aalst, “Trace Clustering in
Management Workshops, vol. 100, F. Process Mining,” in Business Process
Daniel, K. Barkaoui, and S. Dustdar, Management Workshops, 2009, pp.
Eds. Berlin, Heidelberg: Springer, 109–120.
2012, pp. 394–397. [43] W. M. P. van der Aalst, V. Rubin, H.
[36] M. Jans, N. Lybaert, K. Vanhoof, and M. W. Verbeek, B. F. Dongen, E. Kin-
J. M. Van Der Werf, “Business Pro- dler, and C. W. Günther, “Process
cess Mining for Internal Fraud Risk Mining: A Two-Step Approach to Bal-
Reduction: Results of a Case Study,” ance Between Underfitting and Over-
2008. fitting,” Software & Systems Model-
ing, vol. 9, no. 1, pp. 87–111, Nov.
[37] N. Müller-Wickop and M. Schultz,
2008.
“Modelling Concepts For Process Au-
dits - Empirically Grounded Extension [44] N. Müller-Wickop, M. Schultz, N.
Of BPMN,” in Proceedings of the 21st Gehrke, and M. Nüttgens, “Towards

248
APPENDIX A: PUBLICATIONS

Automated Financial Process Audit- net-oriented approach. Cam-bridge,


ing: Aggregation and Visualization of Mass.: MIT Press, 2011.
Process Models,” in Proceedings of [53] K. Jensen and L. M. Kristensen, Col-
the Enterprise Modelling and Infor- oured petri nets. Springer, 2009.
mation Systems Architectures, Ger-
many, 2011. [55] S. T. March and G. F. Smith, “Design
and natural science research on in-
[45] International Federation of Account- formation technology,” Decision sup-
ants, “ISA 320 Materiality in Planning port systems, vol. 15, no. 4, pp. 251–
and Performing an Audit.” 2009. 266, 1995.
[46] M. B. Romney and P. J. Steinbart, Ac- [56] C. Riege, J. Saat, and T. Bucher, “Sys-
counting Information Systems, 11th tematisierung von Evaluationsmetho-
Revised edition (REV). Prentice Hall, den in der gestaltungsorientierten
2008. Wirtschaftsinformatik,” Wissen-
[47] G. Greco, A. Guzzo, L. Pontieri, and D. schaftstheorie und gestaltungsorien-
Sacca, “Mining expressive process tierte Wirtschaftsinformatik, pp. 69–
models by clustering workflow 86, 2009.
traces,” in Advances in Knowledge [57] J. Venable, J. Pries-Heje, and R. Bas-
Discovery and Data Mining, Springer, kerville, “A comprehensive frame-
2004, pp. 52–62. work for evaluation in design science
[48] A. Rozinat and W. M. P. van der Aalst, research,” Design Science Research in
“Conformance Checking of Processes Information Systems. Advances in
Based on Monitoring Real Behavior,” Theory and Practice, pp. 423–438,
Information Systems, vol. 33, no. 1, 2012.
pp. 64–95, Mar. 2008. [58] yWorks GmbH, “yEd - Graph Editor,”
[49] M. Weske, Business process manage- 2013.:
ment concepts, languages, architec- [Link]
tures. Berlin; New York: Springer, ucts_yed_about.html.
2012. [59] University of Hamburg, “Renew - The
[50] N. Mueller-Wickop and M. Schultz, Reference Net Workshop,” 2013.
“ERP Event Log Preprocessing: [Link]
Timestamps vs. Accounting Logic,” in [60] M. Werner, N. Gehrke, and M.
Design Science at the Intersection of Nüttgens, “Towards Automated Anal-
Physical and Virtual Design, Berlin, ysis of Business Processes for Finan-
Heidelberg, 2013, vol. 7939, pp. 105– cial Audits,” in Proceedings of the
119. 11th International Conference on
[51] S. X. Sun and J. L. Zhao, “Formal Wirtschaftsinformatik, Leipzig, 2013.
workflow design analytics using data [61] S. Gregor and A. R. Hevner, “Position-
flow modeling,” Decision Support ing and Presenting Design Science Re-
Systems, vol. 55, no. 1, pp. 270–283, search for Maximum Impact,” MIS
Apr. 2013. Quarterly, vol. 37, no. 2, pp. 337–
[52] W. M. P. van der Aalst and C. Stahl, 355, 2013.
Modeling business processes : a petri [62] M. Werner, N. Gehrke, and M.
Nüttgens, “Business Process Mining

249
APPENDIX A: PUBLICATIONS

and Reconstruction for Financial Au-


dits,” in Proceedings of the 45th Ha-
waii International Conference on Sys-
tem Sciences, Maui, 2012, pp. 5350–
5359.

250
APPENDIX B: PUBLICATION LIST FROM LITERATURE REVIEW

11 Appendix B: Publication List from Literature Review


Title Authors Publication Outlet
a-Algorithm: Structured Workflow Pro- Kwanghoon Kim,Clarence Advances in Knowledge
cess Mining Through Amalgamating A. Ellis Discovery and Data Min-
Temporal Workcases ing
A business process mining application Mieke J. Jans,Jan Martijn Expert Systems with Ap-
for internal transaction fraud mitiga- van der Werf,Nadine Ly- plications
tion baert,Koen Vanhoof
A Case Study on the Suitability of Pro- Maria Leitner,Anne Business Process Man-
cess Mining to Produce Current-State Baumgrass,Sigrid Sche- agement Workshops
RBAC Models fer-Wenzl,Stefanie Rin-
derle-Ma,Mark Strem-
beck
A comprehensive investigation of the Filip Caron,Jan Van- Computers in Industry
applicability of process mining tech- thienen,Bart Baesens
niques for enterprise risk management
A family of case studies on business Ricardo Perez-Cas- The Journal of Systems
process mining using MARBLE tillo,Jose A Cruz- and Software
Lemus,Ignacio Garcia-Ro-
driguez de Guz-
man,Mario Piattini
A genetic programming approach to C. J. Turner,Ashutosh Ti- Genetic and evolutionary
business process mining wari,Jorn Mehnen computation
A GP Process Mining Approach from a Anhua Wang,Weidong Artificial Intelligence and
Structural Perspective Zhao,Chongchen Computational Intelli-
Chen,Haifeng Wu gence
A Grid-Based Multi-relational Ap- Antonio Turi,Annalisa Ap- Database and Expert Sys-
proach to Process Mining pice,Michelangelo tems Applications
Ceci,Donato Malerba
A Heuristic Genetic Process Mining Al- Jiafei Li,Jihong Computational Intelli-
gorithm Ouyang,Mingyong Feng gence and Security
A Hybrid Approach for Dynamic Busi- Ning Li,Jianchu Kang,Wei- E-Business Engineering
ness Process Mining Based On Recon- feng Lv
figurable Nets and Event Types
A Hybrid Approach for Process Mining: Eren Esgin,Pinar Hybrid Artificial Intelli-
Using From-to Chart Arranged by Ge- Senkul,Cem Cimenbicer gence Systems
netic Algorithms
A Hybrid Approach to Process Mining: Eren Esgin,Pinar Senkul International Conference
Finding Immediate Successors of a on Machine Learning and
Process by Using From-To Chart Applications
A Method of Adaptive Process Mining Mei-Hong Shi,Shou-Shan Intelligent System Design
Based on Time-Varying Sliding Window Jang,Yong-Gang and Engineering Applica-
and Relation of Adjacent Event De- Guo,Liang Chen,Kai-Duan tion
pendency Cao
A multi-dimensional quality assess- Jochen De Weerdt,Manu Information Systems
ment of state-of-the-art process dis- De Backer,Jan Vanthie-
covery algorithms using real-life event nen,Bart Baesens
logs

251
APPENDIX B: PUBLICATION LIST FROM LITERATURE REVIEW

Title Authors Publication Outlet


A New Method for Business Process Hua Hu,Jianen Xie,Hai- Web Information Systems
Mining Based on State Equation yang Hu Engineering – WISE 2010
Workshops
A New Process Mining Algorithm Dongyi Wang,Jidong Dependable, Autonomic
Based on Event Type Ge,Hao Hu,Bin Luo and Secure Computing
A New Process Mining Algorithm of Shenghui He,Tao Industrial and Infor-
Workflow Lv,Binggui Huang mation Systems
A novel approach for process mining Lijie Wen,Jianmin Journal of Intelligent In-
based on event types Wang,Wil M. P. van der formation Systems
Aalst,Biqing Huang,Jiagu-
ang Sun
A Novel Approach of Process Mining Hui Zhang,Ying Liu,Chun- Knowledge-Based and In-
with Event Graph ping Li,Roger Jiao telligent Information and
Engineering Systems
A policy-based process mining frame- Jiexun Li,Harry Jiannan Information Systems and
work: mining business policy texts for Wang,Zhu Zhang,J. Leon eBusiness Management
discovering process models Zhao
A Principled Approach to the Analysis Phil Weber,Behzad Intelligent Data Engineer-
of Process Mining Algorithms Bordbar,Peter Tino ing and Automated Learn-
ing - IDEAL 2011
A Process Mining Approach to Rede- N. R. T. P. van Symposium on Symbolic
sign Business Processes - A Case Study Beest,Laura Maruster and Numeric Algorithms
in Gas Industry for Scientific Computing
A process mining based approach to Ming Li,Lu Liu,Lu Yin,Yan- Information Systems
knowledge maintenance qiu Zhu Frontiers
A process-mining framework for the Wan-Shiou Yang,San-Yih Expert Systems with Ap-
detection of healthcare fraud and Hwang plications
abuse
A process-oriented methodology for Ronny S. Mans,Hajo Rei- Information Systems
evaluating the impact of IT: A proposal jers,Daniel Wismeijer,Mi-
and an application in healthcare chiel van Genuchten
A review of business process mining: Ashutosh Tiwari,C. J. Business Process Man-
state-of-the-art and future trends Turner,Basim Majeed agement Journal
A Study of Quality and Accuracy Trade- Zan Huang,Akhil Kumar INFORMS Journal on
offs in Process Mining Computing
A Workflow Process Mining Algorithm Xing-Qi Huang,Li-Fu Journal of Computer Sci-
Based on Synchro-Net Wang,Wen Zhao,Shi-Kun ence and Technology
Zhang,Chong-Yi Yuan
Abstractions in Process Mining: A Tax- R. P. Jagadeesh Chandra Business Process Man-
onomy of Patterns Bose,Wil M. P. van der agement
Aalst
Agent-Based Analysis and Detection of Agnes Werner-Stark,Ti- Agent and Multi-Agent
Functional Faults of Vehicle Industry bor Dulai Systems. Technologies
Processes: A Process Mining Approach and Applications
Algorithms for anomaly detection of Fabio Bezerra,Jacques Information Systems
traces in logs of process aware infor- Wainer
mation systems

252
APPENDIX B: PUBLICATION LIST FROM LITERATURE REVIEW

Title Authors Publication Outlet


An Effective Algorithm for Business Gun-Woo Kim,Seung International Conference
Process Mining Based on Modified FP- Hoon Lee,Jae Hyung on Communication Soft-
Tree Algorithm Kim,Jin Hyun Son ware and Networks
An empirical comparison of static and Ricardo Pérez-Castillo,Ig- Symposium on Applied
dynamic business process mining nacio Garcia-Rodriguez Computing
de Guzman,Mario Piat-
tini,Barbara Weber,Ange-
les S. Places
An empirical evaluation of process Jianmin Wang,Shijie Symposium on Applied
mining algorithms based on structural Tan,Lijie Wen,Raymond Computing
and behavioral similarities K. Wong,Qinlong Guo
An experimental evaluation of genetic C. J. Turner,Ashutosh Ti- Genetic and evolutionary
process mining wari computation
An Exploration of Genetic Process Min- C. J. Turner,Ashutosh Ti- Applications of Soft Com-
ing wari puting
An Extended CompDepend-Algorithm Wang Xiaohui,Zhao Wen- International Conference
to Discovery Duplicate Tasks in Work- qing,Guo Fengjuan on Research Challenges in
flow Process Mining System Computer Science
An improved simulated annealing al- Dianfang Gao,Qiang Liu Conference on Computer
gorithm for process mining Supported Cooperative
Work in Design
An Incremental Process Mining Ap- Andre Cristiano Enterprise Distributed
proach to Extract Knowledge from Leg- Kalsing,Gleison Samuel Object Computing Con-
acy Systems do Nascimento,Cirano ference
Lochpe,Lucineia Heloisa
Thom
An iterative approach to synthesize Ahmed Awad,Rajeev Information Systems
business process templates from com- Gore,Zhe Hou,James
pliance rules Thomson,Matthias
Weidlich
An Outlook on Semantic Business Pro- A. K. A. de Medeiros,C. On the Move to Meaning-
cess Mining and Monitoring Pedrinaci,Wil M. P. van ful Internet Systems
der Aalst,J. 2007: OTM 2007 Work-
Domingue,Minseok shops
Song,Anne Rozinat,B.
Norton,L. Cabral
Analysis of Multi-Agent Interactions Lawrence Cabac,Nicolas Multiagent System Tech-
with Process Mining Techniques Knaak,Daniel nologies
Moldt,Heiko Rölke
Analyzing Multi-agent Activity Logs Us- Anne Rozinat,Stefan Zick- Distributed Autonomous
ing Process Mining Techniques ler,Manuela Veloso,Wil Robotic Systems 8
M. P. van der Aalst,Colin
McMillen
Analyzing Resource Behavior Using Joyce Nakatumba,Wil M. Business Process Man-
Process Mining P. van der Aalst agement Workshops
Analyzing Vessel Behavior Using Pro- Fabrizio M. Maggi,Arjan Situation Awareness with
cess Mining J. Mooij,Wil M. P. van der Systems of Systems
Aalst

253
APPENDIX B: PUBLICATION LIST FROM LITERATURE REVIEW

Title Authors Publication Outlet


Anomaly Detection Using Process Min- Fabio Bezerra,Jacques Enterprise, Business-Pro-
ing Wainer,Wil M. P. van der cess and Information Sys-
Aalst tems Modeling
Application of Process Mining in Ronny S. Mans,M. H. Biomedical Engineering
Healthcare – A Case Study in a Dutch Schonenberg,Minseok Systems and Technologies
Hospital Song,Wil M. P. van der
Aalst,P. J. M. Bakker
Applying Clustering in Process Mining Daniela Luengo,Marcos Business Process Man-
to Find Different Versions of a Busi- Sepulveda agement Workshops
ness Process That Changes over Time
Applying Inductive Logic Programming Evelina Lamma,Paola Inductive Logic Program-
to Process Mining Mello,Fabrizio ming
Riguzzi,Sergio Storari
Applying Petri Net to Analyze a Multi- C. Ou-Yang,Yeh-Chun Global Perspective for
Agent System Feasibility - a Process Juan,C. S. Li Competitive Enterprise,
Mining Approach Economy and Ecology
Applying process mining approach to C. Ou-Yang,Yeh-Chun Journal of Systems Sci-
support the verification of a multi- Juan ence and Systems Engi-
agent system neering
Applying Process Mining in SOA Envi- Ateeq Khan,Azeem Service-Oriented Compu-
ronments Lodhi,Veit Köppen,Gamal ting. ICSOC/ServiceWave
Kassem,Gunter Saake 2009 Workshops
Approaching Process Mining with Se- Diogo R. Ferreira,Mari- Business Process Man-
quence Clustering: Experiments and elba Zacarias,Miguel Mal- agement
Findings heiros,Pedro Ferreira
Auditing 2.0: Using Process Mining to Wil M. P. van der IEEE Computer Society
Support Tomorrow's Auditor Aalst,Kees M. van Press
Hee,Jan Martijn van der
Werf,Marc Verdonk
Beyond Process Mining: From the Past Wil M. P. van der Advanced Information
to Present and Future Aalst,Maja Pesic,Minseok Systems Engineering
Song
Book Review: Process Mining: Discov- Vojtech Huser Journal of Biomedical In-
ery, Conformance and Enhancement formatics
of Business Processes
Bridging Abstraction Layers in Process Thomas Baier,Jan Business Process Man-
Mining by Automated Matching of Mendling agement
Events and Activities
Bridging Abstraction Layers in Process Thomas Baier,Jan Enterprise, Business-Pro-
Mining: Event to Activity Mapping Mendling cess and Information Sys-
tems Modeling
Business alignment: using process min- Wil M. P. van der Aalst Requirements Engineer-
ing as a tool for Delta analysis and con- ing
formance testing
Business process analysis in healthcare Alvaro Rebuge,Diogo R. Information Systems
environments: A methodology based Ferreira
on process mining

254
APPENDIX B: PUBLICATION LIST FROM LITERATURE REVIEW

Title Authors Publication Outlet


Business Process Mining and Recon- Michael Werner,Nick Hawaii International Con-
struction for Financial Audits Gehrke,Markus Nüttgens ference on System Sci-
ences
Business Process Mining and Rules De- Dafne A. Rosso-Pe- Mexican International
tection for Unstructured Information layo,Raul A. Trejo- Conference on Artificial
Ramirez,Miguel Gonza- Intelligence
lez-Mendoza,Neil Her-
nandez-Gress
Business Process Mining Based on Sim- Wei Song,ShaoZhuo International Conference
ulated Annealing Liu,Qiang Liu for Young Computer
Business Process Mining by Means of Dafne A. Rosso-Pe- Mexican International
Statistical Languages Model layo,Raul A. Trejo- Conference on Artificial
Ramirez Intelligence
Business Process Mining from E-Com- Nicolas Poggi,Vinod Mu- Business Process Man-
merce Web Logs thusamy,David Car- agement
rera,Rania Khalaf
Business process mining from group Joao Carlos de A. R. Gon- Conference on Computer
stories calves,Flavia Maria San- Supported Cooperative
toro,Fernanda Araujo Work in Design
Baiao
Business process mining: An industrial Wil M. P. van der Information Systems
application Aalst,Hajo Reijers,A. J. M.
M. Weijters,B. F. van
Dongen,A. K. A. de Me-
deiros,Minseok Song,H.
M. W. Verbeek
Business Process Workarounds: What Nesi Outmazgin,Pnina Enterprise, Business-Pro-
Can and Cannot Be Detected by Pro- Soffer cess and Information Sys-
cess Mining tems Modeling
Case of Process Mining from Business Joonsoo Bae,Young Ki Intelligent Decision Tech-
Execution Log Data Kang nologies
Case Study in Process Mining in a Mul- Paul Taylor,Marcello Data-Driven Process Dis-
tinational Enterprise Leida,Basim Majeed covery and Analysis
Classification and evaluation of timed Hua Duan,Qingtian The Journal of Systems
running schemas for workflow based Zeng,Huaiqing and Software
on process mining Wang,Sherry X.
Sun,Dongming Xu
Clustering and Operation Analysis for Dongha Lee,Jaehun Asia Pacific Business Pro-
Assembly Blocks Using Process Mining Park,Iq Reviessay Pul- cess Management
in Shipbuilding Industry shashi,Hyerim Bae
Combination of Process Mining and Santiago Aguirre,Carlos Data-Driven Process Dis-
Simulation Techniques for Business Parra,Jorge Alvarado covery and Analysis
Process Redesign: A Methodological
Approach
Combining Process Mining and Statisti- Michael Leyer,Jürgen Business Process Man-
cal Methods to Evaluate Customer In- Moormann agement Workshops
tegration in Service Processes

255
APPENDIX B: PUBLICATION LIST FROM LITERATURE REVIEW

Title Authors Publication Outlet


Comprehensive rule-based compliance Filip Caron,Jan Van- Decision Support Systems
checking and risk management with thienen,Bart Baesens
process mining
Configurable Services in the Cloud: Wil M. P. van der Aalst On the Move to Meaning-
Supporting Variability While Enabling ful Internet Systems:
Cross-Organizational Process Mining OTM 2010
Conformance checking of processes Anne Rozinat,Wil M. P. Information Systems
based on monitoring real behavior van der Aalst
Consistent Process Mining over Big Antonia Azzini,Paolo Congress on Big Data
Data Triple Stores Ceravolo
Context-aware process mining frame- Zerari Mounira,Boufaida Information Integration
work for business process flexibility Mahmoud and Web-based Applica-
tions & Services
Continuous Quality Improvement of IT Kerstin Gerke,Gerrit AMCIS 2009 Proceedings
Processes based on Reference Models Tamm
and Process Mining
Coupling Case Based Reasoning and Sameh Triki,Narjes Bella- Workshops on Enabling
Process Mining for a Web Based Crisis mine Ben Saoud,Julie Technologies: Infrastruc-
Management Decision Support System Dugdale,Chihab Hanachi ture for Collaborative En-
terprises
Data and Process Mining: Minitrack In- Selwyn Piramuthu,H. Mi- Hawaii International Con-
troduction chael Chung ference on System Sci-
ences
Data Flow-Oriented Process Mining to Thomas Stocker Service-Oriented Compu-
Support Security Audits ting - ICSOC 2011 Work-
shops
Data Improvement to Enable Process Reinhold Dunkl Computer Aided Systems
Mining on Integrated Non-log Data Theory - EUROCAST 2013
Sources
Data Mining and Process Mining: Busi- H. Michael Chung,Selwyn Hawaii International Con-
ness Impact and Application Chal- Piramuthu ference on System Sci-
lenges ences
Data Transformation and Semantic Log Linh Thao Ly,Conrad Indi- Advanced Information
Purging for Process Mining ono,Jürgen Mangler,Ste- Systems Engineering
fanie Rinderle-Ma
Data-aware process mining: discover- Massimiliano de Le- Symposium on Applied
ing decisions in processes using align- oni,Wil M. P. van der Computing
ments Aalst
Decision Support Based on Process Wil M. P. van der Aalst Handbook on Decision
Mining Support Systems 1
Decision-Process Mining: A New Razvan Petrusel Studia Universitatis
Framework for Knowledge Acquisition Babes-Bolyai
of Business Decision-making Processes
Declarative Process Mining Marco Montali Specification and Verifica-
tion of Declarative Open
Interaction Models
Decomposing Petri nets for process Wil M. P. van der Aalst Distributed and Parallel
mining: A generic approach Databases

256
APPENDIX B: PUBLICATION LIST FROM LITERATURE REVIEW

Title Authors Publication Outlet


Decomposing Process Mining Prob- Wil M. P. van der Aalst Application and Theory of
lems Using Passages Petri Nets
Definition and Validation of Process Irina Ailenei,Anne Ro- Business Process Man-
Mining Use Cases zinat,Albert Eckert,Wil M. agement Workshops
P. van der Aalst
Deriving Event Graphs through Process Hui Zhang,Ying Liu,Chun- Modelling and Manage-
Mining for Runtime Change Manage- ping Li,Roger Jiao ment of Engineering Pro-
ment cesses
Design and Implementation of Process Yong Ling,Liqun International Conference
Mining System Based on a-Algorithm Zhang,Meng Gao on Innovative Computing
Information and Control
Development of a co-operative distrib- G. T. S. Ho,H. C. W. Lau,S. International Journal of
uted process mining system for quality K. Kwok,C. K. M. Lee,W. Production Research
assurance Ho
Development of a process mining sys- H. C. W. Lau,G. T. S. Ho,Y. International Journal of
tem for supporting knowledge discov- Zhao,N. S. H. Chung Production Economics
ery in a supply chain network
Development of an OLAP-Fuzzy Based G. T. S. Ho,H. C. W. Lau Intelligent Information
Process Mining System for Quality Im- Processing III
provement
Development of Distance Measures Joonsoo Bae,Ling International Journal of
for Process Mining, Discovery, and In- Liu,James Caverlee,Liang- Web Services Research
tegration Jie Zhang,Hyerim Bae
Discovering Business Rules through Raphael Crerie,Fernanda Enterprise, Business-Pro-
Process Mining Araujo Baiao,Flavia Maria cess and Information Sys-
Santoro tems Modeling
Discovering Changes of the Change Jana Samalikova,Jos J. M. Software Process Im-
Control Board Process during a Soft- Trienekens,Rob J. provement
ware Development Project Using Pro- Kusters,A. J. M. M.
cess Mining Weijters
Discovering simulation models Anne Rozinat,Ronny S. Information Systems
Mans,Minseok Song,Wil
M. P. van der Aalst
Discovery and analysis of e-mail-driven Marco Stuit,Hans Wort- Information Systems
business processes mann
Distributed Genetic Process Mining Us- Carmen Bratosin,Natalia Parallel Computing Tech-
ing Sampling Sidorova,Wil M. P. van nologies
der Aalst
Divide-and-Conquer Strategies for Pro- Josep Carmona,Jordi Cor- Business Process Man-
cess Mining tadella,Michael Kishinev- agement
sky
Does Process Mining Add to Internal Mieke J. Jans,Benoit De- Enterprise, Business-Pro-
Auditing? An Experience Report paire,Koen Vanhoof cess and Information Sys-
tems Modeling
DPMine/P: modeling and process min- Sergey A. Shershakov Software Engineering
ing language and ProM plug-ins Conference in Russia
Dynamic context-aware Business Pro- Mounira Ze- International Journal of
cess flexibility: an artefact-based ap- rari,Mahmoud Boufaida Business Intelligence and
proach using process mining Data Mining

257
APPENDIX B: PUBLICATION LIST FROM LITERATURE REVIEW

Title Authors Publication Outlet


EMiT: A Process Mining Tool B. F. van Dongen,Wil M. Applications and Theory
P. van der Aalst of Petri Nets 2004
Exploiting Inductive Logic Program- Federico Chesani,Evelina Transactions on Petri
ming Techniques for Declarative Pro- Lamma,Paola Nets and Other Models of
cess Mining Mello,Marco Mon- Concurrency II
tali,Fabrizio Riguzzi,Ser-
gio Storari
Exploring regulatory processes during Cornelia Schoor,Maria Computers in Human Be-
a computer-supported collaborative Bannert havior
learning task using process mining
Exploring the CSCW spectrum using Wil M. P. van der Aalst Advanced Engineering In-
process mining formatics
Finding Structure in Unstructured Pro- Wil M. P. van der Aalst,C. Application of Concur-
cesses: The Case for Process Mining W. Günther rency to System Design
From Local Patterns to Global Models: Nikola Trcka,Mykola Intelligent Systems Design
Towards Domain Driven Educational Pechenizkiy and Applications
Process Mining
Generating event logs from non-pro- Ricardo Perez-Cas- Enterprise Information
cess-aware systems enabling business tillo,Barbara We- Systems
process mining ber,Jakob Ping-
gera,Stefan Zugal,Ignacio
Garcia-Rodriguez de Guz-
man,Mario Piattini
Genetic Process Mining Wil M. P. van der Aalst,A. Applications and Theory
K. A. de Medeiros,A. J. M. of Petri Nets 2005
M. Weijters
Genetic Process Mining: A Basic Ap- A. K. A. de Medeiros,A. J. Business Process Man-
proach and Its Challenges M. M. Weijters,Wil M. P. agement Workshops
van der Aalst
Genetic process mining: an experi- A. K. A. de Medeiros,A. J. Data Mining and
mental evaluation M. M. Weijters,Wil M. P. Knowledge Discovery
van der Aalst
Getting a Grasp on Clinical Pathway Jochen De Weerdt,Filip Emerging Trends in
Data: An Approach Based on Process Caron,Jan Van- Knowledge Discovery and
Mining thienen,Bart Baesens Data Mining
Goal-Heuristic Analysis Method for an Su-Jin Baek,Jong-Won Proceedings of the Inter-
Adaptive Process Mining Ko,Gui-Jung Kim,Jung- national Conference on IT
Soo Han,Young-Jae Song Convergence and Security
2011
Handling Concept Drift in Process Min- R. P. Jagadeesh Chandra Advanced Information
ing Bose,Wil M. P. van der Systems Engineering
Aalst,Indre Zliobaite,My-
kola Pechenizkiy
Healthcare Process Mining with RFID Wei Zhou,Selwyn Piramu- Business Process Man-
thu agement Workshops

258
APPENDIX B: PUBLICATION LIST FROM LITERATURE REVIEW

Title Authors Publication Outlet


Impact of Data Aggregation on the Sig- Kerstin Gerke Workshop on Enabling
nificance of Process Mining Results: An Technologies: Infrastruc-
Experimental Evaluation ture for Collaborative En-
terprises
Improving Business Process Models Fernando Szimanski,Celia Enterprise, Business-Pro-
with Agent-Based Simulation and Pro- G. Ralha,Gerd Wagner,Di- cess and Information Sys-
cess Mining ogo R. Ferreira tems Modeling
Improving process models by discover- Sharmila Subrama- Information Systems
ing decision points niam,Vana Kalogeraki,Di-
mitrios Gunopulos,Fabio
Casati,Malu Castella-
nos,Umeshwar
Dayal,Mehmet Sayal
Incremental Declarative Process Min- Massimiliano Smart Information and
ing Cattafi,Evelina Knowledge Management
Lamma,Fabrizio
Riguzzi,Sergio Storari
Insuring Sensitive Processes through Jorge Munoz-Gama,Isao Ubiquitous Intelligence
Process Mining Echizen and Computing
Integrating Computer Log Files for Pro- Jan Claes,Geert Poels Advanced Information
cess Mining: A Genetic Algorithm In- Systems Engineering
spired Technique Workshops
Intra- and Inter-Organizational Process Wil M. P. van der Aalst The Practice of Enterprise
Mining: Discovering Processes within Modeling
and between Organizations
Leveraging Process-Mining Techniques Geetika T. Laksh- IT Professional Magazine
manan,Rania Khalaf
Logic-Based Incremental Process Min- Stefano Ferilli,Berardina Recent Trends in Applied
ing in Smart Environments De Carolis,Domenico Artificial Intelligence
Redavid
Merging Computer Log Files for Pro- Jan Claes,Geert Poels Business Process Man-
cess Mining: An Artificial Immune Sys- agement Workshops
tem Technique
Model-based Business Process Mining Jon Espen Ingvaldsen,Jon Information Systems
Atle Gulla Management
Multidimensional process mining: a Thomas Vogelgesang,H.- EDBT/ICDT 2013 Work-
flexible analysis approach for health Jürgen Appelrath shops
services research
Multi-phase Process Mining: Building B. F. van Dongen,Wil M. Conceptual Modeling –
Instance Graphs P. van der Aalst ER 2004
Net Components for the Integration of Lawrence Cabac,Nicolas Transactions on Petri
Process Mining into Agent-Oriented Denz Nets and Other Models of
Software Engineering Concurrency I
New methodology for modeling large Pamela Viale,Claudia International Conference
scale manufacturing process: Using Frydman,Jacques Pinaton on Computer Systems
process mining methods and experts' and Applications
knowledge

259
APPENDIX B: PUBLICATION LIST FROM LITERATURE REVIEW

Title Authors Publication Outlet


Obtaining Thresholds for the Effective- Ricardo Perez-Cas- Empirical Software Engi-
ness of Business Process Mining tillo,Laura Sanchez-Gon- neering and Measure-
zalez,Mario Piattini,Felix ment
Garcia,Ignacio Garcia-Ro-
driguez de Guzman
On Recommendation of Process Min- Jianmin Wang,Raymond International Conference
ing Algorithms K. Wong,Jianwei on Web Services
Ding,Qinlong Guo,Lijie
Wen
On the exploitation of process mining Rafael Accorsi,Thomas Symposium on Applied
for security audits: the conformance Stocker Computing
checking case
On the exploitation of process mining Rafael Accorsi,Thomas Symposium on Applied
for security audits: the process discov- Stocker,Günter Müller Computing
ery case
On the Representational Bias in Pro- Wil M. P. van der Aalst Workshops on Enabling
cess Mining Technologies: Infrastruc-
ture for Collaborative En-
terprises
Online Techniques for Dealing with Josep Carmona,Ricard Advances in Intelligent
Concept Drift in Process Mining Gavalda Data Analysis XI
Ontological approach to enhance re- Wirat Jareevong- Business Process Man-
sults of business process mining and piboon,Paul Janecek agement Journal
analysis
Outlier Detection Techniques for Pro- Lucantonio Foundations of Intelligent
cess Mining Applications Ghionna,Gianluigi Systems
Greco,Antonella
Guzzo,Luigi Pontieri
Performance Analysis of GRID Middle- Anastas Misev,Emanouil Computational Science –
ware Using Process Mining Atanassov ICCS 2008
Petri-net integration – An approach to C. Ou-Yang,H. Winarjo Expert Systems with Ap-
support multi-agent process mining plications
Preprocessing Support for Large Scale Jon Espen Ingvaldsen,Jon Business Process Man-
Process Mining of SAP Transactions Atle Gulla agement Workshops
Probabilistic Declarative Process Min- Elena Bellodi,Fabrizio Knowledge Science, Engi-
ing Riguzzi,Evelina Lamma neering and Management
Process Cubes: Slicing, Dicing, Rolling Wil M. P. van der Aalst Asia Pacific Business Pro-
Up and Drilling Down Event Data for cess Management
Process Mining
Process diagnostics using trace align- R. P. Jagadeesh Chandra Information Systems
ment: Opportunities, issues, and chal- Bose,Wil M. P. van der
lenges Aalst
Process Diagnostics: A Method Based Melike Bozkaya,Joost Ga- Information, Process, and
on Process Mining briels,Jan Martijn van der Knowledge Management
Werf
Process Mining Wil M. P. van der Aalst Communications of the
ACM

260
APPENDIX B: PUBLICATION LIST FROM LITERATURE REVIEW

Title Authors Publication Outlet


Process Mining Rafael Accorsi,Meike Ull- Informatik-Spektrum
rich,Wil M. P. van der
Aalst
Process Mining and Petri Net Synthesis Ekkart Kindler,Vladimir Business Process Man-
Rubin,Wilhelm Schäfer agement Workshops

Process Mining and Security: Detecting Wil M. P. van der Aalst,A. Electronic Notes in Theo-
Anomalous Process Executions and K. A. de Medeiros retical Computer Science
Checking Process Conformance
Process Mining and Security: Visualiza- Viet H. Huynh,An N. T. Le Intelligence and Security
tion in Database Intrusion Detection Informatics
Process Mining and Simulation Moe Wynn,Anne Ro- Modern Business Process
zinat,Wil M. P. van der Automation
Aalst,Arthur ter Hof-
stede,Colin Fidge
Process Mining and the ProM Frame- Jan Claes,Geert Poels Business Process Man-
work: An Exploratory Survey agement Workshops
Process Mining and Verification of Wil M. P. van der Aalst,H. On the Move to Meaning-
Properties: An Approach Based on T. de Beer,B. F. van Don- ful Internet Systems
Temporal Logic gen 2005: CoopIS, DOA, and
ODBASE
Process Mining Applied to the BPI R. P. Jagadeesh Chandra Business Process Man-
Challenge 2012: Divide and Conquer Bose,Wil M. P. van der agement Workshops
While Discerning Resources Aalst
Process mining applied to the test pro- Anne Rozinat,I. S. M. De IEEE Transactions on Sys-
cess of wafer scanners in ASML Jong,C. W. Günther,Wil tems
M. P. van der Aalst
Process Mining Approach for Traffic Kirill Krinkin,Eugene Ka- Internet of Things, Smart
Analysis in Wireless Mesh Networks lishenko,S. P. Shiva Spaces, and Next Genera-
Prakash tion Networking
Process Mining Approach to Promote Mehdi Ghazanfari,Mo- Networked Digital Tech-
Business Intelligence in Iranian Detec- hammad Fathian,Mo- nologies
tives’ Police stafa Jafari,Saeed Rou-
hani
Process Mining as First-Order Classifi- Stijn Goedertier,David Business Process Man-
cation Learning on Logs with Negative Martens,Bart Bae- agement Workshops
Events sens,Raf Haesen,Jan
Vanthienen
Process Mining Based on Clustering: A A. K. A. de Medeiros,An- Business Process Man-
Quest for Precision tonella Guzzo,Gianluigi agement Workshops
Greco,Wil M. P. van der
Aalst,A. J. M. M. Weijters,
B. F. van Dongen,Dome-
nico Sacca
Process Mining Based on Regions of Robin Bergenthum,Jörg Business Process Man-
Languages Desel,Robert Lorenz,Se- agement
bastian Mauser

261
APPENDIX B: PUBLICATION LIST FROM LITERATURE REVIEW

Title Authors Publication Outlet


Process Mining Based on Specification Su-Jin Baek,Jinhyang International Conference
Slicing for Dynamic Reconfiguration Kim,Young-Jae Song on Computational Science
and Its Applications
Process Mining by Measuring Process Joonsoo Bae,James Cav- Business Process Man-
Block Similarity erlee,Ling Liu,Hua Yan agement Workshops
Process Mining for Electronic Data In- Robert Engel,Worarat E-Commerce and Web
terchange Krathu,Marco Zaple- Technologies
tal,Christian Pichler,Wil
M. P. van der Aalst, Han-
nes Werthner
Process Mining for Job Nets in Inte- Shinji Kikuchi,Yasuhide Enterprise Information
grated Enterprise Systems Matsumoto,Motomitsu Systems
Adachi,Shingo Moritomo
Process Mining for the multi-faceted Jochen De Weerdt,Anne- Computers in Industry
analysis of business processes - A case lies Schupp,An Vander-
study in a financial services organiza- loock,Bart Baesens
tion
Process Mining for Ubiquitous Mobile A. K. A. de Medeiros,B. F. Ubiquitous Mobile Infor-
Systems: An Overview and a Concrete van Dongen,Wil M. P. van mation and Collaboration
Algorithm der Aalst,A. J. M. M. Wei- Systems
jters
Process Mining Framework for Soft- Vladimir Rubin,C. W. Software Process Dynam-
ware Processes Günther,Wil M. P. van ics and Agility
der Aalst,Ekkart Kind-
ler,B. F. van Dongen,Wil-
helm Schäfer
Process Mining from a Basis of State Marc Sole,Josep Carmona Applications and Theory
Regions of Petri Nets
Process Mining in Auditing: From Cur- Mieke J. Jans Business Process Man-
rent Limitations to Future Challenges agement Workshops
Process Mining in Healthcare: A Con- Silvana Quaglini Business Process Man-
tribution to Change the Culture of agement Workshops
Blame
Process Mining in Healthcare: Data Ronny S. Mans,Wil M. P. Process Support and
Challenges When Answering Fre- van der Aalst,Rob J. B. Knowledge Representa-
quently Posed Questions Vanwersch,Arnold J. Mo- tion in Health Care
leman
Process Mining Manifesto Wil M. P. van der Business Process Man-
Aalst,and other agement Workshops
Process Mining Meets Abstract Inter- Josep Carmona,Jordi Cor- Machine Learning and
pretation tadella Knowledge Discovery in
Databases
Process Mining of RFID-Based Supply Kerstin Gerke,Alexander Commerce and Enterprise
Chains Claus,Jan Mendling Computing
Process Mining Put into Context Wil M. P. van der IEEE Internet Computing
Aalst,Schahram Dustdar
Process Mining Software Repositories Wouter Poncin,Alexander Conference on Software
Serebrenik,Mark van den Maintenance and Reengi-
Brand neering

262
APPENDIX B: PUBLICATION LIST FROM LITERATURE REVIEW

Title Authors Publication Outlet


Process Mining Techniques in Con- Zbigniew Paszkiewicz Business Information Sys-
formance Testing of Inventory Pro- tems Workshops
cesses: An Industrial Application
Process Mining towards Semantics A. K. A. de Medeiros,Wil Advances in Web Seman-
M. P. van der Aalst tics I
Process Mining Versus Intention Min- Ghazaleh Khodabande- Enterprise, Business-Pro-
ing lou,Charlotte Hug,Re- cess and Information Sys-
becca Deneckere,Camille tems Modeling
Salinesi

Process Mining, Discovery, and Inte- Joonsoo Bae,Ling Conference on Web Ser-
gration using Distance Measures Liu,James Caverlee,Wil- vices
liam B. Rouse
Process Mining: A Block-Structured Yan-Liang Qu,Tie-Shi Communication Systems
Mining Approach Zhao and Information Technol-
ogy
Process mining: a research agenda Wil M. P. van der Aalst,A. Computers in Industry
J. M. M. Weijters
Process mining: a two-step approach Wil M. P. van der Software & Systems Mod-
to balance between underfitting and Aalst,Vladimir Rubin,H. eling
overfitting M. W. Verbeek
Process Mining: Algorithm for S-Cover- Jianchun She,Dongqing Workshop on Knowledge
able Workflow Nets Yang Discovery and Data Min-
ing
Process Mining: Discovering Direct Laura Maruster,A. J. M. Discovery Science
Successors in Process Logs M. Weijters,Wil M. P. van
der Aalst,Antal van den
Bosch
Process Mining: Extending α-Algorithm Jiafei Li,Dayou Liu,Bo Advances in Web and
to Mine Duplicate Tasks in Process Yang Network Technologies,
Logs and Information Manage-
ment
Process mining: from theory to prac- C. J. Turner,Ashutosh Ti- Business Process Man-
tice wari,Richard agement Journal
Olaiya,Yuchun Xu
Process Mining: Fuzzy Clustering and B. F. van Dongen,Arya Business Process Man-
Performance Visualization Adriansyah agement Workshops
Process Mining: Overview and Outlook B. F. van Dongen,A. K. A. Transactions on Petri
of Petri Net Discovery Algorithms de Medeiros,L. Wen Nets and Other Models of
Concurrency II
Process Mining-Driven Optimization of Arjel D. Bautista,Lalit Business Process Man-
a Consumer Loan Approvals Process Wangikar,Syed M. Kumail agement Workshops
Akbar
Process-Aware Information Systems: Wil M. P. van der Aalst Transactions on Petri
Lessons to Be Learned from Process Nets and Other Models of
Mining Concurrency II
Process-Mining-Based Workflow Sherry X. Sun,Qingtian IEEE Transactions on Sys-
Model Fragmentation for Distributed Zeng,Huaiqing Wang tems
Execution

263
APPENDIX B: PUBLICATION LIST FROM LITERATURE REVIEW

Title Authors Publication Outlet


Projection approaches to process min- Josep Carmona Data Mining and
ing using region-based techniques Knowledge Discovery
Reality mining via process mining O. M. Hassan,M. S. Applied Informatics and
Farag,M. M. MohieEl-Din Communications
Redesigning business processes: a Laura Maruster,N. R. T. P. Knowledge and Infor-
methodology based on simulation van Beest mation Systems
andprocess mining techniques
Regelbasierte Steuerung von Ge- Heinz Lothar Grob,Frank WIRTSCHAFTSINFORMATI
schäftsprozessen – Konzeption eines Bensberg,Andre Coners K
Ansatzes auf Basis von Process Mining

Relation-Centric Task Identification for Jiexun Li,Harry Jiannan ICIS 2008 Proceedings
Policy-Based Process Mining Wang,Zhu Zhang,J. Leon
Zhao
Requirements towards Effective Pro- Matthias Lohrmann,Alex- On the Move to Meaning-
cess Mining ander Riedel ful Internet Systems:
OTM 2012 Workshops
Rule-Based Business Process Mining: Filip Caron,Jan Van- Management Intelligent
Applications for Management thienen,Bart Baesens Systems
Sequence partitioning for process min- Michal Walicki,Diogo R. Data & Knowledge Engi-
ing with unlabeled event logs Ferreira neering
Similarity-based behavior and process Shusaku Tsumoto,Haruko Future Generation Com-
mining of medical practices Iwata,Shoji Hirano,Yuko puter Systems
Tsumoto
Simplifying discovered process models Dirk Fahland,Wil M. P. Information Systems
in a controlled manner van der Aalst
Skeletal Algorithms in Process Mining Michal R. Przybylek Computational Intelli-
gence
Source Code Partitioning Using Process Koki Kato,Tsuyoshi Business Process Man-
Mining Kanai,Sanya Uehara agement
The case for process mining in audit- Mieke J. Jans,Michael Al- International Journal of
ing: Sources of value added and areas les,Miklos Vasarhelyi Accounting Information
of application Systems
The Need for a Process Mining Evalua- Anne Rozinat,A. K. A. de Business Process Man-
tion Framework in Research and Prac- Medeiros,C. W. Gün- agement Workshops
tice ther,A. J. M. M. Weij-
ters,Wil M. P. van der
Aalst
The Process Mining Manifesto—An in- Gottfried Vossen Information Systems
terview with Wil van der Aalst
The ProM Framework: A New Era in B. F. van Dongen,A. K. A. Applications and Theory
Process Mining Tool Support de Medeiros,H. M. W. of Petri Nets 2005
Verbeek,A. J. M. M. Weij-
ters,Wil M. P. van der
Aalst
The Research of Process Mining As- Zhenyu Wang,Qing International Conference
sessment Used in Business Intelligence Yao,Yuqing Sun on Computer and Infor-
mation Science

264
APPENDIX B: PUBLICATION LIST FROM LITERATURE REVIEW

Title Authors Publication Outlet


The Research on the Usage of Business Xie Yi-Wu,Li Xiao- Network and Parallel
Process Mining in the Implementation Wan,Chen Yan Computing
of BPR
Time prediction based on process min- Wil M. P. van der Information Systems
ing Aalst,M. H. Schonen-
berg,Minseok Song
Time-interval process model discovery Chieh-Yuan Tsai,Henyi Applied Intelligence
and validation--a genetic process min- Jen,Yi-Ching Chen
ing approach
Towards Cross-Organizational Process J. C. A. M. Buijs,B. F. van Business Process Man-
Mining in Collections of Process Mod- Dongen,Wil M. P. van der agement Workshops
els and Their Executions Aalst

Towards Improving the Representa- Wil M. P. van der Aalst,J. Data-Driven Process Dis-
tional Bias of Process Mining C. A. M. Buijs,B. F. van covery and Analysis
Dongen
Trace Alignment in Process Mining: R. P. Jagadeesh Chandra Business Process Man-
Opportunities for Process Diagnostics Bose,Wil M. P. van der agement
Aalst
Trace Clustering in Process Mining Minseok Song,C. W. Gün- Business Process Man-
ther,Wil M. P. van der agement Workshops
Aalst
Translating Message Sequence Charts Kristian Bisgaard Las- Transactions on Petri
to other Process Languages Using Pro- sen,B. F. van Dongen Nets and Other Models of
cess Mining Concurrency I
Using classification methods to label Scott Buffett,Liqiang Journal of Software
tasks in process mining Geng Maintenance and Evolu-
tion
Using Genetic Process Mining Technol- Chieh-Yuan Tsai,I-Ching Next-Generation Applied
ogy to Construct a Time-Interval Pro- Chen Intelligence
cess Model
Using Mapreduce to Scale Events Cor- Hicham Reguieg,Farouk Business Process Man-
relation Discovery for Business Pro- Toumani,Hamid Reza agement
cesses Mining Motahari-Nezhad,Boua-
lem Benatallah
Using minimum description length for Toon Calders,C. W. Gün- Symposium on Applied
process mining ther,Mykola Computing
Pechenizkiy,Anne Rozinat
Using process mining metrics to meas- Chris Thomson,Marian Evaluation and Assess-
ure noisy process fidelity Gheorghe ment in Software Engi-
neering
Using Process Mining to Bridge the Wil M. P. van der Aalst Computer
Gap between BI and BPM
Using process mining to business pro- Faramarz Safi Esfa- Symposium on Applied
cess distribution hani,Masrah Azrifah Azmi Computing
Murad,Md. Nasir Sulai-
man,Nur Izura Udzir

265
APPENDIX B: PUBLICATION LIST FROM LITERATURE REVIEW

Title Authors Publication Outlet


Using Process Mining to Generate Ac- Wil M. P. van der Aalst Business Information Sys-
curate and Interactive Business Pro- tems Workshops
cess Maps
Using process mining to identify coor- Theresa M. Edgington,T. Decision Support Systems
dination patterns in IT service manage- S. Raghu,Ajay S. Vinze
ment
Using process mining to identify mod- Peter Reimann,Jimmy Computer Supported Col-
els of group decision making in chat Frerejean,Kate Thomp- laborative Learning
data son
When Process Mining Meets Bioinfor- R. P. Jagadeesh Chandra IS Olympics: Information
matics Bose,Wil M. P. van der Systems in a Diverse
Aalst World
Workflow simulation for operational Ying Liu,Hui Zhang,Chun- Decision Support Systems
decision support using event graph ping Li,Roger Jiao
through process mining

266
APPENDIX C: FINAL DECLARATION

12 Appendix C: Final Declaration


Hiermit erkläre ich,
Michael Werner, geboren am 10. Februar 1979,
an Eides statt, dass ich die Dissertation mit dem Titel:
„Business Process Analysis Automation for Financial Audits - A design science-
oriented approach to support internal and external auditors in process audits
by using process mining techniques “
selbständig und ohne fremde Hilfe verfasst habe.
Andere als die von mir angegebenen Quellen und Hilfsmittel habe ich nicht benutzt. Die
den herangezogenen Werken wörtlich oder sinngemäß entnommenen Stellen sind als sol-
che gekennzeichnet.

Hamburg, den 7. Oktober 2014 ___________________________


Michael Werner

267

Das könnte Ihnen auch gefallen