SLR Example
SLR Example
Modeling and predicting player behavior is of the utmost importance in developing games. Experience has
proven that, while theory-driven approaches are able to comprehend and justify a model’s choices, such
models frequently fail to encompass necessary features because of a lack of insight of the model builders.
In contrast, data-driven approaches rely much less on expertise, and thus offer certain potential advantages.
Hence, this study conducts a systematic review of the extant research on data-driven approaches to game
player modeling. To this end, we have assessed experimental studies of such approaches over a nine-year
period, from 2008 to 2016; this survey yielded 46 research studies of significance. We found that these studies
pertained to three main areas of focus concerning the uses of data-driven approaches in game player model-
ing. One research area involved the objectives of data-driven approaches in game player modeling: behavior
modeling and goal recognition. Another concerned methods: classification, clustering, regression, and evolu-
tionary algorithm. The third was comprised of the current challenges and promising research directions for
data-driven approaches in game player modeling.
CCS Concepts: • Information systems → Massively multiplayer online games; Data mining; • Com-
puting methodologies → Machine learning; Artificial intelligence; • Applied computing → Computer
games;
Additional Key Words and Phrases: Game player modeling, data-driven approaches, computational models,
systematic literature review (SLR)
ACM Reference format:
Danial Hooshyar, Moslem Yousefi, and Heuiseok Lim. 2018. Data-Driven Approaches to Game Player Mod-
eling: A Systematic Literature Review. ACM Comput. Surv. 50, 6, Article 90 (January 2018), 19 pages.
[Link] 90
1 INTRODUCTION
Player modeling entails describing a game player’s traits and tendencies within a model. Such
characteristics may include a player’s actions and behavior, dispositions and style, motivations, and
aims (Bakkes et al. 2012; Yannakakis et al. 2013). These models can allow a game to tailor its content
or goals automatically to suit the needs of a specific player. Adaptation does not necessarily require
modeling, since games can respond merely to alterations in the world of play, or to a player’s
biometric data (Cowley et al. 2016). Nevertheless, there are multiple advantages to developing a
This work is supported by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIP)
(No. 2016R1A2B2015912).
Author’s addresses: D. Hooshyar and H. Lim, Lyceum, Department of Computer Science and Engineering, Korea Univer-
sity, Anam-ro, Seongbuk-gu, Seoul, the Republic of Korea; emails: [Link]@[Link], limhseok@[Link]; M.
Yousefi, School of Civil, Environmental and Architectural Engineering, Korea University, Anam-ro, Seongbuk-gu, Seoul,
the Republic of Korea; email: [Link]@[Link].
Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee
provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the
full citation on the first page. Copyrights for components of this work owned by others than the ACM must be honored.
Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires
prior specific permission and/or a fee. Request permissions from permissions@[Link].
© 2018 ACM 0360-0300/2018/01-ART90 $15.00
[Link]
ACM Computing Surveys, Vol. 50, No. 6, Article 90. Publication date: January 2018.
90:2 D. Hooshyar et al.
player model to mediate these responses. First, a player model offers comprehension of a player’s
behavior, and thus can justify certain adaptations. Second, it makes adaptations generalizable to
other games.
There are three domains of data conducive to the building of player models. First among these
is gameplay data, that is, data collected immediately from interactions between player and game
within the game world. Second is subjective data—that is, data collected via questionnaires (in-
volving, for instance, demographics, emotional states, personality tests, or psychometrics). Ob-
jective data constitute a third domain, culled from biometrical observations. There are two basic
approaches that are employed in building player models, one theory-driven and one data-driven.
A theory-driven approach relies on the social sciences, particularly experimental psychology; it
entails advancing a model derived from literature and domain experience, then subsequently using
empirical methods to validate the model (Lucas et al. 2013). A data-driven approach, in contrast, re-
lies on the natural sciences and computer science; it first gathers a significant quantity of measure-
ments before applying computational methods to derive models, either fully or semi-automatically
(Yannakakis et al. 2013).
Experience has proven that, while theory-driven approaches are able to comprehend and justify
a model’s choices, such models are frequently missing relevant features because their architects
lack sufficient insight. Games, as inherently open-ended constructs, tend to create a massive space
of actions. For this reason, data-driven approaches show promise because they do not rely on
expert domain knowledge. In addition to being less dependent on expert guidance, data-driven
approaches can likewise collect huge quantities of low-level user actions prior to any modular de-
scription of those actions (Min et al. 2014). Furthermore, this allows for the granular examination of
game player behavior, leading to further understanding. Player movement errors can be examined,
for instance, in order to gain insight into the misconceptions that caused them. Such knowledge
can then generate more sophisticated cognitive models and a more comprehensive understanding
of user behavior.
For these reasons, we are convinced of the promise of data-driven approaches to game player
modeling, a belief shared by a number of works (Cowley et al. 2009; Galway et al. 2008; Lee et al.
2014; Machado et al. 2011; Yannakakis 2012; Zook et al. 2012). However, the increasing interest in
this field has yet to lead to any effort to link its concepts and techniques. Such an organization
of knowledge is vitally important, and requires a systematic review of the current empirical evi-
dence regarding the benefits, challenges, and application of data-driven approaches to game player
modeling. To our knowledge, to date there has been neither such a review nor even a summary of
empirical evidence. In executing a systematic review of the empirical evidence, this article aims to
make several contributions. First, it seeks to offer a review, both systematic and comprehensive,
detailing how data-driven approaches have been implemented in game player modeling. Second,
it intends to provide a classification of the assorted data-driven approaches to game player mod-
eling. In so doing, it will (third) offer an understanding of the current challenges and promising
future directions in the field.
The structure of the article is as follows: Section 2 details the methodology of our review;
Section 3 outlines the results and analyzes important studies; Section 4 elaborates on the
discussion and considers future directions, limitations, and conclusions.
2 RESEARCH METHOD
A systematic literature review (SLR) necessitates a comprehensive and impartial search plan, so
as to ensure the completeness of the search of materials under consideration. The recent upswing
in interest in the field of data-driven approaches to game player modeling has yet to produce
ACM Computing Surveys, Vol. 50, No. 6, Article 90. Publication date: January 2018.
Data-Driven Approaches to Game Player Modeling 90:3
any such comprehensive effort to survey its concepts, methods, and problems; our aim here is to
employ Keele’s methodology in a much-needed SLR (Keele 2007).
ACM Computing Surveys, Vol. 50, No. 6, Article 90. Publication date: January 2018.
90:4 D. Hooshyar et al.
Inclusion criteria
IC1 Peer-reviewed articles published in journals or full-length articles
published in International Conference/Workshop Proceedings
IC2 Presenting full results
IC3 Dated between 2008 and 2016
IC4 English language studies
Third, the full texts of these 94 articles were screened for final eligibility, with matters of inclu-
sion weighed according to the inclusion criteria and quality attributes of Table 2. The 94 articles
were distributed to the three reviewers, who carefully reviewed each one in an effort to correct
for any bias. The researchers applied the four criteria to assess quality of the articles that had
ACM Computing Surveys, Vol. 50, No. 6, Article 90. Publication date: January 2018.
Data-Driven Approaches to Game Player Modeling 90:5
been filtered through the second phase. They used a scale of 1 to 5 (with 5 denoting “very high,”
4 “high,” 3 “medium,” 2 “low,” and 1 “very low” on the given category); total scores thus could
span from 20 (highest quality, without much risk of bias) to 4 (poor quality, potentially significant
bias). Any conflicting opinions arising between the reviewers in the process of selection, assess-
ment, and data extraction were addressed in meetings to reach a consensus. The result of this
exhaustive screening process was to winnow the 94 full-text research articles down to 46 research
studies of the application of data-driven approaches to game player modeling that were conducted
from 2008 to 2016. Thus, the suitability of the articles finally selected for inclusion was assured.
Figure 2 presents the number of articles initially derived from each source in proportion to the
articles included in the final tally.
The reviewers carefully evaluated the full text of all the studies that passed the second phase,
and then catalogued relevant material using a pre-established form for the extraction of data ac-
cording to research strategy used (Keele 2007). The form included research discipline, data gath-
ering, research objectives, data acquisition, analysis technique, and results. The demographics and
overview of these studies are presented in Appendix A.
ACM Computing Surveys, Vol. 50, No. 6, Article 90. Publication date: January 2018.
90:6 D. Hooshyar et al.
ACM Computing Surveys, Vol. 50, No. 6, Article 90. Publication date: January 2018.
Data-Driven Approaches to Game Player Modeling 90:7
Paper ID in
Appendix A Game Purpose Reference
A1 Alice in AreaLand learning (Falakmasir et al. 2016)
A2 Combat with learning (Zook et al. 2012)
Monsters
A3 The Hunter entertainment (Ramirez-Cano et al. 2010)
A4 Forza Motorsport 5 entertainment (Harpstead et al. 2015)
A5 World of Warcraft entertainment (Harrison and Roberts 2011)
A6, CRYSTAL ISLAND: learning (Baikadi et al. 2011; Ha et al. 2012;
A13,A24,A25, OUTBREAK/ Min et al. 2016a, 2016b, 2014)
A41 CRYSTAL ISLAND:
UNCHARTED
DISCOVERY
A7, A17 Infinite Mario entertainment (Shaker et al. 2010; Weber et al. 2011)
A8,A26,A30, StarCraft entertainment (Bisson et al. 2015; Kabanza et al. 2010;
A32 Synnaeve and Bessiere 2011; Weber and
Mateas 2009)
A9 The Legend of Zelda entertainment (Gold 2010)
A10 Rush Football entertainment (Laviers and Sukthankar 2011)
ACM Computing Surveys, Vol. 50, No. 6, Article 90. Publication date: January 2018.
90:8 D. Hooshyar et al.
Number of
Research objectives (goals) papers Reference
Player behavior modeling/ 34 (Anagnostou and Maragoudakis
Experience modeling/Procedural 2009; Bellotti et al. 2009; Borbora
content generation and Srivastava 2012, Butler et al.
2015; Cowley et al. 2009, 2013,
2014; Drachen et al. 2009;
Etheredge et al. 2013; Falakmasir
et al. 2016; Gao et al. 2016; Gow
et al. 2012; Grappiolo et al. 2011;
Harpstead et al. 2015; Harrison and
Roberts 2011; Lee et al. 2014; Liu
et al. 2013; Luo et al. 2016;
Mahlmann et al. 2010; Missura and
Gärtner 2009; Pedersen et al. 2010,
Ramirez-Cano et al. 2010; Shaker
et al. 2010; Sharma et al. 2010;
Spronck and den Teuling 2010;
Tamassia et al. 2016; Valls-Vargas
et al. 2015; Weber et al. 2011a,
2011b; Xie et al. 2014; Yannakakis
and Hallam 2009; Yu and Riedl
2012; Zook et al. 2012; Zook and
Riedl 2012)
Goal/plan recognition 12 (Baikadi et al. 2011; Bisson et al.
2015; Gold 2010; Ha et al. 2012;
Kabanza et al. 2010; Laviers and
Sukthankar 2011; Min et al. 2014,
2016a, 2016b, 2014; Orkin et al.
2010; Synnaeve and Bessiere 2011;
Weber and Mateas 2009)
characterize the AI methods that are primary or secondary in each approach. Primary methods
denote the techniques most frequently applied in the literature, whereas secondary methods
denote ones that appear in a significant number (but not a majority) of studies.
In reviewing the studies given in Table 5, we discovered that a number of classification
algorithms are employed in the behavioral modeling of game players. These include Hidden
Markov Models (HMM), Bayesian Networks (BN), Markov Logic Networks (MLN), Support Vector
Machines (SVM), Neural Networks (NN), Long Short-Term Memory (LSTM), Recursive Neural
Networks (RNN), K-Nearest Neighbor (KNN), Decision Trees (DT), and Deep Learning (DL). The
classification algorithms employed most often for behavior modeling were the NN and its variant,
as well as MLN and HMM, which thus garnered increased attention concerning their modeling
potential (see A16, A17, A20-A22, A28, A34, A36 and A43 in Appendix A). After these, the next
in popularity were SVM and DT (A19, A37, A38 and A40). Note that classification methods were
sometimes paired with clustering or evolutionary algorithms in the studies under review; in such
ACM Computing Surveys, Vol. 50, No. 6, Article 90. Publication date: January 2018.
Data-Driven Approaches to Game Player Modeling 90:9
cases, we labeled them under classification, since the data-driven methods those articles advanced
depend primarily on the methods of classification.
Liu et al. (2013), for instance, worked out a method to predict player behavior and movement
in an educational game. Their algorithm selects from a combination of methods—Markov models,
state aggregation, and player heuristic search—according to whichever offers the largest amount
of data. Their approach promises to minimize the burden of hand-designing a cognitive model and
system-specific features. Lee et al. (2014) likewise developed a framework for predicting player
movements within a pair of puzzle games, which, they showed, greatly outperformed both games’
baselines. They proposed a data-driven approach, a mixture of a one-depth heuristic search model
and a data-driven Markov model with no knowledge of individual history other than the current
game state, to learning individual behavior. They added to the prior framework by constructing
a state-action graph and employing methods of feature selection to minimize the number of fea-
tures for each state. They then applied this new framework to another game (DragonBox) to as-
certain its applicability for extension, and reported similar success. The studies undertaken by Lee
et al. (2014) and Liu et al. (2013) hold the potential to address the two main issues of purely data-
driven approaches—their failure to offer any semantically meaningful interpretation of outputs,
on the one hand, and their shortcomings in algorithmic efficiency on the other. However, unlike
the method proposed by Lee et al. (2014), Liu et al. (2013) assume that players do not change over
ACM Computing Surveys, Vol. 50, No. 6, Article 90. Publication date: January 2018.
90:10 D. Hooshyar et al.
time and define a user reward as a combination of pre-defined heuristics. Thus, the quality of the
learned policy is strongly dependent on these heuristics, which are often system-specific and time
consuming to refine. Lee et al. (2014) overcome this issue by incorporating temporal variations in
player performance, thus allowing for modeling change in a player’s skill over time.
Pedersen et al. (2010) took a different tack, developing a neural network which charts
three elements—behavioral characteristics of players, their self-reported emotions, and level
parameters—in order to be able to successfully predict from in-game behavior and content what
sort of affective experiences (enjoyment, frustration, effort, etc.) the players might report. This
approach begins by studying play style, then constructs a model of users’ preferences from which
it produces levels to meet those desires. In so doing, they assume that a user’s preferences remain
fixed within a game session. However, if this assumption is suspect (or if a player’s style of play
might undergo variation), then a more adaptable model is needed to observe and describe how a
player’s behavior changes over time. Valls-Vargas et al. (2015) addressed this issue in proposing
a new framework for player modeling, one which does not depend on the notion of fixed player
style but instead attempts to predict a dynamic state of play. This framework captures fluctua-
tions in player behavior through the use of episodic information and time interval models within
a sequential machine learning method capable of learning several models progressively (such as
a Support Vector Machine or Decision Tree). Testing various strategies of trace segmentation for
predicting player style, they found that sequential machine learning approaches are superior to
non-sequential ones when they include predictions from prior segments. Moreover, they found
that predictive performance declines when the trace segmentation is either too detailed (minute-
by-minute) or too imprecise (entire traces).
Factor analysis is also frequently employed in learning individual behavior. Zook and Riedl
(2012) developed a tensor factorization technique for anticipating a player’s performance in skill-
based computer games. This data-driven tensor factorization method offers more precise adapta-
tions of challenges to individual players because it is capable of forecasting alterations in a player’s
game mastery in time. They show via an empirical study (involving game users engaging a basic
role-playing combat game) that tensor factorization models are both effective and easily scalable.
But while their approach can even incorporate temporal behavior by including time factors, such
approaches do not easily fit games which require predicting a fine granularity of action instead of
predicting user’s single valued performance.
After classification algorithms, various regression techniques coupled with clustering algo-
rithms have been used in a pair of studies (see A1, A3, A4, A7, A12, A15, A27, A30, A31, A33,
A35, A45, and A46 in Appendix A). These approaches primarily use clustering to obtain appro-
priate groupings of players and then perform a regression method within each group. Hence, we
find the primary and secondary methods of player behavior modeling to be classification first,
regression algorithms generally coupled with clustering techniques second.
Mahlmann et al. (2010), for example, considered whether one can forecast specific aspects of
player behavior, examining by way of supervised learning the commercial game Tomb Raider: Un-
derworld (TRU). Specifically, they sought to anticipate the moment at which a player will cease
playing, or conversely how long the player would take to finish a game. The manufactured predic-
tors focus on metrical data derived from game play in the two initial levels of TRU. The outcomes
clearly demonstrate the efficacy of linear regression techniques as well as nonlinear approaches
to classification.
Mere evolutionary algorithms were the least commonly applied data-driven approach for
learning player behavior (see A2 and A39 in Appendix A). For instance, Bellotti et al. (2009)
devised an engine for learning games derived from evolutionary computation approaches such
as genetic algorithm and reinforcement learning (which separate expertise in game design from
ACM Computing Surveys, Vol. 50, No. 6, Article 90. Publication date: January 2018.
Data-Driven Approaches to Game Player Modeling 90:11
domain expertise). Their method approaches the design of a learning game in terms of task
authoring: domain experts identify and gloss domain knowledge to manifest within gameplay
as tasks and task selection, while game designers establish the manner in which tasks are
represented and chosen. A game experience module would then, at the start of a game, determine
from a player’s profile what subset of tasks to generate. In this manner, the game experience
module takes on a double role: it becomes at once a domain expert (ascertaining a player’s current
capabilities and knowledge level) and a designer (selecting and staging suitable game tasks).
3.3.2 Goal Recognition. Abductive reasoning seeks to discern the user’s intentions by study-
ing his or her actions (Carberry 2001). This is referred to as goal recognition, and it views the
user as enacting goal-directed behavior (or, broadly speaking, understands that he/she is aiming
at realizing a specific state in the world). What goal recognition seeks to accomplish is to derive,
from a user’s prior actions and knowledge of the domain, an understanding of the state that the
user seeks to realize in the world. Goal recognition is closely related to the challenges involved in
other forms of recognition—specifically, of actions and plans. The former, also known as activity
recognition (Turaga et al. 2008), deduces a user’s action from sensory information (for instance,
computer vision); whereas plan recognition (Kautz 1987) tackles the broader, more difficult chal-
lenge of predicting both a user’s goal and the precise sequence of actions by which he or she will
pursue it. In method, goal recognition can be categorized into alternate approaches—one rooted in
planning systems, another rooted in models of probability. Each type of goal recognition possesses
its own technological limitations, but that is a topic beyond the focus of this study.
Notably, player modeling significantly overlaps with goal/plan recognition, as such recognition
entails a problem in which action sequences are interpreted in terms of the goal they are most
probably attempting to realize. Plan recognition algorithms capture a set of observations and from
it generates a set of goals that can account for the observed content. The smallest set of goals that
can account for a sequence of actions is generally deemed superior to larger sets, since it probably
holds more significant explanatory value (Carberry 2001). These methods of plan recognition are
also applicable to forecasting player behavior: if a player’s goal can be ascertained, it follows that
the player will execute the intervening steps necessary for completing those goals. Player modeling
differs from plan recognition, however, in that it sidesteps the process of predicting goals and
deducing behavior from them and instead predicts behavior immediately.
Table 4 demonstrates the importance of goal recognition in player modeling, as goal recognition
constitutes the research objective of 12 studies. In these articles, the authors pursue the prediction
or recognition of goals, plans, and actions (see A6, A8-A11, A13, A24-A26, A30, A32, and A41 in
Appendix A). Several AI techniques have been studied and used for predicting goals, plans, and
actions by the authors of reviewed research studies (see Table 5), such as classification, clustering,
regression, and evolutionary algorithms. For each approach, as with behavior modeling, we have
determined the primary or secondary AI techniques employed. In reviewing the studies given in
Table 5, we discovered that a number of classification algorithms are used in goal/plan recognition.
These include HMM, MLN, SVM, BN, LSTM, RNN, and DL. The classification algorithms employed
most often for goal recognition were the MLN, HMM, and NN and its variant, which thus garnered
increased attention concerning their modeling potential (see A6, A9, A13, A25, A26, and A41 in
Appendix A). After these, the next in popularity were SVM (A10), BN (A32), and DL (A24). After
classification algorithms, regression techniques along with clustering algorithms have also been
used in a pair of studies (see A8 and A30 in Appendix A). Evolutionary algorithms, however, were
not applied for goal recognition.
Ha et al. (2011), for instance, apply MLNs to develop a successful goal recognition framework for
players. This framework employs model parameters derived from a corpus of player interactions
ACM Computing Surveys, Vol. 50, No. 6, Article 90. Publication date: January 2018.
90:12 D. Hooshyar et al.
within a nonlinear game and enables an automatic goal recognition system superior to a number
of baseline models. MLNs are also put to work in a similar goal recognition approach, by Baikadi
et al. (2011), wherein problem-solving goals are linked to discovery events by means of probabilis-
tic inference and first-order logical reasoning. When empirically tested, models employing dis-
covery events-based models surpassed previous best approaches on all fronts. However, in using
this approach—a combination of hand-authored logic formulae and machine-learned weights—
some labor-intensive feature engineering efforts need to be eliminated (by utilizing, for example,
multi-level feature abstraction techniques).
Another method is put forward in a study by Gold (2010), in which an Input-Output Hidden
Markov Model (IOHMM) is trained to identify a player’s goal within an action-adventure game.
The goals—the hidden states in the IOHMM—were Fight, Explore, and Return to Town, and the
observation model learned through players being directed toward specific goals and counting ac-
tions. Initially, there was little difference in performance between models trained to specific first-
time players and the model trained to the experimenter. Subsequent models trained to these same
players’ ensuing gameplay vastly improved, however, over both the first-time player- and the
experimenter-trained models. In other words, it appears that game goal recognition systems work
best with players who have developed a specific playing style. It should be noted that the proposed
approach by Gold (2010) was compared to a hand-authored finite state machine, a common com-
putational framework used in commercial games, where the results showcase the effectiveness of
the proposed approach; computational approaches based on deep learning for goal recognition,
however, outperform the proposed approach (Min et al. 2014).
Furthermore, Min et al. (2014) developed a goal recognition framework grounded in stacked
denoising autoencoders (a form of deep learning), in which the goal recognition models are
trained through a corpus of player interactions. These models featured superior performance—
significantly outperforming the best of the MLN-based goal recognition frameworks—and, more-
over, present a profound advantage in that they require no labor-intensive engineering of features.
Laviers and Sukthankar (2011) have developed a technique for the real-time adaptation of plan
repair policies, by means of Upper Confidence Bounds applied to Trees (UCT). Their research
shows how such policies can, in a game of American football, be paired with plan recognition
to generate a fully autonomous offense that can counter surprising shifts in defensive strategy.
The offense produced by this real-time UCT approach easily surpassed both the baseline game
and domain-specific heuristics in offensive efficiency, measured in average yardage and turnover
avoidance. They cannot model goal alterations, however—that is, instances wherein a player
adopts a new goal and forsakes the previous one. Moreover, this approach is unable to speak to-
wards the relative superiority of training individual players versus aggregating player behavior, or
the question of whether the usefulness of models depend on players having in-game experience.
Alternate approaches involve employing LSTM to learn these frameworks—goal recognition,
multiple-goal recognition, task recognition, and plan repair policies. Regarding the use of latter
in adversarial games, it is advisable to pair it with plan repair, for the reason that such games
frequently require players to undertake a plan prior to being able to gauge the intentions of the
opponent. This is especially true of multi-agent domains, in which it is not computationally pos-
sible to undertake total replanning. Hence, Min et al. (2016b) have approached goal recognition
as a sequence labeling task by means of LSTM. Such an approach to goal recognition outper-
forms previous techniques (e.g., Markov’s logic network-based models and n-gram encoded feed
forward neural networks pre-trained with stacked denoising autoencoders) in accuracy testing.
In this respect, LSTM-based approaches apparently solve a major problem in game player mod-
eling, at once offering superior accuracy in goal recognition and also eliminating the exhaustive
engineering previously required for such modeling.
ACM Computing Surveys, Vol. 50, No. 6, Article 90. Publication date: January 2018.
Data-Driven Approaches to Game Player Modeling 90:13
— Insufficient or Inaccessible Data: Regardless of their specific area of focus in game player
modeling, researchers struggle first and foremost with a dearth of publicly available, robust
data. An expansive multimodal body of data and descriptions of gameplay and players is
particularly urgent. Such a body of data must exhaustively cover gameplay data for a few
highly populated multiplayer games—their actions and locations, events and timestamps.
Moreover, researchers must be able to draw on information concerning players’ demograph-
ics, questionnaire responses, and formal interviews (Of course, such data need not extend
to every player in the database). The database should also incorporate the few significant
databases of gameplay available at the present.
— Heterogeneous Data: While some machine learning methods, such as logistic regression, and
ad hoc graphical models do promise certain advantages in analyzing and mining particular
games data, they demand too much expert adjustment to be more widely applicable. There
is also the challenge of interpreting a scarcity of raw data for sophisticated player modeling.
This difficulty necessitates a comprehensive data-driven method, one which is not reliant on
costly domain knowledge engineering and which produces a coherent and understandable
model that is accurate both in its account of the data and in its prediction of game outcomes.
— Problem of Generalizability: The problem of generalizability describes a gaming situation
in which an agent adapts only to a limited set of states in a game’s world and is unable to
generalize beyond them. It causes poor performance as soon as the game states change, or
the agent engages new states. The standard technique in applied machine learning for pre-
empting the tendency towards generalizability—a set of validation examples—might prove
insufficient in a game environment, on account of both the time required to learn a set of
validation examples and the difficulty of extracting such a set from a scarcity of training
data.
— Algorithmic Efficiency: In highly populated and open-ended games, the massive amount
of data collected can rapidly overwhelm players’ PCs or gaming consoles. The solution
to this technical difficulty requires a separation of modeling and predicting into two non-
coterminous phases. A model can be developed offline, for instance, unconstrained by any
untenable efficiency requirements, leaving only the prediction phase to be executed online
with the efficiency necessary to prevent gameplay delays.
— Data Scarcity Problem: For a number of reasons, it is difficult to develop a set of training
examples for games with a large search space, and particularly for games where a relevant
example might be bound up with a series of delayed, interleaving, or compounding results.
In such cases, player-generated material might have to provide the training examples, unless
concrete targets can be removed from the learning process altogether. This scarcity of data
can be offset, though, by joining methods of collaborative filtering with external information
offered by content-based approaches.
— Knowledge Engineering: Multiple data-driven methods (e.g., Knowledge tracing/Bayesian
network models) are unable to connect modeled player types to game events on their
own, and must be supplemented by knowledge engineering annotations. Hence, these
ACM Computing Surveys, Vol. 50, No. 6, Article 90. Publication date: January 2018.
90:14 D. Hooshyar et al.
4 DISCUSSIONS
4.1 Research Question 3
RQ3: Looking forward, what are promising future directions in data-driven game player model-
ing?
Here we outline four promising future directions in the application of data-driven approaches
in game player modeling (see A1, A5, A14, A19-A21, A25, A27-A30, A38, and A39 in Appendix A).
— Data Mining Techniques for Individual Prediction: While data-mining techniques have been
shown to offer, via careful application, insight regarding the behaviors of groups, their util-
ity in predicting individual behavior is still limited. Hence, player models generally offer
only broad and fuzzy indications for ways in which a game should tailor itself to specific
players (Yannakakis et al. 2013). One potential solution to this problem is to define a few
player models and then categorizes individual players accordingly. In subsequent gameplay,
the model can undergo small changes to better match the player. In other words, rather than
viewing the model as a fixed representation of the player, by which the game makes adapta-
tion decisions, it is seen as a dynamic representation of a group of players that modulates to
match a specific player’s characteristics and thus initiates dynamic adaptation in the game.
We believe that this second approach to game adaptation via player modeling offers prac-
tical advantages: it can quickly produce effective results in developing educational games
that offer a personalized form of engagement to the player.
— Hybrid Player Modeling Approaches: In analyzing model-free and model-based methods of
player modeling, we observe that model-free approaches manifest none of the argumen-
tation and interpretation concerning the model’s choices that model-based approaches
inevitably involve. And yet, model-based methods frequently overlook relevant features
because their builders did not have sufficient insight, whereas model-free methods will au-
tomatically detect such features (Yannakakis 2012). The error of model-free approaches,
however, is a tendency to deduce connections between user attributes, experience, and con-
text that are entirely meaningless. These points to a broader present issue particularly in
open-ended games: such games present the chance to collect and measure a set of player
behavior features far more extensive than our current understanding of what these features
might signify. Given this situation, model-free methods seem preferable, even though mean-
ingful player models require domain-specific knowledge and the extraction and selection
of features. We believe that one must characterize model-based and model-free approaches
ACM Computing Surveys, Vol. 50, No. 6, Article 90. Publication date: January 2018.
Data-Driven Approaches to Game Player Modeling 90:15
as points at either end of a continuum, along which one can place various hybrid efforts to
understand player behavior in games.
— Data-Driven Approaches to Conceptualizing Log Data: While log files generated by games
allow the researcher to learn behavior of players as they play the game, there are a number
of practical issues associated with analyzing such data. To name a few, log files represent
prohibitively large quantities of data; it is hard to interpret them since the responses of
individual player are highly context dependent. It is also difficult to determine which actions
represent key features of player performance given that log files are generally designed
to capture all player actions relevant to game play, and not until after analysis can one
know which of those actions were relevant to learning. For this reason, we hold that a data-
driven approach offers significant promise to researchers, insofar as it does not depend
on expensive domain knowledge engineering, and insofar as it can select and extract the
essential conceptual features of player performance from games’ log data.
— Data-Driven Approaches to Modeling Players in Infinitely Open and Replayable Worlds:
Interest has never been higher in the promise procedural content generation methods hold
for optimized game design, in both commercial and independent game development. It is
now assumed that new games will have more user-generated or [Link]
content than manual content, particularly because costly and bottlenecked content
creation are major impediments in the game development process. But as automatically
generated games become more common, it becomes increasingly difficult to model players
in endlessly open worlds of infinite replayability value. Here also we are of the mind that
researchers would significantly benefit from a data-driven approach that is not dependent
on expensive domain knowledge engineering and can model players within such infinitely
open and replayable worlds.
4.2 Limitations
We have constrained our references with the criteria of inclusion. Additionally, we have included,
in applying assessment criteria, studies that addressed the work that initially generated the specific
lines of research discussed. Readers must be advised that it is impossible to exhaustively review,
in a single article, the myriad aspects of data-driven methods in game player modeling. In fact,
certain topics (such as procedural content generation) are deserving of an entire review on their
own. Rather, our aim here has been to give a representative selection of current research within
each learning method.
4.3 Conclusions
In spite of the increased interest in data-driven approaches to game player modeling, there has
yet to be any effort to review current empirical evidence regarding the benefits, challenges, and
applications of such approaches. This article has thus focused on conclusions from empirical stud-
ies, thereby offering a systematic overview of the evidence pertaining to data-driven approaches
in game player modeling. By examining the available literature for representative, rigorous, and
frequently cited case studies from game player modeling domains, we cast light on the data-driven
approaches that have been employed and in so doing demonstrate the potential held by this new
field in game research. In Table 6, we present our findings concerning the aims of data-driven ap-
proaches to player modeling in games, the algorithms employed in this type of player modeling,
and the present difficulties and future directions of game player modeling.
While previous methods relied on questionnaires to gain insight into perceptions and attitudes,
now every movement and “click” in an electronic learning environment might encode useful infor-
mation that can be tracked and interpreted. Computational methods offer the potential to isolate,
ACM Computing Surveys, Vol. 50, No. 6, Article 90. Publication date: January 2018.
90:16 D. Hooshyar et al.
identify, and categorize any simple or complex action within a game, placing it within meaningful
patterns. Interactions of all sorts can be coded into behavioral patterns and interpreted to guide
decision-making. This exciting juncture lies at the crossroads of computer and learning science,
psychology, and pedagogy. What remains is to grasp the deeper learning processes by means of
breaking them down into simpler and discrete mechanisms. It is our hope and belief that this active
area of research will continue to provide invaluable contributions to the development of dynamic,
powerful, and accurate games, both for designers and for players.
APPENDIX A
Demographic data and overview of the selected studies:
ACKNOWLEDGMENTS
Special thanks to the anonymous reviewers for their insightful comments that helped us make this
article better.
REFERENCES
Kostas Anagnostou and Manolis Maragoudakis. 2009. Data mining for player modeling in videogames. In Proceedings of
the 13th Panhellenic Conference on Informatics (PCI’09). IEEE, 30–34.
Alok Baikadi, Jonathan Rowe, Bradford Mott, and James Lester. 2011. Generalizability of goal recognition models in
narrative-centered learning environments. In Proceedings of the International Conference on User Modeling, Adaptation,
and Personalization. Springer International Publishing, 278–289.
Sander Bakkes, Pieter H. M. Spronck, and Giel van Lankveldl. 2012. Player behavioural modeling for video games. Enter-
tainment Computing, 3, 3, 71–79. DOI: [Link]
ACM Computing Surveys, Vol. 50, No. 6, Article 90. Publication date: January 2018.
Data-Driven Approaches to Game Player Modeling 90:17
Francesco Bellotti, Riccardo Berta, Alessandro De Gloria, and Ludovica Primavera. 2009. Adaptive experience engine for
serious games. IEEE Transactions on Computational Intelligence and AI in Games 1, 4, 264–280.
Francis Bisson, Hugo Larochelle, and Froduald Kabanza. 2015. Using a recursive neural network to learn an agent’s decision
model for plan recognition. In Proceedings of the 25th International Joint Conference on Artificial Intelligence (IJCAI). 918–
924.
Zoheb Borbora and Jaideep Srivastava. 2012. User behavior modeling approach for churn prediction in online games. In
Proceedings of the 2012 International Conference on Privacy, Security, Risk and Trust (PASSAT) and Proceedings of the 2012
International Conference on Social Computing (SocialCom). IEEE, 51–60.
Eric Butler, Erik Andersen, Adam M. Smith, Sumit Gulwani, and Zoran Popović. 2015. Automatic game progression design
through analysis of solution features. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing
Systems. ACM, 2407–2416.
Sandra Carberry. 2001. Techniques for plan recognition. User Modeling and User-Adapted Interaction 11, 1, 31–48.
Ben Cowley, Darryl Charles, Michaela Black, and Ray Hickey. 2009. Analyzing player behavior in Pacman using feature-
driven decision theoretic predictive modeling. In Proceedings of the IEEE Symposium on Computational Intelligence and
Games (CIG 2009). IEEE, 170–177.
Benjamin Cowley, Darryl Charles, Michaela Black, and Ray Hickey. 2013. Real-time rule-based classification of player types
in computer games. User Modeling and User-Adapted Interaction 23, 5, 489–526.
Benjamin Cowley, Marco Filetti, Kristian Lukander, Jari Torniainen, Andreas Henelius, Lauri Ahonen, Oswald Barral 2016.
The psychophysiology primer: A guide to methods and a broad review with a focus on human–computer interaction.
Foundations and Trends® Human–Computer Interaction 9, 3–4, 151–308.
Benjamin Cowley, Ilkka Kosunen, Petri Lankoski, J. Matias Kivikangas, Simo Järvelä, Inger Ekman, Jaakko Kemppainen,
and Niklas Ravaja. 2014. Experience assessment and design in the analysis of gameplay. Simulation & Gaming 45, 1,
41–69.
Anders Drachen, Alessandro Canossa, and Georgios N. Yannakakis. 2009. Player modeling using self-organization in Tomb
Raider: Underworld. In Proceedings of the IEEE Symposium on Computational Intelligence and Games (CIG 2009). IEEE,
1–8.
Marlon Etheredge, Ricardo Lopes, and Rafael Bidarra. 2013. A generic method for classification of player behavior. In
Proceedings of the 2nd Workshop on Artificial Intelligence in the Game Design Process. 1–7.
Mohammad Falakmasir, José P. González-Brenes, Geoffrey J. Gordon, and Kristen E. DiCerbo. 2016. A data-driven approach
for inferring student proficiency from game activity logs. In Proceedings of the 3rd ACM Conference on Learning@ Scale.
ACM, 341–349.
Leo Galway, Darryl Charles, and Michaela Black. 2008. Machine learning in digital games: A survey. Artificial Intelligence
Review 29, 2, 123–161.
Chen Gao, Haifeng Shen, and M. Ali Babar. 2016. Concealing jitter in multi-player online games through predictive be-
haviour modeling. In Proceedings of the IEEE 20th International Conference on Computer Supported Cooperative Work in
Design (CSCWD). IEEE, 62–67.
Kevin Gold. 2010. Training goal recognition online from low-level inputs in an action-adventure game. In Proceedings of
AIIDE.
Jeremy Gow, Robin Baumgarten, Paul Cairns, Simon Colton, and Paul Miller. 2012. Unsupervised modeling of player style
with LDA. IEEE Transactions on Computational Intelligence and AI in Games 4, 3, 152–166.
Corrado Grappiolo, Yun-Gyung Cheong, Julian Togelius, Rilla Khaled, and Georgios N. Yannakakis. 2011. Towards player
adaptivity in a serious game for conflict resolution. In Proceedings of the 3rd International Conference on Games and
Virtual Worlds for Serious Applications (VS-GAMES). IEEE, 192–198.
Eunyoung Ha, Jonathan P. Rowe, Bradford W. Mott, and James C. Lester. 2011. Goal recognition with Markov logic networks
for player-adaptive games. In Proceedings of AIIDE.
Erik Harpstead, Thomas Zimmermann, Nachiappan Nagapan, Jose J. Guajardo, Ryan Cooper, Tyson Solberg, and Dan
Greenawalt. 2015. What drives people: Creating engagement profiles of players from game log data. In Proceedings of
the 2015 Annual Symposium on Computer-Human Interaction in Play. ACM, 369–379.
Brent Harrison and David L. Roberts. 2011. Using sequential observations to model and predict player behavior. In Proceed-
ings of the 6th International Conference on Foundations of Digital Games. ACM, 91–98.
Froduald Kabanza, Philipe Bellefeuille, Francis Bisson, Abder Rezak Benaskeur, and Hengameh Irandoust. 2010. Opponent
behaviour recognition for real-time strategy games. Plan, Activity, and Intent Recognition 10, 5.
Henry Kautz. 1987. A Formal Theory of Plan Recognition. Ph.D. Dissertation, Bell Laboratories.
Keele Staffs. 2007. Guidelines for Performing Systematic Literature Reviews in Software Engineering. EBSE Technical Report,
Ver. 2.3. EBSE. sn.
Kennard Laviers and Gita Sukthankar. 2011. A real-time opponent modeling system for rush football. In Proceedings-of the
International Joint Conference on Artificial Intelligence (IJCAI). 22, 3, 2476.
ACM Computing Surveys, Vol. 50, No. 6, Article 90. Publication date: January 2018.
90:18 D. Hooshyar et al.
Jae Lee Seong, Yun-En Liu, and Zoran Popovic. 2014. Learning individual behavior in an educational game: a data-driven
approach. In Proceedings of the Conference on Educational Data Mining 2014. 114–121.
Yun-En Liu, Travis Mandel, Eric Butler, Erik Andersen, Eleanor O’Rourke, Emma Brunskill, and Zoran Popovic. 2013.
Predicting player moves in an educational game: A hybrid approach. In Proceedings of the Conference on Educational
Data Mining 2013.
Simon Lucas, Michael Mateas, Mike Preuss, Pieter Spronck, Julian Togelius, Peter I. Cowling, Michael Buro, Michal Bida,
Adi Botea, and Bruno Bouzy. 2013. Artificial and computational intelligence in games. Schloss Dagstuhl-Leibniz-Zentrum
fuer Informatik.
Linbo Luo, Haiyan Yin, Wentong Cai, Jinghui Zhong, and Michael Lees. 2016. Design and evaluation of a data-driven
scenario generation framework for game-based training. IEEE Transactions on Computational Intelligence and AI in
Games 99, 1–1.
Marlos Machado, Eduardo P. C. Fantini, and Luiz Chaimowicz. 2011. Player modeling: Towards a common taxonomy. In
Proceedings of the 16th IEEE International Conference on Computer Games (CGAMES). IEEE, 50–57.
Tobias Mahlmann, Anders Drachen, Julian Togelius, Alessandro Canossa, and Georgios N. Yannakakis. 2010. Predicting
player behavior in Tomb Raider: Underworld. In Proceedings of the 2010 IEEE Symposium on Computational Intelligence
and Games (CIG). IEEE, 178–185.
Wookhee Min, Alok Baikadi, Bradford Mott, Jonathan Rowe, Barry Liu, Eun Young Ha, and James Lester. 2016a. A gen-
eralized multidimensional evaluation framework for player goal recognition. In Proceedings of the 12th Annual AAAI
Conference on Artificial Intelligence and Interactive Digital Entertainment.
Bradford Mott, Jonathan Rowe, Barry Liu, and James Lester. 2016b. Player goal recognition in open-world digital games
with long short-term memory networks. In Proceedings of the 25th International Joint Conference on Artificial Intelligence.
2590–2596.
Wookhee Min, Eunyoung Ha, Jonathan P. Rowe, Bradford W. Mott, and James C. Lester. 2014. Deep learning-based goal
recognition in open-ended digital games. In Proceedings of AIIDE.
and Thomas Gärtner. 2009. Player modeling for intelligent difficulty adjustment. In Proceedings of the International Confer-
ence on Discovery Science. Springer, Berlin, 197–211.
Jeff Orkin, Tynan Smith, Hilke Reckman, and Deb Roy. 2010. Semi-automatic task recognition for interactive narratives
with EAT & RUN. In Proceedings of the Intelligent Narrative Technologies Iii Workshop. ACM.
Christopher Pedersen, Julian Togelius, and Georgios N. Yannakakis. 2010. Modeling player experience for content creation.
IEEE Transactions on Computational Intelligence and AI in Games 2, 1, 54–67.
Daniel Ramirez-Cano, Simon Colton, and Robin Baumgarten. 2010. Player classification using a meta-clustering approach.
In Proceedings of the 3rd Annual International Conference Computer Games, Multimedia & Allied Technology. 297–304.
Noor Shaker, Georgios N. Yannakakis, and Julian Togelius. 2010. Towards automatic personalized content generation for
platform games. In Proceedings of AIIDE.
Manu Sharma, Santiago Ontañón, Manish Mehta, and Ashwin Ram. 2010. Drama management and player modeling for
interactive fiction games. Computational Intelligence 26, 2, 183–211.
Pieter Spronck and Freek den Teuling. 2010. Player modeling in civilization IV. In Proceedings of AIIDE.
Gabriel Synnaeve and Pierre Bessiere. 2011. A Bayesian model for plan recognition in RTS games applied to StarCraft.
arXiv preprint arXiv:1111.3735.
Marco Tamassia, William Raffe, Rafet Sifa, Anders Drachen, Fabio Zambetta, and Michael Hitchens. 2016. Predicting player
churn in destiny: A hidden Markov models approach to predicting player departure in a major online game. In Proceed-
ings of the IEEE Conference on Computational Intelligence and Games (CIG). IEEE, 1–8.
Pavan Turaga, Rama Chellappa, Venkatramana S. Subrahmanian, and Octavian Udrea. 2008. Machine recognition of human
activities: A survey. IEEE Transactions on Circuits and Systems for Video Technology 18, 11, 1473–1488.
Josep Valls-Vargas, Santiago Ontanón, and Jichen Zhu. 2015. Exploring player trace segmentation for dynamic play style
prediction. In Proceedings of the 11th AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment.
93–99.
Ben George Weber and Michael Mateas. 2009. A data mining approach to strategy prediction. In Proceedings of the IEEE
Symposium on Computational Intelligence and Games (CIG 2009). IEEE, 140–147.
Ben George Weber, Michael Mateas, and Arnav Jhala. 2011a. Using data mining to model player experience. In Proceedings
of the FDG Workshop on Evaluating Player Experience in Games.
Ben George Weber, Michael John, Michael Mateas, and Arnav Jhala. 2011b. Modeling player retention in Madden NFL 11.
In Proceedings of IAAI. 1701-1706.
Hanting Xie, Daniel Kudenko, Sam Devlin, and Peter Cowling. 2014. Predicting player disengagement in online games. In
Proceedings of the Workshop on Computer Games. Springer International Publishing, 133–149.
Georgios N. Yannakakis and John Hallam. 2009. Real-time game adaptation for optimizing player satisfaction. IEEE Trans-
actions on Computational Intelligence and AI in Games 1, 2, 121–133.
ACM Computing Surveys, Vol. 50, No. 6, Article 90. Publication date: January 2018.
Data-Driven Approaches to Game Player Modeling 90:19
Georgios N. Yannakakis, Pieter Spronck, Daniele Loiacono, and Elisabeth André. 2013. Player modeling. In Dagstuhl Follow-
Ups. Schloss Dagstuhl-Leibniz-Zentrum fuer Informatik.
Georgios N. Yannakakis 2012. Game AI revisited. In Proceedings of the 9th Conference on Computing Frontiers. ACM, 285-292.
Hong Yu and Mark O. Riedl. 2012. A sequential recommendation approach for interactive personalized story generation. In
Proceedings of the 11th International Conference on Autonomous Agents and Multiagent Systems. International Foundation
for Autonomous Agents and Multiagent Systems, 71-78.
Alexander Zook, Stephen Lee-Urban, Michael R. Drinkwater, and Mark O. Riedl. 2012. Skill-based mission generation: A
data-driven temporal player modeling approach. In Proceedings of the 3rd Workshop on Procedural Content Generation
in Games. ACM, 6.
Alexander Zook and Mark O. Riedl. 2012. A temporal data-driven player model for dynamic difficulty adjustment. In Pro-
ceedings of AIIDE.
ACM Computing Surveys, Vol. 50, No. 6, Article 90. Publication date: January 2018.