Conference Paper Title*
*
Note: Sub-titles are not captured in Xplore and should not be used
1st Given Name Surname
2nd Given Name Surname 3rd Given Name Surname
dept. name of organization (of Aff.)
dept. name of organization (of Aff.) dept. name of organization (of Aff.)
name of organization (of Aff.)
name of organization (of Aff.) name of organization (of Aff.)
City, Country
City, Country City, Country
email address or ORCID
email address or ORCID email address or ORCID
4th Given Name Surname
5th Given Name Surname 6th Given Name Surname
dept. name of organization (of Aff.)
dept. name of organization (of Aff.) dept. name of organization (of Aff.)
name of organization (of Aff.)
name of organization (of Aff.) name of organization (of Aff.)
City, Country
City, Country City, Country
email address or ORCID
email address or ORCID email address or ORCID
Abstract—Automated machine learning (AutoML) is the a brief overview of advancements in the field, all of this at a
branch of machine learning focused on automating, to a certain surface but broad scope focus. The rest of the chapter is
degree, all phases of a machine learning system design. For structured as follows. In Section 2 the basics of AutoML are
supervised learning, AutoML is interested in feature extraction, presented, namely definitions, AutoML system components and
pre-processing, model design, and post processing. Significant concepts. Subsequently, in Section 3 a short review of the most
contributions and developments in AutoML have been paradigmatic AutoML approaches is introduced. Afterwards, in
happening over the last decade. We are thus well-placed to Section 4, a short review on AutoML challenges and AutoML
reflect on and appreciate what we have gained. This chapter is challenges impact on the formation of the field is introduced.
designed to capture the key results in the initial years of Then in Section 5 open issues and research opportunities are
AutoML. Specifically, in this chapter an overview of AutoML for emphasized. At last, in Section 6 a summary of the chapter and
supervised learning is offered and a historical overview of take-home messages are offered.
advancements in this area is given. Similarly, the Principal
paradigms for AutoML are outlined and the opportunities for
research are presented.
I. INTRODUCTION II. FUNDAMENTALS OF AUTOML
Automated4 Machine Learning (AutoML) is the research area
Automated Machine Learning or AutoML is the
nomenclature introduced by the machine learning community to concerned with techniques that focus on minimizing the
denote approaches which try to automate the design and requirement of user intervention in machine learning system and
construction of machine learning systems and applications. application design. The subject has been researched primarily in
Under the scope of supervised learning, AutoML tries to the supervised learning context, though unsupervised and semi
eliminate the user in the loop requirement from all steps in supervised learning activities are also on the rise. This chapter is
supervised learning system design (i.e., any system dependent concerned with AutoML in the supervised learning context.
upon models for classification, recognition, regression,
prediction, etc.). This is a concrete requirement now, since data 1 Supervised learning:
are being produced enormously and in essentially any
environment and situation, yet the amount of available machine Supervised learning is likely the best-researched subject of
learning professionals to do this kind of analysis is overseeded. machine learning since it has broad usage. Spam filtering
AutoML for supervised learning has been under research for algorithms, face detection programs, handwriting digit
over a decade now1, and excellent progress has been made to recognition methods and text classification algorithms are but a
date, look for example at the practical AutoML techniques in
the most widely used machine learning toolboxes, and the few examples of traditional uses depending on supervised
AutoML systems in big scale platforms (e.g., Azure 2 or learning. The unique aspect of supervised learning algorithms is
H2O.ai3). Indeed, AutoML is currently a trendy topic in that they have to learn how to map objects into labels, using a
machine learning that is getting a lot of attention from industry, subset of labeled information (i.e., the supervision). More
academy and even the general public. precisely, within the supervised learning framework we dispose
examples, xi ∈ R^d, and labels5 yi ∈ {−1, 1}, i.e.: D = {(xi,
With such advancements and community interest, it is of a set of data D constituted by N pairs of d-dimensional
required to revisit the basics and key results obtained during
the past decade. This is the objective of the current chapter, yi)}i∈1,.,N. The examples xi codify objects of interest (e.g.,
which is intended to summarize the most significant texts, images or videos) by means of a set of numerical
advancements during the past few years, describe the basics of descriptors, whereas labels yi decide the class of the objects
AutoML, and point out open problems and research directions
in the field. This chapter is supplemental to great surveys and (e.g., spam and no-spam). The general aim of supervised
reviews in the literature that appear in. In comparison to these learning is to discover a function f : R^d n → {−1, 1}
sources, this chapter provides an introduction to AutoML, and transforming inputs to outputs, i.e., yj = f(xj ), which will be able
to generalize outside D. Where choices for the function f • β−level: Learning algorithm search. Refers to the problem of
include linear models, decision trees, instance based finding the optimal learning algorithm for a specific task.
classifiers among others. In all cases for f, learning simplifies Including methods that:
to identifying the f which best describes dataset D. Typically,
– Search the space of all estimators of a specific class, e.g.,
D is divided into training and validation sets, and thus
hyperparameter tuning of a support vector machine (SVM)
learning objectives are to learn f from D so that label
classifier. The shape of f in this scenario is given as: f(x) =
predictions are possible for any other example drawn from the
sign(PN j=1 δjyjk(xj, x) + b), where δ is the variables that
same distribution as D. Let us refer to T as the test set,
correspond to the Lagrange multipliers and k a suitable kernel
constituted by examples originating from the same
function. β−level approaches in this context may attempt to find
distribution as D but not found in said set. We can use T to
good kernel functions k and other hyperparameters for f (e.g.,
test the generalization ability of f. The reader is referred to
regularizer term); likewise, these procedures should still
definitions and thorough treatments of supervised machine
determine the model's parameters, i.e., δ, w and b values.
learning.
– Investigate the area of all estimators which can be
2 Concepts of AutoML:
constructed from a collection of learning algorithms and/or
Having laid out the supervised learning context, we can connected processes such as feature selection, normalization if
loosely define AutoML as the process of searching for the f variables, etc. These kinds of β−level techniques are those which
which generalizes better in any potential T with less possible automatically create classification pipelines such as: PSMS and
human intervention. Where f can be the composition of Auto-WEKA. These approaches are able to decide on what kind
several functions which might project the input space, of function f (e.g., select an SVM or a decision tree classifier),
subsample data, aggregation of several predictors, etc. For but also they can define other procedures to be applied to the
instance, f may be of the type: f(x) = νθν(ΦθΦ (x), where ν is model and/or the training data, prior to, during or following f
a model of classification (e.g., a random forest classifer) and being learned. For example, standard operations might include:
Φ is a feature transformation technique (e.g., feature feature extraction/selecting, assembling with partial solutions
standardization and principal component analysis) with and refining the outputs of models (for example, as per class
hyperparameters θν and θΦ, respectively, and where each of imbalance ratios). β−level techniques are also responsible for
these models might be constructed in turn by a number of setting hyperparameters related to any part of the complete
other functions/models. Functions of the type f(x) = νθν(ΦθΦ model.
(x) are referred to as full models [11, 12] or pipelines, since
• γ−level: Look for meta-learning algorithms. This level
they include all of the steps that must be applied to data to get
concerns techniques that attempt to leverage a knowledge base
a supervised learning model. AutoML can be viewed as the
of tasks-solutions to learn to recommend/select β−level
search for functions ν and Φ, with their respective
techniques based on a new task. This level encompasses
hyperparameters θν and θΦ from D. In the following we show
techniques ranging from the original meta-learning techniques
traditional definitions of AutoML, but the intuitive notion is
for recommending an algorithm among a set of alternatives, to
broad enough to encompass much of what has already been
portfolio optimization techniques, to surrogate models applied
defined, but it should also be more explicit for new people to
by current AutoML frameworks, to state-of-the-art few-shot
the field.
meta-learning approaches. Representative γ−level AutoML
2.1 Levels of automation in AutoML approaches include AutoSklearn using metalearning as warm
start of the optimisation, and first AutoML systems that involve
There are a number of conceptions of AutoML for
surrogates. The specific strength of γ−level methods is that they
supervised learning as far back as 2006 (see the Full model
are based on making use of task-level data and apply it to any
selection definition^6 in), one of the most widely adopted
phase of the AutoML task. Under Liu et al.’s notion, most
being that of Feurer et al. Such definition, though, only the
methodologies aiming to automate the design of machine
automatic pipeline generation problem, while other related
learning systems can be covered. From the (manual)
problems in supervised learning have been referred to as
optimization of parameters for a fixed model, to the automation
AutoML at various points in time. In fact, any task attempting
of any aspect of the design process. A remarkable feature of the
to automate part of the machine learning design process can
above notion is that authors consider budgets (in time and space)
be viewed as AutoML. For example, algorithm selection,
for the different levels. Also, one should note that this notion is
hyperparameter tuning, meta-learning, full model selection,
transverse to the classification7 of model selection methods into
Combined Algorithm Selection and Hyperparameter
filters, wrappers and embedded methods by Guyonet al. For
optimization (CASH), neural architecture search, etc. Since all
more information and examples of tasks/methods belonging to
these activities are interconnected to each other, we cite the
each of these categories, please see.
unifying perspective presented by Liu et al. instead. Liu et al.
identify at least three levels of automation in which AutoML 3. Separating AutoML approaches
systems can be divided, these are summarized below:
The domain of AutoML has expanded quite fast over the past
• α−level: Estimator/predictor search. This level indicates couple of years and as a result a huge number of solutions are
the activity of specifying / setting a function mapping inputs available. To facilitate the reader to differentiate between
regressor for solving a specific task (here yi ∈ R), or
to outputs, such as manually fixing the weights of a linear various AutoML methods, in this section we present the most
important elements of any AutoML approach. In the author's
producing hard-coded classifiers (e.g., based on if-then rules). view, three major elements can be differentiated, namely:
Optimizer, Meta-learner and data-model processing techniques.
This differentiation is visually represented in Figure 1. this is not a strict classification, most γ−level AutoML solutions
follow it. Also, interaction between these elements is highly
The optimizer is the heart of the AutoML process and it
flexible, for example, there are AutoML techniques that employ
includes the optimization algorithm itself, along with the
the meta-learner prior to the optimization step whereas others
objective function (in most cases, a loss function for
employ it during the search. This will be more evident in the
supervised learning). Resource controlling mechanisms are
subsequent section where the most used AutoML methodologies
typically linked to the optimizer, and the aim is to address the
are explained.
optimization problem while satisfying time and memory
budget requirements. While generic optimization techniques
III AUTOML METHODOLOGIES:
(e.g., evolutionary and bio-inspired algorithms, pattern search,
etc.) have been conventionally employed for this central
module of AutoML, ad-hoc optimization methods specific to As noted above, advancements in AutoML have led to
the AutoML context are preferred. These include on a budget, number of methodologies that automate the design and
anytime, and derivative free based approaches. Similarly, development of supervised learning systems at various levels.
multi objective methods and approaches that are capable of It is beyond the scope of this chapter to present a full review
working on intricate structures have the potential to contribute of current methodologies, rather in this section the most
positively towards the overall performance of AutoML representative AutoML methodologies currently available are
approaches. Some of these approaches are discussed in other presented. The reader is referred to recent surveys in AutoML
chapters of this book. for full account of the methodologies available.
1. First wave: 2006-2010:
Particle Swarm Model Selection (PSMS) is one of the
earliest existing AutoML approaches addressing the full
pipeline generation problem [11, 12]. Authors defined
the so called full model selection problem, which is to
find the optimal combination of data preprocessing,
feature selection/extraction and classification models, as
well as the optimization of all of the related
hyperparameters. An heterogenous vector based
representation was introduced to represent models as
vectors and Particle Swarm Optimization (PSO) was
Figure 1. Graphical diagram of the main components of employed to address the issue. Several data sampling
an (γ−level) AutoML method. methods were employed to render the method
A meta-learner is defined as any estimator utilized during manageable. Similarly, Gorissen et al. suggested a
the process of optimization in AutoML, this could be a meta- similar evolutionary algorithm to find surrogates, where
learning method for generating recommendations on possibly any large range of data preprocessing, feature
helpful models, or any other estimator (e.g., of expected selection/extraction and model postprocessing methods
performance or running time) employed by the optimizer. could be explored to construct the model. To the
Meta-learners are included in any γ−level method. The meta- author's best knowledge, this was the initial work
learner is usually paired with the optimizer (e.g., in Auto- suggesting constructing ensembles as a constituent of
WEKA and AutoSklearn)). One must observe that meta- the process of AutoML. This thought motivated other
learning alone can be regarded as an AutoML approach: Early methodologies such as Ensemble PSMS, where models
methods were employed to produce rough model
of ensembles of partial solutions discovered during the
recommendations to address supervised learning issues. Such
a type of algorithm recommendation/choice was available
PSMS search process were returned as a solution.
prior to the initial AutoML designs emerging. Nevertheless, Ensembles are nowadays a component in most
most early attempts at meta-learning were aimed at suggesting successful AutoML implementations such as
a classification model, rarely they also recommended AutoSkLearn. The final AutoML method from early
hyperparameters. Thus, end-to-end pipelines were not in focus days of AutoML that we would like to mention is GPS,
in metalearning in the beginning. Currently, meta-learning is a Quan et al. approached the full model selection problem
topical issue by itself, see, and it has been synergistically with a quite novel formulation: in a first step, authors
utilized in AutoML systems. The reader is referred to for a looked for a promising template for a classification
comprehensive review on meta-learning. pipeline and in a second stage authors optimized
Data processing mechanisms are those which alter, arrange hyperparameters for the selected template. In the
data as per the requirement of AutoML methods. These are author’s view this was a form of warm-starting the
data splitting and sampling for solution assessment (e.g., AutoML process, something that is common in
successive halving algorithms). Lastly, model processing contemporaneous AutoML solutions. As seen from the
techniques are those that improve the model with ad hoc above, it is surprising that a number of key contributions
mechanisms for enhancing AutoML solutions. For example, commonly applied in today's state of the art AutoML
constructing ensembles with partial solutions such as in While systems were made in the first wave. Specifically, an
initial formulation of the AutoML problem, the a SMBO strategy known as SMAC that used a random
concept of constructing ensembles with data obtained forest estimator. The strategy amplified research on
from the AutoML process and early thoughts on SMBO for AutoML that now is the preeminent
warm-starting the search process. optimization method in the field. Remarkably, non-
2. Second wave: 2011-2016: probabilistic methods in contrast to such a formulation
A second wave of AutoML began in the early 2010s were also presented. For instance, Rosales et al.
with the emergence of Bayesian Optimization/ proposed an AutoML method that was grounded on
Sequential Model-Based Optimization (SMBO) based multi-objective optimization and surrogate models. A
models for hyperparameter tuning and algorithm regressor (and subsequent classifier) were employed to
selection. The basic intuition behind these approaches approximate the performance of solutions so that only
is to employ a kind of surrogate model to approximate the high-potential ones were approximated with the
the relationships between performance and expensive objective function. This approach is similar in
hyperparameters, and using this approximation to spirit to SMBO but works the problem differently.
inform the optimization process through an Subsequent to these works, SMBO-based solutions have
acquisition function. In 2013 Thornton et al. presented been proposed in the subsequent years, most
Auto-WEKA, an AutoML technique founded upon prominently AutoSklearn. This technique emerged out
SMBO that can construct classification pipelines of an academic competition (see Section 4).
within the widely used WEKA software [28]. The AutoSklearn is an SMBO with the unique aspect that the
authors defined the CASH problem, which is search is initialized first with a meta-learner with the
analogous to similarities with the task of choosing the goal of guiding the search towards promising models.
full model. Auto-WEKA used a SMBO technique Additionally, this technique produces an ensemble of
named SMAC with a random forest estimator. This solutions visited throughout the search process.
strategy enhanced research in SMBO for AutoML that AutoSklearn had a series of AutoML challenges with
today is the prevalent optimization methodology in wide margin at certain points, and even beat humans that
the domain. Surprisingly, other methodologies that do were trying to fine tune a model8. AutoSklearn was
not comply with a probabilistic formulation were also released for public use and it is widely used today. One
suggested. For instance, Rosales et al. created an of the new AutoML approaches published following
AutoML methodology utilizing multi-objective AutoSklearn is TPOT (Tree-based Pipeline
optimization and surrogate models. A regressor (and Optimization Tool), an evolutionary computation-based
then a classifier) was employed to predict the approach with a unique twist. The key difference of
performance of solutions so that only the good ones TPOT from initial attempts that are evolutionary
were assessed with the expensive objective function. computation based is that TPOT employs genetic
This solution has the same spirit as SMBO but works programming as the optimizer, and models are
in a different manner. In subsequent years, SMBO- represented as syntactic trees constructed from
based solutions have been introduced, most primitives corresponding to models. Every tree
significantly AutoSklearn. This approach emerged in corresponds to a complete classification pipeline and
the framework of an academic competition (see these are evolved in order to optimize performance and
Section 4). AutoSklearn is an SMBO-based one with decrease the complexity (number of primitives utilized)
the characteristic that the search process is initialized of the pipeline. Codifying pipelines as trees is the
with a meta-learner with the aim of guiding the search natural solution not attempted elsewhere. In summary,
towards promising models. Furthermore, this the second waive can be attributed by the development
approach creates an ensemble of solutions that were of Bayesian Optimization as the defacto optimizer in
explored during A second wave of AutoML began in AutoML, the majority of AutoML offerings today use
the early 2010s with the advent of models based on such modeling framework but vary in either how the
Bayesian Optimization / Sequential Model-Based estimators are parameterized or used and combined with
Optimization (SMBO) for algorithm selection and other processes. This waive also saw the rise of meta-
hyperparameter tuning. The basic intuition behind learning as an integral step towards auto-selecting the
these approaches is to employ a kind of surrogate classification models. Similarly, advances in
model to approximate the relationships between hyperparameter optimization led to methods (e.g.,
performance and hyperparameters, and employing this multifidelity methods) that have accelerated research
approximation to inform the optimization process into AutoML, see for a recent review on advances in this
through an acquisition function. In 2013 Thornton et area.
al. presented Auto-WEKA, an AutoML approach
based on SMBO that can construct classification 3. Third wave: 2017 and on:
pipelines in the widely used WEKA platform. The The recent trend in AutoML is that of methods for
authors constructed the CASH problem, a similarity to Neural Architecture Search (NAS) [64, 10, 9, 56, 73].
full model selection problem. Auto-WEKA employed The spectacular success of deep learning in so many
domains, coupled with the huge complexity that it attempts of the previous competitions, were gathered
takes to manually fine-tune a model in order to years later through a sequence of challenges that were
achieve the desired performance in a specific dataset fundamental to increase the interest of the community
has pushed AutoML into the deep learning paradigm towards the AutoML area: the ChaLearn AutoML series
(actually, some authors employ as synonym AutoML [27, 22, 21]. ChaLearn11 directed the organization of a
with NAS). NAS addresses the issue of finding the sequence of competitions with the objective of creating
optimal architecture and hyperparameters of deep the hoped-for AutoML black-box within a 5-stage
learning models. Since this is a very difficult issue evaluation protocol. To begin with, participants worked
due to the number of involved parameters (of order on binary classification tasks, then more difficult
billions) and the size of datasets needed by these supervised learning tasks (regression, multiclass and
models to be able to perform well. Moreover, another multi-label classification) were added in later phases.
particular difficulty that needs to be overcome by Five fresh data sets were published at each phase where
NAS methods is the comparison between participants didn't have any information about the data,
heterogeneous structures. in fact, data was private stayed in the cloud until testing,
(i.e., deep learning models). NAS is beyond the so no code saw the data in advance. This was one of the
purview of this chapter, however the reader is most new-age aspect of the AutoML competition
requested to refer to [9, 10, 64, 56] for current surveys compared to contemporaneous competitions: participant
on this rapidly evolving and dynamic area. NAS along solutions to the AutoML competition were run
with few shot learning, and applying reinforcement automatically in the cloud by anyone without human
learning for AutoML operations make up the third interference. All AutoML solutions were run in exactly
wave, grand progress will be made in these areas as the same manner and under exactly the same
these subjects are at the forefront of the machine circumstances, it was in this challenge that budget limits
learning community. This section has given a general were officially accounted for. Competition also enabled
overview on the history of AutoML over the past the comparison between pure AutoML approaches and
decade. Though the overview is not comprehensive, it regular offline manually tuned solutions. It was realized
provides the reader with a good idea on how the field that there was still some difference between fully
has evolved and, most notably, presents the building autonomous vs. tuned approaches, prompting further
blocks of AutoML. In the rest of this chapter we research for the next editions. It is worth noting that it
outline the role that challenges have played in was within the frame of this challenge that a well-known
AutoML's evolution and we outline open questions and highly effective AutoML approach emerged:
and research opportunities in the field AutoSklearn [17]. The initial AutoML challenge series
targeted mid size tabular data related to supervised
IV AUTOML CHALLENGES: learning problems. The level of complexity in the tasks
undertaken was heightened in the later series targeting
It is a fact that competitions have served to push the more realistic environments and harder problems. As an
state of the art and to solve very hard problems that example, Life Long AutoML challenge required
otherwise would have taken a long time, even students to come up with solutions able to learn
centuries, see e.g., the Longitude Act^10. incessantly in datasets of a larger scale that derived from
For AutoML, they have been instrumental, and, while real deployments [15] and employed standard data
debatable, in the authors' view, AutoML was models (e.g., relational and temporal data12). Due to the
conceived in the heart of academic difficulties. The sheer size of these datasets, solutions to the challenge
2006 Prediction challenge [20] and the 2007 agnostic emphasized efficiency, thus, other features of AutoML
learning vs. prior knowledge competition [25, 26] were targeted by participants (e.g., exhaustive search or
posed a challenge to the participants to come up with overfitting prevention mechanisms) Moreover, the life
methodologies that, using the lesser possible domain long setting encouraged participants to come up with
knowledge, could address generic classification incremental solutions, actually top ranked participants
problems. This challenge spawned several early based on boosting ensembles of trees, see e.g., [71]. One
AutoML approaches, see e.g. [47, 7, 57, 70, 13]. of the most significant results of the challenge was that
While most of them addressed the problem of efficiency in AutoML has not been given its due share
hyperparameter tuning, for the first time were of attention by the community. Additionally, there was
evaluated the benefits of incorporating domain proved the inability of state-of-the-art techniques to
knowledge vs. creating fully agnostic methods when manage non-tabular data. For a detailed description of
constructing generic classifiers. The results of the the challenge please refer to [15]. The newest member
challenge shed light in that it was indeed possible to of the AutoML challenge series is the AutoDL13
build an independent black box that could solve competition [43]. Here, there is a requirement to develop
numerous classification problems. For in-depth AutoML approaches capable of operating directly on
analyses on the results of the challenges and the raw data, where data may be heterogeneous (e.g., text,
solutions developed, please see [26, 25]. The first
images, time series, videos, speech signals, etc.). can make a significant impact are enumerated.
Although the competition focus is on deep learning
methodologies, any kind of method can be submitted. 1. Explainable AutoML models: AutoML solutions
As in previous editions, the solutions are are typically black boxes that are working towards
automatically tested in the CodaLab14 platform with navigating the space of models that are possible to
no intervention on the part of the user, methods are construct with a set of primitives. Successful
not allowed to have access to the data until it is AutoML solutions exist which can be applied by
evaluated, and budgetary constraints exist. In early any user without any training in machine learning.
stages of evaluation in AutoDL which have been Even though the progress has been made, an area
conducted in one modality of the data (for example, which has not been approached by the community
images), highly powerful and efficient deep learning is that of creating transparent AutoML approaches.
architectures already exist [44]. Top ranked According to the author, AutoML models must be
candidates of these evaluation stages have created supplemented with explainability and
effective auto augmentation methods [41] and have interpretability mechanisms. Such an improvement
used light architectures utilized as warmstart for the may have significant advantages for popularizing
AutoML process (e.g., on mobile net [31]). For more AutoML. While this venue remains unexplored so
information on the already established methodologies far, we are not too distant from having explainable
solving recognition tasks from raw data, without AutoML methods since AutoML is in itself a
human intervention and by using reasonable search-intensive process that produces enormous
resources, see [43, 44, 46]. This section has given an amounts of data that can be leveraged to produce
overview of academic competitions addressing the interpretable and explainable AutoML models.
AutoML problem under very disparate and difficult 2. AutoML in feature engineering: While data
conditions. By offering data, resources and correct processing, i.e., feature extraction and selection,
evaluation protocols, AutoMLchallenges have have been thought of as blocks of pipelines
increased research in various fronts of machine synthesized by AutoML methods, the process of
learning. From the 2006 prediction challenge to the feature engineering alone has been largely
2020 AutoDL competition the AutoML field has overlooked by the community. It is only now that
experienced its growth within the machine learning attempts to directly process raw data are on the rise,
community. A number of viable ways to solve the see [36, 51, 43, 44, 48]. We think this research area
problem have been suggested thus far, some of which will be crucial for full automation of the AutoML
are used on a large scale today. AutoML is a good pipeline.
example of what challenges can accomplish, the area 3. AutoML for non-tabular data: In line with the
is developing with significant share of the machine above point, AutoML techniques for handling non
learning community being engaged in it. AutoML tabular data, such as raw data (e.g., text, images,
challenges have also contributed to establishing the etc.) and structured data (e.g., graphs, networks,
foundation for equitable and comparable evaluations. etc.) are increasingly needed, thus this topic could
Such evaluations are highly needed in AutoML, being also be a promising research area.
a data driven and resource consuming process, 4. Large scale AutoML: Large scale problems
ensuring autonomy, and delivering finding solutions remain an open issue for state of the art AutoML
within a reasonable timeframe is imperative. solutions. This was attested in recent AutoML
challenges were only few solutions were able to
V . OPENISSUES AND RESEARCH OPPORTUNITIES: execute search intensive AutoML processes [15].
This is an open issue which requires more
In the last decade, AutoML has achieved a community attention. Similarly, in deep learning
tremendous progress in trying to automate model models, AutoML must be efficient and there
design and development, mainly in the context of already exist a number of efficient implementations
supervised learning. From initial efforts in trying to of NAS models.
approac the problem with straightforward black-box 5. Transfer learning in AutoML: As noted earlier,
optimizers based on vector representations, to the existing AutoML solutions produce huge and rich
most recent studies aiming to compare graphs, and information that can be beneficial for various
adopting meta-learning schemes. With such progress, purposes. One of the promising purposes of
the reader can be fooled that the AutoML problem is leveraging such information is to conduct transfer
solved (at least for tasks such as classification), this is learning to improve the performance of AutoML
still a long way to go. As there are a number of models. Meta-learning methods already do a form
challenges which the community deserves to pay of transfer learning, but an interesting area for
attention to. In the following research is to transfer information on information
some of the most promising issues on which research regarding the optimization process (e.g.,
transferring information on the optimization become feasible to push automation of deep learning.
process dynamics from task to task).
Benchmarking and reproducibility for AutoML :
As AutoML is an optimization over data, work on
providing platforms and frameworks for
evaluation and unbiased comparison across
AutoML approaches is an open problem whose
interest is increasing for the ML community (cf.,
e.g., [42]). While challenges provide such a
platform, they tend to be made obsolete very
quickly due to the rate at which AutoML is
expanding. Similarly, code sharing and processes
that facilitate the reproducibility of results in
AutoML could play a tremendous positive role in
the matury of the field. Interactive
6. AutoML techniques: While the general purpose
of AutoML is to automate operations and taking
as much away from design loop as possible to the
user, interactive AutoML techniques might carry
the performance of AutoML models far beyond
its present stage. Techniques for incorporating
prior knowledge into AutoML process might play
a beneficial role.
VI . CONCLUSIONS:
Automated machine learning aims at helping users with the
design of machine learning systems. From the optimization
of hyperparameters of fixed models, to model type
selection and full model/pipeline generation to the
automated design of deep learning structures, AutoML is
now a mature discipline with broad applicability in the data
science age. Significant advances have been made in the
initial years, with extremely effective methodologies freely
available for use by those with limited machine learning
expertise. Similarly, solutions facilitating the design task
even for machine learning professionals. This chapter has
given an overview of the significant achievements in this
first decade, the most representative methods were shown
and the basis of the AutoML task were given. Maybe the
most significant conclusion that can be obtained from this
early years of progress is that today we have evidence that
AutoML is an achievable task, this is a very significant
outcome because during the early years the machine
learning community was extremely sceptical regarding the
future of the field. Thus, even when the black-box solution
to all problems is still far from being achieved, nowadays
we can build on AutoML methods in order to get close to
problems that used to demand much effort. And it is certain
the contribution that AutoML challenges have made
towards the creation of the discipline. With the
development experienced in this first decade, much is
anticipated from AutoML in the coming years. Specifically,
it will be fascinating to learn of techniques that can get
close to the open problems outlined in the above section.
Furthermore, it is also interesting to what degree will it
VII . References: Optimization, pages 3–33. Springer International
[1] James Urquhart Allingham. Unsupervised automatic dataset Publishing, Cham, 2019.
repair. Master’s thesis, Computer Labo [17] Matthias Feurer, Aaron Klein, Katharina Eggensperger, Jost
ratory, University of Cambridge, 2018. Tobias Springenberg, Manuel Blum, and Frank Hutter. Auto-
[2] Ethem Alpaydin. Introduction to machine learning. Adaptive sklearn: Efficient and Robust Automated Machine Learning, pages
computation and machine learning. MIT Press, 3rd edition, 2014. 113–134. Springer International Publishing, Cham, 2019.
[3] P. J. Angeline, G. M. Saunders, and J. B. Pollack. An [18] Dirk Gorissen, Tom Dhaene, and Filip De Turck. Evolutionary
evolutionary algorithm that constructs recurrent neural networks. model type selection for global surrogate modeling. J. Mach. Learn.
Trans. Neur. Netw., 5(1):5465, January 1994. Res., 10:2039–2078, 2009.
[4] C. Bishop. Pattern Recognition and Machine Learning. [19] Dirk Gorissen, Luciano De Tommasi, Jeroen Croon, and Tom
Springer, 1st edition, 2006. Dhaene. Automatic model type selection with heterogeneous
[5] Bernhard E. Boser, Isabelle M. Guyon, and Vladimir N. evolution: An application to RF circuit block modeling. In
Vapnik. A training algorithm for optimal margin classifiers. In Proceedings of the IEEE Congress on Evolutionary Computation,
Proceedings of the Fifth Annual Workshop on Computational CEC 2008, June 1-6, 2008, Hong Kong, China, pages
Learning Theory, COLT92, page 144152, New York, NY, USA, 989–996, 2008.
1992. Association for Computing Machinery. [20] Isabelle Guyon, Amir Reza Saffari Azar Alamdari, Gideon
[6] L. Breiman. Random forests. Machine learning, 45:5–32, Dror, and Joachim M. Buhmann. Perfor manceprediction challenge.
2001. In Proceedings of the International Joint Conference on Neural
[7] Gavin C. Cawley and Nicola L. C. Talbot. Agnostic learning Networks, IJCNN2006, partof the
versus prior knowledge in the design of kernel machines. In IEEEWorldCongressonComputationalIntelligence, WCCI 2006,
Proceedings of the International Joint Conference on Neural Vancouver, BC, Canada, 16-21 July 2006, pages 1649–1656, 2006.
Networks, IJCNN [21] Isabelle Guyon, Kristin P. Bennett, Gavin C. Cawley, Hugo
2007, Celebrating 20 years of neural networks, Orlando, Florida, Jair Escalante, Sergio Escalera, Tin Kam Ho, N´ uria Maci` a,
USA, August 12-17, 2007, pages Bisakha Ray, Mehreen Saeed, Alexander R. Statnikov, and Evelyne
1732–1737, 2007. Viegas. Design of the 2015 chalearn automl challenge. In 2015
[8] Tobias Domhan, Jost Tobias Springenberg, and Frank Hutter. International Joint Conference on Neural Networks, IJCNN 2015,
Speeding up automatic hyperparameter Killarney, Ireland, July 12-17, 2015, pages 1–8, 2015.
optimization of deep neural networks by extrapolation of [22] Isabelle Guyon, Imad Chaabane, Hugo Jair Escalante, Sergio
learning curves. In Proceedings of the 24th Escalera, Damir Jajetic, James Robert
International Conference on Artificial Intelligence, IJCAI15, Lloyd, N´ uria Maci` a, Bisakha Ray, Lukasz Romaszko, Mich`ele
page 34603468. AAAI Press, 2015. Sebag, Alexander R. Statnikov, S´ ebastien Treguer, and Evelyne
[9] Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. Viegas. A brief review of the chalearn automl challenge: Any-time
Neural architecture search: A survey, 2018. any-dataset learning without human intervention. In Proceedings of
[10] Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. the 2016 Workshop on Automatic
Neural Architecture Search, pages 63–77 Springer International Machine Learning, AutoML 2016, co-located with 33rd
Publishing, Cham, 2019. International Conference on Machine Learning (ICML 2016), New
[11] Hugo Jair Escalante. Results on the model selection game: York City, NY, USA, June 24, 2016, pages 21–30, 2016.
Towards a particle swarm model selection algorithm. NIPS2016 [23] Isabelle Guyon and Andr´e Elisseeff. An introduction to
Multi-level Inference Workshop and Model Selecion Game, variable and feature selection. J. Mach. Learn. Res.,
2006. 3(null):11571182, March 2003.
[12] Hugo Jair Escalante, Manuel Montes, and Luis Enrique [24] Isabelle Guyon, Amir Saffari, Gideon Dror, and Gavin
Sucar. Particle swarm model selection. J. Mach. Learn. Res., Cawley. Model selection: Beyond the bayesian/frequentist divide. J.
10:405–440, June 2009. Mach. Learn. Res., 11:6187, March 2010.
[13] Hugo Jair Escalante, Manuel Montes-y-G´ omez, and Luis [25] Isabelle Guyon, Amir Saffari, Gideon Dror, and Gavin C.
Enrique Sucar. PSMS for neural networks on the IJCNN 2007 Cawley. Agnostic learning vs. prior knowledge challenge. In
agnostic vs prior knowledge challenge. In Proceedings of the Proceedings of the International Joint Conference on Neural
International Joint Conference on Neural Networks, IJCNN 2007, Networks, IJCNN 2007, Celebrating 20 years of neural networks,
Celebrating 20 years of neural networks, Orlando, Orlando, Florida, USA, August 12-17, 2007, pages 829–834, 2007.
Florida, USA, August 12-17, 2007, pages 678–683, 2007. [26] Isabelle Guyon, Amir Saffari, Gideon Dror, and Gavin C.
[14] Hugo Jair Escalante, Manuel Montes-y-G´ omez, and Luis Cawley. Analysis of the IJCNN 2007 agnostic learning vs. prior
Enrique Sucar. Ensemble particle swarm model selection. In knowledge challenge. Neural Networks,21(2-3):544–550, 2008.
International Joint Conference on Neural Networks, IJCNN 2010, [27] Isabelle Guyon, Lisheng Sun-Hosoya, Marc Boull´ e, Hugo
Barcelona, Jair Escalante, Sergio Escalera, Zhengying Liu, Damir Jajetic,
Spain, 18-23 July, 2010, pages 1–8, 2010. Bisakha Ray, Mehreen Saeed, Mich`ele Sebag, Alexander R.
[15] Hugo Jair Escalante, Wei-Wei Tu, Isabelle Guyon, Daniel Statnikov, Wei-Wei Tu, and Evelyne Viegas. Analysis of the
L. Silver, Evelyne Viegas, Yuqiang Chen, Wenyuan Dai, and automl challenge series 2015-2018. In Automated Machine
Qiang Yang. Automl @ neurips 2018 challenge: Design and Learning- Methods, Systems, Challenges, pages 177–219. 2019.
results. In Sergio Escalera and Ralf Herbrich, editors, The [28] Mark Hall, Eibe Frank, Geoffrey Holmes, Bernhard
NeurIPS ’18 Competition, pages 209–229, Cham, 2020. Springer Pfahringer, Peter Reutemann, and Ian H. Witten. The weka data
International Publishing. mining software: An update. SIGKDD Explor. Newsl., 11(1):1018,
[16] Matthias Feurer and Frank Hutter. Hyperparameter November 2009.
[29] T. Hastie, R. Tibshirani, and J. Friedman. The Elements of Escalante, Wei-Wei Tu, Zhen Xu, and Sebastien Treguer. AutoCV
Statistical Learning. Springer, 2nd edition, Challenge Design and Baseline Results. In CAp 2019- Conf´erence
2009. sur l’Apprentissage Automatique, Toulouse, France, July 2019.
[30] Xin He, Kaiyong Zhao, and Xiaowen Chu. Automl: A [44] Zhengying Liu, Zhen Xu, Sergio Escalera, Isabelle Guyon,
survey of the state-of-the-art, 2019. Julio C S Jacques Junior, Meysam Madadi, Adrien Pavao, Sebastien
[31] Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Treguer, and Wei-Wei Tu. Towards Automated Computer Vision:
Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, Analysis of the AutoCV Challenges 2019. working paper or
and Hartwig Adam. Mobilenets: Efficient convolutional neural preprint, November 2019.
networks for mobile vision applications. CoRR, abs/1704.04861, [45] Zhengying Liu, Zhen Xu, Meysam Madadi, Julio Jacques
2017. Junior, Sergio Escalera, Shangeth Rajaa, and Isabelle Guyon.
[32] Frank Hutter, Holger H. Hoos, and Kevin Leyton-Brown. Overview and unifying conceptualization of automated machine
Sequential model-based optimization for general algorithm learning. In Proc. of Automating Data Science Workshop @ECML-
configuration. In Carlos A. Coello Coello, editor, Learning and PKDD, 2019.
Intelligent Optimization, pages 507–523, Berlin, Heidelberg, [46] Zhengying Liu, Zhen Xu, Shangeth Rajaa, Meysam Madadi,
2011. Springer Berlin Heidelberg. Julio C. S. Jacques Junior, Sergio Escalera, Adrien Pavao, Sebastien
[33] Frank Hutter, Lars Kotthoff, and Joaquin Vanschoren, Treguer, Wei-Wei Tu, and Isabelle Guyon. Towards automated
editors. Automated Machine Learning- Methods, Systems, deep learning: Analysis of the autodl challenge series 2019
Challenges. The Springer Series on Challenges in Machine zhengying liu. In Proceedings of Machine
Learning. Springer, 2019. Learning Research, volume 123, page 242252, 2020.
[34] Kevin G. Jamieson and Ameet Talwalkar. Non-stochastic [47] Roman W. Lutz. Logitboost with trees applied to the WCCI
best arm identification and hyperparameter optimization. In 2006 performance prediction challenge datasets. In Proceedings of
Arthur Gretton and Christian C. Robert, editors, Proceedings of the International Joint Conference on Neural Networks, IJCNN
the 19th International Conference on Artificial Intelligence and 2006, part of the IEEE World Congress on Computational
Statistics, AISTATS 2016, Cadiz, Spain, May 9-11, 2016, Intelligence, WCCI 2006, Vancouver, BC, Canada,
volume 51 of JMLR Workshop and Conference Proceedings, 16-21 July 2006, pages 1657–1660, 2006.
pages 240–248. [Link], 2016. [48] Jorge G. MadridandHugoJairEscalante. Meta-learning of text
[35] Haifeng Jin, Qingquan Song, and Xia Hu. Auto-keras: An classification tasks. In Ingela Nystr¨ om, Yanio Hern´ andez
efficient neural architecture search system. In Proceedings of the Heredia, and Vladimir Mili´an N´ u˜ nez, editors, Progress in
25th ACM SIGKDD International Conference on Knowledge Pattern Recognition, Image Analysis, Computer Vision, and
Discovery & Data Mining, KDD 19, page 19461956, New York, Applications- 24th Iberoamerican Congress, CIARP 2019, Havana,
NY, USA, 2019. Association for Computing Machinery. Cuba, October 28-31, 2019, Proceedings, volume 11896 of Lecture
[36] Udayan Khurana. Automating feature engineering in Notes in Computer Science, pages 107–119. Springer, 2019.
supervised learning. In Feature Engineering for Machine [49] Geoffrey F. Miller, Peter M. Todd, and Shailesh U. Hegde.
Learning and Data Analytics, pages 221–244. 2018. Designing neural networks using genetic algorithms. In Proceedings
[37] Erin LeDell. H2o automl: Scalable automatic machine of the Third International Conference on Genetic Algorithms, page
learning. In Proceedings of the AutoML Work 379384, San Francisco, CA, USA, 1989. Morgan Kaufmann
shop at ICML 2020, 2020. Publishers Inc.
[38] Bin Li and Steven C. H. Hoi. Online portfolio selection: A [50] Michinari Momma and Kristin P. Bennett. A Pattern Search
survey. ACM Comput. Surv., 46(3), January Method for Model Selection of Support Vector Regression, pages
2014. 261–274. SIAM, 2002.
[39] Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin [51] Fatemeh Nargesian, Horst Samulowitz, Udayan Khurana,
Rostamizadeh, and Ameet Talwalkar. Hyperband: Elias B. Khalil, and Deepak Turaga. Learning feature engineering
Anovelbandit-based approach to hyperparameter optimization. J. for classification. In Proceedings of the 26th International Joint
Mach. Learn. Res., 18(1):67656816, Conference on Artificial Intelligence, IJCAI’17, page 25292535.
January 2017. AAAI Press, 2017.
[40] Yu-Feng Li, Hai Wang, Tong Wei, and Wei-Wei Tu. [52] Randal S. Olson and Jason H. Moore. TPOT: A tree-based
Towards automated semi-supervised learning. In The Thirty- pipeline optimization tool for automating machine learning. In
Third AAAI Conference on Artificial Intelligence, AAAI 2019, Proceedings of the 2016 Workshop on Automatic Machine
The Thirty-First Innovative Applications of Artificial Intelligence Learning, AutoML 2016, co-located with 33rd International
Conference, IAAI 2019, The Ninth AAAI Symposium on Conference on Machine Learning (ICML 2016), New York City,
Educational Advances in Artificial Intelligence, EAAI 2019, NY, USA, June 24, 2016, pages 66–74, 2016.
Honolulu, Hawaii, USA, January 27- February 1, 2019, pages [53] Randal S. Olson and Jason H. Moore. TPOT: A Tree-Based
4237–4244. AAAI Press, 2019. Pipeline Optimization Tool for Automating Machine Learning,
[41] Sungbin Lim, Ildoo Kim, Taesup Kim, Chiheon Kim, and pages 151–160. Springer International Publishing, Cham, 2019.
Sungwoong Kim. Fast autoaugment. CoRR, [54] Nelishia Pillay, Rong Qu, Dipti Srinivasan, Barbara Hammer,
abs/1905.00397, 2019. and Kenneth Sorensen. Automated design of machine learning and
[42] Marius Lindauer and Frank Hutter. Best practices for search algorithms [guest editorial]. Comp. Intell. Mag., 13(2):1617,
scientific research on neural architecture search. ArXiv preprint, May 2018.
1909.02453, 2019. [55] Esteban Real, Sherry Moore, Andrew Selle, Saurabh Saxena,
[43] Zhengying Liu, Isabelle Guyon, Julio Jacques Junior, Yutaka Leon Suematsu, Jie Tan, Quoc V. Le, and Alexey Kurakin.
Meysam Madadi, Sergio Escalera, Adrien Pavao, Hugo Jair Large-scale evolution of image classifiers. In Proceedings of the
34th International Conference on Machine Learning- Volume 70, under concept drift. In Sergio Escalera and Ralf Herbrich, editors,
ICML17, page 29022911. [Link], 2017. The NeurIPS ’18 Competition, pages 317–335, Cham, 2020.
[56] Pengzhen Ren, Yun Xiao, Xiaojun Chang, Po-Yao Huang, Springer International Publishing.
Zhihui Li, Xiaojiang Chen, and Xin Wang. Acomprehensive [72] Quanming Yao, Mengshuo Wang,Yuqiang Chen ,Wenyuan
survey of neural architecture search: Challenges and solutions, Dai,Yu-Feng Li, Wei-Wei Tu, QiangYang, and Yang Yu. Taking
2020. human out of learning applications: A survey on automated
[57] Juha Reunanen. Model selection and assessment using machine learning, 2018.
cross-indexing. In Proceedings of the International Joint [73] Marc-Andr´e Z¨ oller and Marco F. Huber. Survey on
Conference on Neural Networks, IJCNN 2007, Celebrating 20 automated machine learning, 2019.
years of neural networks, Orlando, Florida, USA, August 12-17, [74] Barret Zoph and Quoc V. Le. Neural architecture search with
2007, pages 2581–2585, 2007. reinforcement learning. CoRR, abs/1611.01578, 2016.
[58] John R. Rice. The algorithm selection problem. In Morris
Rubinoff and Marshall C. Yovits, editors, Advances in
Computers, volume 15, pages 65– 118. Elsevier, 1976.
[59] Alejandro Rosales-P´ erez, Jesus A. Gonzalez, Carlos A.
Coello Coello, Hugo Jair Escalante, and Carlos A. Reyes Garc´
ıa. Multi-objective model type selection. Neurocomputing,
146:83–94, 2014.
[60] Bernhard Scholkopf and Alexander J. Smola. Learning with
Kernels: Support Vector Machines, Regularization, Optimization,
and Beyond. MIT Press, Cambridge, MA, USA, 2001.
[61] Kate A. Smith-Miles. Cross-disciplinary perspectives on
meta-learning for algorithm selection. ACM Comput. Surv.,
41(1), January 2009.
[62] Jasper Snoek, Hugo Larochelle, and Ryan P. Adams.
Practical bayesian optimization of machine learning algorithms,
2012.
[63] Quan Sun, Bernhard Pfahringer, and Michael Mayo. Full
model selection in the space of data mining
operators. In Proceedings of the 14th Annual Conference
Companion on Genetic and Evolutionary Computation, GECCO
12, page 15031504, New York, NY, USA, 2012. Association for
Computing Machinery.
[64] El-Ghazali Talbi. Optimization of deep neural networks: a
survey and unified taxonomy. Working paper or preprint, June
2020.
[65] C. Thornton, F. Hutter, H. H. Hoos, and K. Leyton-Brown.
Auto-WEKA: Combined selection and hyperparameter
optimization of classification algorithms. In Proc. of KDD-2013,
pages 847–855,
2013.
[66] Lukas Tuggener, Mohammadreza Amirian, Katharina
Rombach, Stefan L¨ orwald, Anastasia Varlet, Christian
Westermann, and Thilo Stadelmann. Automated machine
learning in practice: State of the art and recent results. CoRR,
abs/1907.08392, 2019.
[67] Joaquin Vanschoren. Meta-learning: A survey. CoRR,
abs/1810.03548, 2018.
[68] Ricardo Vilalta and Youssef Drissi. A perspective view and
survey of meta-learning. Artif. Intell. Rev., 18(2):77–95, 2002.
[69] Yaqing Wang and Quanming Yao. Few-shot learning: A
survey. CoRR, abs/1904.05046, 2019.
[70] J¨org D. Wichard. Agnostic learning with ensembles of
classifiers. In Proceedings of the International Joint Conference
on Neural Networks, IJCNN 2007, Celebrating 20 years of neural
networks, Orlando, Florida, USA, August 12-17, 2007, pages
2887–2891. IEEE, 2007.
[71] Jobin Wilson, Amit Kumar Meher, Bivin Vinodkumar
Bindu, Santanu Chaudhury, Brejesh Lall, Manoj Sharma, and
Vishakha Pareek. Automatically optimized gradient boosting
trees for classifying large volume high cardinality data streams