0% found this document useful (0 votes)
4 views22 pages

Variable Selection in Functional Regression

This paper reviews methodological advances in variable selection for functional regression models, highlighting the intersection of Functional Data Analysis (FDA) and High-Dimensional Data Analysis (HDS). It discusses the challenges of infinite-dimensional functional objects and the need for effective variable selection techniques, focusing on both finite-dimensional regression methods and their adaptations for functional settings. The review aims to foster collaboration between FDA and HDS, presenting various methodologies and their implications for future research in functional regression.

Uploaded by

Manuel Castro
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views22 pages

Variable Selection in Functional Regression

This paper reviews methodological advances in variable selection for functional regression models, highlighting the intersection of Functional Data Analysis (FDA) and High-Dimensional Data Analysis (HDS). It discusses the challenges of infinite-dimensional functional objects and the need for effective variable selection techniques, focusing on both finite-dimensional regression methods and their adaptations for functional settings. The review aims to foster collaboration between FDA and HDS, presenting various methodologies and their implications for future research in functional regression.

Uploaded by

Manuel Castro
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Variable selection in functional regression models: a

review
Germán Aneirosa∗ Silvia Novob Philippe Vieuc
a Department of Mathematics, MODES, CITIC, ITMATI, Universidade da Coruña, A Coruña, Spain
b
arXiv:2401.14867v1 [[Link]] 26 Jan 2024

Department of Mathematics, MODES, CITIC, Universidade da Coruña, A Coruña, Spain


c Institut de Mathématiques, Université Paul Sabatier, Toulouse, France

Abstract

Despite of various similar features, Functional Data Analysis and High-Dimensional Data Anal-
ysis are two major fields in Statistics that grew up recently almost independently one from each
other. The aim of this paper is to propose a survey on methodological advances for variable selection
in functional regression, which is typically a question for which both functional and multivariate
ideas are crossing. More than a simple survey, this paper aims to promote even more new links
between both areas.

Keywords: Functional Data Analysis; Regression; Variable selection

1 Introduction
Nowadays, Functional Data Analysis (FDA) is among the main fields in Statistics. The rich produc-
tion is confirmed in various surveys (see, for instance, Goia and Vieu (2016), Aneiros et al. (2019a)).
In the beginning, the presence of functional data in applications was rare. However, with the devel-
opment of modern technology most applied sciences have to treat datasets containing one, or more,
functional object. For the same reasons one has to treat High-(but finite) Dimensional Data, and
High-Dimensional Statistics (HDS) grew up at the same time with FDA. The main common feature
of both fields is that they take part of the recent enfatuation for Big Data Analysis. At the beginning,
both fields developed in the statistical community in rather independent ways but the benefits that
one could get by crossing ideas from both fields have been highlighted in the last decade as well in
the HDS community (see Ahmed (2017), Sangalli (2018), Vieu (2018)) as in the FDA community (see
Aneiros et al. (2019b), Bongiorno et al. (2014)). Following its 50 years long tradition of publishing
top level innovative methodological advances on multidimensional data analysis, the Journal of Multi-
variate Analysis has played a leading role in the last decade for bridging gaps between FDA and HDS.
This is, for instance, attested by various special issues aiming to promote methodological advances
by linking both fields (see Goia and Vieu (2016), Aneiros et al. (2019a), Aneiros et al. (2022)). This


Corresponding author email address: ganeiros@[Link]

1
paper aims to celebrate this 50th birthday by proposing a review on variable selection methods within
a functional framework which is a topic where both fields are crossing in a natural way.
When dealing with a regression problem with one (or more) functional predictor, sometimes with
some non-functional multivariate predictor, one has many concerns. First of all one has to take into
account the fact that we are dealing with infinite-dimensionality of functional objects (see, for instance,
Cuevas (2014)) and to keep in mind the necessity of building models balancing flexibility, dimension
reduction properties and interpretability (Vieu (2018)). Secondly, one also has to worry about the
quantity of information to be included into the model: this concerns the number of predictors as well as
the number of discretizations that one has at hand for each functional predictor. In a pragmatic way,
as it is the case in HDS with non-functional high-dimensional predictors, one would like to determine a
smaller subset of variables that exhibits the strongest effects on the response (see Hastie et al. (2009)).
In the last decade, there has been a rather large production on sparse modelling and variable selection
techniques in functional setting, and this article is aiming to review the state of art on this topic.
Our paper is organized as follows. Because most of the variable selection procedures in functional
setting were extended from finite-dimensional regression, we start in Section 2 with a selected review
on the techniques employed in HDS with main attention on penalized methods. In the exposition we
will present the procedures by splitting them according to the nature of the model: linear, grouped and
additive regression. Of course, Section 2 is not supposed to be an exhaustive review of the very wide
set of contributions existing in a multivariate setting, but only a presentation of those contributions
which have been adapted for FDA. In Section 3 we will go through the functional setting. The
rich production and the variability in types of models, variables included and tools, led us to make
distinction between four types of methodologies. Subsections 3.1, 3.2 and 3.3 are dedicated to scalar
response models which are most often studied in the literature: firstly we will study selection of scalar
variables in models which contain some functional predictor, secondly we will deal with the selection
of scalar variables derived from the discretization of a functional object, and thirdly with the selection
of functional objects. Finally, Subsection 3.4 concerns the regression models with functional response.
To conclude the paper, in Section 4 we will present some ideas about how variable selection could
behave in functional regression in the next following years.
To make simpler the exposition of all the methodologies, we will assume without loss of generality
that the involved variables (functional or not) are centred to have zero mean. In the same way, we
will not show the assumptions (neither on the random errors nor the covariates) used in the different
methodologies (note that such assumptions could change from one methodology to other one).

2 Variable selection in finite-dimensional regression models


In the finite-dimensional setting, there is an extensive literature in variable selection tools (see, for
instance, Fan and Lv (2010) or Desboulets (2018) for recent reviews). In Section 2, we will present and
briefly comment some of these techniques, paying main attention to those that have been extended
to the functional framework. In that way, we could refer to them in the next section dedicated to
functional models.
The production in variable selection procedures for finite-dimensional regression started with naive
ideas such as stepwise regression (backward (Efroymson (1960)), forward (Weisberg (1980)) or both),

2
forward-stagewise regression or best subset regression (Furnival and Wilson (1974)). However, these
methods are computationally intensive, unstable (see Breiman (1996) or Fan and Li (2001)) and it is
hard to derive sampling properties. They are “discrete procedures” (variables are either selected or
discarded), so they often exhibit high variance, and therefore, in some cases they do not reduce the
prediction error of the full model.
For that, other techniques appeared, like shrinkage methods, also known as regularization, penalty-
based or penalized methods. Shrinkage procedures are more continuous, and do not suffer as much
from high variability (see Hastie et al. (2009)). Most of these procedures attempt to select variables
automatically and simultaneously (a notorious exception is bridge regression for Lq norms with q >
1; see Frank and Friedman (1993) and Fan and Li (2001)). These methods are based on adding a
penalization term in the estimation task which, under suitable conditions, generates a sparse solution,
in the sense that some estimated coefficients are zero. Penalized methods are highly developed,
specially in the case of linear modelling. Specifically, the well-known linear model is given by the
P
expression Yi = pj=1 βj Xij + εi , i ∈ {1, . . . , n}, where Yi is a scalar response, X i = (Xi1 , . . . , Xip )⊤
is a vector of scalar covariates, β = (β1 , . . . , βp )⊤ is a vector of unknown real coefficients and εi is the
random error. The penalized estimator of the vector of unknown parameters, is the solution of the
optimization problem
β = arg min (ℓ(β
β̂ β ) + nPλ (β
β )) , (1)
β ∈R p

where ℓ(·) is a real-valued function which depends on the model and on its estimation procedure; if
the estimation is made though penalized least squares, then ℓ(β Y − X β )⊤ (Y
β ) = (Y Y − X β ), where Y =
(Y1 , . . . , Yn )⊤ and X = (X
X 1 , . . . , X n )⊤ . Pλ (·) is a penalty function which depends on a regularization
parameter λ > 0. The parameter λ controls the amount of penalty and, in the case that (1) gives rise
to sparse solutions, also controls the sparseness of the resulting vector (as noted above, not all the
estimators verifying (1) give rise to sparse solutions).
The penalty function employed has a big influence to the properties of the derived estimator (see
Fan and Li (2001)). In the literature there are several proposals for this penalization term, but we are
going to comment briefly the ones most used in the functional setting. Among penalty functions, the
majority of them are based on norms. Probably the most famous shrinkage method, based on norms,
was proposed in Tibshirani (1996), where L1 penalty was used:
p
X
β) = λ
Pλ (β |βj | . (2)
j=1

He gave the name least absolute shrinkage and selection operator (LASSO) method to the combination
of this penalty with the least squares procedure. However, several objections emerged about this
penalty. On the one hand, LASSO estimators do not satisfy oracle properties (see Fan and Li (2001)).
On the other hand, Meinshausen and Bühlmann (2006) showed that in LASSO the optimal λ for
prediction gives inconsistent variable selection results. This problem was also found by Leng et al.
(2006), who, in particular, showed that for any sample size n, when there are non relevant variables in
the model and the design matrix is orthogonal, the probability that LASSO correctly identify the true
set of important variables is less than a constant (not depending on n) smaller than one. For that,
other penalties were studied. Zou (2006) proposed adaptive LASSO (adaLASSO), where the penalty

3
term has the form
p
X
β) = λ
Pλ (β wj |βj |, (3)
j=1

and wj , j ∈ {1, . . . , p}, are known weights. They showed that if the weights are data-dependent and
cleverly chosen, then the adaptive LASSO estimators can have the oracle properties. Another famous
proposal is the elastic-net penalty (see Zou and Hastie (2005)) which is a compromise between L1
P P
and L2 penalties: Pλ (β β ) = λ2 pj=1 |βj | + λ1 pj=1 βj2 . In a general way, Huang et al. (2008) studied
P
bridge penalties Pλ (β β ) = λ pj=1 |βj |q (related with the Lq norm) and showed that they verify the
oracle property for 0 < q < 1. In that paper they consider the more general context in which the
number of covariates, say pn , may increase to infinity with n (pn → ∞ as n → ∞). In addition, a
robust approach was studied in Wang et al. (2007a), where instead of least squares estimation, they
P P
used least absolute deviation (LAD) with ℓ(β β ) = ni=1 |Yi − pj=1 βj Xij | combined with L1 penalty
(LAD-LASSO).
Probably the main competitor of penalties based on norms is the proposal in Fan and Li (2001):
the smoothly clipped absolute deviation penalty (SCAD) defined, for a > 2, as


 λ |βj | , |βj | < λ,


X p  2 2
 (a − 1)λ − (|βj | − aλ) 2
β) =
Pλ (β Pλ (βj ), Pλ (βj ) = , λ ≤ |βj | < aλ, (4)
 2(a − 1)
j=1 

 2
 (a + 1)λ ,

|βj | ≥ aλ
2
(Fan and Li (2001) suggested to take a = 3.7). SCAD penalty improves properties of L1 penalty,
satisfying the oracle property. For that, it was very often used in works related with generalized linear
models (GLM), in which was assumed that Yi is a real variable verifying E(Yi |X X i ) = g−1 (ηi ) with
ηi = X ⊤i β (i ∈ {1, . . . , n}) and where g(·) is a known injective continuous link function. Fan and Li
(2001) studied GLM and proposed obtaining a penalized log-likelihood estimator using SCAD. That
is, the estimator derived from (1) when ℓ(β β ) denotes the conditional log-likelihood of Yi and Pλ (β β)
is the SCAD (4). They studied properties of this estimator for fixed number of covariates p, while
Fan and Peng (2004) studied them when the number of covariates p = pn diverges (pn → ∞ as
n → ∞).
The extension of shrinkage methods to the context of grouped models (or multifactor analysis-of-
variance (ANOVA) models; see Yuan and Lin (2006)) follows ideas that will be used later in functional
variable selection. For that, we are going to include them in this brief revision. In these models each
explanatory factor is represented by a group of derived input variables. Specifically, the grouped linear
model is given by the relationship
M
X
Yi = X⊤
imβ m + εi , i ∈ {1, . . . , n}, (5)
m=1

where regressors are divided into M groups, so X im = (Xim1 , . . . , Ximvm )⊤ (the case v1 = · · · = vM = 1
gives standard linear regression) and β m = (βm1 , . . . , βmvm )⊤ , m ∈ {1, . . . , M }. Therefore, in this
case, the interest is not in selecting variables individually, but in choosing important factors and each

4
one is in correspondence with a group of covariates. Therefore, the following optimization problem
should be solved:

β = arg
β̂ v
min v (ℓ∗ (β β ∗ )) ,
β ∗ ) + nPλ∗ (β (6)
∗ β ∈R 1 ×···×R M

where β∗ β 1 , . . . , β M ) and
= (β Pλ∗ (·) : Rv1
× · · · × RvM → R denotes the penalty function. Note
P  P 2
that to obtain a penalized least squares estimator ℓ∗ (ββ ∗ ) = ni=1 Yi − M ⊤
m=1 X miβ m should be
considered. The question now is how to choose the penalty function for selecting groups of covariates.
Yuan and Lin (2006) proposed the group LASSO penalty defined as
M q
X
β ∗) = λ
Pλ∗ (β β⊤
m Kmβ m , (7)
m=1

where Km is a positive definite matrix, m ∈ {1 . . . , M }. Penalty (7) is intermediate between the L1


and the L2 penalties. A derived problem is the selection of the matrices Km ; Yuan and Lin (2006)
used Km = vm Ivm with m ∈ {1 . . . , M } where Ivm is the identity matrix of size vm . The adaptation
of the LASSO gave the way to other extensions, like the group SCAD penalty
M vm
!
X X
∗ ∗ 2
β )=λ
Pλ (β Pλ βmr , (8)
m=1 r=1

where Pλ (·) was defined in (4). This penalty was proposed in Wang et al. (2007b) in the context of
the varying coefficients models with functional response, that we will discuss later. These authors
also proved oracle properties for this penalty. Another general proposal suggests to use composite
absolute penalties (CAP) studied in Zhao et al. (2009). CAP depend on a vector of norm parameters,
(γ0 , γ1 , . . . , γM ); these penalties are given by the expression
M
X
Pλ∗ (β
β ∗) = λ β m ||γm |γ0 ,
|||β (9)
m=1

where || · ||γ denotes the Lγ norm. The parameter γ0 determines how groups relate to each other while
γm dictates the relationship of the coefficients within group m. Therefore, this family of penalties
allows grouped selection and the hierarchical variable selection is reached by defining groups with
particular overlapping patterns.
So far we have studied penalized methods for linear models. However, these procedures can be
employed even in nonlinear regression. An interesting case for the relations with functional regression
P
is additive model given by the expression Yi = M m=1 fm (Xmi ) + εi , i ∈ {1, . . . , n}, where fm (·) with
m ∈ {1, . . . , M } are smooth univariate real-valued functions which should be estimated. In this case
the optimization problem is carried out in a space of functions, say F, since the target functions are
the solution of

f̂f = arg min (ℓ∗∗ (ff ∗ ) + nPλ∗∗ (ff ∗ )) , (10)
∗ f ∈F

where f∗ = (f1 , . . . , fM )⊤
and Pλ∗∗ (·)
: F → R is a penalization term. For these models, Meier et al.
Pn P 2
(2009) proposed ℓ∗ (ff ∗ ) = i=1 Yi − M m=1 f m (Xmi ) and, as penalty function, the sparsity-smoothness
penalty that simultaneously controls smoothing of functions fm (·) and sparseness,
v
M u n Z
X u1 X
∗∗ ∗
Pλ (ff ) = Pλ1 ,λ2 (fm ), Pλ1 ,λ2 (fm ) = λ1 t fm (Xim )2 + λ2 (fm ′′ (x)dx)2 . (11)
n
m=1 i=1

5
Two tuning parameters λ1 and λ2 control the amount of penalization: λ1 is a sparseness/tuning
parameter and λ2 is a smoothing/tuning parameter, since the second term in Pλ1 ,λ2 (·) controls the
smoothness of functions fm (·) with m ∈ {1, . . . , M }. To solve the optimization problem (10) in
practice, Meier et al. (2009) use cubic B-spline basis expansion of functions fm (·), that is, fm (x) =
PV ⊤
r=1 βmr bmr (x), where bmr (·) are B-spline basis functions and β m = (βm1 , . . . , βmV ) is the parameter
vector of fm (·). In this way, Meier et al. (2009) reduced the optimization problem (10) to (6), with
vm = V , since number of functions in the B-spline basis is independent from m, and penalty (11)
adopts the form of the group LASSO penalty (7). Huang et al. (2010) also studied an adaptive group
LASSO procedure for additive modelling.
Although we have focused the exposition on shrinkage methods, other different procedures have
been proposed in the literature to select relevant variables. In the context of linear modelling
Efron et al. (2004) proposed a Least Angle Regression (LARS) algorithm, a refined version of the
forward stagewise procedure that uses a simple mathematical formula to accelerate the computations.
This method is computationally efficient and it has LASSO (LARS-LASSO) and forward stagewise
methods as variants. A different idea is the Dantzig selector proposed in Candès and Tao (2007),
based on linear programming, which is able to deal with the case p ≫ n (that is, p is much larger
than n). Another important contribution was the sure independence screening procedure proposed
in Fan and Lv (2008), based on correlations. The enumeration of methods could go on; see, for in-
stance, Li et al. (2012) for a distance correlation method, Ke et al. (2014) for sparse models where
signals are both rare and weak, Mielniczuk and Teisseyre (2014) for a random subspace method and
O’Hara and Sillanpää (2009) for a review of Bayesian approaches.

3 Variable selection in functional regression models


We have presented variable selection methods in the finite-dimensional context. Here we are going to
study their extension to the infinite-dimensional setting and Section 3 is the main part of our paper.
Because variable selection may occur from various points of view, we will divide the exposition into four
subsections. The first three subsections are dealing with the scalar response: in Section 3.1 we are going
to revise works dealing with scalar variable selection when models also contain functional predictors;
in Section 3.2 we will study variable selection of scalar covariates originated from the discretization of
a curve; in Section 3.3 we are going to deal with functional covariate selection. Finally, in Section 3.4
we are going to treat models with functional response.

3.1 Selection of scalar covariates


As commented in the introduction, the combination of scalar and functional predictors in applications
becomes a frequent question in many applied sciences problems. One has situations where, in addition
to a very large number of covariates, pn , there is also some functional predictor involved. Aneiros et al.
(2015) dealt with this reality in the case of a scalar response, proposing a sparse partial linear model
with functional covariate, which allows pn → ∞ as n → ∞. The model that they studied is given by
the expression
Yi = X ⊤
i β + m(ζi ) + εi , i ∈ {1, . . . , n}, (12)

6
where Yi is the scalar response, X i = (Xi1 . . . , Xipn )⊤ are real random covariates, ζi = ζi (t) is the
functional random covariate valued in a semi-metric space and β = (β1 , . . . , βpn )⊤ ∈ Rpn is the vector
of unknown parameters, m(·) is the nonlinear unknown link operator and εi is the random error. The
strategy that they propose is to carry out variable selection in the linear component by transforming
model (12) into a linear one. For that, the effect of the functional covariate should be extracted from
the response and the other scalar predictors. That is, one should consider the model
pn
X
Yi − E(Yi |ζi ) = βj (Xij − E(Xij |ζi )) + εi , i ∈ {1, . . . , n}. (13)
j=1

For estimating the conditional expectations E(·|ζi ) in the expression (13), functional nonparametric
regression can be employed (see Ferraty and Vieu (2006)). Once the model is transformed (in an
approximate way) into a linear one, penalized estimation (1) can be applied. In Ferraty and Vieu
(2006) the SCAD penalty (4) was used.
In the same context Novo et al. (2021a) assumed semiparametric effect for the functional predictor.
That is, m(ζi ) = g(hθ, ζi i), where ζi belongs to a separable Hilbert space with inner product denoted
by h·, ·i, θ = θ(t) is an unknown functional parameter and g(·) is a real-valued smooth function
to estimate. In this case, for the transformation into a linear model, conditional expectations in
expression in (13) were estimated using functional single-index regression (see Ait-Saı̈di et al. (2008)).
Penalized least squares estimation with SCAD penalty (4) were applied to the resulting model.

3.2 Selection of scalar covariates with functional origin


Infinite-dimensionality of functional objects has constantly been of concern in the FDA literature.
When FDA was still not developed, predictive modelling in applied areas consisted in considering the
discretized functional object X (t), that is, scalar variables with functional origin, X (t1 ), . . . , X (tp ).
Then, existing techniques were applied (like principal components regression or partial least squares)
to reduce the dimension (see Frank and Friedman (1993) for a review). Since FDA emerged, other
techniques were proposed in order to reduce dimensionality of the functional predictor, but taking into
account its continuous nature. The concept of “sparseness” in functional regression was usually not
assumed with respect to the coefficients of the model. The common practise was to rewrite the model
using a “sparse” expansion of X (t) (see, for instance, Ramsay and Silverman (2005)). At this stage it
is worth to stress that the word “sparse” is used in FDA for different purposes (see Aneiros and Vieu
(2016) for a discussion): here we are meaning sparsity in the model and not for the curve data itself.
However, in some recent publications the interpretability of the results led authors back to con-
sider discretized functional objects. In addition, they realized that discretized values of the curves
X (t1 ), . . . , X (tp ) may contain information which is not reachable through the continuous curve X (t),
and conversely. Then different modelling options emerge (such as McKeague and Sen (2010)), many
of them combined with the sparse concept in finite-dimensional regression as we will see in this section.
In this case, new proposed procedures for variable selection are designed to deal with the very strong
dependency between resulting scalar variables (taking into account the continuous origin) and with
the very-high-dimension of the resulting vector. In addition, these sparse ideas are combined with
either parametric, nonparametric or semiparametric regression modelling.

7
In Ferraty et al. (2010), authors follow a nonparametric approach. Specifically, suppose that Yi is
a scalar response variable and Xi (t) is a functional random predictor with t ∈ I and I is a compact
subset of the real line. The functional nonparametric model (FNM) is given by the expression:

Yi = m(Xi ) + εi , i ∈ {1, . . . , n}, (14)

where m(·) is a smooth functional and εi denotes the random error (for details on this model, see
Ferraty and Vieu (2006)). In Ferraty et al. (2010), they studied how to select the most predictive
design points of the curve Xi (t), say t1 , . . . , ts , using a procedure based on local linear regression
(properties and references about local linear regression can be found in Fan and Gijbels (1996)). For
that, they consider the discretized Xi (t) and transform the functional model (14) into the underlying
multivariate nonparametric model:

Yi = g(Xi (t1 ), . . . , Xi (tp )) + εi , i ∈ {1, . . . , n}, (15)

where g(·) : Rp → R. Then, they transform the estimation of the most predictive design points into
a multivariate function estimation problem. For that, they propose a two stages algorithm based on
the cross-validation (CV) function:
n
1X
cv(tt∗ , h) = gh,−i (Xi (tt∗ )))v(Xi (tt∗ )), i ∈ {1, . . . , n},
(Yi − b (16)
n
i=1

where b gh,−i (·) is the leave-one-out local linear estimator of g(·), t ∗ is a vector of design points and
Xi (tt∗ ) denotes the discretized values of the functional object at these design points, h is a vector of
smoothing parameters and v(·) is a nonnegative, integrable function of s variables, which allows cases
of marked heterocedasticity (under homocedasticy, v(Xi (tt∗ )) is set to 1).

• In the first stage, called forward addition, the algorithm adds the most predictive design points
(in correspondence with a criterion based on (16)) step by step, while the addition of such
points diminishes the value of the PCV function (a penalized version of the CV function, which
penalizes the number of selected points).

• In the second stage, named backward deletion, the algorithm deletes the least predictive points
(in correspondence with a criterion based on (16)) step by step, while the elimination of such
points diminishes the value of the PCV. Note that the second step allows to enlarge the number
of possible combinations and therefore to find lower values of the cross-validation criterion.

As exposed in Ferraty et al. (2010), the algorithm could be combined with variable selection penalty
methods (like LASSO or LARS) to preselect design points and accelerate calculations. In addition, a
boosting step could be added to improve the predictive performance. In fact, authors conclude that
the incorporation of both discrete and continuous aspects of functional predictors could benefit the
prediction power.
In Kneip and Sarda (2011), authors follow a parametric approach. They also work with a scalar
response variable Yi and a discretized functional covariate Xi (t). In this case they consider the linear
model
Xp
Yi = βj Xi (tj ) + εi , i ∈ {1, . . . , n}, (17)
j=1

8
where X i = (Xi (t1 ), . . . , Xi (tp ))⊤ is the discretized functional predictor, β = (β1 , . . . , βp )⊤ is the vector
of unknown parameters and εi is the random error. They studied variable selection in a linear factor
model by assuming that the predictor X i can be decomposed into a sum of two uncorrelated random
components in Rp ,
X i = W i + Z i , i ∈ {1, . . . , n},

where W i is intended to describe high correlations of the Xij = Xi (tj ) while the components Zij of Z i ,
j ∈ {1, . . . , p}, are uncorrelated. Kneip and Sarda (2011) assume that the components Wij = Wi (tj )
of W i as well as Zij represent nonnegligible parts of the variance of Xij (common variability and
specific variability, respectively). Taking the decomposition above into account, model (17) can be
expressed as
Xp p
X
∗∗
Yi = βj Wi (tj ) + βj Xi (tj ) + εi , i ∈ {1, . . . , n}, (18)
j=1 j=1
Pp P
where βj∗∗ = βj∗ − βj and Yi = j=1 βj∗ Wi (tj ) + pj=1 βj Zij + εi , i ∈ {1, . . . , n}.
In order to estimate and select relevant variables in model (18), Kneip and Sarda (2011) use
different techniques to delete dependency between variables.

1. On the one hand, variables Wi (tj ), j ∈ {1, . . . , p}, are heavily correlated. However, the term
Pp ∗∗
j=1 βj Wi (tj ) could represent an important, common effect of all covariates. For avoiding the
effect of the dependence, Kneip and Sarda (2011) propose to employ principal components to
P
rewrite W i = pj=1 ψ ⊤ j W iψ j . Then, they assume that the effect of W i can be described with a
suitable small number of components, say d (d ≤ p). In that way, model (18) takes the form
d
X p
X
Yi = ψ⊤
αr (ψ r W i) + βj Xi (tj ) + εi , i ∈ {1, . . . , n}, (19)
r=1 j=1
Pp ∗∗
where αr = j=1 βj ψjr .

2. Before estimating and carrying out variable selection in model (19), the dependence between
ψ⊤r W i and Xi (tj ) should be deleted. For that, Kneip and Sarda (2011) used a projected model.
They consider as predictor, instead of X i , the projection of X i onto the orthogonal space of the
space spanned by the eigenvectors corresponding to the k largest eigenvalues of the covariance
matrix of X i . Then, a variable selection procedure (such as LASSO or the Dantzig selector) is
applied to the resulting model to select relevant variables simultaneously in both components.

Kneip et al. (2016) study a slightly different approach. They consider a generalization of the
classical functional linear regression model assuming that there exists an unknown number of “points
of impact”, that is, discrete observation times, where the corresponding functional values possess some
significant influences on the response variable. Specifically, they consider the model
Z s
X
Yi = α(t)Xi (t)dt + βk Xi (tk ) + εi , i ∈ {1, . . . , n}, (20)
I k=1

where Xi (t) is a curve with domain in the interval I and α(t) is a function parameter, while t1 , . . . , ts ∈
I are the points of impact where the curve has influence on the response. For estimating the model,

9
they should identify which discretized times of Xi (t) enter into the second component of the model,
a topic related with variable selection of scalar covariates with functional origin. For estimating the
number and location of impact points, they extract local variations from the functional covariate (to
diminish correlations between Xi (tk ), k ∈ {1, . . . , s}), that is, define Z(X , t) = X (t) − (1/2)(X (t −
δ) + X (t + δ)) for δ > 0 with [t − δ, t + δ] ∈ I and choose as impact points those time points where
there is a special high correlation between the response and Z(Xi , t).
A different idea for selection of impact points in (17) was proposed in Berrendero et al. (2019).
They assume that in the associated functional linear model, the function parameter, say α(t), belongs
to a Reproducing Kernel Hilbert Space (RKHS) instead of the more usual L2 space. Using the
properties derived from such an assumption, they define an optimality criterion for selecting impact
points which only depends on the covariance function of X (t) at each pair of time points and on the
covariance between Xi (tj ) and Yi , j ∈ {1, . . . p}, i ∈ {1, . . . n}. Based on the optimality criterion, they
introduce a recursive expression that is used to carry out the selection.
Aneiros and Vieu (2014) follow a different approach for dealing with both dependency in the
discretized curve and sparse linear modelling. In fact, their idea is to build a specific method for the
case in which scalar covariates have continuous origin. They work with the linear model (17)
where Yi is a scalar response and assume that Xi (t) is a random curve observed at the grid
a ≤ t1 ≤ · · · ≤ tpn ≤ b, βj with j ∈ {1, . . . , pn } are the unknown coefficients and εi the random error.
Note that in this case p = pn , that is, it is allowed that the discretization size tends to infinity with the
sample size (pn → ∞ as n → ∞). Therefore, for selecting relevant variables and estimating model (17)
they propose the so-called partitioning variable selection (PVS) procedure. This two-stage algorithm
relies on the idea that the values X (tj ) and X (tk ) with tj and tk very close will contain very similar
information of the response.

• In the first stage, a reduced linear model is considered, with only very few covariates, say wn ,
cover the entire discretization interval for Xi (t). That is, assuming without loss of generality
that pn = qn wn , the wn variables taken into account are Xi (t(2k−1)qn /2 ), k ∈ {1, . . . , wn }. The
rest of the pn variables are directly discarded. Then, a standard variable selection procedure is
applied to this reduced model, such as penalized least squares (1) with L1 penalty (2) or SCAD
penalty (4). In this way, dependence between covariates is reduced before the application of the
procedure for variable selection.

• In the second stage, a linear model is built when considering the selected variables in the first
step and those in their neighbourhood. That is, if Sb1 = {k ∈ {1, . . . , wn }, βbk 6= 0}, the following
set of variables is considered in the second step: ∪k∈Sb1 {Xi (t(k−1)qn +1 ), . . . , Xi (tkqn )}. In this way,
relevant information which was missed at the first step is taken into account. After that, the
same standard variable selection procedure is applied again to this resulting model.

The algorithm requires a division of the sample to be carried out in the two stages. The natural choice
is to use half of the sample in the first step and the other half in the second step (in some applications
that cannot be the optimal option).
The PVS idea has the advantage of being able to extended it to more complex models with a linear
component. In Aneiros and Vieu (2015), the PVS procedure was extended to the bi-functional partial

10
linear model, which is defined as
pn
X
Yi = βj Xi (tj ) + m(ζi ) + εi , i ∈ {1, . . . , n}, (21)
j=1

where ζ denotes a random variable valued on some semimetric space, and m(·) is an unknown smooth
functional (the other notation in model (17) remains). The idea to apply the PVS procedure to select
relevant variables in the linear component of the model (21), is to transform it into a linear model as
in (13). Then, the estimation of coefficients of the resulting linear model can be obtained by the PVS
procedure combined with penalized least squares (1) with SCAD penalty (4).
One of the main advantages of model (21) is that it allows the inclusion of both, pointwise and
continuous effects of functional predictors (which was found as profitable in Ferraty et al. (2010) or
in Kneip et al. (2016)). However, the presence of the nonparametric component could bring inter-
pretability and dimensionality problems in some applications. For that, in Novo et al. (2021b), the
PVS procedure was extended to a complete semiparametric model, which replaces the nonparametric
component of the model (21) by a functional single-index structure m(ζi ) = g(hθ, ζi i), where ζi belongs
to some separable Hilbert space with inner product h·, ·i, θ is an unknown functional parameter and
g(·) is a real-valued smooth function to estimate.
The problem of variable selection in this model has two additional difficulties in comparison with
models (17) and (21): the estimation of the functional parameter θ is computationally expensive and
needs a relatively big sample size. Therefore, in addition to the PVS procedure, Novo et al. (2021b)
studied the behaviour of the method that uses only the first step of the PVS procedure and obtained
good results (from a theoretical and practical point of view). The reason behind this proposal is
to reduce computational cost in situations of very large pn and improve the behaviour of the PVS
procedure in situations of small sample size (with only one stage the division of the sample is not
needed).
Up to now, the PVS procedure was applied to models where the discretized functional objects have
linear effect in the response. But it can be applied also to select variables in sparse nonparametric
functional modelling. Aneiros and Vieu (2016) studied its application in the model given by the
expression
pn
X
Yi = fj (Xi (tj )) + εi , i ∈ {1, . . . , n},
j=1

where fj (·) are unknown smooth real-valued functions and εi is the random error. To estimate models
in both stages of the PVS procedure and simultaneously select relevant variables in them, authors
use a pilot multivariate additive model procedure for variable selection, such as the one proposed in
Huang et al. (2010), based on approximation of the additive components by truncated series expan-
sions with B-splines bases, and then apply adaptive group LASSO.

3.3 Selection of functional covariates


In models with scalar response, Yi , i ∈ {1, . . . , n}, and under linear relationship between response
and functional predictors, the selection of functional variables requires dealing with the functional
nature of the predictors, say (Xi1 (t), . . . , XiM (t)) and its corresponding coefficient functions, say

11
(α1 (t), . . . , αM (t)). In the case of non-relevant variables, the corresponding coefficient function should
be estimated as constant 0 for all t in its domain I.
The application of shrinkage methods for selecting relevant functional covariates involves opti-
mization in a functional space, so the problem can not be directly addressed. Testing procedures for
variable selection also require a reduction of the dimension of functional predictors. For dealing with
this inconvenience, some authors follow a group modelling strategy, in advance for sake of brevity,
GM strategy, which basically consists in transforming the given model into a grouped linear model
(5). Specifically, this includes the following steps.

1. Firstly, it is assumed that functional predictors and its corresponding coefficient functions be-
long to a separable Hilbert space, so they can be expressed via countable orthonormal basis.
Therefore, we can obtain basis expansions of the functional objects. These basis expansions
can be truncated in order to contain a finite number of basis functions, and still offer a good
approximation of the functional object. The truncation parameter will be denoted as um , since
it will depend on each functional predictor. To sum up, the functional elements of the model
can be expanded in the following way,

X um
X
Xim (t) = Wimr φmr (t) ≈ Wimr φmr (t) = W ⊤
imφ m (t),
r=1 r=1
∞ u (22)
X X m

αm (t) = βmr φmr (t) ≈ βmr φmr (t) = β⊤


mφ m (t), i ∈ {1, . . . , n}, m ∈ {1, . . . , M }
r=1 r=1

where W im = (Wim1 , . . . , Wimum )⊤ , β m = (βm1 , . . . , βmum )⊤ and φ m (t) = (φm1 (t), . . . , φmum (t))⊤
(note that both W im and β m are vectors of coefficients, while φ m (t) are vectors of basis func-
tions). For building basis, authors use different approaches, some of them involve known basis
functions and others involve empirical basis functions. The first ones include Fourier, B-spline
and wavelet bases (see, Ramsay and Silverman (2005) for a general presentation) or Gaussian
radial bases (see Ando et al. (2008)); the last ones include bases based on functional principal
components (FPC) (see, for instance, Ramsay and Silverman (2005) or Hall et al. (2006)).

2. Therefore, each functional covariate Xm (t) is in correspondence with a finite set of coefficients
β m , m ∈ {1, . . . , M }, that should be treated together in order to select or discard a functional
predictor. Then, the model is transformed into a grouped linear model (5).

Combining this strategy in functional linear modelling with penalized methods, the optimization
problem is reduced to (6) and group penalties or sparsity-smoothing penalties can be applied. This
combination can be seen in several papers. For instance, Matsui and Konishi (2011) studied the
selection of functional covariates in a multiple functional linear model given by the expression
XM Z
Yi = αm (t)Xim (t)dt + εi , i ∈ {1, . . . , n}, (23)
m=1 I

where εi denotes the random error. Matsui and Konishi (2011) used Gaussian bases for constructing
basis expansions (22). After that, expression (23) is transformed:
M Z
X M
X
Yi = W⊤ mi φ m φ
(t)φ m (t)⊤
β m dt + εi = X⊤
imβ m + εi , i ∈ {1, . . . , n}, (24)
m=1 I m=1

12
R
where X im = (W W⊤ ⊤
im Jφm ) and Jφm = I φ m (t)φ φ m (t)⊤ , i ∈ {1, . . . , n}, m ∈ {1, . . . , M }. Therefore,
an optimization problem of type (6) is obtained. So using notation β ∗ = (β β 1 , . . . , β M )⊤ , they take
P P
β ∗ ) = ni=1 (Yi − M
ℓ∗ (β ⊤ 2
m=1 X imβ m ) and penalty function (8). Other examples using the GM strategy
and penalization methods for selecting relevant functional predictors in model (23) are Lian (2011)
(who studied selection of relevant variables as Matsui and Konishi (2011), but using FPC basis expan-
sions in (22)) and Huang et al. (2016), who studied robust estimation of model (23) using penalized
LAD method, combined with FPC expansions in (22) and CAP penalty (9) with γ0 = 1 and γm = ∞
for all m ∈ {1, . . . , M }. But the GM strategy was also used to select variables via penalization in more
complex models, such as generalized multiple functional linear models (Zhu and Cox (2009) studied
the selection of functional predictors in a model which also included scalar covariates, using a FPC
basis and group LASSO regularization (7); Gertheiss et al. (2013) used B-spline basis expansions and
sparsity-smoothness penalties (11)) or multiclass logistic regression for functional data (see Matsui
(2014) and Matsui (2019)).
Since datasets containing both functional and non-functional predictors are very frequent, Kong et al.
(2016) studied simultaneous variable selection of both functional (following the GM strategy) and
scalar covariates. Specifically, they worked with the model given by the expression
M Z
X

Z
Yi = i + γ αm (t)Xim (t)dt + εi , i ∈ {1, . . . , n}, (25)
m=1 I

where Z i = (Zi1 , . . . , Zip )⊤ is a vector of scalar covariates, γ = (γ1 , . . . , γp )⊤ is a vector of coefficients


and εi denotes the random error. They use FPC to obtain expansions (22). In order to obtain
sparse estimators of coefficients in both components of the model, the optimization problem that they
solved was a combination of (6), due to the functional predictors, and (1) due to the scalar predictors.
Specifically, using notation β ∗ = (β β 1 , . . . , β M )⊤ the optimization problem to solve is

β , γ̂γ ) = arg
(β̂ min (ℓ∗ (β
β ∗ , γ ) + nPλ∗ (β
β ∗ ) + nPλ (γγ )) , (26)
β ∗ ∈Ru1 ×···×RuM , γ ∈Rp
P P
β ∗ , γ ) = ni=1 (Yi − M
where ℓ∗ (β ⊤ ⊤
m=1 β mX m − Z i γ ), using notations in (24); Kong et al. (2016) use
as penalties group SCAD (8) and SCAD (4). Ma et al. (2019) extended the procedure in Kong et al.
(2016) to quantile regression for functional partially linear regression. In the model they studied,
the conditional quantile τ of the response variable follows expression (25) and the response variable
has linear relation with the conditional quantile. For selecting relevant functional and non-functional
variables simultaneously, Ma et al. (2019) follow the same techniques as Kong et al. (2016).
The idea of using various penalization terms can be extended to other contexts, for instance, when
we have two different groups of functional predictors. Feng et al. (2021) studied model (23) but added
a term for taking into account interaction effects between functional predictors. So their purpose
was to identify relevant main effects and corresponding interactions associated with the response
variable. For that, they carried out variable selection in both terms of the model (main effects and
interactions) separately. Firstly, they followed the GM strategy, obtaining basis expansions of the
functional predictors, coefficient functions corresponding with main effects and those corresponding
with interaction effects (they used FPC for the basis). Then, they applied least squares estimation
combined with two adaptive group lasso penalties (that is, penalties of form (7) but added weights in
the sum as in (3)), one for main effects and the other for interactions.

13
In the previously commented papers, it is assumed that there is linear relationship between the
response (or a known function of the response) and the functional predictors. However, this assumption
can be restrictive in some contexts and more accurate fits can often be produced by modelling a
nonlinear relationship. For that, Fan et al. (2015) went further using the GM strategy combined with
shrinkage methods: they studied a model with scalar response but nonlinear relationship with the
functional predictors. Specifically, they studied sparse functional additive regression (FAR), given by
the relationship
M
X
Yi = fm (Xim (t)) + εi , i ∈ {1, . . . , n}, (27)
m=1

where the fm (·) are general nonlinear functionals of Xim (t), m ∈ {1, . . . , M }. To fit model (27)
and simultaneously select relevant predictors, an optimization problem of type (10) should be solved.
P  P 2
Denoting f ∗ = (f1 , . . . , fM )⊤ , in this case ℓ∗∗ (ff ∗ ) = ni=1 Yi − M m=1 f m (X im (t)) , and instead of
∗∗ ∗
using a particular penalization Pλ (ff ), Fan et al. (2015) explore general penalizations of the form
v 
M u n
X u 1 X
Pλ∗∗ (ff ∗ ) = ρλ t fm (Xim )2  (28)
m=1
n
i=1

(that is, involving l2 -norm of the vectors of values of fm (·) evaluated in the sample), where ρλ (·)
is a concave function. However, given the functional nature of Xim (t), solving problem (10) requires
knowing the form of the functionals fm (·). For that, they specialize the methodology to the linear case,
using model (23) already widely studied, and to the nonlinear case, studying the multiple functional
single-index model taking
Z 
fm (Xmi (t)) = gm αm (t)Xim (t)dt , m ∈ {1, . . . , M }, (29)
I

where gm (t) are smooth nonparametric functions.


In this case, the usage of the GM strategy needs the expansions (22) (Fan et al. (2015) use or-
thonormal basis expansions, with dimension independent from the predictor um = u, but depending
on the sample size u = un ) and also requires obtaining basis expansions for functions gm (·), that is,
gm (x) ≈ h (x)⊤δ m where h (x) is a basis of dimension V = Vn . Then the optimization problem (10)
is transformed into (6) but in this case, the optimization is carried out by minimizing the objective
function ℓ∗ (β β ∗ , δ ∗ ) with respect to β ∗ = (β
β ∗ , δ ∗ ) + nPλ∗ (β β 1 , . . . , β M ) and δ ∗ = (δδ 1 , . . . , δ M )⊤ , where
!2 v
n M M u n
X X X u X
∗ ∗ ∗
ℓ (ββ ,δ ) = Yi − h (W ⊤ ⊤
W imβ m ) δ m ) ∗ ∗ ∗
and Pλ (ββ ,δ ) = t
ρλ ( (1/n) h (WW⊤ ⊤
imβ m ) δ m ).
i=1 m=1 m=1 i=1

So far we have presented procedures based on penalization techniques for selecting relevant func-
tional variables, but in the literature there exist other options for selecting relevant functional pre-
dictors, such as testing procedures. In fact, there is a connection between model testing and vari-
able selection: dropping a variable from the model is equivalent to not reject the null hypothesis
that its corresponding parameter is 0. In this case, the GM strategy can be applied in order to
reduce dimensionality of covariates and function-parameters and apply finite-dimensional techniques.
Collazos et al. (2016) worked with model (23) applying significance testing of the functional predictors

14
Xim (t), m ∈ {1, . . . , M }. For that, they firstly obtained basis expansions (22), and then formulated
the following test, for each m ∈ {1, . . . , M }:

H0 : β m = 0 , H1 : β m 6= 0 . (30)

The test can be solved via a likelihood ratio-test which compares the residual sum of squares (RSS)
obtained under the elimination of the predictor m from the model (H0 ), with the RSS obtained with
the complete model. Since the test (30) is performed for each functional predictor, the M p-values
obtained should be corrected using Bonferroni correction or the false discovery rate. Then, one selects
as influential predictors those covariates, the corrected p-value of which leads to the rejection of the
null hyphotesis in (30).
Other procedures for variable selection, not based on penalization, were presented in Smaga and Matsui
(2018). They worked with model (23), and obtained basis expansions (22). After that, they built two
algorithms based on the random subspace method proposed in Mielniczuk and Teisseyre (2014), which
explores subsets of variables and measure variable importance using t-statistics.
Another way of selecting relevant functional predictors without using a penalization term is the
Bayesian approach. Zhu et al. (2010) studied variable selection in a Bayesian functional hierarchical
model for classification, to deal with situations when functional predictors are contaminated by random
batch effects. This model uses expression (25) with latent response in order to classify observations
from a binary random variable. For the estimation procedure, they assume Gaussian processes for
the priors for the parameter-functions and introduce an hyperparameter in them that indicates if the
functional variable is selected or not. They used orthonormal basis expansions (22) (in particular, they
used FCA basis) to reduce dimensionality of functional objects and transformed the functional poste-
rior sampling problem into a multivariate one, and then applied a hybrid Metropolis–Hastings/Gibbs
sampler (see, for instance, George and McCulloch (1997)) to obtain posterior samples of the parame-
ters and estimate them (selecting relevant variables at the same time).
A different proposal, not involving penalized methods, can be found in Febrero-Bande et al. (2019).
They work in the context of general additive regression with scalar response, that is, they work with
model (27) but with the difference that predictors can be of different nature (functional, scalar, mul-
tivariate, directional, etc.); in addition, the effect of each predictor, fm (·), can be linear or nonlinear.
For selecting relevant variables in that model, they construct an algorithm based on distance cor-
relation R(·, ·) proposed in Székely et al. (2007), which only depends on distances among data. If
Xi denotes a predictor of any nature, R(Xi , Yi ) = 0 characterizes independence with the response.
Therefore, the algorithm starts with a null model and sequentially selects new variables to be incor-
porated into the model accordingly with R(·, ·). Specifically, the covariate, that provides the high
value for the distance correlation with the current residual, is chosen and, then, a test of independence
based on R(·, ·) is carried out: the variable is candidate to be included into the model only if the
null hypotheses of independence is rejected. Each candidate variable is added to the model fixing the
effects (linear/nonlinear) of the previously added variables and setting as the effect of the candidate
variable the one (linear/nonlinear) that gives rise to the best contribution; then, the model is checked
again to see if the addition of this covariate is relevant or not (by means of a generalized likelihood
ratio test). The procedure ends when no more variables can be added to the model because the set of
remaining candidates is empty or all the remaining variables accept the independence null hypothesis

15
of the distance correlation test.

3.4 The functional response case


It also could be the case that the functional variable is the one that we want to predict, and for
that we have a set of scalar covariates, but only few of them are really related with the response.
In this framework Wang et al. (2007b) proposed a model with functional response and time-varying
coefficients. Specifically, the model is given by the relationship
M
X
Yi (t) = αm (t)Xim + εi (t), i ∈ {1, . . . , n}, (31)
m=1

where Yi (t) is a functional variable with domain I, Xi1 , . . . , XiM are real covariates, αm (t) is a function-
parameter and εi (t) is a stochastic process corresponding to the random error. For estimating function-
parameters, they follow the GM strategy: they expand these coefficient functions using B-spline basis
(22) with um = u, and the same basis functions for all the coefficients; then they use penalized
P P
β ∗ ) = ni=1 Tj=1 (Yi (tj ) −
least squares (with the discretized response Yi (t1 ), . . . , Yi (tT ), that is, l∗ (β
PM P u 2
m=1 ( r=1 βmr φr (tj ))Xim ) ) and propose group SCAD penalty (8).
Mingotti et al. (2013) also studied model (31). In order to estimate coefficient functions they ap-
plied a penalized least squares procedure, obtaining previously B-spline expansions for the coefficient-
functions (22) and for the response variable using um = u and the same basis for all the parameter
P P R P
functions and for the response, that is, Yi (t) ≈ ur=1 ari φr (t) and l∗ (β β ∗ ) = ni=1 I ( ur=1 ari φr (t) −
PM P u 2
m=1 ( r=1 βmr φr (t))Xim ) dt. In this case, the penalty term proposed, called functional LASSO, is
given by the expression
M Z X
X u
∗ ∗
β
Pλ (β ) = nλ βmr φr (t) dt.
m=1 I r=1

Mingotti et al. (2013) got advantage of B-spline properties in computations of integrals (see de Boor
(2001)).
Hong and Lian (2011) generalize variable selection with LASSO penalty for multiple functional
linear model for the case in which both response variable and covariates are functional, but parameters
of the model are scalar.

4 Conclusions and future perspectives


As discussed along this review there exists a rich production in variable selection methods for functional
regression. Many different techniques have been developed, all of them incorporate in some sense
previous ideas from variable selection in mutivariate regression models (most of them, ideas related
to LASSO, despite the criticism received by this selector; it is expected that ideas based on other
selectors will be adapted in the near future to the functional case). This is an evidence of the needs
for bridging gaps between FDA and HDS, and also of the benefits one could get from such crossing of
the ideas. The functional production is still very small compared to the dramatically high number of
research papers related with this topic in finite-dimensional models, and undoubtly the next few years

16
should lead to many new advances in the domain, and the extension of techniques in the functional
setting is expected to continue.
As we have commented in the Introduction, the complexity of the data continues to increase day
by day. Müller (2016) talked about “next generation” functional data, and provided some speculative
notions about the challenges that the area has to overcome in the future (functional data irregularly
and sparsely observed, the mix between big data and functional data, the multivariate time domain in
biostatistical modelling, etc). In this context, the dimension and the increasing number of predictors
become a more serious problem. Undoubtly, there will be the necessity of building new models and,
consequently, new variable selection techniques to treat efficiently this new kind of functional objects.

Acknowledgments
The authors are confident on the fact that this volume will contribute (as JMVA did since 50 years)
to promote further researches in high (and infinite) dimensional statistics, and they wish to express
their sincere gratitude to Professor Dietrich von Rosen for having taken the initiative of this Jubilee
Volume and for having invited us in presenting a contribution.
This research was supported by MICINN grant PID2020-113578RB-I00 and by the Xunta de
Galicia (Grupos de Referencia Competitiva ED431C-2020-14 and Centro de Investigación del Sistema
Universitario de Galicia ED431G 2019/01), all of them through the ERDF. The second author also
thanks the financial support from the Xunta de Galicia and the European Union (European Social
Fund - ESF), the reference of which is ED481A- 2018/191.

References
Ahmed, S. E., editor (2017). Big and complex data analysis, Contributions to Statistics. Springer.

Ait-Saı̈di, A., Ferraty, F., Kassa, R., and Vieu, P. (2008). Cross-validated estimations in the single-
functional index model. Statistics, 42(6):475–494.

Ando, T., Konishi, S., and Imoto, S. (2008). Nonlinear regression modeling via regularized radial basis
function networks. Journal of Statistical Planning and Inference, 138(11):3616–3633.

Aneiros, G., Cao, R., Fraiman, R., Genest, C., and Vieu, P. (2019a). Recent advances in functional
data analysis and high-dimensional statistics. Journal of Multivariate Analysis, 170:3–9.

Aneiros, G., Cao, R., Fraiman, R., and Vieu, P. (2019b). Editorial for the special issue on functional
data analysis and related topics. Journal of Multivariate Analysis, 170:1–2.

Aneiros, G., Ferraty, F., and Vieu, P. (2015). Variable selection in partial linear regression with
functional covariate. Statistics, 49(6):1322–1347.

Aneiros, G., Horova, I., Huskova, M., and Vieu, P. (2022). On functional data analysis and related
topics. Journal of Multivariate Analysis, 189:104861.

Aneiros, G. and Vieu, P. (2014). Variable selection in infinite-dimensional problems. Statistics &
Probability Letters, 94:12–20.

17
Aneiros, G. and Vieu, P. (2015). Partial linear modelling with multi-functional covariates. Computa-
tional Statistics, 30(3):647–671.

Aneiros, G. and Vieu, P. (2016). Comments on: Probability enhanced effective dimension reduction
for classifying sparse functional data. TEST, 25:27–32.

Berrendero, J. R., Bueno-Larraz, B., and Cuevas, A. (2019). An RKHS model for variable selection
in functional linear regression. Journal of Multivariate Analysis, 170:25–45.

Bongiorno, E. G., Goia, A., Salinelli, E., and Vieu, P. (2014). An overview of IWFOS’2014. In
Contributions in Infinite-Dimensional Statistics and Related Topics, pages 1–5. Esculapio, Bologna.

Breiman, L. (1996). Heuristics of instability and stabilization in model selection. The Annals of
Statistics, 24(6):2350 – 2383.

Candès, E. and Tao, T. (2007). The Dantzig selector: statistical estimation when p is much larger
than n. Annals of Statistics, 35:2392–2404.

Collazos, J. A., Dias, R., and Zambom, A. Z. (2016). Consistent variable selection for functional
regression models. Journal of Multivariate Analysis, 146:63–71.

Cuevas, A. (2014). A partial overview of the theory of statistics with functional data. Journal of
Statistical Planning and Inference, 147:1–23.

de Boor, C. (2001). A Practical Guide to Splines. Applied Mathematical Sciences. Springer-Verlag,


New York.

Desboulets, L. D. D. (2018). A review on variable selection in regression analysis. Econometrics, 6(4).

Efron, B., Hastie, T., Johnstone, I., and Tibshirani, R. (2004). Least angle regression. Annals of
Statistics, 32:407–499.

Efroymson, M. A. (1960). Multiple regression analysis. In Ralston, A. and Wilf, H. S., editors,
Mathematical Methods for Digital Computers, New York. Wiley.

Fan, J. and Gijbels, I. (1996). Local Polynomial Modelling and its Applications. Monographs on
Statistics and Applied Probability 66. Routledge.

Fan, J. and Li, R. (2001). Variable selection via nonconcave penalized likelihood and its oracle
properties. Journal of the American Statistical Association, 96(456):1348–1360.

Fan, J. and Lv, J. (2008). Sure independence screening for ultrahigh dimensional feature space. Journal
of the Royal Statistical Society: Series B (Statistical Methodology), 70(5):849–911.

Fan, J. and Lv, J. (2010). A selective overview of variable selection in high dimensional feature space.
Statistica Sinica, 20(1):101–148.

Fan, J. and Peng, H. (2004). Nonconcave penalized likelihood with a diverging number of parameters.
Annals of Statistics, 32(3):928–961.

18
Fan, Y., James, G. M., and Radchenko, P. (2015). Functional additive regression. The Annals of
Statistics, 43(5):2296–2325.

Febrero-Bande, M., González-Manteiga, W., and de la Fuente, M. O. (2019). Variable selection in


functional additive regression models. Computational Statistics, 34:469–487.

Feng, S., Zhang, M., and Tong, T. (2021). Variable selection for functional linear models with strong
heredity constraint. Annals of the Institute of Statistical Mathematics.

Ferraty, F., Hall, P., and Vieu, P. (2010). Most-predictive design points for functional data predictors.
Biometrika, 97(4):807–824.

Ferraty, F. and Vieu, P. (2006). Nonparametric Functional Data Analysis, Theory and Practice.
Springer Series in Statistics. Springer-Verlag, New York.

Frank, I. E. and Friedman, J. H. (1993). A statistical view of some chemometrics regression tools.
Technometrics, 35(2):109–135.

Furnival, G. M. and Wilson, R. W. (1974). Regressions by leaps and bounds. Technometrics, 16(4):499–
511.

George, E. I. and McCulloch, R. E. (1997). Approaches for Bayesian variable selection. Statistica
Sinica, 7(2):339–373.

Gertheiss, J., Maity, A., and Staicu, A. M. (2013). Variable selection in generalized functional linear
models. Stat, 2(1):86–101.

Goia, A. and Vieu, P. (2016). An introduction to recent advances in high/infinite dimensional statistics.
Journal of Multivariate Analysis, 146:1–6.

Hall, P., Müller, H.-G., and Wang, J.-L. (2006). Properties of principal component methods for
functional and longitudinal data analysis. The Annals of Statistics, 34(3):1493–1517.

Hastie, T., Tibshirani, R., and Friedman, J. (2009). Linear methods for regression. In The Elements
of Statistical Learning, pages 43–99. Springer Series in Statistics, New York.

Hong, Z. and Lian, H. (2011). Inference of genetic networks from time course expression data us-
ing functional regression with lasso penalty. Communications in Statistics - Theory and Methods,
40(10):1768–1779.

Huang, J., Horowitz, J. L., and Ma, S. (2008). Asymptotic properties of bridge estimators in sparse
high-dimensional regression models. The Annals of Statistics, 36(2):587–613.

Huang, J., Horowitz, J. L., and Wei, F. (2010). Variable selection in nonparametric additive models.
The Annals of Statistics, 38(4):2282–2313.

Huang, L., Zhao, J., Wang, H., and Wang, S. (2016). Robust shrinkage estimation and selection
for functional multiple linear model through lad loss. Computational Statistics & Data Analysis,
103:384–400.

19
Ke, Z. T., Jin, J., and Fan, J. (2014). Covariate assisted screening and estimation. The Annals of
Statistics, 42(6):2202–2242.

Kneip, A., Poß, D., and Sarda, P. (2016). Functional linear regression with points of impact. The
Annals of Statistics, 44(1):1–30.

Kneip, A. and Sarda, P. (2011). Factor models and variable selection in high-dimensional regression
analysis. The Annals of Statistics, 39(5).

Kong, D., Xue, K., Yao, F., and Zhang, H. H. (2016). Partially functional linear regression in high
dimensions. Biometrika, 103(1):147–159.

Leng, C., Lin, Y., and Wahba, G. (2006). A note on the Lasso and related procedures in model
selection. Statistica Sinica, 16(4):1273–1284.

Li, R., Zhong, W., and Zhu, L. (2012). Feature screening via distance correlation learning. Journal
of the American Statistical Association, 107(499):1129–1139.

Lian, H. (2011). Shrinkage estimation and selection for multiple functional regression. Statistica
Sinica, (23):51–74.

Ma, H., Li, T., Zhu, H., and Zhu, Z. (2019). Quantile regression for functional partially linear model
in ultra-high dimensions. Computational Statistics & Data Analysis, 129:135–147.

Matsui, H. (2014). Variable and boundary selection for functional data via multiclass logistic regression
modeling. Computational Statistics & Data Analysis, 78:176–185.

Matsui, H. (2019). Sparse group lasso for multiclass functional logistic regression models. Communi-
cations in Statistics - Simulation and Computation, 48(6):1784–1797.

Matsui, H. and Konishi, S. (2011). Variable selection for functional regression models via the L1
regularization. Computational Statistics & Data Analysis, 55(12):3304–3310.

McKeague, I. W. and Sen, B. (2010). Fractals with point impact in functional linear regression. The
Annals of Statistics, 38(4):2559–2586.

Meier, L., van de Geer, S., and Bühlmann, P. (2009). High-dimensional additive modeling. The Annals
of Statistics, 37(6B):3779–3821.

Meinshausen, N. and Bühlmann, P. (2006). High-dimensional graphs and variable selection with the
Lasso. The Annals of Statistics, 34(3):1436 – 1462.

Mielniczuk, J. and Teisseyre, P. (2014). Using random subspace method for prediction and variable
importance assessment in linear regression. Computational Statistics & Data Analysis, 71:725–742.

Mingotti, N., Lillo Rodrı́guez, R. E., and Romo Urroz, J. (2013). Lasso variable selection in functional
regression. Statistics and Econometrics Series 13 Working paper 13–14, Universidad Carlos III de
Madrid.

20
Müller, H.-G. (2016). Peter Hall, functional data analysis and random objects. The Annals of Statistics,
44(5):1867–1887.

Novo, S., Aneiros, G., and Vieu, P. (2021a). Sparse semiparametric regression when predictors are
mixture of functional and high-dimensional variables. TEST, 30:481–504.

Novo, S., Vieu, P., and Aneiros, G. (2021b). Fast and efficient algorithms for sparse semiparametric
bi-functional regression. Australian and New Zealand Journal of Statistics, 63:606–638.

O’Hara, R. B. and Sillanpää, M. J. (2009). A review of Bayesian variable selection methods: what,
how and which. Bayesian Analysis, 4(1):85–117.

Ramsay, J. O. and Silverman, B. (2005). Functional Data Analysis. Springer Series in Statistics.
Springer-Verlag, New York, 2nd edition.

Sangalli, L. M. (2018). The role of statistics in the era of big data. Statistics & Probability Letters,
136:1–3.

Smaga, L. and Matsui, H. (2018). A note on variable selection in functional regression via random
subspace method. Statistical Methods & Applications, 27:455–477.

Székely, G. J., Rizzo, M. L., and Bakirov, N. K. (2007). Measuring and testing dependence by
correlation of distances. The Annals of Statistics, 35(6):2769–2794.

Tibshirani, R. (1996). Regression shrinkage and selection via the Lasso. Journal of the Royal Statistical
Society: Series B, 58:267–288.

Vieu, P. (2018). On dimension reduction models for functional data. Statistics & Probability Letters,
136:134–138.

Wang, H., Li, G., and Jiang, G. (2007a). Robust regression shrinkage and consistent variable selection
through the lad-lasso. Journal of Business & Economic Statistics, 25(3):347–355.

Wang, L., Chen, G., and Li, H. (2007b). Group SCAD regression analysis for microarray time course
gene expression data. Bioinformatics, 23(12):1486–1494.

Weisberg, S. (1980). Applied Linear Regression. Wiley, New York.

Yuan, M. and Lin, Y. (2006). Model selection and estimation in regression with grouped variables.
Journal of the Royal Statistical Society: Series B (Statistical Methodology), 68(1):49–67.

Zhao, P., Rocha, G., and Yu, B. (2009). The composite absolute penalties family for grouped and
hierarchical variable selection. The Annals of Statistics, 37(6A):3468–3497.

Zhu, H. and Cox, D. D. (2009). A functional generalized linear model with curve selection in cervical
pre-cancer diagnosis using fluorescence spectroscopy. In Rojo, J., editor, Optimality: The Third
Erich L. Lehmann Symposium, volume 57, pages 173–189.

Zhu, H., Vannucci, M., and Cox, D. D. (2010). A Bayesian hierarchical model for classification with
selection of functional predictors. Biometrics, 66(2):463–473.

21
Zou, H. (2006). The adaptive Lasso and its oracle properties. Journal of the American Statistical
Association, 101(476):1418–1429.

Zou, H. and Hastie, T. (2005). Regularization and variable selection via the elastic net. Journal of
the Royal Statistical Society: Series B (Statistical Methodology), 67(2):301–320.

22

You might also like