Chapter 1
Chapter 1
[Link]
November 2024
1
In particular, conjuncture forecasts do not respond to the same needs than
long term forecasts. They assist monetary and budgetary government policies,
and for the firms their decisions about prices, stocks and cash flows. Most other
organizations are also concerned about prediction and forecasts in their own
way.
Even basic prediction or forecast can be useful to a firm, if they yield better
results than its competitors.
Prediction seems to be the general term, while forecast rather refer to a
time framework. Cases in which predictions are used without referring to time
or to the future are frequent. For example, when part of a population is not
observed, the policy maker or the administrations must predict what are the
characteristics of the individuals that are not observed.
However, the performance of the typical economic forecast is poor, to say
the least. Some of the stated forecasts by experts and forecasting institutions
are often belied by the occurrence of contrary events. Some of the best time
series models for forecasting are sometimes beaten by a mere random walk.
Other indicators are better forecast such as population projections and foreign
trade deficits, while not always. Prediction errors and forecast errors are often
originated from emotions, with predicting agents hoping or fearing too much.
Blind extrapolations are also common sources of prediction errors.
Moreover, it is usually hard to draw clear conclusions about the economic
mechanisms at play from the published forecasts. Besides, this is consistent with
forecasts not needing to be related to causal relationships. As a consequence,
The notions of exogeneity and endogeneity of variables are somewhat blurred
in predictions. Endogenous variables, in terms of an explanatory models, can
often be used as valid predicting variables, while past exogenous variables are
also massively exploited.
Predictions and structural models are not necessarily related. There are
many examples where causal models do not provide the best forecasts. Some
forecasters even completely neglect economic theoretical bases for design forecast
methods. Still, having in mind an economic model should help listing predicting
factors and having ideas about how they interact. Then, one can search data
about these factors. In that case, it is also necessary to think about other factors
that are not included in the theory, but may be relevant.
Another issue, with the current explosion of data availability and comput-
erized power of calculus, is that traditional asymptotics in statistics does not
handle well the situations where they are more available variables than observa-
tions. Moreover, the usual goodness-of-fit indicators of statistics do not respond
well to evaluating the future. Standard errors of predictions and forecasts are
still useful for dealing with inaccuracies in parameter estimation and specifica-
tion search, but they do not help for other types of uncertainties occurring with
predictions. Besides, overfitting is a scourge of many prediction attempts. The
nature of the data used is also unstable, with macroeconomic indicators being
routinely revised over time, for example. In all these respect, more and different
statistical methods are needed when predicting. However, traditional statistics
must be well understood and mastered first.
2
An issue with macroeconomic forecasts, is the Lucas critique: The way the
economic system respond to changes and to policies evolve over time and is
affected by exogenous variables and policies themselves. This challenges the
idea that a stable model of agents’ behaviour could provide precise predictions.
More specifically, the publication of forecasts can itself influence people, and
made them believe that they will be true, and thus make them self-fulfilling
when people follow these results by supporting them, or self-defeating when
people fight them for example because of protection motives.
Finally, predictions and forecasts should be integrated in the decisions that
they are supposed to guide. This is rarely the case from an analytical point of
view. For example, forecasting can be used for many different decisions: policy
design and evaluation, tax and budget management, control of processes, scien-
tific research and model evaluation, risk and portfolio management, investment
and saving decisions, insurance decisions, planning of operations, etc.
However, any economic agent engaged in decision-making is better try to
know where it goes, that is to predict and forecast. Our challenge in this course is
to try to better understand predictions and forecast, and make progress on their
methods. Statistical decision theory and time series techniques are among the
tools that will assist us. Some ways of summarizing the data, with parsimonious
while flexible models should be developed. Some ways of investigating out-of-
the-sample situations should be used.
Rational decision-making often requires evaluation of risks and the careful
weighting of the potential positive and negative consequences of risky choices.
We may willingly embrace a certain level of risk based on our assessment of
specific contexts.
All calculations about the future involve inductive reasoning, but they do
not involve the future itself. These reasonings are only reasonings. For exam-
ple, strictly speaking, it remains a possibility that it might snow in Marseille
tomorrow, despite inductive reasonings based on past records suggesting other-
wise. Moreover, fundamentally all that we discuss about the uncertainty of the
future can be transposed to any situation of ignorance, not exclusively about
future uncertainties. So, predictions can deal with the future or with current
ignorance.
Uncertainty results from the fact that all the factors affecting an agent’s
actions cannot be determined or predicted. One can distinguish two cases:
(1) The uncertainty that is generated by interferences with the actions of
other persons, and
(2) The uncertainty that directly stems from random events, whether these
are natural or social events.
In contrast, regarding the uncertainty emerging from pure randomness, log-
ical tools that describe the random events must be constructed. When these
events can be associated with probabilities - often probabilities of states of na-
ture that also need be defined, we are able to specify probability spaces, which
are associated with a set of random events.
Another, technical, distinction of how uncertainty is modeled, is whether
it is probabilistic, as for financial risks; measurable, when random events can
3
be specified but not their probabilities; and more evanescent than this. This
distinction is somewhat related to the notions of predictive uncertainty that we
will dealt with, and unpredictive uncertainty that cannot be grasped precisely
but must be accounted for.
Among the error terms that partly account for uncertainty in econometric
models, some are motivated by diverse envisaged sources: innovation shocks,
measurement errors, optimization errors by agents, individual heterogeneity,
environment heterogeneity, efficiency error, changes in behavior, sampling error,
policy-generated randomness, etc.
In that case, Y and u are random n-vectors and X can be a fixed or random
n × k matrix. The model can also be written separately for each observation i:
and
E(ui uj |xi , xj ) = σ 2 δij
for all i, j = 1, ..., n,
where δij is the Kronecker index (δij = 1 if i = j, and 0 otherwise).
In all these expressions, the expectations are conditional on X because these
variables can be random. In general, there is a constant in the X ′ s.
These notations may also directly denote random variables for Y and u, and
a k-random vector for X that includes k distinct random variables. This is a
perspective before a sample has been drawn, and before the sample size is known.
It has the advantage of focussing on the fundamental notion of randomness
associated to a given distribution for each variable.
Sometimes, in these two cases, the variables in X can be considered as fixed
or chosen by the analyst, as is the case for some experiments. This view has the
4
advantage of making some calculations more concrete, for example the formula
of the OLS estimator β̂ = (X ′ X)−1 X ′ Y can be discussed as a mere algebraic
calculus based on some given observed numbers.
When distinguishing the observations of each individuals i, one can denote
Yi , Xi and ui , respectively.
With time series instead, the indices i and t can be substituted. However,
in that case, the sequentiality of the periods somehow changes the theoretical
perspective. First, serial correlations will play a central role in the econometric
analyses. Second, non-stationarity issues when they arise will invalidate many
methods used in cross sections, and new methods must be used. Third, since
forecasting is the objective here, on can systematically take advantage of the
available information about the past to predict the future, and this in a way
that account for the distinct time distance of the past observed periods to the
future.
Typical equations could read like:
(yi1 , ..., yiT , x1i , ..., xiT , ηi ), i = 1, ..., N . Then, one writes the model:
5
where vi = (vi1 , ..., viT )′ and xi = (xi1 , ..., xiT )′ , with vi et ηi unobserved.
This is called the strict exogeneity hypothesis. One says that xi = (xi1 , ..., xiT )′
are strictly exogenous conditionnally to the unobserved effects ηi .
6
(1) The probabilities of events in the sense of observable frequencies (about
our observations of the world, and extrapolations from these observations).
They can be called objective probabilities, notably because they generally do
not vary across individuals. For example, the number of rainy days within a
year is a frequency number that provides a measure of the probability of rain
over the year; Another example is the occurrence of the results of a dice, which
is 1/6 for each face and corresponds approximately to the observed frequencies
when throwing it enough times. Objective probabilities are the ones typically
used in time series models for forecasting.
(2) The probabilities of events in the sense of likelihood (in our judgements
and feelings). They can be called: subjective probability, notably because they
generally vary across individuals. Some degrees of likelihood may become mea-
surable by referring to a fictional story or to personal feelings. One can start
from prior subjective probabilities of events and use implicit or explicit calcu-
lus of probabilities to refine them. For example, by looking at the sky in the
morning, one may develop a feeling about the probability of rain in the after-
noon, then by seeing how the clouds gather overtime in the sky, one may adjust
this probability number intuitively. Subjective probabilities are often used by
experts for predictions when data is poor or missing.
Let s be a vector of possible values given to each of the uncertain elements in
the given context. These values are typically called the states of nature, random
states, or the states of the world. For example, if we are deciding to take an
umbrella with us or not when leaving home, rainy and sunny days can define
two distinct states of nature in which we are interested. Another example is,
for a given financial asset, the infinite set of possible monetary returns on the
next day. Let us denote Ω as the set of all possible states s, or universe of
eventualities, or of possibilities.
Let us now formally define the states of nature and the events. Often,
economic theories are specified in terms of the states of nature. However, the
policies or practical decisions rather refer to the set of events, E, that is a way to
summarize all the useful information for decision-making. These two notions are
related. For example, two events of interest for a financial asset are: its return
tomorrow is positive; and, respectively, negative. Therefore, moving from Ω
to E is not merely a question of information aggregation but also of modeling
choice. When predicting or forecasting, choosing what are the relevant variables,
and the relevant random events, looks like a preliminary stage of the reflexion.
For example, when forecasting GDP for a given year, using the past observed
values of GDP as a basis for the prediction is such a restrictive while inevitable
convention.
Definition:
Axioms for a set of random events E and a measurable space (Ω, E):
E ⊂ P(Ω ), Ω ∈ E;
E is stable by countable intersection ( ∩), countable union ( ∪) and comple-
menting ( ).
7
Ω is the set of random states, while E is denoted the σ-field of events.1
An uncertain or random event is defined as a subset, say E, of Ω, belonging
to a given σ-field E of Ω.
This is the harmonious definition of random events, allowing for their com-
binations, that makes the space ‘measurable’. A useful kind of measure are the
probability functions that will fit well the description of the events. We discuss
these functions below. We note the distinction between random states that be-
long to the universe Ω, for the complete set of events E, and random events
that belong to the set of subsets of Ω, i.e., to P(Ω ). That is, for an event E:
E ⊂ Ω and E ∈ E. Each of the events in E is a set of states from Ω.
Indeed, to define events consistently, a σ- field structure must be imposed on
the set E of the events. This allows, for example, to specify the coincidence of two
events as a new event. For example, if Ω = {rain, snow, sun}, then one can spec-
ify E = {{rain} , {snow} , {sun} , {rain, snow} , {rain, sun} , {snow, sun} , {rain, snow, sun} , }.
This is for the analyst who chooses the most appropriate specification for
E, while Ω may be somewhat imposed on her by a generally admitted theory.
Defining a σ-field E of Ω often amounts to defining the kind of information used.
This is one reason a careful specification of the possible events in E matters.
Definition:
(i) Let be a measurable space (Ω, E). A probability function on this space is
such that
an (Ai )i∈N all incompatible (i.e., disjoint) events from E, let the
and for P
∞
event A ≡ i=1 Ai [whichP∞is the occurrence of any of the events Ai , and can
be denoted in that case i=1 Ai ], then
P∞ P∞ P∞
P( i=1 Ai ) = P ( i=1 Ai ) = i=1 P (Ai ) .
An example is the usual regression model with normal errors for a given
observation i:
yi = x′i β + ui , where u follows a Gaussian distribution N (0, σ 2 ).
This yields a probability density function with density
1 In this context, a σ-algebra is a set of set of states that satisfies the above-described
8
(1/(2πσ))exp(−(yi − x′i β)2 /(2σ)).
The corresponding random events are generated from intersections and unions
of intervals of values for yi . The set of the states of the world can be anything
imagined that can generate the random events and values of yi . useful to the
analysis.
Definition:
(i) In a probability space, a measurable function of the random states is
called a random variable, and it can be endowed with a corresponding ‘image
probability function’.
A real-valued random variable X is a measurable function from a measurable
space (Ω, E) to a real measurable space (R, R), and describes the consequences
in each state s in Ω in terms of real numbers: X(s) ∈ R. That is:
For any event B in R, there exists an event A in E such that X −1 (B) ⊂ A.
The σ-field on real consequence events, R, is typically the ‘ Borel set’ of
events that can be constructed by using intersections and unions of the intervals
in R.
(ii) One can also define real-valued random vectors that are obtained by
stacking several real-valued random variables (e.g., Y = (X1 , ..., Xn )′ ).
9
Definition:
(i) Two events A and B are independent if and only if P (A∩B) = P (A)P (B).
(ii) The events in a set E are independent (‘among themselves’) if the probabil-
ity of the intersection of any subset of events from E is equal to the product of
the probabilities of these events in the subset.
Definition:
The conditional probability of an event B conditional on an event A , which
is such that P (A) ̸= 0, is:2
P (B|A) = P (A ∩ B)/P (A).
In the case that A and B are independent, then the knowledge that A
happens does not influence the perceived probability of B:
Proposition
For any event A and any complete finite set of exclusive events Bi , i =
1, ..., I, PI
P (A) = P (A|B)P (B) + P (A|B̄)P (B̄) = i=1 P (A|Bi )P (Bi ),
PI
Proof: Use P (A) = i=1 P (A ∩ Bi ), [since the A ∩ Bi are exclusive events
and sum to A] and substitute P (A|Bi )P (Bi ) for P (A ∩ Bi ).
2 The joint occurrence of two events A and B, in A ∩ B, brings some information about the
likelihood of B happening if A has already been observed. One can take advantage of this
intuition to develop the notion of conditional probability. Let us consider any two events A
and B, which are not independent, and such that A is nonnegligible (i.e., P (A) ̸= 0). One’s
subjective feeling about the likelihood that B occurs should be affected by the knowledge that
A has taken place. If A and B are incompatible (A ∩ B =), this likelihood is necessarily zero.
If A ⊂ B, the probability of occurrence of B is clearly one. In the other cases, one could
assume that this likelihood is proportional to P (A ∩ B). This is supported by noticing that
in the Venn diagram of A and B, the measured area A ∩ B may be seen as depicting the
information of A about B, and vice versa versa. This suggests a definition of the conditional
probability of B conditional on A: P (B|A) = kP (A ∩ B), with k a positive scalar constant to
determine. The normalization of this newly constructed probability function P (.|A), based on
a nonnegligible event A (P (A) ̸= 0) yields the formula. Indeed, given that A has happened,
A is certain and therefore 1 = P (A | A) = kP (A).
10
Defining a conditional probability P (B|A) can be seen as a tentative partial
prediction of event B with event A.
The proposition shows that the conditional probabilities can be used to
decompose the probability number of a given event into a sum of products of
probability numbers. The Bayes’s formula in the next subsection is a specific
use of these decompositions.
Learning about some uncertainty often amounts to learning about the prob-
ability numbers of events of interest. This learning process can be modelled
as updating ex ante probabilities by using information about the occurrence of
some observable events.
P (A|B).P (B)
P (B|A) = P (A|B).P (B)+P (A|B̄).P (B̄)
.
(ii) Consider a probability space with a finite number S of states of the world:
s = 1, ..., S; and a given probability function P (.) Let E be a random event
whose realization brings some information that can be exploited for updating
the probability function P (.). If all the singletons of the states of the world
{{s = j} , j = 1, ..., S} can be considered as random events, then any ex ante
probability number pj ≡ P ({s = j}) of a state s = j, can be updated as follows
into an ex-post probability number pj∗ ≡ P ({s = j} | E):
P (E | {s=j}).pj
pj∗ = PS = w j pj ,
P (E | {s=k}).pk
k=1
P (E | {s=j})
with the updating factors wj = PS .
P (E | {s=k}).pk
k=1
P (E | Bj ).P (Bj )
P(Bj | E) = PK
P (E | Bk ).P (Bk )
k=1
11
The intuition behind the formula in (ii) is that the ex post probability num-
ber of a state is a rescaling of its ex ante probability numbers by the likelihood
of the realized event E conditional on the considered state, normalized by the
denominator so that all ex-post probabilities sum to one. This implies that, due
to knowledge updating, some probability numbers of events should increase,
while other ones should diminish.
Proposition
(a) For a discrete distribution of a random variable X:
∞ ∞
X X P ({Xi } ∩ A)
E (X | A) = P ({Xi } | A) Xi = Xi
i=1 i=1
P (A)
(b) Let be the, respectively, joint, conditional and marginal density func-
tions,( f , fY |X and fX , respectively), of two random variables X and Y , with
any respective realizations x and y, then
f (x , y) = f Y |X (y | x ).f X (x ).
12
allows us to examine the role of the unobserved errors in the models, and the
problems that these may generate for the estimation Unclear. Third, they are
suggestive of modelling approaches based on regressions. Fourth, and not least,
conditoning expectations is often a good way to construct predictors.
Y = E(Y | X) + u,
Let be a k-dimensional real vector X and a real random variable Y , the con-
ditional expectation of Y given X is a numerical mesurable function E(Y | X)
defined on Rk , such that
2 2
E [Y − E(Y | X)] = min E [Y − f (X)]
f
Proposition:
(i) The conditional expectation E(Y | X) is not unique, in general. However,
it is almost surely unique.
(ii) The FOC that are here necessary and sufficient for a solution is that for
any measurable and square integrable function f ,
′
E [Y − E(Y | X)] [f (X)] = 0.
(iii) When X and Y are jointly Gaussian, then the conditional expectation
is equal to the linear regression and one can write:
′
E(Y | X) = β0 +β1 X 1 +...+βk X k , where X = X 1 , ..., X k and β0 , β1 , ..., βk
are real coefficients corresponding to the OLS formula.
Other sets of function f can be used to, and yield slightly different properties
of the conditional expectation.
13
The Law of Iterated Expectations
We will use the law of iterated expectations for calculating conditional ex-
pectation of the dependent variables on independent variables; and on diverse
random events. This property may not be as convenient for models for truncated
or censored data that are highly nonlinear.
The law of iterated expectations is what makes feasible many results of the
linear regression model. In particular, it allows the description of a (dependent)
variable as the mean over the covariates of the regression function (conditional
expectation conditional on covariates).
Proposition:
Let Y and X be two random variables and PX the probability function of
X, then (Law of Iterated Expectations):
Z
EY = E (Y | X = x )dP X (x ) = E PX [E (Y | X = x)] ,
Proposition:
Let W be a random vector that determines X through a functional rela-
tionship, X = f (W ). Typically the dimension of X is smaller than that of W .
Then,
E(Y | X) = E [E (Y | W ) | X] .
E(Y | X) = E [E (Y | X) | W ]
Y = Xβ + u,
14
where u is an unobserved residual. Since this model is only an approxima-
tion, u ̸= Y − E(Y | X) is likely. As a consequence, and the properties that u
is centered, and of u independent of the X ′ s may not be satisfied. That is: this
linear functional form may suffer from endogeneity issues for its estimation.
Let us show that, under the expected utility hypothesis, information is useful
to decisions, and that its usefulness can be measured in terms of VNM utility
levels.
Let us express the idea of information through the conditioning of the expec-
tation operator on a given informative event. Indeed, a sigma-field of events is a
way to describe relevant and consistent information about the world. Consider
the following example with two random events A and B, inspired from Gollier
(2004) a changer. Let P (A) = 1/2, P (B) = 1/4, P (A ∩ B) = 1/8. Assume the
expected utility hypothesis and that the VNM utility function, u, under the act
X, has a level equal to u = 1 for any random state in A ∩ B̄ or in Ā ∩ B, and
that its level is equal to u = −1 otherwise. Because of the cardinality property
of the expected utility, it is always possible to choose a VNM utility function
that satisfies these normalizations.
example to changeExample 1 : The event A could be ‘rain tomorrow’ and
the event B could be ‘be given an umbrella tomorrow’, while the act X could
consist in deciding to use the umbrella if it is given, whether it rains or not. The
umbrella prevents one from getting wet when it rains, but is uselessly heavy to
carry when it does not rain.
It is easy to calculate: E(u(X)) = 0, E(u(X)|A) = 1/2 and E(u(X)|Ā) =
−1/2.
Indeed, Eu(X) = P (Ā ∩ B̄).u(Ā ∩ B̄) + P (A ∩ B̄).u(A ∩ B̄) + P (A ∩ B).u(A ∩
B) + P (Ā ∩ B).u(Ā ∩ B)
= P (Ā ∩ B̄).(−1) + P (A ∩ B̄).1 + P (A ∩ B).(−1) + P (Ā ∩ B).1.
Moreover, P (Ā∩B) = P (B)−P (A∩B) = 1/8, P (A∩B̄) = P (A)−P (A∩B) =
3/8, P (Ā ∩ B̄) = 1 − P (A) − P (Ā ∩ B) = 1 − 1/2 − 1/8 = 3/8. Therefore:
Eu(X) = −3/8+3/8−1/8+1/8 = 0. The conditional expectations E(u(X)| A)
and E(u(X)| Ā), respectively under the information that event A occurs or
that event A does not occur, can be calculated similarly with first computing
the appropriate conditional probabilities using the formula P (E | A) = P (E ∩
A)/P (A) for any event E.
The words that we use when we read P (E | A), as ‘Probability of E knowing
A’ fit well the interpretation of the conditioning on A as referring to information
15
about the occurrence of A. This can also be related to Bayes formula used to
update some information quantified by using probabilities.
Consider now another act X ′ such that the VNM utility level is now equal
to -1 for any random state in A ∩ B̄ or in Ā ∩ B, and that it is equal to 1 oth-
erwise. That is: exactly the opposite signs to above. In that case: E(u(X ′ )) =
0, E(u(X ′ )|A) = −1/2 and E(u(X ′ )|Ā) = 1/2, with a symmetric calculus as
above.
Therefore, the agent should be indifferent between the two acts X and X ′
since in both cases Eu(X) = 0 and Eu(X ′ ) = 0.
However, if the agent knew whether A is true or not, he would be able to
choose act X if A is true since E(u(X)|A) = 1/2 while E(u(X ′ )|A) = −1/2;
and instead to choose act X ′ if A is not true, since E(u(X ′ )|Ā) = 1/2 (that is:
‘Eu knowing non-A’, while E(u(X)|Ā) = −1/2. Therefore, the information on
the occurrence of A has a utility value of 1/2.
Therefore, the gain from information can not only be manifested by com-
paring the levels of E(u(X)|Ā) and E(u(X)), but also concretely by making
distinct decisions, according to whether the agent is or is not informed, through
assessing the consequences of the selected act vs. those of the non-selected act.
Note however, that this utility gain has only a relative, not an absolute, mean-
ing, because the cardinality of the VNM utility function makes it unsufficient
to identify the gain exactly.
The utility value of information is always nonnegative in this model. This
may not be the case with other ‘non expected utility’ models.
These results, under the expected utility hypothesis, justify the interest of
searching for more information for predicting or forecasting. However, in that
case, new complications may arise because of searching costs that should be
included in the decision model. Moreover, ex ante information search may itself
be risky because the agent does not know which observations he will make during
his search.
Definition:
Let (Xn )n∈N be a sequence of real random variables and X is a real random
variable.
(i) Convergence in Probability:
P
Xn X as n −→ +∞ if and only if
−→
16
(ii) Almost Sure Convergence:
a.s.
Xn X as n −→ +∞ if and only if
−→
∀ε > 0, P sup |Xm − X| > ε −→ 0 as n −→ +∞.
m≥n
The following laws of large numbers and central limit theorems are useful to
prove the consistency and the asymptotic normality of the estimators that will
be considered, including the OLS estimators.
P∞ σi2
and i=1 i2 < ∞, Pn
1
then X̄ − µ̄ converges almost surely to zero, where µ̄ ≡ n i=1 µi .
√ d
then n X̄ − µ N (0, σ 2 ).
→
17
(vi) Central-Limit Theorem of Lindeberg-Feller:
If Xi , i = 1, ..., n are independent with E(Xi ) = µi < ∞ and V (Xi ) = σi2 <
∞.
1
Pn 1
Pn
Let µ̄ = n i=1 µi and σ̄n2 = n i=1 σi2 .
max(σi )
If lim nσ̄n = 0 and if lim σ̄n2 = σ 2 < ∞,
n−→∞ n−→∞
√ d
then n X̄ − µ̄ N (0, σ̄ 2 ).
→
Proposition (derivatives):
(vii) Leibniz formula: Let be a n-times differentiable function f of two vari-
ables x and y,
n
∂ n−k ∂ k
∂ ∂
dn f = dx + dy f =nk=0 f dxn−k dy k .
∂x ∂y ∂xn−k ∂y k
(viii) Matrix derivatives: The matrix at the denominator is the one that
indicates
′the orientation of the elements in the obtained matrix of derivatives.
∂f ∂f
∂A = ∂A ′.
∂ (a′ x) ∂ (a′ x)
∂x = a and ∂x′ = a′ .
∂(Ax) ∂(Ax)
∂x = A′ , where A is n × k and x is k × 1; and ∂x′ = A.
∂ (x′ Ax)
∂x = 2Ax, when A is symmetric; = (A + A′ )x when A is squared k × k
but not symmetric.
References:
18