0% found this document useful (0 votes)
36 views28 pages

Tobacco Consumption Analysis in Australia

The document presents a zero-inflated ordered probit (ZIOP) model to analyze tobacco consumption in Australia, addressing the issue of excessive zero observations in traditional ordered probit models. This model distinguishes between two types of zero observations: non-participants and zero-consumption participants, which allows for more accurate policy implications. Monte Carlo simulations demonstrate the model's effectiveness, and it is applied to a dataset of nearly 29,000 individuals, revealing significant insights into tobacco consumption patterns.

Uploaded by

gado sema
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
36 views28 pages

Tobacco Consumption Analysis in Australia

The document presents a zero-inflated ordered probit (ZIOP) model to analyze tobacco consumption in Australia, addressing the issue of excessive zero observations in traditional ordered probit models. This model distinguishes between two types of zero observations: non-participants and zero-consumption participants, which allows for more accurate policy implications. Monte Carlo simulations demonstrate the model's effectiveness, and it is applied to a dataset of nearly 29,000 individuals, revealing significant insights into tobacco consumption patterns.

Uploaded by

gado sema
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

See discussions, stats, and author profiles for this publication at: [Link]

net/publication/228232802

Analysis of Tobacco Consumption in Australia Using a Zero-Inflated Ordered


Probit Model

Article · June 2007

CITATIONS READS

6 37

2 authors, including:

Mark N. Harris
Curtin University
196 PUBLICATIONS 2,870 CITATIONS

SEE PROFILE

All content following this page was uploaded by Mark N. Harris on 13 February 2020.

The user has requested enhancement of the downloaded file.


8:07f=WðJul162004Þ Prod:Type:FTP ED:ArulselviU:
þ model
ECONOM : 2910 pp:1227ðcol:fig::NILÞ PAGN:Sakthi SCAN:

ARTICLE IN PRESS

3 Journal of Econometrics ] (]]]]) ]]]–]]]


[Link]/locate/jeconom
5

7
A zero-inflated ordered probit model, with an
9
application to modelling tobacco consumption
11
Mark N. Harris, Xueyan Zhao
13
Department of Econometrics and Business Statistics, Building 11, Wellington Road, Clayton Campus, Monash

F
University, Vic. 3800, Australia
15

O
17

O
19 Abstract
PR
Data for discrete ordered dependent variables are often characterised by ‘‘excessive’’ zero
21 observations which may relate to two distinct data generating processes. Traditional ordered probit
models have limited capacity in explaining this preponderance of zero observations. We propose a
23 zero-inflated ordered probit model using a double-hurdle combination of a split probit model and an
D
ordered probit model. Monte Carlo results show favourable performance in finite samples. The
25 model is applied to a consumer choice problem of tobacco consumption indicating that policy
TE

recommendations could be misleading if the splitting process is ignored.


27 r 2007 Elsevier B.V. All rights reserved.
EC

29 JEL classification: C3; D1; I1


Keywords: Ordered outcomes; Discrete data; Tobacco consumption; Zero-inflated responses
31
R

33
R

1. Introduction and Background


O

35
Often in empirical economics interest lies in modelling a discrete random variable that is
C

37 inherently ordered. Examples include survey responses on opinions, employment status


levels, bond ratings and job classifications by skill levels. Typically the empirical strategy
N

39 employed would involve estimation of an ordered probit (OP) or logit model (see, for
example, Zavoina and McElvey, 1975; Marcus and Greene, 1985). However, often data for
U

41
Corresponding author. Tel.: +61 39905 9414; fax: +61 39905 5474.
43
E-mail addresses: [Link]@[Link] (M.N. Harris), [Link]@[Link]
(X. Zhao).
45
0304-4076/$ - see front matter r 2007 Elsevier B.V. All rights reserved.
47 doi:10.1016/[Link].2007.01.002

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
2 M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]]

1 such ordered random variables are characterised by excessive observations in the choice at
one end of the ordering or, typically, ‘‘zeros’’. For example, in a survey corresponding to
3 illicit drug use, answers to a question such as ‘‘how often do you use drug X?’’ are likely to
be characterised by an excess of zero observations when discrete choices of consumption
5 levels including ‘‘never/not recently’’ ðy ¼ 0Þ are presented.
Traditional OP models have limited capacity in explaining such a preponderance of zero
7 observations, especially when the zeros may relate to two distinct sources. In the case of
discrete levels of reported drug consumption, zeros will be recorded for non-participants
9 who abstain due to health or legal concerns and who pay no regard to drug prices, for
example, in their decision making. However, there may also be zeros who are the corner
11 solution of a standard consumer demand problem and who may become consumers if the
price is lower or income is higher. Thus, it is likely that these two types of ‘‘zeros’’ are
13 driven by different systems of consumer behaviour. Here, zero consumption potential users
are likely to possess characteristics similar to those of the users and are likely to be

F
15 responsive to standard consumer demand factors such as prices and income. On the other

O
hand, genuine non-participants are likely to have perfectly inelastic price and income
17 demand schedules, and are driven by a separate process relating to sociological, health and

O
ethical considerations. If such underlying processes are modelled incorrectly, it could
19 invalidate any subsequent policy implications. Additionally, even the same explanatory
PR
variable could have different effects on the two decisions. One example is the effect of
21 income on drug consumption. Higher income, acting as an indicator for social class and
health awareness, may increase the chance of genuine non-participation. However if, for
23 participants, tobacco consumption is a normal good, higher incomes will be associated with
D
lower chances of zero consumption for these participants. An OP model generated by a
25 single latent equation cannot allow for the differentiation between the two opposing
TE

effects.
27 In a manner analogous to the zero-inflated/augmented Poisson (ZIP/ZAP) models in the
count data literature (see, for example, Mullahey, 1986; Heilbron, 1989; Lambert, 1992;
EC

29 Greene, 1994; Pohlmeier and Ulrich, 1995; Mullahey, 1997) and double-hurdle models in
the limited dependent variable literature (see, for example, Cragg, 1971), this paper
31 proposes an extension to the OP model to take into account of the possibility that the zeros
R

can arise from two different aspects of individual behaviour. Unlike the Poisson and
33 negative binomial regression framework, the ultimate data generating process here can be
R

seen as coming from two separate underlying latent variables. We propose a zero-inflated
O

35 ordered probit (ZIOP) model that involves a system of a probit ‘‘splitting’’ model and an
OP model which relate to potentially differing sets of covariates. We also further allow the
C

37 error terms of the two latent equations to be correlated (denoted a ZIOPC model), along
the lines of a Heckman-selection-type model (Heckman, 1979).
N

39 Monte Carlo experiments are conducted under various true models to examine the finite
sample performance. We also report performances of various specification tests and model
U

41 selection criteria for choosing between the OP, ZIOP and ZIOPC models. The model is
then applied to a unit record data set from Australia on tobacco consumption, which
43 involves an estimation sample of nearly 29,000 individuals and 76% of zero observations.
The application clearly illustrates the extra insights provided by the ZIOP/ZIOPC model in
45 analysing the effects of some important explanatory factors on individuals’ tobacco
consumption patterns.
47

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]] 3

1 2. The econometric framework

3 2.1. A zero-inflated ordered probit model (ZIOP)

5 We start by defining a discrete random variable y that is observable and assumes the
discrete ordered values of 0; 1; . . . ; J. A standard OP approach would map a single latent
7 variable to the observed outcome y via so-called boundary parameters, with the latent
variable being related to a set of covariates. Here we propose a ZIOP model that involves
9 two latent equations: a probit selection equation and an OP equation. This splits the
observations into two regimes that relate to potentially two different sets of explanatory
11 variables. Consider the drug consumption example. Here an individual user is modelled as
having to overcome two hurdles: whether to participate, and then, conditional on
13 participation, how much to consume which also includes zero consumption. The two types
of zero-consumption observations relate to those non-participants with perfectly inelastic

F
15 demand to prices and income, and those zero consumption participants who report zero

O
consumption at the time but who may consume once the price is right, for example.1 The
17 former may relate to personal demographics and socioeconomic status, whilst the latter

O
group may exhibit behaviour similar to other non-zero users and be more responsive to
19 economic factors such as prices and income.

21
PR
Let r denote a binary variable indicating the split between Regime 0 (with r ¼ 0 for non-
participants) and Regime 1 (with r ¼ 1 for participants). r is related to a latent variable r
via the mapping: r ¼ 1 for r 40 and r ¼ 0 for r p0. The latent variable r represents the
23 propensity for participation and is defined as
D
r ¼ x0 b þ , (1)
25
TE

where x is a vector of covariates that determine the choice between the two regimes, b a
27 vector of unknown coefficients, and  a standard-normally distributed error term.
Accordingly, the probability of an individual being in Regime 1 is given by (Maddala,
EC

29 1983)
Prðr ¼ 1jxÞ ¼ Prðr 40jxÞ ¼ Uðx0 bÞ, (2)
31
R

where Uð:Þ is the cumulative distribution function (c.d.f.) of the univariate standard normal
33 distribution.
R

Conditional on r ¼ 1, consumption levels under Regime 1 are represented by a discrete


variable yeðe
y ¼ 0; 1; . . . ; JÞ that is generated by an OP model via a second underlying latent
O

35
variable ye :
C

37 ye ¼ z0 c þ u, (3)
N

with z being a vector of explanatory variables with unknown weights c, and u an error term
39
following a standard normal distribution. The mapping between ye and ye is given by
U

8 
41 < 0 if ye p0;
>
ye ¼ j if mj1 oe y pmj ðj ¼ 1; . . . ; J  1Þ; (4)
43 >
: J if m pe 
J1 y ;
45 where m are boundary parameters to be estimated in addition to c, and we assume

1
47 Prices may, however, be more important for younger individuals in the participation decision.

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
4 M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]]

1 throughout the paper that m0 ¼ 0. Note that, importantly, Regime 1 also allows for zero
consumption. Also, there is no requirement that x ¼ z. Under the assumption that u is
3 standard Gaussian, the OP probabilities are given by (Maddala, 1983)
8
> y ¼ 0jz; r ¼ 1Þ ¼ Fðz0 cÞ;
Prðe
5 <
yÞ ¼ Prðe
Prðe y ¼ jjz; r ¼ 1Þ ¼ Fðmj  z0 cÞ  Fðmj1  z0 cÞ ðj ¼ 1; . . . ; J  1Þ; (5)
>
: Prðe
7 y ¼ Jjz; r ¼ 1Þ ¼ 1  FðmJ1  z0 cÞ:

9 While r and ye are not individually observable in terms of the zeros, they are observed via
the criterion
11 y ¼ re
y. (6)
That is, to observe a y ¼ 0 outcome we require either that r ¼ 0 (the individual is a non-
13
participant) or jointly that r ¼ 1 and ye ¼ 0 (the individual is a zero consumption

F
participant). To observe a positive y, we require jointly that the individual is a participant
15
ðr ¼ 1Þ and that ye 40. Under the assumption that  and u identically and independently

O
follow standard Gaussian distributions, the full probabilities for y are given by
17 (

O
Prðy ¼ 0jz; xÞ ¼ Prðr ¼ 0jxÞ þ Prðr ¼ 1jxÞ Prðey ¼ 0jz; r ¼ 1Þ
19 PrðyÞ ¼

21
8
>
>
<
Prðy ¼ jjz; xÞ ¼ Prðr ¼ 1jxÞ Prðe
Prðy ¼ 0jz; xÞ ¼ ½1  Fðx0 bÞ þ Fðx0 bÞFðz0 cÞ
0 0 0
PR
y ¼ jjz; r ¼ 1Þ ðj ¼ 1; . . . ; JÞ

¼ Prðy ¼ jjz; xÞ ¼ Fðx bÞ½Fðmj  z cÞ  Fðmj1  z cÞ ðj ¼ 1; . . . ; J  1Þ ð7Þ


23 >
>
D
: Prðy ¼ Jjz; xÞ ¼ Fðx0 bÞ½1  Fðm  z0 cÞ:J1
25
TE

In this way, the probability of a zero observation has been ‘‘inflated’’ as it is a


combination of the probability of ‘‘zero consumption’’ from the OP process plus the
27
probability of ‘‘non-participation’’ from the split probit model. Note that this specification
EC

is analogous to the zero-inflated/augmented count models, and that there may or may not
29
be overlaps with the variables in x and z. Moreover, the model is also directly comparable
to the double-hurdle limited dependent variable models (see, for example, Cragg, 1971).
31
R

Once the full set of probabilities has been specified and given an i:i:d: sample of size N
from the population on ðyi ; xi ; zi Þ, i ¼ 1; . . . ; N, the parameters of the full model h ¼
33
R

ðb0 ; c0 ; l0 Þ0 can be consistently and efficiently estimated using maximum likelihood (ML)
criteria, yielding asymptotically normally distributed maximum likelihood estimates
O

35
(MLEs).2 The log-likelihood function is
C

37 N X
X J
‘ðhÞ ¼ hij ln½Prðyi ¼ jjxi ; zi ; hÞ, (8)
N

i¼1 j¼0
39
U

where the indicator function hij is


41 
1 if individual i chooses outcome j
hij ¼ ði ¼ 1; . . . ; N; j ¼ 0; 1; . . . ; JÞ (9)
43 0 otherwise:

45
2
An anonymous referee pointed out that an interesting line of future research would be to estimate such a model
47 by Bayesian Monte Carlo Markov Chain techniques.

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]] 5

1 2.2. Generalising the model to correlated disturbances (ZIOPC)

3 As described above, the observed realisation of the random variable y can be viewed as
the result of two separate latent equations, Eqs. (1) and (3), with uncorrelated error terms.
5 However, these correspond to the same individual so it is likely that the two stochastic
terms  and u will be related. We now extend the model to have ð; uÞ follow a bivariate
7 normal distribution with correlation coefficient r, whilst maintaining the identifying
assumption of unit variances. The full observability criteria are thus
9 8 
< 0 if ðr p0Þ or ðr 40 and ye p0Þ;
 
>
11 y ¼ rey ¼ j if ðr 40 and mj1 oe y pmj Þ ðj ¼ 1; . . . ; J  1Þ; (10)
>
: J if ðr 40 and m oe 
J1 y Þ;
13
which translate into the following expressions for the probabilities:
8

F
15 >
> Prðy ¼ 0jz; xÞ ¼ ½1  Fðx0 bÞ þ F2 ðx0 b; z0 c; rÞ;
>

O
>
< Prðy ¼ jjz; xÞ ¼ F2 ðx0 b; mj  z0 c; rÞ  F2 ðx0 b; mj1  z0 c; rÞ
17 PrðyÞ ¼ (11)
> ðj ¼ 1; . . . ; J  1Þ;

O
>
>
>
: Prðy ¼ Jjz; xÞ ¼ F2 ðx0 b; z0 c  m ; rÞ;
19 J1

21
correlation coefficient l between the two univariate random elements.
PR
where F2 ða; b; lÞ denotes the c.d.f. of the standardised bivariate normal distribution with

ML estimation would again involve maxmisation of Eq. (8) replacing the probabilities of
23
(7) with those of (11) and re-defining h as h ¼ ðb0 ; c0 ; l0 ; rÞ0 . A Wald test of r ¼ 0 is a test for
D
independence of the two error terms and thus a test of the more general model given by Eq.
25
TE

(11) against the null of a simpler nested model of Eq. (7).3


27
2.3. Marginal effects
EC

29
There are several sets of marginal effects that may be of interest in this model. For
example, one may be interested in the marginal effects of an explanatory variable on the
31
R

probability of ‘‘participation’’, Prðr ¼ 1Þ; or the probabilities for the levels of consumption
conditional on participation, Prðe y ¼ jjr ¼ 1Þ; or on the overall probabilities for different
33
R

levels of consumption, Prðy ¼ jÞ. In particular, the marginal effect on the overall
probability of observing zero consumption, Prðy ¼ 0Þ, is the sum of the effects on the
O

35
probabilities of two types of zeros; that is, the probability of non-participation and the
probability of zero-consumption arising from participants who are infrequent or potential
C

37
consumers.
N

The marginal effect of a dummy variable can be calculated as the difference in the
39
probability of interest with the relevant dummy variable turned ‘‘on’’ and ‘‘off’’,
U

conditional on given values of all other covariates. Note that the explanatory variable of
41
interest may appear in only one of x or z, or in both. For a continuous variable xk , the
marginal effect on the participation probability in Eq. (2), which only relates to
43
explanatory variables in x, is given by
45 3
An OP model in Eqs. (3) and (4) with ye  y will yield starting values for c and l. Those for b can be obtained
from estimating a binary probit model defined by (1) and (2) with Prðr ¼ 1jxÞ  Prðy40jxÞ. With fixed b hZIOP a
47 grid-search over ð0:9; 0:9Þ would provide a start value for r.

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
6 M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]]

1 q Prðr ¼ 1Þ
ME ¼ ¼ fðx0 bÞbk . (12)
Prðr¼1Þ qxk
3
To derive the marginal effects on the overall probabilities for the general model of
5 ZIOPC, we partition the explanatory variables and the associated coefficients as
  !   !
w bw w cw
7 x¼ ; b¼ e ; z¼ and c ¼ , (13)
e
x b ez ec
9
where w represents the common variables that appear in both x and z, with the associated
11 coefficients bw and cw for the participation and consumption equations, respectively. e x and
ez denote those distinct variables that only appear in one of the latent equations, with e b and
13 ec as their associated coefficients for the two equations.
Denote the unique explanatory variables for the whole model as x ¼ ðw0 ; ex0 ; ez0 Þ0 , and set
0 e0 0 0

F
   0 0 0 0
15 the associated coefficient vectors for x as b ¼ ðb w ; b ; 0 Þ and c ¼ ðcw ; 0 ; e c Þ . The

marginal effects of the explanatory variable vector x on the full probabilities in Eq. (11)

O
17 are given by

O
" ! # !
z0 c þ rx0 b x0 b  rz0 c
19 ME ¼ F pffiffiffiffiffiffiffiffiffiffiffiffiffi  1 fðx bÞb  F pffiffiffiffiffiffiffiffiffiffiffiffiffi fðz0 cÞc ,
0 

21
Prðy¼0Þ
"
1  r2
z0 c þ rx0 b
ME ¼ F pffiffiffiffiffiffiffiffiffiffiffiffiffi
! #
x 0
b 
PR
1  r2

 1 fðx0 bÞb  F pffiffiffiffiffiffiffiffiffiffiffiffiffi fðz0 cÞc


rz 0
c
!

Prðy¼0Þ 1  r2 1  r2
23 " ! !#
D
0 x0 b  rz0 c 0 x0 b þ rðm1  z0 cÞ
25 þ fðz cÞF pffiffiffiffiffiffiffiffiffiffiffiffiffi  fðm1  z cÞF pffiffiffiffiffiffiffiffiffiffiffiffiffi c ,
TE

1  r2 1  r2
" ! !#
27 m2  z0 c þ rx0 b m1  z0 c þ rx0 b
ME ¼ F pffiffiffiffiffiffiffiffiffiffiffiffiffi F pffiffiffiffiffiffiffiffiffiffiffiffiffi fðx0 bÞb
1  r2 1  r2
EC

Prðy¼2Þ
29 " ! !#
0 x0 b þ rðm1  z0 cÞ 0 x0 b þ rðm2  z0 cÞ
þ fðm1  z cÞF pffiffiffiffiffiffiffiffiffiffiffiffiffi  fðm2  z cÞF pffiffiffiffiffiffiffiffiffiffiffiffiffi c ,
1r 2 1r 2
31
R

..
33 .
R

!
z0 c  mJ1  rx0 b
ME ¼ F pffiffiffiffiffiffiffiffiffiffiffiffiffi fðx0 bÞb þ fðz0 c  mJ1 Þ
O

35 Prðy¼JÞ 1  r2
!
C

37 x0 b  rðz0 c  mJ1 Þ 
F pffiffiffiffiffiffiffiffiffiffiffiffiffi c , ð14Þ
1  r2
N

39
where fð:Þ is the p:d:f : of the standard univariate normal distribution. Note that the
U

41 marginal effect on Prðy ¼ 0Þ can be decomposed into the marginal effects on the
probabilities of the two types of zeros. Marginal effects for the ZIOP model are obtained
43 as above but with r ¼ 0. Standard errors of the marginal effects can be obtained by the
Delta method (see, for example, Greene, 2003, pp. 674–675). An alternative method is to
45 use simulated asymptotic sampling techniques. Specifically, randomly draw h from
MVNðb h; Var½bhÞ a large number of times. For each random draw, calculate the marginal
47 effects using either the analytical expressions of Eq. (14) or the numerical derivatives of the

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]] 7

1 probability expressions. The empirical standard deviation of the simulated marginal effects
is a valid asymptotic estimate of their standard errors.4
3

5 2.4. Hypothesis testing and model selection issues

7 Testing between the ZIOP and ZIOPC models can be based on a simple t-test of r ¼ 0,
using the standard errors from the estimated Hessian. With regard to the ZIOP (or
9 ZIOPC) versus the OP model, they are not nested in the usual sense of parameter
restrictions. The ZIOP (or ZIOPC) model becomes an OP model when Prðr ¼ 1jxÞ  1, or
11 x0 b ! 1 in Eq. (2), implying all individuals are in Regime 1 and there is no ‘‘zero-
splitting’’ process. Although having two non-nested models in this case, a generalised
13 likelihood ratio (LR)-type statistic could be used, with degrees of freedom being given by
the number of additional parameters estimated in the more general model. The LR test is

F
15 known to have good properties in non-standard testing problems (see, for example,

O
Andrews and Ploberger, 1995; Chesher and Smith, 1997). Similarly lacking theoretical
17 underpinnings in this situation but a useful general specification test in many situations,

O
the Hausman specification test (Hausman, 1978) could also be considered, with the degrees
19 of freedom being the number of common parameters estimated in the competing models.

21
PR
As indicated in the Monte Carlo results in Section 3, both tests actually perform quite well.
A more theoretically based approach is the Vuong test (Vuong, 1989) for testing between
two non-nested models, which has been suggested in the related context of testing a zero-
23 inflated Poisson versus a simple count model (Greene, 2003). Denote f h ðyi jxi ; zi Þ as the
D
predicted probability using Model h (h ¼ 1 and 2 for OP and ZIOP/ZIOPC) that yi equals
25 observed y and let
TE

 
f ðy jxi ; zi Þ
27 mi ¼ log 1 i . (15)
f 2 ðyi jxi ; zi Þ
EC

29 To test the null hypothesis that Eðmi Þ ¼ 0; or that there is no difference in the
probabilities of correct prediction using the two models, the Vuong statistic is given by
31 pffiffiffiffiffi
R

P
N ð1=N N i¼1 mi Þ
33 u ¼ qffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
PN ffi, (16)
R

1=N i¼1 ðmi  m̄Þ2


O

35
which has a standard normal limiting distribution. The test statistic is bidirectional in the
sense that jujo1:96 lends support to neither model, whereas uo  1:96 favours Model 2
C

37
and u41:96 favours Model 1 (Vuong, 1989).
N

Finally, in such a non-nested situation, information based model-selection criteria, such


39
as AIC, BIC and consistent AIC ðCAICÞ, are appropriate for choosing between alternative
U

models. These are given by AIC ¼ 2‘ðhÞ þ k, BIC ¼ 2‘ðhÞ þ ðln NÞk, and CAIC ¼
41
2‘ðhÞ þ ð1 þ ln NÞk (see, for example, Cameron and Trivedi, 1998, p. 183), where k is the
total number of parameters estimated and ‘ðhÞ the maximised log-likelihood function. The
43
preferred model is that with the smallest value.
45 4
The Delta and the simulation methods represent two distinct asymptotic approximations. Nearly identical
results were obtained from the two approaches, indicating the adequacy of the underlying asymptotic theory in
47 this case.

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
8 M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]]

1 3. Finite sample performance

3 Although the model satisfies the regularity conditions for ML estimation (see Greene,
2003), results of some Monte Carlo experiments are presented in this section to provide
5 some evidence on the finite sample performance of the ML estimator and the model
selection criteria. We consider, in turn, the cases when the data generating process (d.g.p.)
7 are ZIOPC, ZIOP and OP. Estimation was undertaken in the Gauss matrix programing
language, using the CML maximum likelihood estimation add-in module.5
9
3.1. Performance under ZIOPC/ZIOP
11
3.1.1. Monte Carlo design
13 R ¼ 1; 000 repeated samples, each with a sample size of N ¼ 1; 000, are generated from a

F
ZIOPC d.g.p., and all three models of ZIOPC, ZIOP and OP are estimated (Experiment 1).
15 This is then repeated for a d.g.p. of ZIOP with r ¼ 0 (Experiment 2). We draw x from

O
x ¼ ð1; x1 ; x2 Þ0 , where x1 ¼ logðUniform½0; 100Þ and x2 ¼ 1fUniform½0;140:25g . Observations
17 for z are generated from z ¼ ð1; z1 Þ0 , where z1  x1 . We have chosen the continuous

O
variable z1  x1 to mimic variables such as age and income, and the dummy variable, x2 to
19 represent qualitative characteristic such as gender or marital status. The set of 1,000 draws
21
of x and z is generated once and subsequently held fixed. PR
Parameter values for both experiments are set as follows: ðb0 ; b1 ; b2 Þ0 ¼ ð1; 0:25; 1Þ0 ,
ðg0 ; g1 Þ0 ¼ ð0:5; 1Þ0 , and ðm1 ; m2 Þ0 ¼ ð4:5; 5:5Þ0 ; for Experiment 1, r ¼ 0:5. Note that we have
23
D
set the coefficients b1 and g1 as having opposite signs allowing for the same explanatory
variable ðx1  z1 Þ to have opposing effects on the two latent variables. The parameter
25
TE

values are also chosen to yield around 70% of zero observations.


27
3.1.2. Monte Carlo results
EC

29 The results are summarised in Tables 1 and 2. As individual coefficients in such discrete
models do not convey much information, we present the results in terms of estimated
31 marginal effects evaluated at sample means of the observed covariates. For each of the
R

j ¼ 0; . . . ; J outcomes, we present the true marginal effects, the estimated marginal effects
33 averaged over the R runs (ME), the root mean square error of the R estimated marginal
R

effects relative to the true MEs (RMSE), as well as the empirical coverage probabilities
O

35 (CP), measured as the percentages of times the true marginal effects fall within the
estimated 95% confidence intervals. We also present the same set of information for the
estimated parameters l and r, though we do not present results for b and c.6
C

37
Some model selection and summary statistics are also presented. RMSE_P refers to the
N

39 root mean square error of the predicted probabilities for all outcomes and observations,
averaged over the R Monte Carlo runs, where for the mth replication
U

qffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
P PJ
41 RMSE m ¼ ð1=NJÞ N ðPbm  Pm Þ2 , ðm ¼ 1; . . . ; RÞ. Correct gives the average
i¼1 j¼0 ij ij
percentage of correct predictions for y based on the maximum probability rule, whilst
43
Time refers to the average estimation time (in minutes). The results for the Wald test of the
45 5
Code is available from the authors on request. Also the model will be available in the next release of Limdep
(version 9.0), NLOGIT (4.0).
6
47 Results for b and c can be found in Harris and Zhao (2004).

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]] 9

1 Table 1
Monte Carlo results under ZIOPC: Experiment 1
3 Prðy ¼ 0jx; zÞ Prðy ¼ 1jx; zÞ

5 TRUE OP ZIOP ZIOPC TRUE OP ZIOP ZIOPC

Marginal effects (ME)


7 x1 ¼ z1 ME 0.080 0.013 0.083 0.082 0.149 0.004 0.147 0.150
RMSE (0.068) (0.018) (0.018) (0.145) (0.015) (0.016)
9 CP 0.003 0.955 0.955 0.000 1.000 1.000
x2 ME 0.320 0.321 0.322 0.165 0.134 0.165
11 RMSE (0.031) (0.031) (0.035) (0.022)
CP 0.957 0.956 1.000 1.000
13
Prðy ¼ 2jx; zÞ Prðy ¼ 3jx; zÞ

F
15 TRUE OP ZIOP ZIOPC TRUE OP ZIOP ZIOPC

O
x1 ¼ z1 ME 0.001 0.004 0.002 0.002 0.071 0.005 0.063 0.070
17
RMSE (0.007) (0.011) (0.012) (0.076) (0.012) (0.010)

O
CP 0.997 1.000 1.000 0.000 1.000 1.000
19

21
x2 ME
RMSE
CP
0.118 0.126
(0.019)
0.998
0.119
(0.019)
1.000
PR
0.037 0.062
(0.027)
1.000
0.038
(0.013)
1.000

23 TRUE OP ZIOP ZIOPC


D
Coefficients
25 m1
TE

m1 4.500 0.446 4.787 4.461


RMSE (4.050) (0.860) (0.830)
27 CP 0.000 0.976 0.947
EC

m2 m2 5.500 0.893 5.872 5.460


29 RMSE (4.610) (0.910) (0.860)
CP 0.000 0.976 0.947
31 r r 0.500 0.482
R

RMSE (0.170)
33 CP 0.927
R

OP ZIOP ZIOPC
O

35
RMSE_P 0.116 0.023 0.020
C

(0.001) (0.005) (0.006)


37 Correct 0.731 0.744 0.744
Time 0.025 0.075 0.493
N

39 Wald 0.235 0.765


Vuong 0.00 1.00
U

41 VuongðCÞ 0.00 1.00


LR 1.00
LRðCÞ 1.00
43 Hausman 1.00
HausmanðCÞ 0.99
45 AIC 0.000 0.074 0.926
BIC 0.000 0.540 0.460
CAIC 0.000 0.589 0.411
47

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
10 M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]]

1 Table 2
Monte Carlo results under ZIOP: Experiment 2
3 Prðy ¼ 0jz; xÞ Prðy ¼ 1jz; xÞ

5 TRUE OP ZIOP ZIOPC TRUE OP ZIOP ZIOPC

Marginal effects
7 x1 ¼ z1 ME 0.080 0.018 0.083 0.083 0.146 0.009 0.146 0.147
RMSE (0.063) (0.019) (0.019) (0.137) (0.017) (0.017)
9 CP 0.003 0.953 0.955 0.000 0.979 0.982
x2 ME 0.320 0.320 0.320 0.206 0.206 0.206
11 RMSE (0.033) (0.032) (0.024) (0.027)
CP 0.951 0.952 1.000 1.000
13
Prðy ¼ 2jz; xÞ Prðy ¼ 3jz; xÞ

F
15 TRUE OP ZIOP ZIOPC TRUE OP ZIOP ZIOPC

O
x1 ¼ z1 ME 0.033 0.005 0.032 0.032 0.033 0.004 0.032 0.032
17
RMSE (0.038) (0.010) (0.010) (0.037) (0.006) (0.006)

O
CP 0.000 0.996 0.997 0.000 1.000 1.000
19

21
x2 ME
RMSE
CP
0.087 0.087
(0.013)
1.000
0.087
(0.015)
1.000
PR
0.028 0.028
(0.007)
1.000
0.027
(0.008)
1.000

23 TRUE OP ZIOP ZIOPC


D
Coefficients
25 m1
TE

m1 4.500 0.667 4.449 4.393


RMSE (3.830) (0.780) (0.780)
27 CP 0.000 0.943 0.939
EC

m2 m2 5.500 1.206 5.454 5.383


29 RMSE (4.290) (0.790) (0.800)
CP 0.000 0.943 0.939
31 r r 0.000 0.005
R

RMSE (0.220)
33 CP 0.927
R

OP ZIOP ZIOPC
O

35
RMSE_P 0.110 0.019 0.019
C

(0.002) (0.006) (0.006)


37 Correct 0.733 0.748 0.748
Time 0.0262 0.0811 0.3184
N

39 Wald 0.927 0.073


Vuong 0.00 1.00
U

41 VuongðCÞ 0.00 1.00


LR 1.00
LRðCÞ 1.00
43 Hausman 1.00
HausmanðCÞ 1.00
45 AIC 0.000 0.670 0.330
BIC 0.000 0.988 0.012
CAIC 0.000 0.991 0.009
47

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]] 11

1 null of ZIOP ðH 0 : r ¼ 0Þ against the alternative of ZIOPC are given as the percentage of
times that the statistic lends support to each model. LR and LRðCÞ and Hausman and
3 HausmanðCÞ are, respectively, the percentage of times the LR and Hausman statistics
favours the ZIOP (ZIOPC) over the OP model. Vuong/Vuong(C) corresponds to Vuong’s
5 (1989) non-nested test as applied to the ZIOP/ZIOPC model versus OP, again expressed as
the percentage of times each model is selected. Finally, the percentage of times each of the
7 AIC; BIC and CAIC selects each of the three models is also reported. All tests are
undertaken at 5% nominal size.
9 As can be seen from Table 1, when the true model is ZIOPC (Experiment 1) and a simple
OP model is estimated, not surprisingly, both the estimated marginal effects and the
11 estimated boundary parameters are severely biased. With the exception of one outcome,
95% empirical CP are essentially zero. On the other hand, estimation of a ZIOP model
13 ignoring the correlation performs quite well. Average estimated marginal effects are very
close to the true ones, and moreover have small RMSEs. Allowing for the (true)

F
15 correlation in estimation (ZIOPC) further improves the results. RMSE measures for

O
ZIOPC are even better with essentially zero biases and coverage probabilities ranging from
17 0.927 to 1.7 The average estimate of r is 0.48 compared to the actual value of 0.5, with an

O
empirical CP of 93%.
19 In terms of correctly estimating probabilities, the OP clearly fares poorly with an
PR
average RMSE of 0.116. Significant improvements are afforded by the ZIOP and ZIOPC
21 models; here the mean RMSE falls sharply to 0.023 and 0.020, respectively. The percentage
of correct predictions is fairly similar across all models. This is the result of a common
23 phenomenon typical in discrete choice models, where models tend to simply predominantly
D
predict the most frequently observed outcome (here zeros), resulting in poor ‘‘predictive’’
25 performance.
TE

For applied researchers, an important issue is a model selection procedure to correctly


27 choose between alternative models. As shown in Table 1, the Wald test of ZIOPC versus
ZIOP correctly rejects the null and selects the correct ZIOPC model in 77% of cases. For
EC

29 choosing between the non-nested models, Vuong’s (1989) statistic correctly selects the
ZIOPC model over the incorrect OP in all cases. The uncorrelated ZIOP model is also
31 preferred to the OP model by the Vuong test in all instances. The LR statistics also
R

correctly reject the OP model all of the time. The Hausman statistics fare similarly well,
33 only incorrectly rejecting in 1% of instances (for the correlated version). In terms of the
R

information criteria, in no instances do any of the criteria incorrectly select the OP model.
O

35 AIC significantly favours the ZIOPC model, whereas BIC and CIAC have an approximate
equal split in choosing between the ZIOP and ZIOPC models. However, as already stated,
C

37 a preferable method of choosing between these two nested models would be a Wald test of
r ¼ 0. Finally, with regard to estimation times, the ZIOPC (with an average of 0.493 min
N

39 per estimation) appears to be more computationally intensive than the simpler OP


(0.025 min) and its uncorrelated counterpart ZIOP (0.075 min).
U

41 Turning to Experiment 2 in Table 2 where the d.g.p. is ZIOP (with r ¼ 0), here we would
expect the OP to fare poorly and the ZIOP and ZIOPC to excel, with the estimates of r in
43 the latter to be ‘‘small’’ and insignificant. Indeed, the OP estimated marginal effects and
boundary parameters are again quite severely biased, with CP of 0.003 or lower, whereas
45
7
Note that all of these CP are based on asymptotic distributions whilst N here is ‘‘small’’ at 1,000. Empirical CP
47 would therefore be much closer to the theoretical 0.95 values for larger sample sizes.

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
12 M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]]

1 both of the ZIOP and ZIOPC ones are essentially identical with small biases with empirical
CP ranging from 0.939 to 1. The average estimate of r is 0.005 and at 5% nominal size one
3 would incorrectly reject the null of r ¼ 0 in 7.3% of cases. The ZIOP and ZIOPC models
clearly dominate the misspecified simple OP model in terms of RMSE_P. Once more, in all
5 instances the Vuong statistic correctly chooses against the simple OP model, as do the LR
and Hausman statistics. With regard to model selection criteria, the OP is never chosen by
7 any of the criteria, and here the information criteria perform much better with regard to
choosing the correct model.
9
3.1.3. Exclusion restrictions
11 It is often the case in such two-part models that precision of parameter estimates is
enhanced if there are explicit exclusion restrictions in the specification of the covariates in
13 the two equations. For example, in the well-known Heckman-selection equation
(Heckman, 1979), although the correlation between the selection and regression equations

F
15 is identified by the nonlinearities involved, due to multicollinearity concerns this

O
correlation is often imprecisely estimated if x  z. Smith (2003) also suggests that
17 identification of the correlation parameter may be stronger in samples with more than 50%

O
zeros. To examine the likely effect of exclusion restrictions, Experiment 1 with a d:g:p: of
19 ZIOPC is re-run assuming x  z. The results are presented in Table 3.
PR
Here all results for the marginal effects are somewhat similar to Experiment 1 where
21 exclusion restrictions were in place. All of the model selection procedures also invariably
correctly select the larger models over OP. However, there is indeed evidence that the
23 correlation coefficient, r, is not properly identified; the average estimate for r is only 0.038
D
compared to the true value of 0.5. Based on the Wald statistic, in only 2.3% of the cases
25 would one correctly select the ZIOPC model over the ZIOP variant. Furthermore,
TE

convergence problems were encountered for particular draws of the random variables
27 within the Monte Carlo experiment. Indeed, Smith (2003) has also suggested that weak
identification can lead to computational problems such as lack of convergence in similar
EC

29 models.
From an empirical point of view however, given that the zeros are assumed to come
31 from two different regimes, a model with x  z is not going to be a scenario that an applied
R

researcher would necessarily entertain.


33
R

3.2. Performance under OP


O

35
We now consider the case when the true model is, in fact, the usual OP model. In this
C

37 case, even though x and b do not feature in the true OP d.g.p., ZIOP and ZIOPC models
were estimated as if they did. We consider several scenarios for the explanatory variables x
N

39 and z. Experiment 4 has partly overlapping x and z as in Experiments 1 and 2, with


x ¼ f1; logðUniform½0; 100Þ; 1fUniform½0;140:25g g and z equal to the first two columns of x.
U

41 Experiment 5 assumes that x and z have explicit exclusion restrictions and no overlapping
variables, with z ¼ f1; jNð0; 4Þjg and x as before. In Experiment 6, we consider a case of
43 complete overlap, with x  z ¼ f1; logðUniform½0; 100Þg. Finally Experiment 7 has x and z
as in Experiment 4 except that x has an additional Nð0; 4Þ variate. Again R ¼ 1; 000 Monte
45 Carlo replications are considered in all scenarios.
The marginal effect results from these experiments are in Table 4, and summary statistics
47 and model selection results in Table 5. Convergence problems were encountered with the

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]] 13

1 Table 3
Monte Carlo results under ZIOPC (x ¼ z): Experiment 3
3 Prðy ¼ 0jz; xÞ Prðy ¼ 1jz; xÞ

5 TRUE OP ZIOP ZIOPC TRUE OP ZIOP ZIOPC

Marginal effects
7 x1 ¼ z1 ME 0.099 0.032 0.102 0.102 0.301 0.005 0.295 0.297
RMSE (0.132) (0.021) (0.021) (0.306) (0.022) (0.021)
9 CP 0.000 0.952 0.959 0.000 0.986 0.999
Prðy ¼ 2jz; xÞ Prðy ¼ 3jz; xÞ
11
TRUE OP ZIOP ZIOPC TRUE OP ZIOP ZIOPC
13
x1 ¼ z1 ME 0.079 0.012 0.071 0.073 0.123 0.015 0.123 0.123

F
RMSE (0.067) (0.019) (0.018) (0.108) (0.012) (0.012)
15 CP 0.000 1.000 0.997 0.000 1.000 1.000

O
TRUE OP ZIOP ZIOPC
17

O
Coefficients
19 m1 m1 4.500 0.694 4.865 4.826

21
m2
RMSE
CP
m2 5.500
(3.810)
0.000
1.315
(0.720)
0.949
5.955
(0.730)
0.986
5.902
PR
23 RMSE (4.190) (0.780) (0.790)
D
CP 0.000 0.949 0.986

25 r r 0.500 0.038
TE

RMSE (0.510)
CP 0.990
27
OP ZIOP ZIOPC
EC

29
RMSE_P 0.122 0.019 0.019
(0.002) (0.006) (0.006)
31 Correct 0.470 0.523 0.523
R

Time 0.027 0.067 0.344


33 Wald 0.977 0.023
R

Vuong 0.00 1.00


VuongðCÞ 0.00 1.00
O

35 LR 1.00
LRðCÞ 1.00
C

37 Hausman 1.00
HausmanðCÞ 0.94
N

39 AIC 0.000 0.976 0.024


BIC 0.000 0.997 0.003
U

CAIC 0.000 0.998 0.002


41

43 ZIOPC model in these experiments. For this reason, only the ZIOP model was estimated.
For the applied researchers, if convergence problems are encountered with the ZIOPC, it
45 may suggest that the data is inconsistent with a zero-splitting process.
Results in Table 4 show that when the true d.g.p. is OP, a ZIOP model actually performs
47 very well. In fact, the average estimated marginal effects from both OP and ZIOP models

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
14 M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]]

1 Table 4
Monte Carlo results under OP: Experiments 4–7—marginal effects
3 Experiment 4 TRUE OP ZIOP TRUE OP ZIOP
Prðy ¼ 0jz; xÞ Prðy ¼ 1jz; xÞ
5
x1 ¼ z1 ME 0.000 0.000 0.000 0.373 0.376 0.376
CP 0.851 0.856 1.000 1.000
7 x2 ME 0.000 0.000 0.000 0.000
CP 0.999 0.999
9 Prðy ¼ 2jz; xÞ Prðy ¼ 3jz; xÞ

x1 ¼ z1 ME 0.216 0.219 0.219 0.157 0.157 0.157


11
CP 0.893 0.994 1.000 1.000
x2 ME 0.000 0.000 0.000 0.000
13 CP 0.999 0.999

F
Experiment 5
15 Prðy ¼ 0jz; xÞ Prðy ¼ 1jz; xÞ

O
x1 ME 0.000 0.0001 0.000 0.0001
17 CP 0.999 0.999

O
z1 ME 0.0412 0.0412 0.0388 0.0171 0.0171 0.0157
19 CP 0.936 0.926 1 1

21
x1 ME
Prðy ¼ 2jz; xÞ

0.000 0
PR Prðy ¼ 3jz; xÞ

0.000 0
CP 1 1
23 z1 0.0227 0.0227 0.0218 0.0014 0.0015 0.0013
D
ME
CP 0.94 0.936 0.817 0.804
25
TE

Experiment 6
Prðy ¼ 0jz; xÞ Prðy ¼ 1jz; xÞ
27 x¼z 0.000 0.000 0.000 0.373 0.375 0.375
ME
EC

CP 0.859 0.906 0.999 1.000


29
Prðy ¼ 2jz; xÞ Prðy ¼ 3jz; xÞ

31 x¼z ME 0.216 0.217 0.218 0.157 0.158 0.157


R

CP 0.924 0.999 1.000 1.000

33
R

Experiment 7
Prðy ¼ 0jz; xÞ Prðy ¼ 1jz; xÞ
O

35 x1 ¼ z1 ME 0.000 0.000 0.000 0.373 0.374 0.374


CP 0.869 0.771 1.000 1.000
C

37 x2 ME 0.000 0.000 0.000 0.000


CP 0.996 0.996
N

x3 ME 0.000 0.000 0.000 0.000


39 CP 0.996 0.996
U

Prðy ¼ 2jz; xÞ Prðy ¼ 3jz; xÞ


41
x1 ¼ z1 ME 0.216 0.218 0.218 0.157 0.156 0.156
CP 0.917 0.997 1.000 1.000
43
x2 ME 0.000 0.000 0.000 0.000
CP 0.996 0.996
45 x3 ME 0.000 0.000 0.000 0.000
CP 0.996 0.996

47

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]] 15

1 Table 5
Monte Carlo results under OP: Experiments 4–7—summary statisticsa
3 Experiment 4 Experiment 5 Experiment 6 Experiment 7

5 OP ZIOP OP ZIOP OP ZIOP OP ZIOP

RMSE_P 0.015 0.018 0.014 0.015 0.014 0.016 0.014 0.019


7 (0.005) (0.005) (0.005) (0.005) (0.005) (0.005) (0.006) (0.006)
Correct 0.594 0.594 0.861 0.861 0.592 0.593 0.593 0.593
9 Vuong 0.00 0.01 0.00 0.00 0.00 0.00 0.00 0.03
LR 0.02 0.04 0.01 0.03
11 Hausman 0.01 0.01 0.01 0.01
AIC 0.791 0.209 0.743 0.257 0.834 0.166 0.724 0.276
BIC 1.000 0.000 0.998 0.002 1.000 0.000 1.000 0.000
13 CAIC 1.000 0.000 0.999 0.001 1.000 0.000 1.000 0.000
Prðr ¼ 1jxÞ – 1.000 – 0.999 – 1.000 – 1.000

F
15 Prðr ¼ 1jxi Þ – 0.993 – 0.998 – 0.994 – 0.993

O
95% range – (0.965,1) – (0.993,1) – (0.968,1) – (0.965,1)
17 a
95% Range refers to the empirical 95% range of Prðr ¼ 1jxi Þ.

O
19
PR
are almost identical, both being very close to the true ones. Table 5 shows why this is the
21 case; when the true data are OP, the zero split of the ZIOP in Eq. (2) rules that almost all
observations are from Regime 1 and Prðr ¼ 1jxÞ ! 1 (x0 b ! 1, Fðx0 bÞ ! 1Þ. As shown in
23 Table 5, the average probability of r ¼ 1 evaluated at sample means, Prðr ¼ 1jxÞ, is
D
between 0.999 and 1 for the four experiments. This probability averaged across all
25 individuals and all replications, Prðr ¼ 1jxi Þ, is also very close to unity, and the empirical
TE

95% range of this latter probability is between 0.965 and 1 for all four experiments. This is
27 a very favorable result indicating that even the (misspecified) ZIOP model is estimated
when the true model is OP, the parameter estimates of b are such that the ZIOP reduces to
EC

29 OP even the two models are not nested in the usual sense.8
In terms of RMSE for the predicted probabilities, OP does slightly better than ZIOP.
31 Again both models have near identical performance in the percentage of correct
R

predictions. The LR statistic again, somewhat surprisingly given its lack of theoretical
33 justification, appears to work very well, with empirical sizes ranging from 1% to 4% (at
R

nominal 5% level). Similarly the Hausman statistic is only marginally undersized with
O

35 empirical sizes of 1%. The information criteria BIC and CAIC correctly choose the OP
model in 99.8% or more of cases, while AIC does so in more than 72.4% of instances. In
C

37 other words, the information criteria and LR and Hausman statistics appear to be able to
choose the correct model when the true data are not from a zero split d.g.p..
N

39 On the other hand, the Vuong statistic appears to be unable to choose between the two
non-nested models. Recall that the test statistic is bidirectional. For the bulk of all
U

41 experiments the test statistic falls in the ‘‘indeterminate’’ region ðjujo1:96Þ, leading to the
conclusion that neither of the two models is preferred. For all experiments, there is not one
43 single case where the true OP is chosen (u41:96), and for around 1–3% of the time the
ZIOP is actually preferred. These results are confirmed in the Q–Q plots in Fig. 1, which
45
8
This result was correctly conjectured by two anonymous referees, one of whom also suggested that as N ! 1
47 so will bb1 .

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
16 M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]]

1 4 4
Theoretical quantiles

Theoretical quantiles
3 2 2

5
0 0

7
-2 -2
9
-4 -4
11 -4 -2 0 2 4 -4 -2 0 2 4
Sample quantiles Sample quantiles
13 experiment 4 experiment 5

F
4 4
15
Theoretical quantiles

Theoretical quantiles

O
2 2
17

O
19 0 0

21 -2 -2
PR
23 -4 -4
D
-4 -2 0 2 4 -4 -2 0 2 4
25 Sample quantiles Sample quantiles
TE

experiment 6 experiment 7
27
Fig. 1. Vuong statistic: theoretical versus empirical quantiles under OP.
EC

29
contains plots of empirical quantiles of the Vuong statistic against theoretical ones (of a
31 standard normal distribution) for the four experiments. For the sample size and parameter
R

settings in these experiments, the plots are clearly not close to the 45 dotted lines where
33 the empirical and theoretical quantiles concur. In fact, for three of the four experiments,
R

OP is unlikely to be chosen as the Vuong statistic seems to stay negative. The empirical 5%
O

35 critical values for claiming the ZIOP model for Experiments 4–7 are 1:65, 0:86, 1:29
and 1:84, compared to the theoretical critical value of 1:96. We also note here that the
C

37 empirical Vuong statistic in Experiments 1–3, where ZIOP models are the true d.g.p., have
large negative values that are below 5. This is shown in the Q–Q plots for Experiments
N

39 1–3 in Fig. 2.
In summary of the Monte Carlo results when the true d.g.p. is ZIOPC/ZIOP, both zero-
U

41 inflated models perform well. The Vuong test, LR and Hausman statistics, as well as the
information criteria should all correctly select the ZIOPC/ZIOP models. In choosing
43 between the ZIOP and ZIOPC models, a standard t-test on the estimated value of r should
be used. When the data has been generated according to an OP process, estimation of a
45 ZIOP model will still yield accurate estimates of the quantities of interest, with the
probability for ‘Regime 1’ tending to unity in the split decision such that ZIOP tends to
47 OP. Moreover, the information criteria are likely to correctly select the smaller model (with

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]] 17

1 4 4 4

Theoretical quantiles

Theoretical quantiles
Theoretical quantiles

2 2 2
3
0 0 0
5
-2 -2 -2
7
-4 -4 -4
-10 -5 0 -10 -8 -6 -4 -2 0 2 4 -10 -5 0
9 Sample quantiles Sample quantiles Sample quantiles
Experiment 1 Experiment 2 Experiment 3
11
Fig. 2. Vuong statistic: theoretical versus empirical quantiles under ZIOP(C).
13

F
15 BIC and CAIC all of the time and AIC over 70% of the time). Similarly, somewhat

O
surprisingly, both of the Hausman and LR tests have empirical sizes very close to nominal
17 ones, with the latter having marginally better performance. There is evidence, however,

O
that the Vuong statistic is heavily biased towards the more heavily parameterised models. It
19 chooses the ZIOP models all of the time when they are true, but fails to choose the true OP
model with the bulk of the values falling in the indeterminate range. Of course, as with any
21 Monte Carlo experiments, all of the above results are conditional on the specific
experimental designs.
PR
23
D
4. An application to tobacco consumption
25
TE

Cigarette smoking has long been acknowledged as a public health issue. Yet a significant
27
proportion of the population in both developed and developing countries smoke. Large
EC

amounts of public funds are spent worldwide on educational programs and promotional
29
campaigns to reduce cigarette consumption. Empirical studies are crucial to help identify
the socioeconomic and demographic factors associated with smoking, providing invaluable
31
R

information to facilitate well-targeted public health policies.


33
R

4.1. The data


O

35
The data we use for the model are from the Australian National Drug Strategy
C

37 Household Survey (NDSHS, 2001). In this data set, neither the monetary expenditures nor
the physical quantities of tobacco consumed are reported. The information on individuals’
N

39 consumption of tobacco is given via a discrete variable measuring the intensity of


consumption. There have been seven surveys conducted through the NDSHS since 1985.
U

41 The surveys collect information from individuals aged 14 and over on attitudes and
consumption of several legal and illegal drugs. Measures have been put in place in the
43 surveys to ensure confidentiality in order to reduce under reporting. In this paper, data
from the three most recent surveys of 1995, 1998 and 2001 are used which involve a total of
45 over 40,000 respondents. After removal of missing values, a sample of 28,813 individuals is
used for estimation. This data set has been used in several previous studies (Cameron and
47 Williams, 2001; Williams, 2003; Zhao and Harris, 2004).

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
18 M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]]

1 Definitions of all variables used in the study are given in the Appendix. In particular, the
information in the data concerning an individual’s consumption of tobacco is collected
3 through the question ‘‘How often do you now smoke cigarettes, pipes or other tobacco
products?’’, where the responses take the form of one of the following choices: not at all
5 ðy ¼ 0Þ; smoking less frequently than daily ðy ¼ 1Þ; smoking daily with less than 20
cigarettes per day ðy ¼ 2Þ; and smoking daily with 20 or more cigarettes per day ðy ¼ 3Þ.
7 Table 6 presents some summary statistics on the observed smoking intensities. On
average around 76% of individuals identify themselves as current non-smokers. With the
9 way the survey questions are asked, these self-identified non-smokers will include genuine
non-smokers, recent quitters, infrequent smokers who are not currently smoking, as well as
11 potential smokers who might smoke when, say, the price falls. It could also be argued that
these observations may include some misreporting respondents who prefer to identify
13 themselves as non-smokers. The choices of consumption intensities are clearly ordered,
thus presenting a good case for the ZIOP(C) model(s) in order to identify the different

F
15 types of zero observations and their potentially different driving factors.

O
The participation decision of Eq. (1) is likely to be driven by factors relating to
17 individuals’ attitudes towards smoking and health concerns. Thus, r is likely to be related

O
to the individuals’ education levels and other standard socio-demographic variables such
19 as income, marital status, age, gender and ethnic background that capture socioeconomic

21
PR
status. There are also studies in the literature suggesting a significant growth in the
smoking prevalence of young females (Boreham and Shaw, 2002). To allow for the recent
rise in participation rates among young females, a dummy variable for young females
23 (defined as females under 25 years of age) is interacted with a time variable and included in
D
x in Eq. (1).
25 In terms of the decision of the levels of consumption conditional on participation,
TE

economists have typically followed a standard consumer demand framework with special
27 characteristics for addictive goods. Much work has been undertaken applying (Becker and
Murphy, 1988) theory of rational addiction to smoking to explain consumer behaviour in
EC

29 terms of an individual’s stock of addiction from past smoking (see Becker and Stigler,
1977; Chaloupka, 1991). Here, for the explanatory variables z in Eq. (3), we include
31 standard demand-schedule variables such as income and own- and cross-drug prices. The
R

related drug prices are included as there is evidence that certain drugs, in particular
33 marijuana and alcohol, act as either compliments or substitutes to tobacco (see, for
R
O

35
Table 6
C

37 Summary of consumption frequencies


N

39 1995 1998 2001 Combined


U

N % N % N % N %
41
Tobacco
Non-smoker 2644 72.4 7047 72.1 20113 78.0 29804 76.0
43
Weekly or less 120 3.3 504 5.2 937 3.6 1561 4.0
Daily, less than 20/day 600 16.4 1472 15.1 3351 13.0 5423 13.8
45 Daily, more than 20/day 286 7.8 749 7.7 1376 5.3 2411 6.2
Total 3650 100 9772 100 25777 100 39199 100
47

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]] 19

1 example, Cameron and Williams, 2001; Zhao and Harris, 2004). Data for marijuana prices
were obtained from information provided by the Australian Bureau of Criminal
3 Intelligence (ABCI, 2002) and the Australian Crime Commission (ACC, 2003). They are
collected quarterly and are based on information supplied by covert police units and police
5 informants. The consumer price indexes for tobacco and alcoholic drinks are obtained
from the Australian Bureau of Statistics (ABS, 2003) for individual states. In addition,
7 standard social demographic factors are also included in z to capture any heterogeneity in
consumption behaviour among smokers.
9 Note that we allow the age factor to enter both equations. The participation decision is
allowed to relate non-linearly to age by including age in natural logarithmic form.
11 However, in the intensity of consumption equation, following a Becker and Murphy (1988)
rational addiction approach, the likelihood that the age-consumption profile will be ‘‘n-
13 shaped’’ is allowed for by including both linear and quadratic terms for age.

F
15

O
17 4.2. The results

O
19 In Table 7 we present some summary statistics from three models: an OP model

21
PR
conditional on z and treating all observed zeros indifferently; a ZIOP model conditional on
both x and z that allows zero observations to come from two distinct sources; and a
ZIOPC that also allows for correlation across the two error terms. Results for some
23 ancillary parameters are also presented. As the magnitudes of the estimated coefficients of
D
b and g are somewhat meaningless, they are not presented here.9
25 The LR statistics clearly reject the OP model here as does the Hausman one for the
TE

correlated version (the statistic for the uncorrelated version yielded a negative value).
27 Furthermore, all of the information criteria, as well as the Vuong test, clearly suggest
superiority of the ZIOP and ZIOPC models over the OP one. With the exception of AIC,
EC

29 the information criteria marginally favour the uncorrelated variant (ZIOP) to the
correlated one (ZIOPC). Moreover, a Wald test on the estimated value of r also suggests
31 that the correlation is not statistically significant.
R

The results are presented as marginal effects on the choice probabilities in Tables 8 and
33 9. For comparison purposes, we also include the results of a simple probit model to
R

compare results on participation (Table 8). We only present those for the ZIOPC model as
O

35 the ZIOP ones were very similar. Note that for variables appearing in both x and z, we
have combined the two parts of the marginal effects, following Eq. (14). In Table 8, we
C

37 present marginal effects on Prðy ¼ 0Þ using a ZIOPC model and compare them with the
results from the probit and OP models. For the ZIOPC model, we also decompose the
N

39 overall marginal effect on Prðy ¼ 0Þ into two parts: the effect on non-participation Prðr ¼
0Þ and the effect on participation with zero consumption Prðr ¼ 1; y~ ¼ 0Þ. In Table 9, we
U

41 present marginal effects on the unconditional probabilities of all three positive levels of
smoking ðy ¼ 1; 2; 3Þ, using an OP model versus the ZIOPC model.
43 9
Details for the estimated coefficients are in Harris and Zhao (2004). Coefficients for several covariates exhibit
opposite signs in b and c: income has a positive coefficient in c for conditional consumption but a negative one in b
45 for participation, while the sole income coefficient in either a probit or an OP model is positive. Note also that,
although the price variables are not included in x on a priori grounds, if included they were individually and
47 jointly insignificant.

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
20 M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]]

1 Table 7
Tobacco consumption: some estimated coefficients and summary statisticsa
3 OP ZIOP ZIOPC

5 m1 0.155 ð0:004Þ 0.273 ð0:011Þ 0.272 ð0:011Þ


m2 0.920 ð0:011Þ 1.387 ð0:031Þ 1.383 ð0:031Þ
r 0.068 (0.222)
7
‘ðhÞ 21,995 21,628 21,626
AIC 44,012 43,292 43,291
9 LR versus OP 734 738
Hausman versus OP 689 1,259
11 BIC 44,206 43,635 43,643
CAIC 44,227 43,672 43,681
Vuong: versus OP 13.830 13.890
13 Time 8.1 25.9 224.8

F
a
15 Significant at 5% ( ) and 10% ( ). Preferred model with regard to each information criteria is indicated with

O
bold.

17

O
19 The marginal effects in Tables 8 and 9 highlight some interesting differences from

21
PR
alternative models for some explanatory factors, such as Ln(Income), Pre-School, School,
Young Female and Study. A key example is the effect of income. This variable clearly acts
as a social class proxy in the participation decision and, accordingly, one would expect a
23 priori it to be positively associated with non-participation. Results based on the ZIOPC
D
model suggest that a 10% increase in personal income results in a 0.0027 rise in the
25 probability of non-participation, but a 0.0017 fall in the probability of participation with
TE

zero consumption. This latter effect indicates that tobacco is a normal good for
27 participants. Overall, there is a 0.001 net positive effect on the probability of observing
zero consumption for a 10% increase in personal income. However, basing policy advise
EC

29 on the probit (or OP) model results, one would conclude that income is positively related
to participation as well as higher consumption.
31 Another example is the marginal effect of Study. Using simple probit and OP models,
R

one would conclude that people who mainly study are more likely to be non-smokers (by
33 0.098 and 0.128, respectively, Table 8). With a single latent equation, we assume, in the
R

case of OP, that there is a homogenous ‘study’ effect that affects an individual moving
O

35 from non-smokers to smokers of higher levels (y ¼ 0; 1; 2; 3Þ in the same direction.


However, when a ZIOPC model is used, we assume that the observed smoking categories
C

37 are the result of two distinct decisions of ‘participation’ and conditional ‘levels of
consumption conditional on participation’, on which Study can have opposite effects.
N

39 Indeed, as shown in Table 8, the ZIOPC estimates that Study has a positive effect on
participation decision but a negative effect on levels of consumption, leading to a
U

41 statistically insignificant total effect of observing a zero ðy ¼ 0Þ outcome when the


opposing effects cancel each other out. This contrasts the positive effects on non-
43 participation by both the probit and OP models. In addition, the resulting marginal effects
on the unconditional probabilities of levels of smoking ðPrðy ¼ jÞ; j ¼ 1; 2; 3Þ in ZIOPC is
45 also the result of two sources: MEs on participation, Prðr ¼ 1Þ, and MEs on levels of
smoking conditional on participation, Prðy ¼ jjr ¼ 1Þ. For example, the 0:032 ME of
47 Study on heavy smoking, Prðy ¼ 3Þ, in Table 9 is the combined result of opposing effects of

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]] 21

1 Table 8
Tobacco consumption: marginal effect for non-participation and zero consumption
3 ZIOPC

5 Probit OP Non-participation Zero consumption Full

Prðy ¼ 0Þ Prðy ¼ 0Þ Prðr ¼ 0Þ Prðr ¼ 1; ye ¼ 0Þ Prðy ¼ 0Þ


7
Young female 0.021 – 0.059 0.023 0.036
9 (0.004)** – (0.025)** (0.010)** (0.015)**
Actual age 0.006 0.004 0.015 0.009 0.006
(0.000)** (0.000)** (0.001)** (0.001)** (0.000)**
11 Ln (Income) 0.014 0.003 0.027 0.017 0.010
(0.004)** (0.004) (0.009)** (0.006)** (0.005)**
13 Male 1 0.018 0.047 0.095 0.023 0.072
(0.006)** (0.005)** (0.013)** (0.009)** (0.007)**

F
15 Married 1 0.082 0.099 0.160 0.039 0.121
(0.006)** (0.005)** (0.013)** (0.010)** (0.007)**

O
Pre-school 1 0.012 0.009 0.054 0.019 0.035
17 (0.008) (0.007) (0.019)** (0.014) (0.009)**

O
Capital 1 0.007 0.012 0.008 0.020 0.012
19 (0.006) (0.005)** (0.012) (0.010)** (0.007)*
Work 1
21 Unemployed 1
0.002
(0.008)
0.129
0.053
(0.008)**
0.044
0.008
(0.017)
0.061
PR 0.049
(0.014)**
0.004
0.041
(0.009)**
0.057
(0.018)** (0.014)** (0.032)* (0.022) (0.018)**
23 Study 1 0.098 0.128 0.194 0.182 0.012
D
(0.010)** (0.012)** (0.061)** (0.033)** (0.032)
25 English 1 0.046 0.057 0.063 0.004 0.067
TE

(0.011)** (0.012)** (0.028)** (0.021) (0.014)**


Degree 1 0.150 0.184 0.080 0.128 0.209
27 (0.006)** (0.008)** (0.022)** (0.019)** (0.010)**
EC

Diploma 1 0.038 0.048 0.027 0.036 0.063


29 (0.007)** (0.007)** (0.014)* (0.012)** (0.008)**
Year 12 1 0.056 0.069 0.020 0.060 0.079
(0.007)** (0.007)** (0.018) (0.014)** (0.010)**
31 School 1
R

0.174 0.170 0.002 0.093 0.091


(0.009)** (0.019)** (0.100) (0.046)** (0.058)
33 LnðPA Þ
R

– 0.322 – 0.300 0.300


– (0.078)** – (0.071)** (0.071)**
35 LnðPM Þ – 0.002 –
O

0.004 0.004
– (0.012) – (0.010) (0.010)
LnðPT Þ – 0.156 – 0.145 0.145
C

37 – (0.020)** – (0.019)** (0.019)**


N

39
U

41 Study on the two decisions. This contrasts the ME of 0:043 from an OP with only one
source of impact.
43
4.3. Model evaluation
45
In Fig. 3 we present the observed sample proportions, average predicted probabilities
47 and the probabilities evaluated at average covariates using a ZIOPC model for the four

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
22 M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]]

1 Table 9
Tobacco consumption: marginal effects for non-zero consumption levels
3 OP ZIOPC OP ZIOPC OP ZIOPC

5 Prðy ¼ 1Þ Prðy ¼ 1Þ Prðy ¼ 2Þ Prðy ¼ 2Þ Prðy ¼ 3Þ Prðy ¼ 3Þ

Young female – 0.006 – 0.021 – 0.008


7 – (0.003)** – (0.009)** – (0.003)**
Actual age 0.001 0.002 0.002 0.004 0.001 0.000
9 (0.000)** (0.000)** (0.000)** (0.000)** (0.000)** (0.000)**
LnðincomeÞ 0.000 0.003 0.002 0.007 0.001 0.000
(0.000) (0.001)** (0.002) (0.003)** (0.001) (0.002)
11
Male 1 0.006 0.010 0.026 0.042 0.016 0.021
(0.001)** (0.001)** (0.003)** (0.004)** (0.002)** (0.005)**
13 Married 1 0.012 0.016 0.054 0.070 0.033 0.035
(0.001)** (0.002)** (0.003)** (0.005)** (0.002)** (0.005)**

F
15 Pre-school 1 0.001 0.006 0.005 0.021 0.003 0.009
(0.002)** (0.006)** (0.004)**

O
(0.001) (0.004) (0.002)
Capital 1 0.001 0.001 0.007 0.005 0.004 0.008
17 (0.001)** (0.001) (0.003)** (0.004) (0.002)** (0.003)**

O
Work 1 0.006 0.003 0.029 0.018 0.018 0.024
19 (0.001)** (0.002) (0.004)** (0.009)** (0.003)** (0.007)**

21
Unemployed 1

Study 1
0.005
(0.002)**
0.015
0.006
(0.004)
0.025
0.024
(0.008)**
0.070
PR0.032
(0.011)**
0.021
0.015
(0.005)**
0.043
0.020
(0.008)**
0.032
(0.002)** (0.007)** (0.007)** (0.023) (0.004)** (0.012)**
23 English 1 0.007 0.006 0.031 0.038 0.019 0.025
D
(0.001)** (0.003)* (0.006)** (0.009)** (0.004)** (0.007)**
25 Degree 1 0.022 0.003 0.101 0.107 0.061 0.099
TE

(0.001)** (0.002) (0.005)** (0.006)* * (0.003)** (0.005)**


Diploma 1 0.006 0.001 0.026 0.032 0.016 0.030
27 (0.001)** (0.002) (0.004)** (0.005)** (0.002)** (0.004)**
EC

Year 12 1 0.008 0.000 0.038 0.040 0.023 0.040


29 (0.001)** (0.002) (0.004)** (0.006)** (0.002)** (0.004)**
School 1 0.020 0.004 0.093 0.044 0.057 0.051
(0.002)** (0.011) (0.011)** (0.035) (0.007)** (0.013)**
31
R

LnðPA Þ 0.038 0.011 0.177 0.144 0.107 0.166


(0.009)** (0.004)** (0.043)** (0.035)** (0.026)** (0.039)**
33 LnðPM Þ
R

0.000 0.000 0.001 0.002 0.001 0.002


(0.001) (0.000) (0.006) (0.005) (0.004) (0.006)
O

35 LnðPT Þ 0.018 0.005 0.085 0.070 0.052 0.081


(0.002)** (0.002)** (0.011)** (0.009)** (0.007)** (0.010)**
C

37
N

39 smoking categories, which we denote zero, low, moderate and high. For Prðy ¼ 0Þ, we also
present the probability of zeros arising from the regime of non-participation ðr ¼ 0Þ as the
U

41 ‘‘selection component’’ . As can be seen, the model fits the data well in terms of mimicking
these sample proportions. Moreover, it is clear that the bulk of the probability mass of
43 Prðy ¼ 0Þ comes from non-participants with r ¼ 0.
In Fig. 4 the overall age-smoking profile is plotted. The expected n-shaped profile is
45 clearly evident for the smokers, and most pronounced for moderate and high levels of
consumption. In terms of the probability of zero consumption, this finds a nadir at around
47 the mid-late 20’s years of age, with the probability reaching nearly 0.9 at age 66.

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]] 23

1 0.8
Sample
3 0.7
Average Probabilities

Average Covariates
0.6
5 Selection Component

0.5
Probability

7
0.4
9
0.3
11
0.2

13 0.1

F
15 0

O
Zero Low Moderate High
17 Tobacco Consumption

O
Fig. 3. Observed and predicted probabilities.
19

21 0.9
Zero Consumption
PR
0.8 Low Consumption
23
D
Moderate Consumption
0.7 High Consumption
25
TE

0.6
Probability

27 0.5
EC

29 0.4

0.3
31
R

0.2
33
R

0.1
O

35 0
16 21 26 31 36 41 46 51 56 61 66
C

37 Age
N

Fig. 4. Predicted probabilities by age.


39
U

41
As previously mentioned, one of the advantages of a ZIOPC is its ability to disentangle
43 the total effect of a covariate on Prðy ¼ 0Þ into those effects on the probabilities of the two
types of zeros: Prðr ¼ 0Þ and Prðe y ¼ 0; r ¼ 1Þ. This is illustrated in Fig. 5 for the effect of
45 age. At younger ages, Prðy ¼ 0Þ is dominated by potential consumers. However, as age
increases and participation rates decline, the bulk of Prðy ¼ 0Þ comes from genuine non-
47 participation.

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
24 M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]]

1 100%

Zero Consumption for Participants


3
80%
5

7
60%
9
% of P(0)

11
40%
13

F
Non-Participation
15
20%

O
17

O
19 0%

21 Age
PR
16 18 20 22 24 26 28 30 32 34 36 38 40 42 44 46 48 50 52 54 56 58 60 62 64 66

Fig. 5. Decomposition of Pðy ¼ 0Þ into component parts.


23
D

25 While the results in Table 7 clearly suggest preference of the ZIOP(C) model(s) over the
TE

OP model, as a specification test, we also experimented with the more flexible models of
27 multinomial logit (MNL) and multinomial probit (MNP). Without imposing restrictions
in the correlation matrix, we encountered convergence problems with MNP using standard
EC

29 econometric software. The MNL model was found inferior on the basis of both BIC and
CAIC. We also compared the log-likelihood values of the ZIOPC and MNL models, which
31 were very close (21; 626 and 21; 601, respectively), while the MNL has 20 additional
R

parameters. In addition, Hausman tests clearly rejected the embodied independence of


33 irrelevant alternatives property of the MNL model, and the smoking levels data here are
R

quite clearly ordered in nature.


O

35
5. Conclusions
C

37
We propose a model for ordered discrete data that allows for the observed zero
N

39 observations to be generated by two different behavioural regimes. Following double-


hurdle and zero-inflated models, we extend the OP model to a zero-inflated OP model
U

41 using a system of two latent equations with potentially different covariates. We also allow
for the likely correlation between the two latent equations. The Monte Carlo experiments
43 suggest that the model performs well in finite samples. Although not strictly valid, both the
LR and Hausman statistics appear to provide useful general specification tests against the
45 simpler OP model. The former may be preferred in terms of ease of computation and better
finite sample performance. On the other hand, the Vuong test does not appear to have
47 favourable small sample properties and tends to favour the ZIOP models, whilst

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]] 25

1 information criteria seem to have good empirical properties with regard to selecting the
correct model. If the d.g.p. is not zero split, the LR and Hausman statistics and the
3 suggested model selection procedures based on information criteria should all correctly
pick the OP model. However, even in this case estimation of a ZIOP model will still
5 provide accurate results in terms of the estimated marginal effects, as the probabilities for
the two regimes tend to zero and unity in the split decision such that ZIOP probabilities
7 tend to OP ones. However, if the d.g.p. is zero split, the evidence is that the ZIOP model
will be correctly selected if any of the LR and Hausman statistics, Vuong test or
9 information criteria are used. With regard to differentiating between the ZIOP and
ZIOPC, a Wald test of r ¼ 0 would be preferred to the information based model selection
11 procedures.
The models are applied to discrete data of tobacco consumption from a national survey
13 from Australia. The empirical application demonstrates the advantages of the ZIOP(C)
model in separating the different behavioural schemes for participants and non-

F
15 participants. In particular, we allow for the split of the observed non-users (‘‘ zeros’’ )

O
into two groups: those of non-participants who choose not to smoke due to health
17 concerns or other non-economic factors and those zero consumption potential users who

O
may be the result of a demand-schedule corner solution and are therefore responsive to
19 economic factors such as prices and income. The example shows that the use of a

21 that have opposing impacts on the two schemes.


PR
conventional OP model would confuse the effects of some important explanatory variables

The ZIOP(C) model has important advantages over the conventional OP model. It can
23 be used to estimate the proportion of zeros coming from each regime, and how this split
D
changes with observed characteristics. Finally, the proposed model allows for the
25 identification of variables that are important in each regime. This is potentially very
TE

important for policy analysis.


27
Acknowledgements
EC

29
We wish to thank John Gweke, three anonymous referees and one anonymous Associate
31 Editor, as well as William Greene, Don Poskitt, Max King, Rob Hyndman, Tim Fry and
R

Brett Inder for constructive comments. The usual caveats apply. Excellent research
33 assistance was provided by Preety Ramful and Xiaohui Zhang. Funding from the
R

Australian Research Council is kindly acknowledged by both the authors.


O

35
Appendix A. Definition of variables
C

37
N

39 y Levels of tobacco consumption; y ¼ 0 if not current smoker, y ¼ 1 if


U

smoking weekly or less, y ¼ 2 if smoking daily with less than 20


41 cigarettes per day, and y ¼ 3 if smoking daily with 20 or more cigarettes
per day.
43 Ln(Age) Logarithm of actual age.
Age actual age divided by 10.
45 Age square Age squared and divided by 10.
Male 1 for male; and 0 for female.
47

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
26 M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]]

1 Married 1 if married or de facto; and 0 otherwise.


Pre-school 1 if the respondent has pre-school aged child/children, and 0 otherwise.
3 Capital 1 if the respondent resides in a capital city, and 0 otherwise.
Work 1 if mainly employed; and 0 otherwise.
5 Unemployed 1 if unemployed; and 0 otherwise.
Study 1 if mainly study; and 0 otherwise.
7 Other 1 if retired, home duty, or volunteer work; and 0 otherwise. This variable
is used as the base of comparison for work status dummies and is
9 dropped in the estimation.
English 1 if English is the main language spoken at home for the respondent, and
11 0 otherwise.
Degree 1 if the highest qualification is a tertiary degree, and 0 otherwise.
13 Diploma 1 if the highest qualification is a non-tertiary diploma or trade certificate,
and 0 otherwise.

F
15 Year 12 1 if the highest qualification is Year 12, and 0 otherwise.

O
School 1 if still studying in school, and 0 otherwise.
17 Noqual 1 if the highest qualification is below Year 12, and 0 otherwise. This

O
variable is used as the base of comparison for education dummies and is
19 dropped in the estimation.
LnðPT Þ
21 LnðPA Þ
LnðPM Þ
PR
Logarithm of real price index for tobacco, divided by 10.
Logarithm of real price index for alcoholic drinks, divided by 10.
Logarithm of real price for marijuana measured in dollars per ounce,
23 divided by 10.
D
Ln(Income) Logarithm of real personal annual income before tax measured in
25 thousands of Australian dollars, divided by 10.
TE

Young female A binary dummy for female aged 25 years or younger, interacted with an
27 annual time trend t ¼ 1; 2; 3.
EC

29

31
R

References
33
R

ABCI, 2002. Australian Illicit Drug Report. Australian Bureau of Criminal Intelligence.
ABS, 2003. Consumer Price Index 14th Series: by region by group, sub-group and expenditure class—alcohol and
O

35 tobacco, Cat. No. 6455.0.40.001, Australian Bureau of Statistics.


ACC, 2003. Australian Illicit Drug Report 2001-02, Australian Crime Commission.
C

37 Andrews, D., Ploberger, W., 1995. Admissibility of the likelihood ratio test when a nuisance parameter is present
only under the alternative. The Annals of Statistics 23 (5), 1609–1629.
N

39 Becker, G., Murphy, K., 1988. A theory of rational addiction. Journal of Political Economy 96 (4), 675–700.
Becker, G., Stigler, G., 1977. De gustibus non est disputandum. American Economic Review 68 (1), 76–90.
U

Boreham, R., Shaw, A., 2002. Drugs Use, Smoking and Drinking Among Young People in England in 2001. The
41 Stationary Office.
Cameron, C., Trivedi, P., 1998. Regression Analysis of Count Data. Econometric Society Monographs.
43 Cambridge University Press, Cambridge, UK.
Cameron, L., Williams, J., 2001. Cannabis alcohol and cigarettes: substitutes or compliments. The Economic
Record 77 (236), 19–34.
45 Chaloupka, F., 1991. Rational addictive behaviour and cigarette smoking. Journal of Political Economy 99 (4),
722–742.
47 Chesher, A., Smith, R., 1997. Likelihood ratio specification tests. Econometrica 65 (3), 627–646.

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002
ECONOM : 2910
ARTICLE IN PRESS
M.N. Harris, X. Zhao / Journal of Econometrics ] (]]]]) ]]]–]]] 27

1 Cragg, J., 1971. Some statistical models for limited dependent variables with application to the demand for
durable goods. Econometrica 39, 829–844.
3 Greene, W., 1994. Accounting for excess zeros and sample selection in Poisson and negative binomial regression
models. Working Paper EC-94-10, Stern School of Business, New York University, Stern School of Business,
New York University.
5 Greene, W., 2003. Econometric Analysis, fifth ed. Prentice-Hall, Englewood Cliffs, NJ, USA.
Harris, M., Zhao, X., 2004. Modelling tobacco consumption with a zero-inflated ordered probit model. Working
7 Paper 14/04, Monash University, Department of Econometrics and Business Statistics, Monash University,
Australia.
Hausman, J., 1978. Specification tests in econometrics. Econometrica 46, 1251–1271.
9 Heckman, J., 1979. Sample selection bias as a specification error. Econometrica 47, 153–161.
Heilbron, D., 1989. Generalized linear models for altered zero probabilities and overdispersion in count data.
11 Discussion paper, University of California, University of California, San Francisco.
Lambert, D., 1992. Zero inflated Poisson regression with an application to defects in manufacturing.
13 Technometrics 34, 1–14.
Maddala, G.S., 1983. Limited Dependent and Qualitative Variables in Econometrics. Cambridge University

F
Press, Cambridge, UK.
15 Marcus, A., Greene, W., 1985. The determinants of rating assignment and performance. Discussion Paper

O
Working Paper CRC528, Center for Naval Analyses.
17 Mullahey, J., 1986. Specification and testing of some modified count data models. Journal of Econometrics 33,

O
341–365.
Mullahey, J., 1997. Heterogeneity, excess zeros and the structure of count data models. Journal of Applied
19

21
Econometrics 12, 337–350.
PR
NDSHS, 2001. Computer Files for the Unit Record Data from the National Drug Strategy Household Surveys.
Pohlmeier, W., Ulrich, V., 1995. An econometric model of the two-part decision-making process in the demand
for health care. Journal of Human Resources 30, 339–361.
23 Smith, M., 2003. On dependency in double-hurdle models. Statistical Papers 44, 581–595.
D
Vuong, Q., 1989. Likelihood ratio tests for model selection and non-nested hypotheses. Econometrica 57,
307–334.
25
TE

Williams, J., 2003. The effects of price and policy on marijuana use: what can be learned from the australian
experience. Health Economics 13 (2), 123–137.
27 Zavoina, R., McElvey, W., 1975. A statistical model for the analysis of ordinal level dependent variables. Journal
of Mathematical Sociology 4, 103–120.
EC

Zhao, X., Harris, M., 2004. Demand for marijuana, alcohol and tobacco: participation frequency, and cross-
29 equation correlation. Economic Record 80 (251), 394–410.
R
R
O
C
N
U

Please cite this article as: Harris, M.N., Zhao, X., A zero-inflated ordered probit model, with an application to
modelling tobacco consumption. Journal of Econometrics (2007), doi:10.1016/[Link].2007.01.002

View publication stats

You might also like