** Ron Kohavi – DRAFT **
[Link]
Please free to comment in the margins using Review->New comment in Word
The Overall Evaluation Criterion (OEC)
If you can't measure it, you can't improve it
-- Peter Drucker
Tell me how you measure me, and I will tell you how I will behave
-- Eliyahu M. Goldratt (1990)
Why you care: Determining the Overall Evaluation Criterion for your experiments is one of the
most important things you can define to clarify the goals and align the organization. It often
requires multiple iterations to adjust and refine the OEC.
When designing controlled experiments, one of the most important questions to ask early is:
what are you optimizing for? Getting agreement on what the goal is and how to measure it is a
huge step forward for the organization.
When we started work on optimizing a support site, we asked the team what they wanted to
optimize, and they were convinced that “time on site” was the key metric to improve. When
asked if more time is better (after all, this is a support site), there was a debate about the
direction: is more time better or worse?
Key Metrics
The good part about software is that online observations are easy to log with high quality. There
are no humans that need to enter observations by hand, and no need to “key-in” the data from a
survey, frequently an error-prone process. While attention needs to be paid to the proper
instrumentation of software, the marginal cost of logging additional observations is minimal.
Combining the observations into metrics is extremely important, and determining the key
metrics, or key performance indicators (KPIs) is critical. There are several great books about
metrics, measurements, and performance indicators (Spitzer, 2007; Parmenter, 2015;
McChesney, Covey, & Huling, 2012). Spitzer (2007) notes that “What makes measurement so
potent is its capacity to instigate informed action—to provide the opportunity for people to
engage in the right behavior at the right time.” In the context of controlled experiments, because
the Treatment is the cause (with high probability for highly statistically significant effects) of the
Ron Kohavi / 2
impact to each metric, formulating the key metrics is an assessment of the value of an idea (the
Treatment) on some axis of interest.
What are characteristics of good key metrics?
1. Key metrics should be hard to game. When given a numerical target, humans can be
quite ingenious, especially when the measures are tied to rewards. Spitzer (2007) has a
chapter “When Measurement Goes Bad” with some great examples
a. Vasili Alexeyev, a famous Russian super-heavyweight weight lifter, was offered
an incentive for every world record he broke. The result of this contingent
measurement was that he kept breaking world records a gram or two at a time to
maximize his reward payout!
b. A manager of a fast-food restaurant striving to achieve an award for attaining a
perfect 100 percent on the restaurant's “chicken efficiency” measure (the ratio of
how many pieces of chicken sold to the number thrown away) did so by waiting
until the chicken was ordered before cooking it. He won the award, but drove the
restaurant out of business because of the long wait times.
c. A company paid bonuses to its central warehouse spare parts personnel for
maintaining low inventory. As a result, necessary spare parts were not available in
the warehouse, and operations had to be shut down until the parts could be
ordered and delivered.
Parmenter (2015) has a chapter on “Unintended Consequence: The Dark Side of
Measures” and gives the following example
d. Managers at a hospital in the United Kingdom were concerned about the time it
was taking to treat patients in the accident and emergency department. They
decided to measure the time from patient registration to being seen by a house
doctor…The nursing staff thus began asking the paramedics to leave their patients
in the ambulance until a house doctor was ready to see them, thus improving the
“average time it took to treat patients.”
The Wikipedia entry on Perverse Incentive (Perverse Incentive, 2018) gives the following
examples:
e. In Hanoi, under French colonial rule, a program paying people a bounty for each
rat tail handed in was intended to exterminate rats. Instead, it led to the farming of
rats. A similar example is mentioned with regards to Cobra snakes, where
presumably the British government offered bounty for every dead cobra in Delhi
and enterprising people began to breed cobras for the income (Cobra Effect,
2018).
f. The Duplessis Orphans: Between 1945 and 1960, the federal Canadian
government paid 70 cents a day per orphan to orphanages, and psychiatric
hospitals received $2.25 per day, per patient. Allegedly, up to 20,000 orphaned
children were falsely certified as mentally ill so the Catholic Church could get
$2.25 per day, per patient.
g. Funding fire departments by the number of fire calls made is intended to reward
the fire departments that do the most work. However, it may discourage them
from fire-prevention activities, which reduce the number of fires.
Ron Kohavi / 3
2. Key metrics should be sensitive. Sensitivity depends on three factors: the statistical
variance of the underlying metric, the effect size (delta between treatment and control in
an experiment), and the number of experiment units (e.g., users). At the extreme, one
could run a controlled experiment and look at the stock price for the company. However,
the ability of routine product changes to impact the stock price during the experiment
period is practically zero, so the stock-price metric will not be sensitive. At the other
extreme, you could measure the existence of the new feature, and that will be very
sensitive, but not informative. Click-throughs on the new feature will be sensitive, but
will not capture the impact on the rest of the page, and possible cannibalization of other
features. A whole-page click-through metric, a measure of “success” (e.g., purchase),
and time to success, are usually good key metrics that are sensitive enough for
experimentation. See Measuring Metrics (Dmitriev & Wu, 2016) for more discussion on
sensitivity.
3. Improvements to the key metrics should have a causal relationship the key organizational
long-term goals. In the case of for-profit organizations, Kaplan and Norton wrote in the
Balanced Scorecard (1996) “Ultimately, causal paths from all the measures on a
scorecard should be linked to financial objectives.” Hauser and Katz (1998) write “the
firm must identify metrics that the team can affect today, but which, ultimately, will
affect the firm’s long-term goals.” Spitzer (2007) wrote that “measurement frameworks
are initially composed of hypotheses (assumptions) of the key measures and their causal
relationships. These hypotheses are then tested with actual data, and can be confirmed,
disconfirmed, or modified.” This characteristic is the hardest to satisfy, as we often do
not know the underlying causal model.
One suggestion that we have is to think of customer lifetime value as a guiding principle
for key metrics. For example, you could increase short-term revenues by raising prices,
or plastering a website with ads, but users will abandon, and customer lifetime value will
decline. A metric that measures ad revenue constrained to some space on the page is a
much better metric, as ad space can hurt the user experience. Think about metrics that
measure user value and avoid vanity metrics that indicate a count of your actions, which
users often ignore (e.g., count of banner ads is a vanity metric, whereas clicks on ads
indicates potential user interest).
Expect to define your key metrics and modify them over time, as experiments shed light on the
strengths and weaknesses of metrics. Hubbard in How to Measure Anything (2014) discusses a
concept called EVI: Expected Value of Information. This captures how additional information
can help you in decision making. Investing time and effort in investigating metrics and
modifying existing ones can have high EVI. It’s not enough to be agile and measure, you need
to make sure the metrics guide you in the right direction.
Ron Kohavi / 4
Combining Key Metrics into an OEC
Assume that you have a few key metrics defined. What do you do with them? Some books,
such as Lean Analytics (Croll & Yoskovitz, 2013) call for focusing on one; an OMTM: One
Metric That Matters. The 4 Disciplines of Execution (McChesney, Covey, & Huling, 2012) calls
for focusing on a WIG: a Wildly Important Goal. These are motivating but oversimplify the
problem.
Except for trivial scenarios, there is usually no single metric that captures what a business is
optimizing for. Kaplan and Norton (1996) give a good example: imagine entering a modern jet
airplane. Is there a single metric that you should put on the pilot’s dashboard? Is it airspeed?
Altitude? Fuel left? You want all of them, and more. When you have an online business, key
metrics are typically measuring user engagement (e.g., active days or sessions per user, clicks per
user) and monetary value (e.g., revenue per user); there is usually no simple single metric to
optimize for: we have multiple objectives.
Ideally, devising a single metric that is a weighted combination of such objectives is highly
desired and recommended (Roy, 2001, pp. 50, 405-429). Sports game do this regularly: you
don’t keep track of close and far basketball shots separately; a successful shot beyond a special
court line counts as three points; others two. Credit scores (FICO) combine multiple metrics into
a single score that ranges from 300 to 850. The ability to have a single summary score is critical
in sports games and can likewise help the business. A single metric makes the exact definition of
success clear, possibly through explicit weights. This approach empowers teams to make
decisions without having to escalate to management. It also opens up the opportunity for
automated searches (parameter sweeps).
If you have multiple metrics, one possibility proposed by Roy (2001) is to normalize each metric
to a predefined range, say 0-1, assign each a weight. Your OEC is the weighted sum of the
normalized metrics.
Coming up with a single weighted combination may be hard initially. In those cases, split the
decisions into four groups:
1. If all key metrics are flat (not stat-sig) or positive (stat-sig), when at least one
positive, then ship the change.
2. If all key metrics are flat or negative, with at least one negative, then don’t ship the
change.
3. If all key metrics are flat, then don’t ship the change and considering either increasing
the experiment power, failing fast, or pivoting.
4. If some key metrics are positive and some key metrics are negative, then decide given
the tradeoffs. When you have accumulated enough of these decisions, you may be
able to assign weights.
If you are unable to combine your key metrics into a single OEC, focus on minimizing the
number of key metrics. Pfeffer and Sutton (1999) warn about the Otis Redding problem, named
Ron Kohavi / 5
after the famous song “Sitting by the Dock of the Bay,” which has this line: “Can’t do what ten
people tell me to do, so I guess I’ll remain the same.” Having too many metrics causes
complexities and the organization is likely to ignore the key metrics. It also helps with the
multiple comparison problems in Statistics. Try to limit your key metrics to five.
Parmenter in Key Performance Indicators (2015) has the following diagram that emphasizes the
importance of aligning key metrics to your strategy.
Figure 1: It is important to align multiple metrics with the overall strategic direction
Depending on the org size and objectives, you may have multiple groups with their own OECs.
Example: OEC for E-mail
When one of us (Kohavi) worked at Amazon, and one of the areas he led was e-mail. A very
nice system was built to send e-mails based on “programs,” which targeted customers who met
certain conditions. For example,
If a new book came out, a program targeted all users who previously bought a book by
the author, e-mailing them that the author of a book they bought has a new book out.
There was a program that used Amazon’s recommendation algorithm. The e-mails start
with something like this: “[Link] has new recommendations for you based on
items you purchased or told us you own.”
There were many programs that were very specific and defined by humans: bought from
this category and that one, how about you buy this new widget.
The question is what OEC should be used for these programs? The initial OEC, or “fitness
function,” as it was called at Amazon, gave credit to a program based on the revenue it generated
from users clicking-through the e-mail.
There is a fundamental problem here: the metric is easy to game, as the metric is monotonically
increasing: spam users more, and at least some will click through, so overall revenue will
increase. This is likely true even if the revenue from the treatment of users who receive the e-
mail is compared to a control group that doesn’t receive the e-mail.
Ron Kohavi / 6
This was recognized as a problem, as users were complaining about being sent too many e-mails.
The initial solution was to put a constraint: a user can only receive an e-mail every X days. An
e-mail traffic cop was built, but the problem was that this becomes an optimization program:
which e-mail should a person receive every X days, if multiple e-mail programs want to target
the person. Is it even clear that some users wouldn’t be open to receiving more e-mails if they
were truly useful?
The key insight is that the click-through revenue OEC is optimizing for short-term revenue
instead of customer lifetime value. Users that are annoyed will unsubscribe, and Amazon then
loses the opportunity to target them in the future. A simple model was used to construct a lower
bound on the lifetime opportunity loss when a user unsubscribes. The OEC was thus
OEC = ∑ Rev i−∑ Rev j −s∗unsubscribe_lifetime_loss
i j
where i ranges over e-mail recipients in Treatment, j ranges over e-mail recipients in Control,
and s is the number of incremental unsubscribes, i.e., unsubscribes in Treatment minus Control
(one could debate whether it should have a floor of zero, or whether it’s possible that the
Treatment actually reduced unsubscribes), and unsubscribe_lifetime_loss was the estimated loss
of not being able to e-mail a person for “life.”
When this new OEC was implemented, with just a few dollars assigned to unsubscribes, more
than half the programs were negative!
More interestingly, the realization that unsubscribes have such a big loss led to a different
unsubscribe page, where the default was to unsubscribe from this “program,” not from all of e-
mails. The cost of an unsubscribe was therefore drastically diminished.
Example: OEC for Search Engines
Search engines (e.g., Google, Bing, Baidu) typically focus on two long-term metrics: query share
and revenue per search. There are significant efforts to increase these metrics, which is why the
following story, initially told in Trustworthy online controlled experiments: Five puzzling
outcomes explained (Kohavi, et al., 2012), is interesting practical. It is a great example where
short-term and long-term objectives diverge diametrically.
When Bing had a ranker bug, which resulted in very poor results being shown to users in
controlled experiment, the two key organizational metrics improved significantly: distinct
queries per user went up over 10%, and revenue per user went up over 30%! What should the
OEC for a search engine be?
Ron Kohavi / 7
This problem was interesting enough that it is now included in Data Science Interviews Exposed
(Huang, You, Wang, Cao, & Gao, 2015)
Clearly these long-term goals do not align with short-term measurements in experiments. If they
did, search engines would intentionally degrade quality to raise query share and revenue!
The degraded algorithmic results (the main search engine results shown to users, sometimes
referred to as the 10 blue links) force people to issue more queries (increasing queries per user)
and click more on ads (increasing revenues).
To understand the problem, we decompose query share. Monthly Query Share is defined as
distinct queries for the search engine divided by distinct queries for all search engines over a
month. Distinct queries per month can be decomposed into the product of three terms:
Users Sessions Distinct queries (1)
× × ,
Month User Session
where the 2nd and 3rd terms in the product are computed over the month, and a session is defined
as user activity that begins with a query and ends with 30 minutes of inactivity on the search
engine.
If the goal of a search engine is to allow users to find their answer or complete their task quickly,
then reducing the distinct queries per task is a clear goal, which conflicts with the business
objective of increasing share. Since this metric correlates highly with distinct queries per session
(more easily measurable than tasks), distinct queries alone should not be used as an OEC for
search experiments.
Given the decomposition of distinct queries shown in Equation 1, let’s look at the three terms
1. Users per month. In a controlled experiment, the number of unique users is going to be
determined by the design. For example, in an equal A/B test, the number of users that
fall into the two variants will be approximately the same. For that reason, this term
cannot be part of the OEC for controlled experiments.
2. Distinct queries per task should be minimized, but it is hard to measure. Distinct queries
per session is a surrogate metric that can be used. This is a subtle metric, however,
because increasing it may indicate that users have to issue more queries to complete the
task, but decreasing it may indicate abandonment. This metric should be minimized
subject to the task being successfully completed.
3. Sessions/user is the key metric to optimize (increase) in experiments, as satisfied users
will come more.
Revenue per user should likewise not be used as an OEC for search and ad experiments without
other constraints. When looking at revenue metrics, we want to increase them without negatively
impacting engagements metrics. A common constraint is to restrict the average number of pixels
that ads can use over multiple queries. Increasing revenue per search given this constraint is a
constraint optimization problem.
Ron Kohavi / 8
Non-Key Metrics
The OEC and key metrics are critical for decision making, but when they move, you will need a
lot of other metrics to help you understand why they moved: debugging.
1. If click-through rate is a key metric, then you might have 20 metrics to indicate clicks
on certain areas of the page.
2. If revenue is a key metric, you might want to decompose revenue into two metrics
o Revenue indicator: a Boolean (0/1) that indicates whether the user purchased
o Conditional Revenue: revenue if the user purchased, and Null otherwise.
When averaged, only the revenue from purchasing users is averaged.
Average overall revenue is the product of these two metrics, but each tells a different
story about revenue. Did it increase/decrease because more/less people purchased, or
because the average purchase price changed?
In addition to debugging, another set of metrics that are commonly defined are “guardrail”
metrics: metrics that should not degrade. These are discussed in @@
[Link]
While a typical scorecard will have an OEC and a few key metrics, it might have hundreds to
thousands of metrics to help debugging and identify surprises. For example, Bing has 6,000
metrics in its main “web” scorecard. These are all metrics that can then be broken down by
segments, such as browsers, markets, etc. Of course, with such a large number of metrics, one
must use lower p-values for detecting surprising impact. @@REF
The Dark Site of Metrics
Charles Goodhart, a British economist, originally wrote the “law” now attributed to him as “Any
observed statistical regularity will tend to collapse once pressure is placed upon it for control
purposes” (Goodhart, 1975; Chrystal & Mizen, 2001). The following phrasing is now more
commonly used: as Goodhart’s law: (Goodhart's law, 2018; Strathern, 1997).
Campbell’s law, named after Donald Campbell, states that “The more any quantitative social
indicator is used for social decision-making, the more subject it will be to corruption pressures
and the more apt it will be to distort and corrupt the social processes it is intended to monitor”
(Campbell's law, 2018; Campbell, 1979).
Ron Kohavi / 9
Lucas Critique (Lucas critique, 2018; Lucas, 1976) observes that relationships observed in
historical data cannot be considered structural, or causal. Policy decisions can alter the structure
of economic models and the correlations that held historically will no longer hold. The Phillips
Curve, for example, showed a historical negative correlation between inflation and
unemployment; over the study period of 1861-1957 in the United Kingdom, when inflation was
high, unemployment was low and vice versa (Phillips, 1958). Raising inflation in hopes that it
would lower unemployment is assuming an incorrect causal relationship. Indeed, in the 1973–
1975 recession, both inflation and unemployment increased. In the long-run, the current belief is
that the rate of inflation has no causal effect on unemployment (Hoover, 2008).
Tim Harford addresses the fallacy of using historical data by using the following example
(Harford, 2014, p. 147):
“Fort Knox has never been robbed, so we can save money by sacking the guards.” You can’t
look just at the empirical data, you need also to think about incentives.
Obviously, such a change in policy would cause robbers to re-evaluate their probability of
success, given the new policy.
What Goodhart’s Law, Campbell’s Law, and the Lucas Critique state is that correlation does not
imply causation. Finding correlations in historical data does not imply that you can pick a point
on a correlational curve by modifying one of the variables and expecting the other to change.
For that to happen, the relationship must be causal, which is why picking metrics for the OEC is
so hard.
Jerry Muller in the Tyranny of Metrics (Muller, 2018) writes about “metric fixation,” which is
the persistent belief that (i) it is possible and desirable to use metrics, (ii) making such metrics
public improves transparency and accountability, and (iii) use of metrics motivates people within
the organization. Muller then warns about the unintended negative consequences when metrics
are put into practice. A key point is not that metrics are bad, but that awareness of the limitations
and unintended negative consequences is important and can lead to the design of better metrics.
The book ends with a useful checklist and the following recommendation “the issue is not one of
metrics versus judgment, but metrics as informing judgment, which includes knowing how much
weight to give to metrics, recognizing their characteristic distortions, and appreciating what can’t
be measured.”
Acknowledgments
Thanks to David Manheim, Colin McFarland, Adil Aijaz, Kevin Anderson, Kris Jack, and Matt
Gershoff for feedback.
Ron Kohavi / 10
Misc Stuff
Add Netflix’s OEC from Designing with Data / Rochelle’s book, and get a quotation from Steve
Urban about insensitivity of churn.
“viewing hours”, or the amount of Netflix content that was consumed, was the strongest proxy
metric for retention for Netflix.
King, Rochelle; Churchill, Elizabeth F; Tan, Caitlin. Designing with Data: Improving the User
Experience with A/B Testing (Kindle Locations 2716-2717). O'Reilly Media. Kindle Edition.
Deming in Out of Crisis: the most important figures that one needs for management are unknown
or unknowable (Lloyd S. Nelson, director of statistical methods for the Nashua corporation), but
successful management must nevertheless take account of them.
[Link]
while so many employees wince at the mere thought of being measured at work, these same
people would be horrified at the prospect of golfing (or, for that matter, bowling, playing
baseball, tennis, football, or any other sport) without keeping score.
From Pavel’s paper: Conceptually, the following four groups of metrics have been recognized as
useful for analyzing experiments in previous research (refs) Success Metrics (the metrics that
feature teams should improve), Guardrail Metrics (the metrics that are constrained to a band and
should not move outside of that band), Data Quality Metrics (the metrics that ensure that the
experiments will be set-up correctly, and that no quality issues happened during an experiment),
and Debug Metrics (the drill down into success and guardrail metrics).
Ulwick describes some ways to measure what customers want (although not specifically for the
web) (Ulwick, 2005).
The team was empowered to continue the experiment and then ship it to 100% of users.
A quantitative OEC allows for an automated search through a space of parameters
Ron Kohavi / 11
Correlated versus causal metrics If two metrics change together, they’re correlated, but if one
metric causes another metric to change, they’re causal. If you find a causal relationship between
something you want (like revenue) and something you can control (like which ad you show),
then you can change the future.
Croll, Alistair; Yoskovitz, Benjamin. Lean Analytics: Use Data to Build a Better Startup Faster
(Lean Series) (Kindle Locations 487-489). O'Reilly Media. Kindle Edition.
- Pirate Metrics — a term coined by venture capitalist Dave McClure. gets its name from the
acronym for five distinct elements of building a successful business. McClure categorizes the
metrics a startup needs to watch into acquisition, activation, retention, revenue, and referral —
AARRR.[16] [Link]
- Avoiding false metrics
[Link]
[car dealers] instead of doing the intended thing--providing a great experience--all they do is
work hard to get people to give them a five when a drone in a call center makes the call. Many of
them will clearly state to a customer, "If anything has happened today that would prevent you
from giving us a five when they call, please tell us right now..."
The system of false metrics doesn't create a better buying experience, it creates a threatened
customer with pressure to give a five.
Ron Kohavi / 12
REFERENCES
Campbell, D. T. (1979). Assessing the impact of planned social change. Evaluation and
Program Planning, 2, 67-90. Retrieved from [Link]
7189(79)90048-X
Campbell's law. (2018). Retrieved from Wikipedia: [Link]
%27s_law
Cheshire, T. (2012, January 5). Test. Test. Test: How wooga turned the games business into a
science. Wired Magazine UK,
[Link]
Chrystal, K., & Mizen, P. D. (2001). Goodhart's Law: Its origins, meaning and implications for
monetary Policy. Prepared for the Festschrift in honour of Charles Goodhart held on 15-
16 November 2001 at the Bank of England. Retrieved from
[Link]
Cobra Effect. (2018). Retrieved from Wikipedia: [Link]
Croll, A., & Yoskovitz, B. (2013). Lean Analytics: Use Data to Build a Better Startup Faster.
O'Reilly Media.
Crook, T., Frasca, B., Kohavi, R., & Longbotham, R. (2009). Seven Pitfalls to Avoid when
Running Controlled Experiments on the Web. (P. Flach, & M. Zaki, Eds.) KDD '09:
Proceedings of the 15th ACM SIGKDD international conference on Knowledge
discovery and data mining, pp. 1105-1114.
Dmitriev, P., & Wu, X. (2016). Measuring Metrics. CIKM: Conference on Information and
Knowledge Management. Indianapolis, In. Retrieved from [Link]
Goldratt, E. M. (1990). The Haystack Syndrome. North River Press.
Goodhart, C. A. (1975). Problems of Monetary Management: The UK Experience. In Reserve
Bank of Australia, Papers in Monetary Economics (Vol. 1).
Goodhart's law. (2018). Retrieved from Wikipedia: [Link]
%27s_law
Harford, T. (2014). The Undercover Economist Strikes Back: How to Run–or Ruin–an Economy.
Riverhead Books.
Hauser, J. R., & Katz, G. (1998, October). Metrics: You Are What You Measure! European
Managemet Journal, 16(5), 516-528. Retrieved from
[Link]
%[Link]
Hoover, K. D. (2008). Phillips Curve. In H. R. David, The Concise Encyclopedia of Economics.
Retrieved from Wikipedia: [Link]
Huang, Y., You, J., Wang, I., Cao, F., & Gao, I. (2015). Data Science Interviews Exposed.
CreateSpace.
Hubbard, D. W. (2014). How to Measure Anything: Finding the Value of Intangibles in Business
(3rd ed.). Wiley.
Kaplan, R. S., & Norton, D. P. (1996). The Balanced Scorecard: Translating Strategy into
Action. Harvard Business School Press.
Ron Kohavi / 13
Kohavi, R., Deng, A., Frasca, B., Longbotham, R., Walker, T., & Xu, Y. (2012). Trustworthy
online controlled experiments: Five puzzling outcomes explained. Proceedings of the
18th Conference on Knowledge Discovery and Data
Mining([Link]/Pages/[Link]). Retrieved
from [Link]
Kohavi, R., Henne, R. M., & Sommerfield, D. (2007, August). Practical Guide to Controlled
Experiments on the Web: Listen to Your Customers not to the HiPPO. The Thirteenth
ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
(KDD 2007), pp. 959-967.
Kohavi, R., Longbotham, R., Sommerfield, D., & Henne, R. M. (2009, February). Controlled
experiments on the web: survey and practical guide. Data Mining and Knowledge
Discovery, 18(1), pp. 140-181.
Lucas critique. (2018). Retrieved from Wikipedia: [Link]
Lucas, R. E. (1976). Econometric Policy Evaluation: A Critique. In K. Brunner, & A. Meltzer,
The Phillips Curve and Labor Markets (Vol. 1, pp. 19-46). New York: Carnegie-
Rochester Conference on Public Policy.
McChesney, C., Covey, S., & Huling, J. (2012). The 4 Disciplines of Execution: Achieving Your
Wildly Important Goals. New York, NY: Free Press.
Muller, J. Z. (2018). The Tyranny of Metrics. Princeton University Press.
Parmenter, D. (2015). Key Performance Indicators: Developing, Implementing, and Using
Winning KPIs (3rd ed.). Hoboken, New Jersey: John Wiley & Sons, Inc.
Perverse Incentive. (2018). Retrieved from Wikipedia:
[Link]
Pfeffer, J., & Sutton, R. I. (1999). The Knowing-Doing Gap: How Smart Companies Turn
Knowledge into Action. Harvard Business Review Press.
Phillips, A. W. (1958, Nov). The Relation between Unemployment and the Rate of Change of
Money Wage Rates in the United Kingdom, 1861-1957. Economica, New Series,
25(100), 283-299. Retrieved from [Link]
Roy, R. K. (2001). Design of Experiments using the Taguchi Approach : 16 Steps to Product and
Process Improvement. John Wiley & Sons, Inc.
Spitzer, D. R. (2007). Transforming Performance Measurement: Rethinking the Way We
Measure and Drive Organizational Success. AMACOM.
Strathern, M. (1997). 'Improving ratings': audit in the British University System. European
Review, 5(3), 305-321. doi:10.1002/(SICI)1234-981X(199707)5:[Link];2-4