0% found this document useful (0 votes)
10 views14 pages

Navigating Software Chaos: Process Evolution

Uploaded by

edcleryton4silva
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views14 pages

Navigating Software Chaos: Process Evolution

Uploaded by

edcleryton4silva
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Process Evolution

in a Mad World
Originally published, 1994

James Bach
Satisfice, Inc.
james@[Link]
[Link]

Introduction flexibility to be applicable even in a chaotic world (Fig. 1).

Then, we’ll examine the question of process evolution and


There is a guaranteed formula for success in show how it can occur within a risk management and
software development. leadership framework, without the need for process control.

1.
2.
Create the right environment.
Get the right people.
Management Strategies
3. Determine the right course of action.
4. Make everyone do it. Leadership
Many books and papers have been written about different
ways of executing this formula.
Risk management
But, what if we can’t afford the ingredients?

This paper is about what to do in cases where we are asked


to accomplish an impossible task, in an impossibly short
Process
time frame. control
We will review the three high level strategies of technical
management, leadership, risk management, and process Impossibility of task
control and see how risk management provides the

Figure 1
Process
evolution
In order to frame the issue of management and evolution,
Process though, we first have to examine the often chaotic nature of
management the world in which software projects occur, and the role of
processes in responding to that world (Fig. 2).

Processes

World
Figure 2

Process evolution in a mad world -1- James Bach, Satisfice, Inc.


Order and chaos in the world Despite very logical arguments from process pundits,
something must be wrong with their premises. None of the
reasons given above seem meaningful when applied to
Software engineering happens within one of two companies that compete effectively in a cut-throat world,
worlds. For want of better terms, we'll call them the and build huge markets out of nothing. Were the creators of
orderly world and the chaotic world. These are not meant to Hypercard sloppy and incompetent? The Hypercard 1.0
be dramatic terms. They refer to specific and persistent project (which enabled widespread end-user programming
characteristics of two profoundly different dynamics of that in turn created an explosion of useful applications for
software development. the Macintosh) was extremely chaotic, yet extremely
successful for Apple Computer.
Orderly world: A working environment for software
development in which problems are visible and controllable, Since chaotic organizations can and do succeed, it indicates
and appropriate resources to solve problems are available. that methodologists who criticize them have oversimplified
the nature of business, the nature of engineering and the
Chaotic world: A working environment for software nature of the software consumer.
development in which problems are ambiguous and
uncontrolled, and resources to solve problems are limited. Let’s take a closer look at each of these three issues.

These worlds represent only the


factors outside the direct control of
a project, not the processes used Chaos is a function of change...
within a project. Chaotic and Change is opportunity.
orderly processes are covered later.

Most of us would rather work in an orderly world. In an information society, chaos is an essential
Much of quality assurance literature is devoted to the condition of business. Tom Peters' landmark analysis
problems of chaos and the virtues of order. Nearly every of the challenges of modern business, Thriving on Chaos,
technique advanced in the literature relies in some sense argues that general advances in technology, the proliferation
upon an assumption of order. of new market channels, and increasingly demanding
customers have forced us into a revolution in the way
As reasonable as that is, the unfortunate side effect of business is done. This revolution reaches into every kind of
focusing on order is that we discover too little about its business, even into the heart of engineering organizations.
opposite. Does chaos have any advantages? Why can’t we
stamp it out? How do we succeed in spite of it? The revolution is about embracing change. The problem for
engineers is that change translates into chaos, especially
Why does chaos persist? when a single error can potentially bring down an entire
system. But, change also translates into opportunity. It’s as
simple as this: if there is time to put a certain amount of
QA methodologists have some of the answers. and functionality into the product easily, then there is time to put
they aren’t flattering to us. Here are the most popular ones: in more functionality at the price of a certain amount of
disruption and risk. Thus does madness creep into our
 We aren't rational. projects— we will tend to take on as much risk as we
 We fear change. possibly can.
 We only think for the short term.
 We are poorly managed. Peters identifies five key strategies to take advantage of this
 We don't value quality. mad situation:
 We don't understand engineering.
 We don't understand management.  Creating total customer responsiveness.
 We are going out of business.  Pursuing fast-paced innovation.
 Achieving flexibility by empowering people.
Now, many of these reasons are true to some extent in many  Building systems for a world turned upside down.
organizations at many times. Yet, there's much more to this  Learning to love change: a new view of leadership.
problem than simple incompetence or sloppiness.
Taken in their truest spirit, the latter four of these strategies
Chaotic organizations have, in fact, produced excellent are the antithetical to the primary tactics of quality
software. At Borland and Apple, where I have personal assurance, which are to measure and control.
experience, plenty of great software has been created. And
having interviewed colleagues and coworkers who at one Measurement and control relies on the establishment of a
time or another have worked at Lotus, SPC, Quarterdeck, stable model of development, stable technologies,
Microsoft, Claris, and WordPerfect—all of them well- established processes that are detailed and specific, and
known and influential companies—it’s clear that these people who adhere to those processes in the course of their
organizations are in the same situation. work.

Process evolution in a mad world -2- James Bach, Satisfice, Inc.


Far from adherence, Peters urges rule breaking, bureaucracy what was going on, whereas my preference was to create
bucking, and skunkworks projects. He urges that we databases and systems to track project information. He
confront uncertainties with very proactive risk taking, rather disliked systems; I disliked walking around. Instead of
than recoiling from those uncertainties. This obviously overlapping, we could work together and leverage off of
plays havoc with measurement and control. each other’s strength.

But let’s say that Peters is wrong. Instead, let’s say that This was a good idea because each of us covered a weakness
uncertainty should be met with processes that enhance order. in the other; but the reason it actually worked was because
Unfortunately, that leaves us with another large problem: we liked each other.
getting the business to go along with us. A lot of
methodologists shrug off this problem. They say, “if you Methodologists assume that everyone will do the right thing
don’t have support from top management for quality, regardless of their unique human strengths and frailties.
nothing can be done.” That’s the reason why so many
management books are dedicated to ways we can convince That is another kind of wishful thinking.
management to give us the authority, staff, tools, and time
we need to do a good job.
Requirements documents don’t buy software,
But what if that fails? What if the future of the business people do. And customers are people with all the
rides on accomplishing a particular project that takes 10 variabilities noted above. In many cases customers are
people to properly test but there is only money for 3 difficult to identify or analyze. We try to predict their
engineers? Do you give up? needs, and we are often wrong. We satisfy some customers
and not others. In general, for mass-marketed software, the
Projects with such constraints are more features we jam in, the larger the number of customers
everywhere. Some succeed, some
fail, many succeed just enough to What are the methods that help us
get by. The key is to find every
possible way to optimize resources.
cope with the mad world?
True, the result isn’t as polished as
it might have been. The point is that we find ourselves in whose needs we are likely to satisfy.
these clutch situations and we could use some tips on how to Then there’s the problem of competition. Competition
cope. changes the customers expectations. In the world of COTS
software development (COTS is a military acronym meaning
Methodologists usually make the assumption that projects “commercial off-the-shelf”), competition is ferocious.
can be protected from shortages, emotions, and internal or Companies like Borland, Microsoft, Novell, Lotus, and
external disruptions. Methodologists assume that the world WordPerfect are either at each other’s throats or making
can be stopped while we settle down to do our requirements alliances. We track each other carefully, and when one
documents and DFD’s, while we design, code and unit test. company releases a product halfway through a competitor’s
project lifecycle, that competitor is forced to make mid-
That’s wishful thinking. stream changes to meet the threat, or risk shipping a product
which is obsolete on the day it is completed.
Processes don’t create software, people do. For As for quality, except for life-critical software, it isn’t as
hard problems like software development, it’s important to important as functionality. Reliability is assumed, by the
understand that the true process for getting to a solution customer, to be in the product until proven otherwise by a
cannot be modeled. If it could be, then we could tell our specific failure. Quality is important, but not as important
problems directly to computers and they would do all the as time to market and getting the right feature set at the right
analysis, design, and coding for us. Instead, we find that price.
only the external attributes of hard processes can be
described and controlled. This leaves the rest up to a Methodologists almost always define quality as adherence to
dangerously variable creature: the human engineer. the specification. In truth, it is adherence to the customer’s
expectation from the time of the sale to the day he buys the
Engineers have personalities. They vary in their talents, next upgrade. That means we need to track customer
ambitions, experiences, and world views. One engineer requirements until the last possible moment before the
might be very good at creating structure charts; another product is released.
might be better at pseudocode. One might be a whiz at
debugging; another at defensive coding.

I once found myself in the position of being a project


coordinator in QA on a project that already had a project
coordinator in the development organization. At first we
were concerned that his duties and mine would overlap.
But, as we talked this through, we discovered that his
preference was to be out and about, keeping people aware of

Process evolution in a mad world -3- James Bach, Satisfice, Inc.


Specific reasons why chaos persists As an example, management can direct a test engineer to
prepare a test plan, and even provide training and a planning
We trade order for opportunity: template. But, there’s no guarantee that the plan will be a
good one. Management can institute a review process, but
 Exploiting a competitor’s weakness. there’s no guarantee the reviewers will be any better at test
 Leveraging a special skill on the team. planning savvy than the original engineer.
 Adapting an existing technology to a new use.
 Hitting a specific marketing window. When I was at Apple, my team conducted a review of 17 test
 Creating a new market. plans in use by our department. Our findings were that none
of the plans were actually in use, at all. Some of them were
Things are often outside our control: almost pure boilerplate. In one case the tester wasn’t aware
that his product had a test plan until we literally opened a
 Engineers are variable drawer in his desk and found it lying where the previous test
 Customers are variable engineer on that product had left it!
 Management interferes.
 Business conditions are poor. Instead of recommending that plans be better managed, we
. suggested that they be abolished. Obviously, if the concept
We aren’t perfect: of written test plans had been a useful one for us, we would
have noticed much sooner that the plans weren’t any good.
 We procrastinate. The test plan study had revealed that test planning was being
 We have low expectations. done in an entirely different way, and by different people
 We don’t know any better. than we had assumed. This is a classic example of how
 We fear change. public and private processes can differ.

Problems can be solved without processes. That’s


For all these reasons and more, chaos cannot simply be called “using your head.” The quality of such non-
laughed off in our discussions of process improvement. processes, as well as private processes, depends on the
There will always be the need to cope with it, reduce it quality of the individual doing them, and the quality of that
where we can, exploit it wherever possible, and endure it individual’s team.
when we must.
Software methodologists usually limit themselves to
Order and chaos in processes discussing the public side of things. They prefer the
objectiveness of public processes, and they assume that
private processes and non-processes can be ignored, as long
Process: A pattern for solving a problem. A standard as the public processes are sufficiently complete and
solution to a standard problem. controlled.

Processes may be mandated from outside the project, but One textbook response to the test plan problem we had at
they are generally under the control of the project team. At Apple would be to redouble efforts to manage test plans,
its heart, any project is just a series of problems which get instead of looking for the private processes that were making
solved. So, the concept of processes helps us talk
about how best to go about that.
Orderly processes maximize accountability.
The concept of processes can be deceptive
Chaotic processes maximize flexibility.
and dangerous. For most of the interesting
problems of software development, especially in a it unnecessary for individual test engineers to use plans in
demanding business context, or wherever humans are the first place. Another textbook response would be to
involved, the processes that we talk about are usually not rigorously define the actual planning process— on the
the ones we actually practice. assumption that rigorous planning processes are cost
effective in every environment.
For one thing, there is a difference between public and
private process. The former is the visible element of Private processes and non-processes are chaotic,
process, i.e. the documentation and ostensible interfaces. public ones are orderly. Either kind can be applied in
The latter is the what might also be called the “true” process, either a chaotic world or an orderly world.
and consists of the actual problem-solving activity,
including the actual interfaces and informal communications Orderly processes: Processes which are quantified and
used in the course of the work. Many private processes controlled by project management, and that are performed in
never get identified, and remain uncontrolled, whereas all a consistent and coordinated fashion by the team. The
non-trivial public processes have private counterparts which dangerous assumption of orderly processes is that
may serve or may subvert them. boilerplate solutions will efficiently solve real problems.

Process evolution in a mad world -4- James Bach, Satisfice, Inc.


Chaotic processes: Processes which are unquantified
and uncontrolled by project management, and that are Leadership: Leadership deals with the problems of
performed in an ad hoc and unilateral fashion. The organization, communication, motivation and conflict
dangerous assumption of chaotic processes is that each team resolution between people. Leadership is a very powerful
member will do the right thing at the right time. discipline, and good software has been created without
resorting either to process control or risk management in any
Of course, there is a spectrum of order and chaos. A systematic way. It encompasses the notion of individual
completely chaotic process is one which no one is aware of, commitment, initiative, creativity, and other human
perhaps not even the person using it. A completely orderly attributes.
process is so ingrained into the fabric of the workplace that
it, too, may not be recognized as a process. In between there Leadership drives risk management and process control.
are many levels of accountability and flexibility.

What ways do we have of managing these processes? Attributes of leadership

 Understanding of people.
Strategies of process management  Identification of goals.
 Deployment of people to achieve goals.
Consider three overall strategies of process management:
 Helping people improve.
process control, risk management and leadership. These are
guiding metaphors of project management, within which
everything else has a place. In fact, each of these three
grand strategies includes the other ones. In operational Process control is more important in an orderly
terms, however, they are all very different. world. When resources are not overburdened and the
problem domain is well understood, it isn’t important to
Process control: Process control seeks to reduce continually relate every task to the specific risk it mitigates.
variability by systematically defining and tracking every That association can be designed into the processes when
process, and taking systematic corrective action as necessary they are publicly defined (e.g. a test outline can be
to keep the project moving toward defined goals. Process constructed to emphasize risk areas, then the process of
control is the epitome of order, and can only track orderly following the test outline automatically takes care of risk).
processes. The issues of leadership are also not as critical, because
there is less communication overhead, and less of a need for
Attributes of process control individual heroism.

Standardization Defined procedures Often, when chaotic processes are applied to an orderly
Defined deliverables world, the result is less efficiency, and higher risk of project
failure. But, when process control is applied in a chaotic
Quantification Quantifiable output world, the very same problem arises.
Measurement systems
The reason for this is the tremendous cost of maintaining
Control Defined management process consistency and coordination in a quantifiable way when the
Corrective action process world is changing all around us. If we go to the effort of
constructing an elaborate network of CASE tools, what
Coordination Integration with other processes happens when our development platform changes? What
happens when we phase out C and adopt C++? Everything
goes to pieces, that’s what.
Risk Management: Risk management seeks to optimize
resources and maximize flexibility by identifying and At a cost of $17,000, I once hired a contract programmer to
prioritizing potential failures and deploying processes, create a program that would automatically check the validity
whether controlled or uncontrolled, to the extent necessary of symbol tables emitted by the Borland C++ compiler.
to avoid the important failures. Risk management is not This was expected to save a lot of testing time. Two weeks
inherently orderly or chaotic. after it was finished, the symbol table format had to be
changed, and the tool became obsolete. Looking hard at the
Risk management drives process control. reason for the change, it became clear that in order to
maintain the tool, I would need to dedicate someone to it
about half-time, because such changes were going to happen
Attributes of risk management again and again.

 Understanding of cause and effect. So much for the big savings. This tool was going to cost us
 Identification of risks. a lot more money than the problem was worth.
 Deployment of necessary process.
 Elimination of unnecessary process. Where process control runs into big trouble is in a chaotic
world where things are changing fast. In order to maintain

Process evolution in a mad world -5- James Bach, Satisfice, Inc.


the controls, substantial effort must go into updating Risk management and chaos
documents, distributing them, modifying measurement
systems, and continuously reestablishing order. Risk management is a parallel effort to normal engineering
tasks (see Fig. 3). It consists of three major activities:
As Figure 1 illustrates, in a world where change is a
constant, it’s important to use risk management and Three levels of risk management
leadership to assure that no more than the essential
processes are in place. As chaos increases, we must resort to Identification What are the risks?
the most flexible techniques we have, because highly Ask “why?” and “what if?”
structured methods break down.
Management How will we reduce risk?
How exactly does risk management help us deal with planning Study project dynamics
change?
Reduction Is the plan working?
Deploy effective processes
Risk management in action
(see Fig. 3)
These correspond to the generic problem solving tasks of
goal setting, task planning, and task completion. Even more
Risk identification...
generically, the levels can be thought of as strategy, tactics,
Every defect in the product increases the risk that the product and execution.
will not successfully satisfy the customer.
Risk management has been routinely underestimated or
Risk management planning... mistaken for “cowboy” methodology in the literature of
quality assurance. Most of the reason for this is the
In a very small project, or one meant for an internal or informal assertion that risk is automatically managed if we simply
customer, the risk represented by the most severe of bugs is apply the “guaranteed formula for success”, mentioned
probably quite low. So it may not be important to have a above.
separate tester assigned to the project, or it may not be
necessary to test at all.
Think of risk management as the art of protecting
On a particular large project, let’s say we plan to perform testing, the project from failure. At each level there is a dialog
track any defects found, and assign them to development
that goes on between risk thinking and success thinking.
engineers.
This often corresponds to an actual conversation between
Development and QA. The dialog can be as simple as a
Dialog with task planning process...
conversation or as complex as a full blown verification
We now have a basic plan for managing that risk. To actually process with supporting documentation, depending on the
carry it out, we need to match our test plan to the development risks involved. Process control falls on both sides of the
plan. A dialog with the development engineers is needed to model. Through the action of the dialog between the two
coordinate this. sides, process control is deployed as needed.

Risk reduction... The other dimension of communication is passing down


assignments (deployment) and passing up information
As the product is delivered, and we find problems, they are regarding success or failure of those assignments. This
supposed to get fixed. escalation of status then feeds back into new plans and
assignments.
Escalation of information for analysis and course
adjustment... This horizontal communication (between success
maximizing tasks and failure prevention tasks) and vertical
If they don’t get fixed, that is a failure of the default risk reduction
communication (between levels of analysis) represent a
strategy. That triggers a dialog at the planning level once again.
Either the plan is altered, or the risk represented by those bugs stimulus and response dynamic. Stimulus and response is
is called into question. In that case failure information is purely situational and inherently flexible.
escalated to the strategic level.
As you can see, risk management need not be a
grand bureaucracy. It can be as simple as a way of
thinking that relates goals with tasks through an
understanding of project dynamics.

Process evolution in a mad world -6- James Bach, Satisfice, Inc.


Risk Management

Goal setting Risk identification


failure
success

Task planning Risk planning


failure

success

Task completion Risk reduction

Figure 3

The risk management cycle is performed not only at the serve as aids to experienced project managers, or as a basis
project plan level but even on small sub-tasks. The model is for training.
recursive.
In the final part of this paper, we examine this list as an
Taking the case of software defects, the risk management artifact of process and a product of evolution.
cycle occurs at project inception, but it also occurs for each
bug, as decisions are made regarding how and when to fix Process evolution
them. The cycle happens again as we adjust the process of
bug review to meet the challenges of efficient management
near the end of the project. Contrary to popular belief, it is possible to evolve
processes even without the aid of a process
The hard part is spotting risks, prioritizing them, control authority. Even if we can’t get help from
and knowing how to eliminate them. The risk management, as long as the environment isn’t actively
management framework is easy to understand and to use, hostile, each of us can individually work toward a better
but it only sets the stage. We need good heuristics to organization. For a dedicated process engineer like me, it’s
actually perform it well and good processes to call upon always nice if the organization both desires improvement
when the risk demands it. How do those heuristics and and commits to working on it. Still, if one or both of these
processes evolve? elements is weak or missing, as they usually are, then a risk-
based, “guerrilla” method of evolution can be employed.
As an example of heuristics for risk management, I’ve
attached the list of key questions that we ask of each defect
at our bug review meetings (see Appendix). These are not
the heuristics, per se, they are only references to them. They

Process evolution in a mad world -7- James Bach, Satisfice, Inc.


The issues involved with this kind of evolution are: 3. Work with the right people
The tendency, when working with processes, is to determine
Evolving a Process what the right process is, and then try to paste that onto the
organization. Instead, it’s much more effective to work in
Engineering it • Use personal initiative harmony with the team. If the team wants an improved
• Focus on risk testing process, don’t force a defect prevention strategy on
• Work with the right people them. Suggest it, demonstrate it, but don’t push it.
• Make it tolerant
At the same time, don't work with more people than
• Experiment informally
necessary. Involving too many people early will bog
everything down. When I have a new process in mind, I
Deploying it • Build on existing processes
may float the idea past a few key influencers in the
• Start small. Score quickly.
company, but I don’t get serious about it until I have some
• Use mentoring, not manuals kind of “straw man” document or tool ready to review.
Even then, I prefer to find a single team that is willing to
Maintaining it • Evaluate frequently give it a try before exposing other teams to it.
• Formalize only if necessary
Whatever the level of actual deployment of the experimental
process, that is the level on which the champion of the
process should involve other people. Also, bear in mind
1. Use personal initiative that success on a particular level may not scale up to the
Processes evolve through direct attention by a champion next level, even though local success does add momentum to
who maintains it until it becomes popular and widely the cause.
understood. Committees aren’t champions, quality manuals
aren’t champions. Evolution is successful when someone If a process is not accepted, recheck the analysis of the
takes the time to experiment with a new way of doing things. importance of the risks that the process is meant to address.
If possible, appoint someone to work full-time as a
coordinator of process improvement. 4. Make it tolerant
For a new process to be successful in a chaotic world, it has
At Borland we invented a role called a “metrics engineer” to tolerate being ignored for a few days or months. It must
or “QA coordinator”. The mission of the metrics engineer is require minimal maintenance or coordination— none, if
to maintain the best possible picture of the state of the possible. It must also tolerate being performed incorrectly.
project. That means maintaining and auditing project
documentation, creating reports including information Everybody’s busy. It’s hard to assimilate new ways of
drawn from various departments, watching the development working.
processes, and helping to evolve them.
5. Experiment informally
A metrics engineer can be the focus for process Again, the temptation is to define and deploy the process
experimentation and documentation, but even if nothing like rapidly. Resist the temptation. Instead, deploy an
that role exists in your organization, it’s always possible to experiment. See if it works. Update it. Don’t appoint
choose an important perennial problem and set aside a committees or task forces. Don’t think it to death. Don’t
portion of time to solve it. Evolution is possible even under design an exotic automated tool. Experiment with different
very difficult circumstances, even if it is often very slow. pilot processes in real situations over a long period. That’s
much less threatening and cumbersome than a new “official”
2. Focus on risk process that comes out of nowhere.
This is very important. If you create processes for imagined
risks rather than clear and present ones, you won’t get The most important reason to experiment rather than define
anywhere. Moreover, the risks must also be accepted as is to allow room for failure. If we invest too much in
such by the team. immediate and complete success, we’re less likely to learn
from failure. Think of experimentation as “failing forward”.
In the Apple test plan example, it would have been
completely useless to improve those test plans in an Another important reason to experiment is to reduce
atmosphere of complacence about planning. That whole interference from people in the organization who may feel
department was phased out, a couple of years later. As it threatened by the new process. If they think they have to
happened, the complacence was due to the unimportance of “speak now or forever hold their peace” they may try to kill
the mission that the department was given. Why go to the the initiative before it’s fully formed. In that situation, I’ve
expense of doing great testing on unimportant software? found that by treating the process as an experiment, it’s
much less threatening. I assure them that before any process
The way to tell if a risk is real is to look at past failures and becomes official, they will be given plenty of time to review
determine how fearful the organization is that a given failure and modify it.
will recur. People learn best from their own painful
experience. Use that as an engine for evolution.

Process evolution in a mad world -8- James Bach, Satisfice, Inc.


6. Build on existing processes These principles, drawn partly from risk
Take stock of existing processes, structures and habits that management and partly about leadership, have
are already in use. Connect the new process into those. proven successful at bringing organizations
There is usually a higher cost to introducing a brand new slowly forward, even in the midst of one crisis
way of thinking than there is to allying with an existing after another.
point of view.

For instance, if there is a formal software build process, but You don’t need top management support to do these. You
no formal acceptance testing process, the natural way to in- don’t need documentation or ISO-9000 certification. These
troduce acceptance testing is to make it a part of the build are techniques like any others, requiring skill and
process, such that no new software is delivered until it determination to accomplish. But, unlike process control-
passes the acceptance test. This will probably be easier to oriented techniques of evolution (i.e. TQM, CMM), risk-
achieve (all things being equal) than asking the developer or based evolution can be practiced by a single person without
the tester to perform it before or after the build occurs. anyone else’s cooperation.

7. Start small.
Score quickly.
In my experience, Failure is a precious resource.
most initiatives that
fail seem to collapse It drives evolution.
under their own
You always have the option to develop in yourself the
weight. That’s why I stress evolution as opposed to
heuristics to spot risk and communicate with others about it.
definition or deployment. The primary sense of the word
You always have the option to be a champion of process
“evolve” is that of starting small and growing in steps.
improvement, especially on the small scale described here.
There are some engineers who can only think in terms of
Even when we are asked to do the impossible, some portion
programmatic solutions, no matter what the problem. In
of that may be possible, and the experience of working
fact, software usually complicates things. A good way to
toward that goal can be harnessed to improve the
kill a new process is to try to develop a dedicated, GUI,
organization.
networked, software platform to drive it!

Instead, think of a completely manual process that can be Example: Bug review meetings
implemented in a few days or a week at most. Use off-the-
shelf software, if necessary. Only develop software after the
process is proven, and even then, create and deploy it in This is an example from my experiences at Borland. Bug
stages. review meetings began in our group after a last minute bug
fix introduced a major new defect which was not discovered
8. Use mentoring, not manuals until after shipment. The bug cost us $250,000 in
Instead of formalizing processes, we should train ourselves remastering, disk reduplication, unpacking and repackaging
and our teams to think situationally. That means focusing expenses. This happened before I joined the company, so I
on cause and effect, and letting processes be provisional and don't know for sure whether it’s a true story. All that
experimental. matters is that a legend was born into our group of the big
bug that got away.
9. Evaluate frequently
Find ways to evaluate the effectiveness of the process. The legend dramatically elevated the sense of risk about
These can be purely qualitative measures, or quantitative. making changes to the product late in the cycle. Bug review
One very good way to do it is to relate the process to past meetings, which we call “bug councils” were introduced as a
failures to show how causes have been eliminated or strategy to reduce the risk.
reduced.
The idea of bug councils is for all the project leaders to look
10. Formalize only if necessary at each and every defect and suggestion on the bug list and
In an orderly world, formalization is great. I’m all for it. In collectively decide what to do about it. It’s a change control
the chaotic world, formalize only when the experimental as well as a quality standard control process.
process is fully understood and accepted by the team. Even
then, plan for the expense of maintenance. With between fifty and one hundred new bugs coming in
every day, and councils puttering at the rate of thirty bugs an
hour, it’s a meeting that lasts several hours each day, every
day, in the final three weeks before sign-off.

Process evolution in a mad world -9- James Bach, Satisfice, Inc.


How do bug reviews look against the list of just to assure that mini-councils are happening), feature
evolution principles? councils (same meeting format applied to the feature set),
schedule councils (reviewing anything that threatens the
1. Use personal initiative schedule), “T” councils (full councils that only look at bugs
Counciling isn’t a very good example of this principle, marked as candidates for deferral), even something dubbed
because it was a technique invented by the team leaders. the “micro-council” (The program manager reviews the list
The bug review questions, however, were recorded on my alone, seeking out QA leaders as necessary to ratify his
initiative. I wanted to help new councilors understand the decision).
thought process. In so doing, the process has been evolved
a notch. The bug review key question list (see Appendix) is the first
documentation of the lore of bug counciling, so the
2. Focus on risk formalization process has begun. This specific practice has
Bug counciling is a risk reduction strategy with regard to not yet spread to other divisions, however.
change control. It is a risk identification strategy with
regard to the quality standard, and drives risk management 6. Build on existing processes.
planning with regard to incremental test planning needed to Bug counciling is a systematic approach to something
verify fixes that have been approved. everybody has to do anyway. We all have to make decisions
about how good is good and whether the risk of changing
Most of all it is fueled by the pain our team has felt when the software exceeds the risk of not changing it. We now do
problems were introduced late in the project cycle, or when it in a meeting, instead of in our offices.
minor known problems were allowed to remain in the
product that later turned out to be major problems. 7. Start small. Score quickly.
Bug counciling got going very quickly after the idea was
3. Work with the right people adopted. It started as an entirely manual process, with
All of the leaders of the project were involved in the ground rules made up in the course of a few minutes.
decision to start the formal bug reviews, and the process was
shaped with their involvement. No one above the first level 8. Use mentoring, not manuals
of management was very much involved, however, and no There is no bug council guide, or rules of order. If ever
one from the other development organizations within there is one, it will probably be used for training purposes
Borland, although it has since propagated to several other only, and not as a rule book. The question list in the
groups. Appendix is just such a tool. Although handy as a reference,
it is still better suited for training.
The clear necessity of counciling, and the success we’ve had
with it, have served to break down most of the resistance to 9. Evaluate frequently
bug counciling. The point of contention now is how early in In the early days of counciling, there were frequent
the project to start doing it. The answer to this question may comments and suggestions regarding its effectiveness. With
lie in better metrics to track critical bugs, or in better each project, we get more confident in the basic technique.
prevention strategies.
One way we evaluate it is to look at how many bugs found
4. Make it tolerant in the field were known to the council beforehand, and how
Councils are tolerant in that they are driven by the bug list. many were caused by fixes authorized by the council. When
If we don’t council for a couple of days, then the list grows a new kind of problem slips through, new heuristics crop up
longer, but we can’t ship without performing the review (see to defend against a repeat disaster. The bug review
#1), so the process is almost self-governing. questions give an idea of how much there is to think about.

At first, we kept the current hot bug list on a big whiteboard. 10. Formalize only if necessary
The board required a lot of maintenance, so a new system The bug council was driven to a certain amount of
was adopted that was driven by the bug tracking database, formalization by the sheer weight of the process, but it
requiring much less coordination. Because our bug database remains today an experimental process, subject to change
is a decentralized system, we were able to modify our and optimization.
version of it to do this without disturbing any other teams.
If a substantial number of leaders leave the company all at
Another way we made it tolerant was to divide the review once, the counciling concept will go with them, but failing
into stages, calling people in when their bugs were up for that, there is little need to further formalize the process.
review. In this way, only a core group of about four leaders
had to be there all the time, instead of the twelve or so Example: Better Bug Metrics
which comprised the full management team.

5. Experiment informally This is a quick history of the evolution of a daily bug


Bug counciling is an institution now, three years later, but metrics collection and distribution system which changed
they are still being tweaked. We’ve tried many variations on the way projects are managed in my particular group within
the them, including mini-councils (a leader each from QA Borland.
and Development), meta-councils (full councils that meet

Process evolution in a mad world - 10 - James Bach, Satisfice, Inc.


The challenges in this area were: both for people trying to find bugs and those trying to fix or
prevent them.
• Determining what to measure.
• Determining how the metrics should be used.
The process is optimized based on experience.
• Packaging the metrics for ease of use.
Automation is introduced. Principles #1, #5, #7,
• Adding a process when we already had too much to do. and #9 apply.
• Getting the team interested in the metrics.
• Overcoming resistance in team members who feared The milestone was reached and I stopped doing the graphs.
the misuse of statistics. But I knew it should be restarted for the next milestone in
order to compare that data against the first run, so I asked a
It took eight months to craft the metrics system, but in small buddy of mine, who was familiar with Paradox database
steps such that I was never away for too long from my other programming, to write a script to automate the data
duties. The system was recognized at our project debriefing collection. While he was at it, I had him throw in a few
as having contributed to the success of the project. Now other basic metrics just for fun.
there are two large teams at Borland who are using it, one
small team, and another that has requested it. Meanwhile, I created a new spreadsheet to provide a more
comprehensive graph set and store the other data for later
The system runs every night, queries several bug tracking use.
databases, and produces a comprehensive text report which
is then imported to a spreadsheet, where key elements are
graphed. The graphs and the report are concatenated and The new process is connected to (and reinforces)
sent via email to the managers of the project, and anyone existing processes. Principles #3 and #6 apply.
else who asks, on a daily basis, and to almost everyone else
on a weekly basis. The program manager called me a couple of weeks later and
said that we needed to increase the sense of urgency about
The system is an example of an experimental, skunkworks shipping, so would I please restart the graphs. I responded
process, growing from the grassroots in response to specific that since the graphs are based on output from the bug
needs. It shows that improvement can happen even without review meetings, he would have to get those meetings
an organizational cattle-prod in the form of a documented restarted first, which he did.
official process, and even without a committee or blue-
ribbon task force.
An attempt at early formalization is resisted. The
process remains identified with the originator.
First, a problem is identified and a solution is Principles #4 and #10 apply.
suggested by a member of the team. Principles #1
and #2 apply. Because I wanted to entrench this new concept I came up
with a name for it: "Hot Convergence Charting". The name
Development on the system started when I witnessed has not caught on, yet. I'm told they're known as "James'
disagreement among the managers in the Languages Graphs". Outperforming the projections suggested by the
department as to how long it would take us to converge from graphs is called "beating James."
a certain number of critical bugs to zero, so that we could
ship. I thought it would be useful to graph the number of I also installed the graph data collector on various other
critical bugs each day after the review meeting, to see if manager's machines, but they almost never ran the scripts.
there were any interesting trends to be seen. That might They found it easier to ask me to do it.
help us improve our ability to schedule project milestones.
A second round of optimization triggers a broader
A solution is hand generated to prove the concept. definition of the process and a higher quality goal
An unexpected side benefit emerges as a result of for the system. Principles #3, #4 and #9 apply.
the experiment. Principles #1, #5, #7, #8, and #9
apply. During the next convergence cycle, I added features to the
graph in response to the way I saw it being used and the
I performed the graphing by hand for 11 days in a row, each questions that came up while discussing it. While doing
time gritting my teeth and thinking of ways to improve the this, I decided to throw out the original Paradox script and
process (which required a total of twelve separate queries of create a brand new, extensible metrics gathering
the bug tracking system with manual data entry to a architecture.
spreadsheet, and took about 20 minutes). I sent the graph to
the managers for their curiosity. Since we have to ship with I commandeered an old Dell 325 (a big heavy beast which
zero critical bugs, the graphs were instantly popular with a fell on a concrete sidewalk while I was wheeling it over on
couple of people who really wanted us to ship soon. They an office chair, but it still booted up when I plugged it in),
began looking for ways to drive the graph to zero. I realized and made that a dedicated metrics collection machine. I put
that the graph could be used as a motivational instrument, all the metrics scripts under version control.

Process evolution in a mad world - 11 - James Bach, Satisfice, Inc.


The new architecture emerged rapidly, since I was pretty The program manager began referring to "James' magic
fired up at this point and trying to get ready for the third formula" in meetings, which piqued some interest and some
convergence cycle. By the time it began, the metrics system alarm among the managers.
had reached it's final (experimental) form, and I was just
tweaking it to add special reports here and there. On the first trial, the formula was correct within 2 days at a
range of three weeks, which was good, given the number of
At that point, the development project which was the guinea variables that affect the actual ship date. On review,
pig for the metrics system was in a state of continuous however, I determined that it was probably luck for it to be
convergence, so I was running the metrics system that good. On the second trial, it was way off. I decided
throughout the day, measuring almost everything I ever that the formula was a failure and scrapped it.
wanted to measure, plus a lot of things other people had
asked for. It had evolved into a generalized defect analysis Another team heard about the formula and I gave them a
system. briefing, noting it's strengths and weaknesses.

An existing process is recognized as The process spreads by popular demand,


anachronistic and discontinued. Principles #2, #5, remaining flexible to meet the needs of it
#9 apply. customers. The effort to automate over the past
months begins to pay off. Principles #2, #3, and #4
The existing, clunky, weekly bug tracking statistics, which apply.
required an hour per week to produce, and were entirely
paper-based, were discontinued because my metrics Another large team expressed interest in the metrics, but
measured everything and more. didn't like some of the reports. So I branched the sources
and created a special version just for them. Doing metrics
for the second team added approximately six minutes of
The process reaches a full-blown experimental
work to my day after the initial startup cost.
form. Its simplicity and un-authoritarian nature
win it support. Principles #1, #2, #3, #5, #6, #8, #9,
and #10 apply. Summary:

The VP began using it to see how we were doing. He also This is an example of people cooperating naturally, solving
occasionally requested special metrics. problems naturally, and improving processes naturally,
without the need for a process control authority.
Decisions were being made on the basis of the reports.
Special effort was made to comb bad data out of the bug No committees or quality circles were needed. No proposals
database in response to concerns that the graphs might be or requirements documents were created or reviewed.
misleading. Many critical questions were posed by some
managers who now found themselves under more intense Should the organization decide to formalize the process,
scrutiny. I tried to treat everyone who approached me as my substantial experience in actually using it will be available
customer and almost always modified the system in response to guide that effort.
to their concerns.

Most of the managers relaxed when they saw that the


metrics were not being used punitively or dogmatically. The
graphs allowed us to spot certain trends and ask certain
informed questions that we would never could have before.
Resistance diminished, as well, because it required no effort
on anybody's part to create them (except me) and since this
wasn't an "official" process, but one that was entirely
voluntary on the part of the people who used the reports.

The process spawns another process, which fails,


and is withdrawn before it does any damage. But,
the failure itself adds to the experience of the
organization. Principles #2, #6, #8, #9, apply.

Meanwhile, in the midst of another disagreement over how


fast we were really converging, the program manager asked
me to prepare a projection based on the convergence charts.
I then created a simple model to project a ship date from
recent trends.

Process evolution in a mad world - 12 - James Bach, Satisfice, Inc.


Appendix: Key Questions For Bug Review

Problem Analysis
1.1. How was the bug found?
Frequency 1.1.1. Was it found by a user?
1.1.2. Is it a natural or contrived case?
1.1.3. Is it a typical or pathological case?
1.1.4. Was the bug caused by a recent fix to another bug?
1.2. How often is it likely to occur?
1.2.1. Is it intermittent or predictable?
1.2.2. Is it a one-time problem or ongoing?
1.3. How soon after the bug was created did we discover it?

2.1. Does the bug cause any data to be lost?


Severity 2.2. Will it cause an additional load for Technical Support?
2.3. How likely is the user to notice it when it occurs?
2.4. Is it the tip of an iceberg?
2.4.1. Will it trigger other problems?
2.4.2. Is it part of a class of bugs that should all be fixed?
2.4.3. Does it represent a basic design deficiency?
2.5. Was this bug shipped in the previous release?
2.5.1. Did Technical Support hear anything about it?
2.5.2. Has anything changed since the last version that would
make it more or less of a problem?
2.6. Are there any international implications?
2.7. Does the bug break test automation?
2.8. Is this bug less severe than others we've deferred? more severe than
others we've fixed?

3.1. Are certain kinds of users more likely to be affected than others?
Publicity 3.1.1. How sophisticated are those users?
3.1.2. How vocal are those users?
3.1.3. How important are those users?
3.1.4. Will it affect the review writers at any major magazines?
3.2. Are our competitors strong or weak in the same functional areas?
3.3. Is this the first release or is there an installed base?
3.4. Is the problem so esoteric that no one will notice before we can update
the product?
3.5. Does it look like a defect to the casual observer, or more like a design
limitation?

Process evolution in a mad world - 13 - James Bach, Satisfice, Inc.


Solution Analysis
4.1. What are the workarounds?
Identification 4.1.1. Are they obvious or esoteric?
4.2. Can we "document around it" instead of fixing it?
4.3. Can Technical Support create a Tech Note to explain the solution?
4.4. Can the solution be postponed until a later milestone?
4.5. Is a fix known?
4.5.1. Are there several possible fixes or just one?
4.5.2. How many lines of code are involved?
4.5.3. Is it complex code or simple code?
4.5.4. Is it familiar code or legacy code?
4.5.5. Is the fix a tweak, rewrite, or substantial new code?
4.5.6. How long will it take to implement the fix?
4.5.7. What components are affected by the fix?
4.5.8. Will it require rebuilds of dependent components?
4.5.9. Does the fix impact doc. in any way? screenshots? help?

5.1. What new problems could the fix cause? what is the worst case?
Verification 5.2. How effectively could we test the fix, if we authorize it?
5.2.1. Was this bug found late in the project? does that indicate a
weakness in the test suite?
5.2.2. Will the test automation cover this case?
5.2.3. Could the fix be sent specially to some or all of the beta
testers?
5.3. How hard would it be to undo the fix, if there's trouble with it?

6.1. How dangerous is it to make changes in this code?


Perspective 6.2. Will a fix to this component be the only reason to rebuild or remaster?
6.3. How does the overall quality compare to previous releases?
6.4. If we think this bug is important, why not slip the schedule by two weeks
and fix more bugs?
6.5. What would be the right thing to do? the safe thing to do?

7.1. Was the problem caused by a previously authorized change?


Prevention 7.2. What was the error that caused the defect?
7.3. Is there any internal error checking or unit test that should be added to
catch bugs of this type?
7.4. Is there any review process that could catch bugs like this before they
get into the build?

Process evolution in a mad world - 14 - James Bach, Satisfice, Inc.

Common questions

Powered by AI

The orderly world in software development is characterized by a working environment where problems are visible, controllable, and appropriate resources are available to solve them. In contrast, the chaotic world is marked by ambiguous, uncontrolled problems and limited resources to address those issues . In an orderly world, quality assurance relies on the assumption of order, making it easier to implement structured processes. However, in a chaotic environment, such order is elusive, necessitating flexibility and adaptability in management practices .

The dialogue between development and quality assurance is crucial for effective risk management because it ensures that risks are identified, communicated, and addressed collaboratively. This interaction facilitates understanding of both risk implications and quality goals, allowing for more informed decision-making. By engaging in constructive discussions, teams can implement risk management measures that effectively balance project goals with quality standards, thus preventing potential failures . The dialogue can vary from simple conversations to complex verification processes, adapting to the risks involved .

Experimental processes at Borland significantl contribute to process improvements by allowing iterative development and refinement in real-world conditions. James Bach illustrates how experimental approaches can lead to rapid adaptation and customized solutions, aligning processes with the evolving needs of teams and projects. These processes enable continuous learning and adjustment, facilitating improvements without the rigidity of formal process controls. As such, they demonstrate an agile methodology that values functionality and applicability over strict adherence to predefined procedures .

Risk management provides flexibility in a chaotic world by optimizing resources and adapting processes to avoid important failures. Its core components include risk identification, planning, and reduction. By identifying potential failures and adapting to changing conditions, risk management helps maintain project viability despite inherent uncertainties . This adaptability is crucial as rigid structures often fail in rapidly changing environments .

Balancing order and chaos in process evolution is significant because it allows organizations to leverage the stability of order while harnessing the adaptability of chaos. James Bach suggests that focusing solely on order can lead to overlooking the benefits of chaos, such as promoting innovation and rapid adaptation. Thus, by integrating risk management and leadership, organizations can evolve processes that are resilient to change, optimizing both structured and flexible approaches .

Relying on a formalized process in a rapidly changing environment can lead to inefficiencies and project failures. This is because maintaining structured processes requires significant effort for updates and coordination as conditions change. Thus, in dynamic settings, too much formality can become a hindrance, as rigid systems struggle to accommodate rapid shifts in technology and requirements . An adaptable approach that prioritizes essential processes while allowing flexibility is more suitable .

Process evolution benefits from a grassroots approach as it encourages organic growth and adaptation based on real user needs and feedback, without the rigidity of formal top-down enforcement. This approach allows processes to develop naturally from the insights and innovations of individuals involved in the work, fostering inclusivity and ownership that can lead to more sustainable and accepted changes . It promotes flexibility, empowering teams to identify practical solutions to issues and experiment with changes that result in effective process improvements .

The evolution of a bug metrics system exemplifies process evolution in a chaotic environment by showcasing incremental developments driven by necessity and feedback. Initially informal and experimental, the system expanded and adapted over time in response to the changing needs of teams, thereby incorporating new features and refinements. This reflects a dynamic adaptation where processes evolve naturally rather than through rigid change mandates, emphasizing flexibility and responsiveness as critical elements in handling chaos .

Leadership in software development involves managing organization, communication, motivation, and conflict resolution. It is essential because it enables teams to achieve goals and improve processes through human attributes like creativity and commitment, without strict reliance on process control or risk management. Leadership drives the other two strategies by fostering an environment where people are motivated to tackle risks and adhere to processes that align with project goals .

Lessons from the failure of a metrics-based projection system highlight the unpredictable nature of software development, where numerous variables can affect outcomes. While initially promising results may offer false confidence, the system's limitations reveal the necessity for cautious interpretation of metrics and the need for iterative adjustments based on feedback and results. This failure underscores the importance of remaining adaptable and reflective when applying automated projections to complex projects, ensuring that metrics are used as one of many tools rather than definitive predictors .

You might also like