Navigating Software Chaos: Process Evolution
Navigating Software Chaos: Process Evolution
in a Mad World
Originally published, 1994
James Bach
Satisfice, Inc.
james@[Link]
[Link]
1.
2.
Create the right environment.
Get the right people.
Management Strategies
3. Determine the right course of action.
4. Make everyone do it. Leadership
Many books and papers have been written about different
ways of executing this formula.
Risk management
But, what if we can’t afford the ingredients?
Figure 1
Process
evolution
In order to frame the issue of management and evolution,
Process though, we first have to examine the often chaotic nature of
management the world in which software projects occur, and the role of
processes in responding to that world (Fig. 2).
Processes
World
Figure 2
Most of us would rather work in an orderly world. In an information society, chaos is an essential
Much of quality assurance literature is devoted to the condition of business. Tom Peters' landmark analysis
problems of chaos and the virtues of order. Nearly every of the challenges of modern business, Thriving on Chaos,
technique advanced in the literature relies in some sense argues that general advances in technology, the proliferation
upon an assumption of order. of new market channels, and increasingly demanding
customers have forced us into a revolution in the way
As reasonable as that is, the unfortunate side effect of business is done. This revolution reaches into every kind of
focusing on order is that we discover too little about its business, even into the heart of engineering organizations.
opposite. Does chaos have any advantages? Why can’t we
stamp it out? How do we succeed in spite of it? The revolution is about embracing change. The problem for
engineers is that change translates into chaos, especially
Why does chaos persist? when a single error can potentially bring down an entire
system. But, change also translates into opportunity. It’s as
simple as this: if there is time to put a certain amount of
QA methodologists have some of the answers. and functionality into the product easily, then there is time to put
they aren’t flattering to us. Here are the most popular ones: in more functionality at the price of a certain amount of
disruption and risk. Thus does madness creep into our
We aren't rational. projects— we will tend to take on as much risk as we
We fear change. possibly can.
We only think for the short term.
We are poorly managed. Peters identifies five key strategies to take advantage of this
We don't value quality. mad situation:
We don't understand engineering.
We don't understand management. Creating total customer responsiveness.
We are going out of business. Pursuing fast-paced innovation.
Achieving flexibility by empowering people.
Now, many of these reasons are true to some extent in many Building systems for a world turned upside down.
organizations at many times. Yet, there's much more to this Learning to love change: a new view of leadership.
problem than simple incompetence or sloppiness.
Taken in their truest spirit, the latter four of these strategies
Chaotic organizations have, in fact, produced excellent are the antithetical to the primary tactics of quality
software. At Borland and Apple, where I have personal assurance, which are to measure and control.
experience, plenty of great software has been created. And
having interviewed colleagues and coworkers who at one Measurement and control relies on the establishment of a
time or another have worked at Lotus, SPC, Quarterdeck, stable model of development, stable technologies,
Microsoft, Claris, and WordPerfect—all of them well- established processes that are detailed and specific, and
known and influential companies—it’s clear that these people who adhere to those processes in the course of their
organizations are in the same situation. work.
But let’s say that Peters is wrong. Instead, let’s say that This was a good idea because each of us covered a weakness
uncertainty should be met with processes that enhance order. in the other; but the reason it actually worked was because
Unfortunately, that leaves us with another large problem: we liked each other.
getting the business to go along with us. A lot of
methodologists shrug off this problem. They say, “if you Methodologists assume that everyone will do the right thing
don’t have support from top management for quality, regardless of their unique human strengths and frailties.
nothing can be done.” That’s the reason why so many
management books are dedicated to ways we can convince That is another kind of wishful thinking.
management to give us the authority, staff, tools, and time
we need to do a good job.
Requirements documents don’t buy software,
But what if that fails? What if the future of the business people do. And customers are people with all the
rides on accomplishing a particular project that takes 10 variabilities noted above. In many cases customers are
people to properly test but there is only money for 3 difficult to identify or analyze. We try to predict their
engineers? Do you give up? needs, and we are often wrong. We satisfy some customers
and not others. In general, for mass-marketed software, the
Projects with such constraints are more features we jam in, the larger the number of customers
everywhere. Some succeed, some
fail, many succeed just enough to What are the methods that help us
get by. The key is to find every
possible way to optimize resources.
cope with the mad world?
True, the result isn’t as polished as
it might have been. The point is that we find ourselves in whose needs we are likely to satisfy.
these clutch situations and we could use some tips on how to Then there’s the problem of competition. Competition
cope. changes the customers expectations. In the world of COTS
software development (COTS is a military acronym meaning
Methodologists usually make the assumption that projects “commercial off-the-shelf”), competition is ferocious.
can be protected from shortages, emotions, and internal or Companies like Borland, Microsoft, Novell, Lotus, and
external disruptions. Methodologists assume that the world WordPerfect are either at each other’s throats or making
can be stopped while we settle down to do our requirements alliances. We track each other carefully, and when one
documents and DFD’s, while we design, code and unit test. company releases a product halfway through a competitor’s
project lifecycle, that competitor is forced to make mid-
That’s wishful thinking. stream changes to meet the threat, or risk shipping a product
which is obsolete on the day it is completed.
Processes don’t create software, people do. For As for quality, except for life-critical software, it isn’t as
hard problems like software development, it’s important to important as functionality. Reliability is assumed, by the
understand that the true process for getting to a solution customer, to be in the product until proven otherwise by a
cannot be modeled. If it could be, then we could tell our specific failure. Quality is important, but not as important
problems directly to computers and they would do all the as time to market and getting the right feature set at the right
analysis, design, and coding for us. Instead, we find that price.
only the external attributes of hard processes can be
described and controlled. This leaves the rest up to a Methodologists almost always define quality as adherence to
dangerously variable creature: the human engineer. the specification. In truth, it is adherence to the customer’s
expectation from the time of the sale to the day he buys the
Engineers have personalities. They vary in their talents, next upgrade. That means we need to track customer
ambitions, experiences, and world views. One engineer requirements until the last possible moment before the
might be very good at creating structure charts; another product is released.
might be better at pseudocode. One might be a whiz at
debugging; another at defensive coding.
Processes may be mandated from outside the project, but One textbook response to the test plan problem we had at
they are generally under the control of the project team. At Apple would be to redouble efforts to manage test plans,
its heart, any project is just a series of problems which get instead of looking for the private processes that were making
solved. So, the concept of processes helps us talk
about how best to go about that.
Orderly processes maximize accountability.
The concept of processes can be deceptive
Chaotic processes maximize flexibility.
and dangerous. For most of the interesting
problems of software development, especially in a it unnecessary for individual test engineers to use plans in
demanding business context, or wherever humans are the first place. Another textbook response would be to
involved, the processes that we talk about are usually not rigorously define the actual planning process— on the
the ones we actually practice. assumption that rigorous planning processes are cost
effective in every environment.
For one thing, there is a difference between public and
private process. The former is the visible element of Private processes and non-processes are chaotic,
process, i.e. the documentation and ostensible interfaces. public ones are orderly. Either kind can be applied in
The latter is the what might also be called the “true” process, either a chaotic world or an orderly world.
and consists of the actual problem-solving activity,
including the actual interfaces and informal communications Orderly processes: Processes which are quantified and
used in the course of the work. Many private processes controlled by project management, and that are performed in
never get identified, and remain uncontrolled, whereas all a consistent and coordinated fashion by the team. The
non-trivial public processes have private counterparts which dangerous assumption of orderly processes is that
may serve or may subvert them. boilerplate solutions will efficiently solve real problems.
Understanding of people.
Strategies of process management Identification of goals.
Deployment of people to achieve goals.
Consider three overall strategies of process management:
Helping people improve.
process control, risk management and leadership. These are
guiding metaphors of project management, within which
everything else has a place. In fact, each of these three
grand strategies includes the other ones. In operational Process control is more important in an orderly
terms, however, they are all very different. world. When resources are not overburdened and the
problem domain is well understood, it isn’t important to
Process control: Process control seeks to reduce continually relate every task to the specific risk it mitigates.
variability by systematically defining and tracking every That association can be designed into the processes when
process, and taking systematic corrective action as necessary they are publicly defined (e.g. a test outline can be
to keep the project moving toward defined goals. Process constructed to emphasize risk areas, then the process of
control is the epitome of order, and can only track orderly following the test outline automatically takes care of risk).
processes. The issues of leadership are also not as critical, because
there is less communication overhead, and less of a need for
Attributes of process control individual heroism.
Standardization Defined procedures Often, when chaotic processes are applied to an orderly
Defined deliverables world, the result is less efficiency, and higher risk of project
failure. But, when process control is applied in a chaotic
Quantification Quantifiable output world, the very same problem arises.
Measurement systems
The reason for this is the tremendous cost of maintaining
Control Defined management process consistency and coordination in a quantifiable way when the
Corrective action process world is changing all around us. If we go to the effort of
constructing an elaborate network of CASE tools, what
Coordination Integration with other processes happens when our development platform changes? What
happens when we phase out C and adopt C++? Everything
goes to pieces, that’s what.
Risk Management: Risk management seeks to optimize
resources and maximize flexibility by identifying and At a cost of $17,000, I once hired a contract programmer to
prioritizing potential failures and deploying processes, create a program that would automatically check the validity
whether controlled or uncontrolled, to the extent necessary of symbol tables emitted by the Borland C++ compiler.
to avoid the important failures. Risk management is not This was expected to save a lot of testing time. Two weeks
inherently orderly or chaotic. after it was finished, the symbol table format had to be
changed, and the tool became obsolete. Looking hard at the
Risk management drives process control. reason for the change, it became clear that in order to
maintain the tool, I would need to dedicate someone to it
about half-time, because such changes were going to happen
Attributes of risk management again and again.
Understanding of cause and effect. So much for the big savings. This tool was going to cost us
Identification of risks. a lot more money than the problem was worth.
Deployment of necessary process.
Elimination of unnecessary process. Where process control runs into big trouble is in a chaotic
world where things are changing fast. In order to maintain
success
Figure 3
The risk management cycle is performed not only at the serve as aids to experienced project managers, or as a basis
project plan level but even on small sub-tasks. The model is for training.
recursive.
In the final part of this paper, we examine this list as an
Taking the case of software defects, the risk management artifact of process and a product of evolution.
cycle occurs at project inception, but it also occurs for each
bug, as decisions are made regarding how and when to fix Process evolution
them. The cycle happens again as we adjust the process of
bug review to meet the challenges of efficient management
near the end of the project. Contrary to popular belief, it is possible to evolve
processes even without the aid of a process
The hard part is spotting risks, prioritizing them, control authority. Even if we can’t get help from
and knowing how to eliminate them. The risk management, as long as the environment isn’t actively
management framework is easy to understand and to use, hostile, each of us can individually work toward a better
but it only sets the stage. We need good heuristics to organization. For a dedicated process engineer like me, it’s
actually perform it well and good processes to call upon always nice if the organization both desires improvement
when the risk demands it. How do those heuristics and and commits to working on it. Still, if one or both of these
processes evolve? elements is weak or missing, as they usually are, then a risk-
based, “guerrilla” method of evolution can be employed.
As an example of heuristics for risk management, I’ve
attached the list of key questions that we ask of each defect
at our bug review meetings (see Appendix). These are not
the heuristics, per se, they are only references to them. They
For instance, if there is a formal software build process, but You don’t need top management support to do these. You
no formal acceptance testing process, the natural way to in- don’t need documentation or ISO-9000 certification. These
troduce acceptance testing is to make it a part of the build are techniques like any others, requiring skill and
process, such that no new software is delivered until it determination to accomplish. But, unlike process control-
passes the acceptance test. This will probably be easier to oriented techniques of evolution (i.e. TQM, CMM), risk-
achieve (all things being equal) than asking the developer or based evolution can be practiced by a single person without
the tester to perform it before or after the build occurs. anyone else’s cooperation.
7. Start small.
Score quickly.
In my experience, Failure is a precious resource.
most initiatives that
fail seem to collapse It drives evolution.
under their own
You always have the option to develop in yourself the
weight. That’s why I stress evolution as opposed to
heuristics to spot risk and communicate with others about it.
definition or deployment. The primary sense of the word
You always have the option to be a champion of process
“evolve” is that of starting small and growing in steps.
improvement, especially on the small scale described here.
There are some engineers who can only think in terms of
Even when we are asked to do the impossible, some portion
programmatic solutions, no matter what the problem. In
of that may be possible, and the experience of working
fact, software usually complicates things. A good way to
toward that goal can be harnessed to improve the
kill a new process is to try to develop a dedicated, GUI,
organization.
networked, software platform to drive it!
Instead, think of a completely manual process that can be Example: Bug review meetings
implemented in a few days or a week at most. Use off-the-
shelf software, if necessary. Only develop software after the
process is proven, and even then, create and deploy it in This is an example from my experiences at Borland. Bug
stages. review meetings began in our group after a last minute bug
fix introduced a major new defect which was not discovered
8. Use mentoring, not manuals until after shipment. The bug cost us $250,000 in
Instead of formalizing processes, we should train ourselves remastering, disk reduplication, unpacking and repackaging
and our teams to think situationally. That means focusing expenses. This happened before I joined the company, so I
on cause and effect, and letting processes be provisional and don't know for sure whether it’s a true story. All that
experimental. matters is that a legend was born into our group of the big
bug that got away.
9. Evaluate frequently
Find ways to evaluate the effectiveness of the process. The legend dramatically elevated the sense of risk about
These can be purely qualitative measures, or quantitative. making changes to the product late in the cycle. Bug review
One very good way to do it is to relate the process to past meetings, which we call “bug councils” were introduced as a
failures to show how causes have been eliminated or strategy to reduce the risk.
reduced.
The idea of bug councils is for all the project leaders to look
10. Formalize only if necessary at each and every defect and suggestion on the bug list and
In an orderly world, formalization is great. I’m all for it. In collectively decide what to do about it. It’s a change control
the chaotic world, formalize only when the experimental as well as a quality standard control process.
process is fully understood and accepted by the team. Even
then, plan for the expense of maintenance. With between fifty and one hundred new bugs coming in
every day, and councils puttering at the rate of thirty bugs an
hour, it’s a meeting that lasts several hours each day, every
day, in the final three weeks before sign-off.
At first, we kept the current hot bug list on a big whiteboard. 10. Formalize only if necessary
The board required a lot of maintenance, so a new system The bug council was driven to a certain amount of
was adopted that was driven by the bug tracking database, formalization by the sheer weight of the process, but it
requiring much less coordination. Because our bug database remains today an experimental process, subject to change
is a decentralized system, we were able to modify our and optimization.
version of it to do this without disturbing any other teams.
If a substantial number of leaders leave the company all at
Another way we made it tolerant was to divide the review once, the counciling concept will go with them, but failing
into stages, calling people in when their bugs were up for that, there is little need to further formalize the process.
review. In this way, only a core group of about four leaders
had to be there all the time, instead of the twelve or so Example: Better Bug Metrics
which comprised the full management team.
The VP began using it to see how we were doing. He also This is an example of people cooperating naturally, solving
occasionally requested special metrics. problems naturally, and improving processes naturally,
without the need for a process control authority.
Decisions were being made on the basis of the reports.
Special effort was made to comb bad data out of the bug No committees or quality circles were needed. No proposals
database in response to concerns that the graphs might be or requirements documents were created or reviewed.
misleading. Many critical questions were posed by some
managers who now found themselves under more intense Should the organization decide to formalize the process,
scrutiny. I tried to treat everyone who approached me as my substantial experience in actually using it will be available
customer and almost always modified the system in response to guide that effort.
to their concerns.
Problem Analysis
1.1. How was the bug found?
Frequency 1.1.1. Was it found by a user?
1.1.2. Is it a natural or contrived case?
1.1.3. Is it a typical or pathological case?
1.1.4. Was the bug caused by a recent fix to another bug?
1.2. How often is it likely to occur?
1.2.1. Is it intermittent or predictable?
1.2.2. Is it a one-time problem or ongoing?
1.3. How soon after the bug was created did we discover it?
3.1. Are certain kinds of users more likely to be affected than others?
Publicity 3.1.1. How sophisticated are those users?
3.1.2. How vocal are those users?
3.1.3. How important are those users?
3.1.4. Will it affect the review writers at any major magazines?
3.2. Are our competitors strong or weak in the same functional areas?
3.3. Is this the first release or is there an installed base?
3.4. Is the problem so esoteric that no one will notice before we can update
the product?
3.5. Does it look like a defect to the casual observer, or more like a design
limitation?
5.1. What new problems could the fix cause? what is the worst case?
Verification 5.2. How effectively could we test the fix, if we authorize it?
5.2.1. Was this bug found late in the project? does that indicate a
weakness in the test suite?
5.2.2. Will the test automation cover this case?
5.2.3. Could the fix be sent specially to some or all of the beta
testers?
5.3. How hard would it be to undo the fix, if there's trouble with it?
The orderly world in software development is characterized by a working environment where problems are visible, controllable, and appropriate resources are available to solve them. In contrast, the chaotic world is marked by ambiguous, uncontrolled problems and limited resources to address those issues . In an orderly world, quality assurance relies on the assumption of order, making it easier to implement structured processes. However, in a chaotic environment, such order is elusive, necessitating flexibility and adaptability in management practices .
The dialogue between development and quality assurance is crucial for effective risk management because it ensures that risks are identified, communicated, and addressed collaboratively. This interaction facilitates understanding of both risk implications and quality goals, allowing for more informed decision-making. By engaging in constructive discussions, teams can implement risk management measures that effectively balance project goals with quality standards, thus preventing potential failures . The dialogue can vary from simple conversations to complex verification processes, adapting to the risks involved .
Experimental processes at Borland significantl contribute to process improvements by allowing iterative development and refinement in real-world conditions. James Bach illustrates how experimental approaches can lead to rapid adaptation and customized solutions, aligning processes with the evolving needs of teams and projects. These processes enable continuous learning and adjustment, facilitating improvements without the rigidity of formal process controls. As such, they demonstrate an agile methodology that values functionality and applicability over strict adherence to predefined procedures .
Risk management provides flexibility in a chaotic world by optimizing resources and adapting processes to avoid important failures. Its core components include risk identification, planning, and reduction. By identifying potential failures and adapting to changing conditions, risk management helps maintain project viability despite inherent uncertainties . This adaptability is crucial as rigid structures often fail in rapidly changing environments .
Balancing order and chaos in process evolution is significant because it allows organizations to leverage the stability of order while harnessing the adaptability of chaos. James Bach suggests that focusing solely on order can lead to overlooking the benefits of chaos, such as promoting innovation and rapid adaptation. Thus, by integrating risk management and leadership, organizations can evolve processes that are resilient to change, optimizing both structured and flexible approaches .
Relying on a formalized process in a rapidly changing environment can lead to inefficiencies and project failures. This is because maintaining structured processes requires significant effort for updates and coordination as conditions change. Thus, in dynamic settings, too much formality can become a hindrance, as rigid systems struggle to accommodate rapid shifts in technology and requirements . An adaptable approach that prioritizes essential processes while allowing flexibility is more suitable .
Process evolution benefits from a grassroots approach as it encourages organic growth and adaptation based on real user needs and feedback, without the rigidity of formal top-down enforcement. This approach allows processes to develop naturally from the insights and innovations of individuals involved in the work, fostering inclusivity and ownership that can lead to more sustainable and accepted changes . It promotes flexibility, empowering teams to identify practical solutions to issues and experiment with changes that result in effective process improvements .
The evolution of a bug metrics system exemplifies process evolution in a chaotic environment by showcasing incremental developments driven by necessity and feedback. Initially informal and experimental, the system expanded and adapted over time in response to the changing needs of teams, thereby incorporating new features and refinements. This reflects a dynamic adaptation where processes evolve naturally rather than through rigid change mandates, emphasizing flexibility and responsiveness as critical elements in handling chaos .
Leadership in software development involves managing organization, communication, motivation, and conflict resolution. It is essential because it enables teams to achieve goals and improve processes through human attributes like creativity and commitment, without strict reliance on process control or risk management. Leadership drives the other two strategies by fostering an environment where people are motivated to tackle risks and adhere to processes that align with project goals .
Lessons from the failure of a metrics-based projection system highlight the unpredictable nature of software development, where numerous variables can affect outcomes. While initially promising results may offer false confidence, the system's limitations reveal the necessity for cautious interpretation of metrics and the need for iterative adjustments based on feedback and results. This failure underscores the importance of remaining adaptable and reflective when applying automated projections to complex projects, ensuring that metrics are used as one of many tools rather than definitive predictors .