Variability and uniformity in designing systems for resilience

“Resilience is the capacity to recover quickly from difficulties, the ability to spring back into shape”

If this is the accepted definition of resilience then it’s not surprising that most resilience thinking is around getting things back to the steady state it was in before whatever shake-up occurred. It’s easy to see how traditional management techniques from a mechanistic worldview would assume that resilience in systems is best achieved by uniformity throughout the system, just as manufacturing ball bearings is improved by removing variation and the tolerance for variability.

I wonder if maybe we don’t want to ‘spring back into shape’ but instead form a different shape, evolve based on our response to the difficulties. From this idea I visualise five layers to a system that are organised by how much variability and uniformity is required to ensure the system is able to adapt to difficulties.

The top layer is for Individuals, it has the most variability and least uniformity, and allows the people in the system to be flexible, creative and solve complex problems in innovative ways.

The second layer is the Team. This is where we start to see some uniformity applied to things like roles and responsibilities but still have more variability to empower the team to change the ways they work with ease.

The third and middle layer is Process. Here there are equal amounts of variability and uniformity. There are standardised approaches to completing step-by-step processes but those processes are open to inspection and adaption. Given that there is an equal intersection of variability and uniformity, Process is where changes that affect the whole system can be made most easily.

The fourth layer is the Applications layer. This has greater uniformity with some variability expressed as fixed functionality in the application that can be used in a variety of ways.

The fifth and final layer is Data. This should be as uniform as possible with minimum variability, using a fixed architecture to . Variability down here prevents reliability and so should be avoided.

Thinking of system design from this intersecting scales of variability and uniformity helps to inform how we can build systems that are resilient to difficulties through being able to adapt to change rather than always seeking to return to a previous steady state, which doesn’t prepare the system for facing future difficulties.

Velocity as a measure for products not teams

I’ve been thinking about velocity as a measure for teams and products. The definition of the word is ‘the speed of something in a given direction‘, not just speed as we often think.

Scrum measures velocity, defined as “the amount of work a Team can tackle during a single Sprint … is calculated at the end of the Sprint by totaling the Points for all fully completed User Stories”, as speed alone. USpS is the MpH of the team, it is ‘output velocity’.

So, in this way of thinking, velocity is a team performance metric. It’s narrow, used to understand only the speed of the team, and doesn’t include the direction element from our dictionary definition. The issues that we see with using USpS to measure team speed alone is that the team could easily be moving quickly in the wrong direction (I guess the assumption in Scrum is that direction is provided in other ways), and that measuring human beings in such a mechanistic way is fraught with all kinds of inequalities, assumptions, and biases to the point where it becomes more damaging to the team than it is helpful.

But that doesn’t mean we have to abandon velocity all together. There are other ways of thinking about it as a useful measure. We could define velocity more broadly as ‘speed in the right direction’. Then, this ‘impact velocity’ could be used more to understanding the performance of the Product as it advances towards its goal state, rather than the team as in Scrum. The same team can measure impact velocity across multiple Products, and compare them, and learn from each other.
So, why measure impact velocity at all? If ‘velocity = speed in the right direction’, then the reasons to measure it are to check direction and course correct, and the sooner this is done because there is pace in achieving goals the more likely the team are to achieve mission.

Quality has to be part of our definition of impact velocity, and something that Scrum seems to be criticised for lacking and the resultant shipping of bugs just to get as many user stories completed in that sprint. Velocity is speed in the right direction, not just speed, so quality along the way; quality thinking, quality customer insight, quality deciding, quality building, quality shipping, quality feedback, provide the team with the ability to correct the course of the products and head in the right direction with more speed.

At least now I have a bit of a working definition that I can use to think about and test was of measuring impact velocity on real products.