In its annual review, the UK’s National Cyber Security Centre called for ‘radical transparency’ around technology. The report’s authors quipped that we know more about what’s in our sausages than our software – and that this allows potential risks to exist that are exploited later. Those risks exist because critical business decisions are made in a fog of technical jargon and ambiguity.
At the heart of this is risk appetite – how much risk is a business willing to accept as part of its operations? Operating at no risk is impossible; if you will not accept any risk, then you would not get out of bed in the morning, let alone run a business with potential profit and loss. It is also not affordable to reduce every single risk to zero. However, operating without any concern for risk is also not possible, as threats and potential losses have to be considered and prevented. The traditional challenge here is how to set that risk appetite effectively so that you can push ahead on investments that make sense and delay or prevent spending where it does not make sense. The issue is that risk appetite has become a boardroom euphemism, a convenient way to justify underinvestment without a rigorous analysis of the consequences. Without specific metrics, teams can state they have a ‘moderate risk appetite’, and the conversation ends without any effective decision made. To achieve the radical transparency our industry needs, we must shift our entire mindset to quantifying what risk surface we have across the organization and how that risk level changes over time.
Removing ambiguity around risk operations
The first step to quantify risk is to remove imprecision around our terms. Risk appetite remains the board’s high-level, strategic statement on the level of risk that the organization is willing to accept to achieve its goals. Under this, risk tolerance is the specific, measurable, and tactical deviation in risk levels that is permissible. Risk tolerances should be defined as numbers – for example: “Our tolerance for unplanned downtime on our e-commerce platform is a maximum of four hours per quarter.”
Alongside risk tolerance, you should also have risk thresholds, which are specific triggers that can be exceeded. In the case of the e-commerce platform above, “If downtime exceeds two hours, notify the CIO. If it exceeds four, the CIO should notify the board”.
To work effectively with your board around risk, you will require a ruthless intolerance for ambiguity around risk on both sides. The quickest way to achieve this is to put all of these discussions into monetary terms. This framework forces IT security to have that quantitative discussion with others across the business. It moves the dialogue from the subjective (“Are we comfortable with this?”) to the objective (“Are we willing to accept a potential loss of £5 million from this system being offline for 24 hours?”). After looking at the most critical issues in your business, you can then expand this into a wider risk operations process. While it might seem difficult to put an accurate price on each and every risk, it does not mean that we should not start. The goal here is to make risk management into a more operational process, rather than remaining hard to define or woolly. The definition of risk will improve in accuracy over time based on real-world experience.
Centralising risk operations
One essential element in this process is how to centralise these conversations and data into one process. Like the Security Operations Centre made it easier to respond to potential intrusions using data from across IT systems, the concept of a Risk Operations Centre is to bring together all the risk signals that an organization might have into one place for decision-making. Whereas a SOC looks backward to help deal with issues and incidents once a fault is discovered, a ROC looks forward to try and prevent faults before they take place. Using monetary data and risk tolerance, it is possible to see those changes in risk levels and then take action to reduce potential impact.
Bringing all this data together into a unified risk view also allows us to report metrics that the board can actually use for strategic oversight. The three most important rates to track are risk arrival, risk survival, and risk departure, otherwise termed burndown.
The risk arrival rate measures how quickly new risks that exceed the defined risk tolerance are discovered. This is a direct measure of the effectiveness of our preventative controls. On the other side, risk departure tracks how quickly those risks are eliminated or removed from the system. It measures the efficiency of the organization’s detection, prioritisation, and remediation processes. It is possible to have lots of risks arriving, but as long as those risks are removed as quickly, then the problem does not grow. The challenge exists when the number of risks coming in exceeds what the team can deal with in a timely manner.
In this scenario, risk survival measures the longevity of risk in our environment, or how long a vulnerability persists from discovery to departure. Risk survival is a direct measure of the speed of remediation and how efficient the team is at removing issues.
A simple chart showing these three trend lines tells the board everything they need to know: are we winning or losing the war on risk? Are our investments in prevention slowing the arrival of new problems? Is our remediation engine keeping pace with new arrivals (departure rate) and are we fixing the most critical things fast enough (survival rate)? Over time, the rate of removing issues should improve to take out more of the risk backlog. The security team should also prioritise those critical issues that might affect systems generating revenue, reducing how long those priority issues exist, while lower-risk threats can be managed or eliminated automatically.
Speaking the language of business around risk helps IT security and resilience teams to improve their approach. By speaking in terms of money, it also makes it easier to get support over time for any investments that are needed. We can move from managing technology to managing risk, based on making better, faster, and more defensible decisions around what risks to eliminate, which ones to manage, and which ones to offload or transfer. Adopting a disciplined risk lexicon makes it easier to achieve the radical transparency our organizations need.
The author
Ivan Milenkovic is Vice President Risk Technology EMEA at Qualys a cloud security company. Ivan leads work with customers on their risk strategies across their operations. Prior to joining Qualys, Ivan held roles as a Global Cyber Consulting Head of Operations with Atos and Global CISO for WebHelp.






