Attack Economics
Where Do All the Attacks Go?
Florencio, Herley 2013 - Economics of Information Security and Privacy III
Why most users are not harmed
- The fact that a majority of Internet users appear unharmed each year is difficult to reconcile with a weakest-link analysis. If security is only as strong as the weakest-link then all who choose weak passwords, re-use credentials across accounts, fail to heed security warnings or neglect patches and updates should be hacked, regularly and repeatedly. Clearly this fails to happen. Two billion people use the Internet; the majority can in no sense be described as secure, and yet they apparently derive more use from it than harm. How can this be? Where do all the attacks go?
- We seek to explain this enormous gap between potential and actual harm. The answer, we find, lies in the fact that an Internet attacker, who attacks en masse (targeted attacks are thus not included in this analysis), faces a sum-of-effort rather than a weakest-link defense. Large-scale attacks must be profitable in expectation, not merely in particular scenarios. For example, knowing the dog’s name may open an occasional bank account, but the cost of determining one million users’ dogs’ names is far greater than that information is worth. The strategy that appears simple in isolation leads to bankruptcy in expectation. Many attacks cannot be made profitable, even when many profitable targets exist.
- The model where a single user Alice faces an attacker Charles fails to capture the anonymous and broadcast nature of web attacks. Indeed, it is numerically impossible: two billion users cannot possibly each have an attacker who identifies and exploits their weakest-link. Instead, we use a cloud threat model where a population of users is attacked by a population of attackers.
- A crowd of users presents a sum-of-effort rather than a weakest-link defense. Many attacks, while they succeed in particular scenarios, are not profitable when averaged over a large population. This is true even when many profitable targets exist and explains why so many attacks types end up causing so little actually observed harm. Thus, how common a security strategy is, matters at least as much as how weak it is. Even truly weak strategies go unpunished so long as the cost of the failures exceeds the gain from the successes.
Threat model
- Threat models often describe the technical capabilities of an attacker. A defender Alice is pitted against an attacker Charles, who can attack in any manner consistent with the threat model.
- There are several things wrong with this threat model.
- It makes no reference of the value of the resource to Alice, or to Charles.
- It makes no reference to the cost of defence to Alice, or of the attack to Charles.
- It makes no reference to the fact that Charles is generally uncertain about the value of the asset and the extent of the defence (i.e., he doesn’t know whether benefit exceeds cost until he attacks successfully).
- It makes no provision for the possibility that exogenous events save Alice, even when her own defence fails (e.g., her bank catches fraudulent transfers).
- It ignores the fact that Charles must compete against other attackers
- It ignores the fact that Charles must compete against other attackers
- It ignores scale: assuming that Internet users greatly outnumber attackers it is simply numerically impossible for every user to have an attacker who identifies and exploits her weakest-link.
Some high-value users may face this threat model, but it is not possible that all do.
- In our threat model a population of Internet users are attacked by a population of hackers. We call the Internet users Alice(i) for i \= 0, 1, · · · , N−1 and the attackers Charles(j) for j \= 0, 1, · · · , M−1. Clearly N >> M : Internet users outnumber attackers.
- Each Alice(i) is subjected to attack by many of the attackers.
- Each attacker goes after as many Internet users as he can reach.
- Cost is the main reason for this approach: it costs little more to attack millions than it does to attack thousands.
-
The attackers’ goal is purely financial. None of the Alice(i)’s are personally known to any of the Charles(j)’s
-
The attacks we study happen in a competitive economic landscape. An attack is not merely a technical exploit but a business proposition. If it succeeds (and makes a profit) it is repeated over and over (and copied by others). If it fails (does not make a profit) it is abandoned and the energy is spent elsewhere.
- Attackers are playing a “numbers game”: they seek victims in the population rather than targeting individuals. For example, if Charles(j) targets Paypal accounts, he isn’t seeking particular accounts but rather any accounts that he happens to compromise.
-
Charles(j) doesn’t know the value, or security investment of any particular Internet user in advance. He discovers this only by attacking.
-
This threat model has some obvious short-comings. It excludes:
- Cases where the attacker and defender are known to each other
-
Cases where the attacker is targeting Alice(i) alone or values her assets beyond their economic value (e.g. APT).
-
There are many possible attacks, call them attack(0), attack(1), · · · , attack(Q− 1) (I think an attack in this context more or less corresponds to a sequence of attack techniques in MITRE ATT\&CK).
- We consider “parallel” attacks:
- Cj (N, k) is the cost to Charles(j) of reaching N users with attack(k).
- Attacking each user with attack(k) individually, one after the other, has much bigger cost:
Expected Loss
- Li is the loss that Alice(i) endures when she succumbs to any attack. This loss is independent of the attack type.
- The effort that Alice(i) devotes to defending against attack(k) is ei(k). Pr{ei(k)} is the probability that she succumbs to this attack (if attacked) at this level of effort. We assume that Pr{ei(k)} is a monotonically decreasing function of ei(k). The greater effort Alice(i) spends on attack(k), the lower her probability of succumbing.
-
There is some chance that, even though she succumbs to attack, Alice(i) suffers no loss because she is saved because of exogenous events. For example, her bank password falls to Charles(j) but her bank detects the fraud and saves her from harm. Here, we use Pr{SP} to denote the probability that her Service Provider saves Alice(i) from harm. (this might correspond also to a backup that recovers everything after a ransomware).
-
The expected loss for Alice is the probability that exogenous events do not save her, times the probability that she succumbs to any of the attacks, times her loss, plus the sum of what she spends defending against all attacks:
Expected Gain
- Gi is the gain (if successful) from Alice(i)
- The expected gain with attack(k) is the probability that exogenous events do not stop his fraud, times the sum of the probable gain over all attacked users, minus the total cost of the attack
-
The summation is over as many users as Charles(j) attacks. We do not assume that all users are attacked, only that the number of attacked users is large enough for statistical arguments to apply.
-
We assume that Gi \<=Li
- Other cases might exist that are not explored here. For example Charles(j) might benefit without harming Alice(i), in which case Gi >> Li
Attack Selection
- Each attacker chooses the attack that maximizes his expected gain.
- The selected attack must satisfy the following condition:
- Thus, if a sufficient number of users have a good defense for attack(k) (average probability of success is “sufficiently low”), then attack(k) is unprofitable.
- Choosing that attack(k) is not a rational choice for the attacker.
-
Different attackers might choose different attacks, for example because they have a different cost structure. However, it is likely that some attacks give the best returns to many attackers and some other attacks are best for almost none.
-
A given attack(k) must be profitable at scale. Not on specific targets, or in specific conditions.
- An attack(k) that is not profitable at scale will be rarely seen because choosing that attack(k) is not economically rational.
Counterintuitive consequences
- A certain attack(k) may not be profitable at scale even if:
- Some users devote zero or very low defense for attack(k) (free rider effect)
- The attacker does not know in advance which are the users where attack(k) will succeed.
- It occasionally works with low cost (e.g., users with very bad passwords or information exposed on the Internet without any protection). Attack cost C(1,k) is small and the probability of success is not zero
- The attacker does not know in advance which are the users where attack(k) will succeed.
- The attack cost will be N*C(1,k), which is bigger than C(N,k).
-
It occasionally gives access to high value assets (e.g., email passwords that unlock high value).
- Most Gi are small, some Gi are large. The attacker does not know in advance which ones.
-
A certain attack(k) may be rarely/never seen even for other reasons:
- There are economically better alternatives (attacks with lower cost; attacks to assets with higher expected value).
- Its probability of success is very low because it is applicable to a tiny fraction of all the users (e.g., rarely used software). Diversity is important.
- See also next paper:
- The problem for the attacker is that profitability is not directly observable. It is not obvious who will succumb to most attacks and who will prove profitable.
- The cost of false positives (unprofitable targets) can consume the gain from true positives. When this happens, attacks that are perfectly feasible from a technical standpoint become impossible to run profitably.
Contrast with threat model individual attacker
- In a different threat model where one attacker faces one target, the attacker selects the attack(k) that maximizes:
- This is maximized by the attack(k) with the highest probability of success/cost ratio:
- In this threat model the target cannot neglect any defense (weakest-link)
- Devoting zero effort to attack(k) means that the probability of success for attack(k) is close to 1. Thus, if the cost for executing that attack is sufficiently low, then that attack(k) will be selected.
When Does Targeting Make Sense for an Attacker?
Herley 2013 - IEEE Security & Privacy ( Volume: 11, Issue: 2, March-April 2013)
- Scalable attacks have costs that grow much slower than linearly in the number, N, of users attacked. Doubling the number attacked causes the costs to increase by far less than a factor of two
- Non-scalable attacks are everything else. Generally they have a linear cost dependence on N. Doubling the number attacked doubles the cost.
-
This segmentation is obviously a simplification. Even spam has a linear cost component (e.g. gathering target addresses, finding enough machines and IP addresses to do the sending). However, first-copy costs dominate, so that doubling the size of the attack has little effect on the overall cost. Equally, attacks that are scalable may need to be followed by a non-scalable component. While phishing may harvest passwords in bulk, the process of cashing out might be non-scalable and one-by-one.
-
Attacks that involve knowledge about the target are not scalable.
- As a consequence, scalable attacks are:
- non-adaptive. Personalization and customization are very limited in a scalable cost model. This is why 419 (Nigerian-style) and lottery scam emails generally begin “Dear Sir/Madam” or “Attn: Beneficiary.” A script can accommodate segments of the population (e.g. a malicious server might attempt different exploits depending on the browser version) but individual-level customization is out (today there are more practical options for customization, due to LLMs).
-
non-selective. They attack anyone and everyone.
-
Because of their cost structure, scalable attacks generally reach orders of magnitude more users.
-
Since the marginal cost per additional user is close to zero, it makes no sense to leave reachable targets un-attacked.
-
Having costs that grow slower than linearly is the exception rather than the rule. The vast majority of attacks do not have this property. On the other hand, there is an almost unlimited suite of non-scalable attacks, but on a per-user basis they are far more expensive.
- Spam-based attacks, for example, seem to be able to attack users for pennies per million. Non-scalable attacks could easily cost six or seven orders of magnitude more than this per attacked user. Their greater cost suggests that non-scalable attacks must be reserved for cases where the value of the target is extremely high.
Competing against Scalable attacks
- Attack revenue is NYV whereN \= number of targets, Y \= fraction of targets that succumb (yield), V \= average value extracted from those that succumb
- The nonscalable attacker has a structural disadvantage because for a given cost N_NS \<\< N_S. It must thus compensate in YV. This can only happen when V_NS >> V_S:
- If V_NS=V_S then it must be Y_NS>>Y_S. But V_S is very low and N_S is orders of magnitudes greater than N_NS. Thus in this case it is unlikely that the nonscalable attack is a rational choice.
-
It must be V_NS >> V_S, so that at least some of the orders of magnitude lost in reach (N_S vs N_NS) are made up in extracted value.
-
A non-scalable attack requires two things to compete successfully.
- First, there must be some users who have much higher extractable value than the average. If the attacker’s cost per user is orders of magnitude higher then he needs to extract orders of magnitude more when he succeeds.
-
Second, it must be observable who those high-value users are. It does the attacker no good to know that they exist, if he doesn’t know where.
-
Concentration and observability of the extractable value are absolute requirements for the non-scalable attacker.
- Suppose that value is uniformly distributed: every user has equal value. Non-scalable attacks make no sense in this case since the scalable attacks gather them at far lower cost.
- Suppose that value is concentrated, but entirely unobservable. Again, the non-scalable attacker can’t compete if there are good victims out there, but they just can’t be found.
- A little better is any distribution where the variance is high, so that at least some users have significantly higher value. Best are heavy-tail distributions where a large portion of the overall value lies with a few individuals. Many phenomena follow power law distributions (e.g. Pareto).
- In these distributions, however, the mean is higher than the median, so most people have below average value. In other words, for the distributions that favor non-scalable attacks, the vast majority of people have below average value. Since the non-scalable attacker needs much higher than average value, he must leave the vast majority of users alone.
Security, cybercrime and scale
Herley 2014, Communications of the ACM, Volume 57, Issue 9, Sep 2014
Financially motivated attacks
- To be sustainable there must first be a supply of profitable targets and a way to find them. Hence, the attacker must do three things:
- decide who and what to attack,
- successfully attack, or get access to, a resource,
- monetize that access
- Turning access into money is much more difficult than it looks.
-
(This is no longer true today, because of cryptocurrencies and ransomware/double extortion)
-
When attacks are financially motivated, the average gain for each attacker, E{G}, must be greater than the cost, C: E{G} – C > 0.
- C must include all costs, including that of finding viable victims and of monetizing access.
-
The gain must be averaged across all attacks, not only the successful ones.
-
The problem for the attacker is that profitability is not directly observable. It is not obvious who will succumb to most attacks and who will prove profitable.
-
The cost of false positives (unprofitable targets) can consume the gain from true positives. When this happens, attacks that are perfectly feasible from a technical standpoint become impossible to run profitably.
-
Most assets escape exploitation not because they are impregnable but because they are not targeted.
- This happens not at random but predictably when the expected monetization value is less than the cost of the attack.
Scalable attacks
- It is clear that the majority of users are regularly attacked by attacks that involve very low cost per attacked user.
-
I find it useful to segment attacks by how their costs grow. Scalable attacks are one-to-many attacks with the property that cost (per attacked user) grows slower than linearly; for example, doubling the number of users attacked increases the cost very little: C(2N) \<\< 2 C(N).
-
Scalable attacks are a minority of attack types. It is the exception rather than the rule that costs have only weak dependence on the number attacked. Anything that cannot be automated completely or involves per-target effort is thus non-scalable.
- Physical side-channel attacks (requiring proximity) are out, as getting close to one million users costs a lot more than getting close to one.
- Labor-intensive social engineering attacks are non-scalable. After an initial scalable spam campaign, the Nigerian scam (and variants) devolves into a non-scalable effort in manipulation.
- Spear-phishing attacks that make use of information about the target are non-scalable. While the success rate on well-researched spear-phishing attacks may be much higher than the scatter-shot (such as “Dear Paypal customer”) approaches, they are non-scalable.
- Attacks that involve knowledge of the target are usually non-scalable; for example, guessing passwords based on knowledge of the user’s dog’s name, favorite sports team, or cartoon character involves significant non-scalable effort. Equally, attacks on backup authentication questions that involve researching where a user went to highschool are non-scalable.
- (Some of these considerations might no longer apply today, due to LLMs and large botnets.)
- ASLR converts scalable attacks to non scalable.
Finding viable targets is hard
- Assume Mallory can estimate a probability, or likelihood, of profit, given everything he observes about a potential target. This is the probability that the target succumbs and access can be monetized (for greater than average cost). Call this P{viable|obs.}.
- We assume the cost of gathering the observables is small relative to the cost of the attack. This makes the problem a binary classification, so receiver operator characteristic (ROC) curves are the natural analysis tool.
-
The paper shows that finding viable targets in a cost-effective way is very hard. A classifier with 99% accuracy is likely to be needed and a population with a large absolute value of viable targets is also needed.
-
The attacker needs concrete observable features to estimate viability. If he does not get it right often enough, he makes a loss. The problem of how viable niches for a particular attack can be identified is worth serious research. If they can be identified, members of these niches (rather than the whole population) are those who must invest extra in the defense.
Which attacks can be neglected?
-
I concentrate on attacks that are financially motivated, where expected gain is greater than cost.
-
Scalable attacks represent an easy case.Their ability to reach vast populations means no one is unaffected. In the question of trade-offs it is difficult to make the case that scalable attacks are good candidates to be ignored. Fortunately, they fall into a small number of types and have serious restrictions, as we saw earlier. Everyone needs to defend against them.
-
Non-scalable attacks present our opportunity; it is here we must look for candidates to ignore. Cybercriminals probably do most damage with attacks they can repeat and for which they can reliably find and monetize targets. I suggest probable harm to the population as a basis for prioritizing attacks:
- Attacks where viable and non-viable targets cannot be distinguished pose the least economic threat. If viability is entirely unobservable, then Mallory can do no better than attack at random.
- Non-scalable attacks with low densities of viable targets are smaller threats than those where it is high (see mathematical treatment in the paper).
- The more difficult an attack is to monetize the smaller the threat it poses.
(note that cryptocurrencies and ransomware/double extortion, not known at the time this paper was written, fit all these criteria)