pProbabilityRULES THEM ALL Back to the experiments

A complete book for beginners

Probability Made Easy
with Taleb.

Start with coins and dice, then learn about conditional probability, expected value, and risk. Each of the 16 chapters has three exercises, with answers at the end. Basic arithmetic is all you need to begin.

16 chapters48 exercisesNo calculus neededExamples, formulas, and worked answers

18 sections in the contents

Contents 16 chapters + appendix

Probability Made Easy with Taleb

Start with a coin

A fair coin has a one-in-two chance of landing heads. Does that mean ten tosses must give five heads? If the last five tosses were heads, is tails now due?

These questions require little mathematics, but they are easy to get wrong. They are where this book begins.

If you can add, subtract, multiply, divide, and read fractions, you can start here. We explain formulas through concrete examples before introducing the symbols. Each chapter has three exercises, with answers at the end. When something is unclear, try it with a coin, a die, or a sheet of paper.

Why learn with Taleb?

Nassim Nicholas Taleb studies how randomness affects our judgments and what those judgments can cost us. Drawing on his approach to risk, this book keeps asking two questions: how likely is an outcome, and what happens if it occurs?

Being right ninety-nine times can still be a bad bargain if one mistake wipes you out. Understanding probability means considering both the chances and the size of the gains or losses.

One game will recur in the book. You start with 10 chips. Each round uses an independent toss of a fair coin: win 1 chip or lose 1 chip. Stop at 20 chips or at 0. Every round is fair, but what is your chance of losing everything before reaching the target? By Chapter 13, you will know how to answer.

This is an independently written introduction. Taleb did not write or review it. The mathematics comes from probability theory; his articles and excerpts collected in the Heron Notes provide a perspective on risk. Sources and reading suggestions appear in the appendix.

Contents

Part I: Start with everyday questions

  1. What probability can tell us: state the conditions before discussing the chances.
  2. List the possible outcomes: when counting is enough.
  3. Why ten tosses need not give five heads: probability versus the record you observe.

Part II: Learn the basic methods

  1. When to add and when to multiply: “or” versus “both.”
  2. How to calculate “at least once”: begin with “not even once.”
  3. New information, new probabilities: why the denominator changes.
  4. Independent or mutually exclusive?: can both happen, and does one change the chance of the other?
  5. A positive test: how likely is illness?: remember the prevalence.

Part III: Clear up common mistakes

  1. Why shared birthdays are not so rare: matching you is different from matching any pair.
  2. Three doors: should you switch?: the host's rules matter.
  3. An average gain is not a promised gain: understand expected value.
  4. Five heads in a row: what next?: independence and the law of large numbers.

Part IV: Use probability to make judgments

  1. Fair rounds can still leave you broke: starting capital and stopping rules.
  2. Probability 1 can allow exceptions: understand “almost surely.”
  3. Rare, large losses: understanding fat tails and what an average may miss.
  4. Before deciding, ask what you can afford to lose: consider probability and consequences together.

How to read this book

On your first reading, aim to understand one problem per chapter. Explain it in your own words, then try the three exercises. If your answer is wrong, check what event you calculated before checking your arithmetic.

The later chapters are harder. Chapter 14, in particular, explains why “probability zero” and “impossible” need not mean the same thing in a continuous model. We begin with a line of length 1. If the idea still feels unfamiliar, set it aside and return to it later.

The games, test parameters, and everyday situations in the exercises are teaching examples. The birthday problem, the three-door problem, and gambler's ruin are classic problems. Every answer depends on the stated assumptions.

What probability can tell us

What does a one-in-two chance mean?

Toss a coin. It may land heads or tails. Suppose the two sides are equally likely, and ignore possibilities such as landing on its edge. We then say the probability of heads is one half.

Probability describes how likely an event is, using a number between 0 and 1. One half can be written as 1/2, 0.5, or 50%. These mean the same thing.

Here, 50% does not tell us which side the next toss will show. Nor does it require two tosses to produce exactly one head and one tail.

First say how the probability is assigned

Together, these assumptions form a probability model. A model specifies the possible outcomes and the probabilities assigned to them.

Our coin has two outcomes, each with probability one half. A biased coin still has those two outcomes, but their chances need not be equal. Calling an object a coin does not establish a 50% chance of heads.

Throughout this book, a “fair coin” means equal chances of heads and tails. A “fair die” means a six-sided die whose six faces are equally likely.

Two examples

One toss of a fair coin has probability 1/2 of heads. One roll of a fair die has probability 1/6 of a six. In both cases, the elementary outcomes are equally likely, so we divide the number of favorable outcomes by the total number of outcomes.

Now ask a different question: will ten tosses of a fair coin guarantee five heads? No. Three, five, or seven heads are all possible. A one-half chance on each toss does not require every group of ten to split evenly.

When reading a probability, get into the habit of asking how it was obtained and what it actually tells you.

Try it

Hover over or tap the links below for hints and answers. Tap again to close.

  1. Roll a fair six-sided die once. What is the probability of a six?
    Hint
    Question 1 · Hint

    Count the equally likely faces, then count how many show a six.

    Show answer
    Question 1 · Answer

    1/6. The six faces are equally likely, and one of them is 6.

  2. A commemorative coin has a heads design on both sides. Ignoring the possibility of landing on its edge, is the probability of seeing heads still one half?
    Hint
    Question 2 · Hint

    Consider each side separately: what do you see whichever side lands upward? Do not assume this behaves like an ordinary coin.

    Show answer
    Question 2 · Answer

    It is 1, not one half. Under the stated conditions, either side landing upward shows a heads design.

  3. “There is a 40% chance of rain here tomorrow.” What further information does this statement need?
    Hint
    Question 3 · Hint

    Where is ‘here’, what time interval is meant, and how much rain counts as rain? Who supplied the estimate?

    Show answer
    Question 3 · Answer

    Specify the location, the time interval, and what counts as rain. Also identify the source of the estimate, such as the forecasting agency.

List the possible outcomes

Start with the six faces

A fair die can show 1, 2, 3, 4, 5, or 6, each equally likely. To find the probability of an even number, count 2, 4, and 6: three of the six outcomes. The answer is 3/6 = 1/2.

The collection of all possible outcomes is the sample space. An event is a selection of those outcomes. For example, the event “an even number” contains 2, 4, and 6.

You do not have to memorize the terms immediately. First learn to list the outcomes and identify which ones the question asks about.

When is counting enough?

You can divide counts directly only when the elementary outcomes are equally likely:

Probability = number of favorable outcomes ÷ total number of outcomes

A bag contains three red balls and one blue ball. If each ball is equally likely to be drawn, the probability of red is three quarters.

There are two colors, but that does not give each color a one-half chance. The four balls are equally likely; three of them share the same color. Be clear about what you are counting.

How many outcomes do two coin tosses have?

Assume a fair coin and independent tosses: the first result does not change the chance on the second. Write the first toss first and the second toss second. There are four equally likely outcomes:

First tossSecond tossShort form
HeadsHeadsHH
HeadsTailsHT
TailsHeadsTH
TailsTailsTT

“Exactly one head” includes HT and TH, so its probability is 2/4 = 1/2.

“At least one head” also includes HH, making its probability 3/4. A small change in wording changes which outcomes count.

A common mistake is to treat “two heads, one head, no heads” as three equally likely outcomes. These categories cover everything, but “one head” contains two orders and is twice as likely as either of the other categories.

Try it

Hover over or tap the links below for hints and answers. Tap again to close.

  1. Roll a fair die once. What is the probability of a number less than 3?
    Hint
    Question 1 · Hint

    List the faces below 3, then compare their count with all six faces.

    Show answer
    Question 1 · Answer

    The outcomes less than 3 are 1 and 2, so 2/6=1/3.

  2. Toss a fair coin independently twice. What is the probability of at least one tail?
    Hint
    Question 2 · Hint

    First calculate the chance of no tails, then subtract from 1.

    Show answer
    Question 2 · Answer

    No tails means HH, with probability 1/4. Therefore at least one tail has probability 1−1/4=3/4.

  3. A bag has three red balls and one blue ball, each equally likely to be drawn. What is the probability of red? Why should you not split the probability equally between the two colors?
    Hint
    Question 3 · Hint

    Each ball is equally likely. Count the red balls and the total number of balls.

    Show answer
    Question 3 · Answer

    3/4. The four balls, not the two colors, are equally likely. Red includes three elementary outcomes; blue includes one.

Why ten tosses need not give five heads

Probability and the observed record are different

Suppose ten tosses of a fair coin produce seven heads. Heads account for 70% of this record, while our model assigns heads probability 50%. These different numbers are not a contradiction.

The 50% is a probability assigned by the model. The 70% is the proportion observed in ten tosses, called the relative frequency. To calculate it, divide the number of occurrences by the total number of trials:

7 ÷ 10 = 70%

With only ten tosses, considerable fluctuation is normal. This record might prompt you to investigate whether the coin is fair, but it does not by itself prove that the coin is biased.

What happens with more tosses?

If the coin remains fair and the tosses are independent, the proportion of heads will usually be close to one half after sufficiently many tosses. Chapter 12 develops this idea.

This does not mean every additional toss brings the proportion closer to one half. Nor must a later run compensate for an earlier imbalance. After seven heads in ten tosses, the next ten could also contain more heads than tails.

More trials are not a universal guarantee. If you change the coin halfway through, or keep only records with many heads, the resulting proportion no longer answers the original question.

Look at the count behind the proportion

Shop A makes two deliveries and one is late: 1/2 = 50%. Shop B makes 100 deliveries and one is late: 1/100 = 1%. The same count of late deliveries means something very different with a different total.

Now compare two surveys. Four of five respondents agree in one; 4,000 of 5,000 agree in the other. Both report 80%. In the small survey, one person changing their answer moves the percentage substantially. The larger survey is less sensitive to one person, but you must still ask how respondents were chosen and whom they represent.

A proportion needs at least two companions: the number of observations and the method used to select them.

Try it

Hover over or tap the links below for hints and answers. Tap again to close.

  1. An outcome occurs 7 times in 20 trials. What is its relative frequency?
    Hint
    Question 1 · Hint

    Relative frequency is the number of occurrences divided by the total number of trials.

    Show answer
    Question 1 · Answer

    7/20=0.35, or 35%.

  2. A fair coin has landed heads four times in a row. Assuming independent tosses, what is the probability of tails on the fifth toss?
    Hint
    Question 2 · Hint

    Notice ‘fair’ and ‘independent’. Can earlier results change the next toss's chance?

    Show answer
    Question 2 · Answer

    1/2. Under the fairness and independence assumptions, the first four results do not change the fifth probability.

  3. Two surveys both report 80% support, but one questioned 5 people and the other 5,000. Why should you not consider the results equally reliable?
    Hint
    Question 3 · Hint

    How much would one changed answer affect each survey? Also ask how respondents were selected.

    Show answer
    Question 3 · Answer

    In a five-person survey, one changed answer moves the result by 20 percentage points. A 5,000-person survey is usually less sensitive to an individual response. But size cannot repair obvious selection bias, such as asking only supporters.

When to add and when to multiply

“Or” combines outcomes

What is the probability of rolling a 1 or a 2 on a fair die? Two of the six outcomes qualify, so the probability is 2/6 = 1/3.

You can also add their probabilities: 1/6 + 1/6 = 1/3.

But addition has a catch. Can you add the probabilities of “an even number” and “a number greater than 3” directly? List them first:

  • Even: 2, 4, 6.
  • Greater than 3: 4, 5, 6.

Both lists contain 4 and 6. Adding directly counts each of these twice, so subtract the overlap once:

3/6 + 3/6 − 2/6 = 4/6 = 2/3

The combined outcomes are 2, 4, 5, and 6. Here “or” means at least one condition is satisfied; satisfying both also counts.

Write it with letters

We often use A and B for events, and P(A) for the probability of A. The calculation becomes:

P(A or B) = P(A) + P(B) − P(A and B)

“And” means both conditions hold. If A and B cannot happen together—for example, one roll showing both 1 and 2—the overlap is empty. Its probability is 0, so simple addition works. Such events are mutually exclusive.

What if both events must happen?

Roll a fair die independently twice and ask for a six on both rolls.

The first roll has probability 1/6 of a six. Whatever the first result, the second still has probability 1/6. Multiply:

1/6 × 1/6 = 1/36

You could instead list all 36 equally likely ordered pairs. Only (6,6) qualifies.

Direct multiplication works here because the rolls are independent. If one result changes the chance of the next, do not copy this shortcut. Drawing two balls without replacement changes the bag before the second draw. That requires the conditional probability introduced in Chapter 6.

Try it

Hover over or tap the links below for hints and answers. Tap again to close.

  1. Roll a fair die once. What is the probability of a 1 or a 6?
    Hint
    Question 1 · Hint

    Can one roll show both 1 and 6? If not, their probabilities can be added directly.

    Show answer
    Question 1 · Answer

    1/6+1/6=1/3. A 1 and a 6 are mutually exclusive on one roll.

  2. What is the probability of an odd number or a number less than 4? Which outcomes would be counted twice by simple addition?
    Hint
    Question 2 · Hint

    List the odd faces and the faces below 4 separately. When merging the lists, keep each face only once.

    Show answer
    Question 2 · Answer

    Odd numbers are 1, 3, and 5; numbers less than 4 are 1, 2, and 3. Simple addition counts 1 and 3 twice. The combined outcomes are 1, 2, 3, and 5, giving 4/6=2/3.

  3. Toss a fair coin independently twice. What is the probability of heads first and tails second?
    Hint
    Question 3 · Hint

    Write the two individual probabilities. Which operation does independence allow?

    Show answer
    Question 3 · Answer

    Independence gives 1/2×1/2=1/4.

How to calculate “at least once”

Start with “not even once”

Toss a fair coin independently three times. What is the probability of at least one head?

“At least one” includes one, two, and three heads. You could count each case, but there is an easier way: first calculate no heads at all, which means three tails.

Three tails have probability 1/2 × 1/2 × 1/2 = 1/8. Every other outcome includes a head, so the answer is:

1 − 1/8 = 7/8

If A is “at least one head,” then “no heads” is its complement, written Aᶜ. The two events cover all outcomes without overlapping:

P(A) = 1 − P(Aᶜ)

From three trials to many

For four independent tosses, the same method gives 1 − (1/2)⁴ = 15/16.

The superscript 4 means multiply 1/2 by itself four times. If the probability of success is p, the probability of failure is 1−p. For n independent trials with the same success probability:

Probability of at least one success = 1 − (1−p)ⁿ

Both conditions matter: independence and the same probability on every trial. If a failure changes the next attempt, or the difficulty varies, calculate using the actual conditions.

A small chance per trial can add up

Suppose an event has only a 1% chance on each of 100 independent trials. What is the probability it happens at least once?

The probability of avoiding it on one trial is 99%. Avoiding it on all 100 has probability 0.99¹⁰⁰, about 36.6%. Subtract from 1: the probability of at least one occurrence is about 63.4%.

So a 1% chance each time does not leave you with only a 1% chance across 100 trials. But 1% × 100 = 100% is not the correct answer either. It double-counts outcomes in which the event happens more than once.

The Heron Notes record a related result. If each trial has probability 1/n and you make n independent trials, the probability of at least one occurrence approaches 63.2% as n grows. The formula is 1 − (1−1/n)ⁿ; its limit is 1−1/e, where e is about 2.718. If limits are unfamiliar, concentrate on the 100-trial example first.

Try it

Hover over or tap the links below for hints and answers. Tap again to close.

  1. Roll a fair die independently twice. What is the probability of at least one six?
    Hint
    Question 1 · Hint

    Find the chance of not rolling a six once, then of avoiding a six twice.

    Show answer
    Question 1 · Answer

    The probability of no six on either roll is (5/6)², so the answer is 1−(5/6)²=11/36.

  2. Each independent attempt has a 20% chance of success. What is the probability of succeeding at least once in three attempts?
    Hint
    Question 2 · Hint

    What is the chance of one failure? How would you calculate three failures in a row?

    Show answer
    Question 2 · Answer

    All three failures have probability 0.8³=0.512. The answer is 1−0.512=0.488, or 48.8%.

  3. An event has probability 1/20 on each of 20 independent trials. How do you calculate the probability of at least one occurrence? Give three decimal places.
    Hint
    Question 3 · Hint

    Put p=1/20 and n=20 into the ‘at least once’ formula. Round only at the end.

    Show answer
    Question 3 · Answer

    1−(19/20)²⁰≈0.642, or about 64.2%. A small probability on each trial can produce a much larger probability of at least one occurrence over many trials.

New information, new probabilities

Why does the denominator change?

A standard deck without jokers has 52 cards, including 13 hearts. Draw a card at random from a well-shuffled deck. Its probability of being a heart is 13/52 = 1/4.

Now someone tells you the card is red. Does the probability of hearts change?

Yes. The 26 black cards are ruled out. Of the 26 red cards left, 13 are hearts. Given that the card is red, its probability of being a heart is 13/26 = 1/2.

The card has not changed; our information has. A probability recalculated given a condition is called a conditional probability.

Reading the formula

Let A mean “the card is a heart” and B mean “the card is red.” P(A|B) reads “the probability of A given B.” The right side of the vertical bar is the known condition; the left side is what we are asking about.

The general formula is:

P(A|B) = P(A and B) ÷ P(B), provided P(B) > 0.

All hearts are red, so P(A and B) = 13/52. Divide by P(B) = 26/52, and the result is again 1/2.

At first, do not worry about memorizing the formula. Rule out outcomes inconsistent with the known condition, then calculate within those that remain.

Two more examples

Roll a fair die and learn that the result is greater than 3. The possibilities left are 4, 5, and 6. Two are even, giving a conditional probability of 2/3.

Now consider a family with two children. In this simplified model, each child is equally likely to be a boy or a girl, independently of the other. Write the older child first. There are four equally likely combinations: BB, BG, GB, and GG.

Choose a family at random from this model and retain only cases with at least one boy. That excludes GG. Of the three remaining combinations, only BB has two boys, so the conditional probability is 1/3.

If the information is instead “the older child is a boy,” only BB and BG remain. The probability of two boys is then 1/2.

This problem deliberately specifies both a simplified model and how information is selected. Meeting one child by chance, or hearing a parent choose which child to mention, can involve a different selection process. Do not assume the same answer applies.

Try it

Hover over or tap the links below for hints and answers. Tap again to close.

  1. Draw a card at random from a standard deck without jokers. Given that it is black, what is the probability it is a spade?
    Hint
    Question 1 · Hint

    Knowing the card is black rules out the red cards. What fraction of the remaining cards are spades?

    Show answer
    Question 1 · Answer

    Of the 26 black cards, 13 are spades, giving 13/26=1/2.

  2. Roll a fair die. Given that the result is odd, what is the probability it is greater than 3?
    Hint
    Question 2 · Hint

    Look only at the odd faces 1, 3, and 5. Which exceed 3?

    Show answer
    Question 2 · Answer

    Among the odd results 1, 3, and 5, only 5 exceeds 3. The answer is 1/3.

  3. In this chapter's two-child model, the younger child is a girl. What is the probability that both children are girls?
    Hint
    Question 3 · Hint

    List combinations in older-younger order and keep only those whose second child is a girl.

    Show answer
    Question 3 · Answer

    1/2. In the independent, equal-probability model, learning that the younger child is a girl does not change the chance that the older child is a girl.

Independent or mutually exclusive?

Mutually exclusive: they cannot happen together

One roll of a die cannot show both 1 and 6. These two events cannot occur together; they are mutually exclusive.

In other words, they share no outcomes. For events A and B, write A∩B = ∅. The middle symbol means intersection, the outcomes they share. The symbol on the right means the empty set, containing no outcomes.

Mutually exclusive events therefore have probability 0 of occurring together. The reverse does not hold in every model: an event with probability 0 can still contain outcomes. Chapter 14 explains this distinction.

Independent: learning one does not change the chance of the other

Toss a fair coin independently twice. Learning that the first toss was heads leaves the second toss's chance of heads at one half. The events “heads first” and “heads second” are independent.

They can happen together: HH does exactly that. Independence does not mean they cannot both happen.

Mathematically, independence means:

P(A and B) = P(A) × P(B)

If B has positive probability, you can also read this as: learning B has occurred does not change A's probability, so P(A|B) = P(A).

Compare events on the same die

On one roll, “a 1” and “a 6” are mutually exclusive. Are they independent? No. Before seeing the roll, the probability of 6 is one sixth. Once you know the result is 1, a 6 is impossible, with probability 0.

Now compare “an even number” and “a 6.” They are not mutually exclusive, since 6 satisfies both. Nor are they independent: learning that the result is even leaves 2, 4, and 6, increasing the chance of 6 from one sixth to one third.

Two events can be neither mutually exclusive nor independent. The terms answer different questions.

Ask in order: can they happen together? Does knowing that one occurred change the probability of the other? This is more reliable than guessing from the words alone.

Try it

Hover over or tap the links below for hints and answers. Tap again to close.

  1. On one toss of a fair coin, are heads and tails mutually exclusive? Are they independent?
    Hint
    Question 1 · Hint

    Ask separately: can both occur? Once you know it is heads, is tails still as likely as before?

    Show answer
    Question 1 · Answer

    Mutually exclusive, but not independent. One toss cannot be both heads and tails. Knowing it is heads changes the probability of tails from 1/2 to 0.

  2. Toss a fair coin independently twice. Are “tails first” and “heads second” mutually exclusive? Are they independent?
    Hint
    Question 2 · Hint

    Write a sequence satisfying both events, then ask whether the first result changes the second chance.

    Show answer
    Question 2 · Answer

    Not mutually exclusive, but independent. TH satisfies both events, and the first toss does not change the second toss's probability.

  3. On one roll of a fair die, are “an even number” and “a number greater than 3” mutually exclusive? Are they independent? List the shared outcomes before calculating.
    Hint
    Question 3 · Hint

    List the shared faces. Does their joint probability equal the product of the two individual probabilities?

    Show answer
    Question 3 · Answer

    They can occur together, sharing outcomes 4 and 6, so they are not mutually exclusive. The joint probability is 1/3, while the product is 1/2×1/2=1/4. These differ, so the events are not independent either.

A positive test: how likely is illness?

Detecting illness is not the same as confirming it

Suppose a test detects, on average, 99 out of every 100 people who have a disease. That sounds accurate. You receive a positive result. Does it mean you have a 99% chance of having the disease?

No. The first number describes how often the test detects people already known to have the disease. Your question runs the other way: among those testing positive, how many actually have it? The group you need to examine has changed.

Let us calculate with hypothetical numbers. This is a probability lesson, not data about any real disease or test.

Divide 10,000 people into two groups

Suppose disease prevalence in a population is 1%. Among 10,000 people, about 100 have the disease and 9,900 do not.

Suppose the test detects 99% of those with the disease, but also falsely reports a positive for 5% of those without it:

Actual conditionPeoplePositive results
Has the disease100100 × 99% = 99
Does not have it9,9009,900 × 5% = 495
Total10,000594

Now look only at the positive results. There are 594, of which 99 belong to people who really have the disease. Under these assumptions, the probability of disease given a positive result is:

99 ÷ 594 ≈ 16.7%

Why less than one in five? People without the disease vastly outnumber those with it. Even a small false-positive percentage can produce more positive results than the genuine cases.

This does not mean the test is useless. Before testing, a randomly chosen person's probability of disease was 1%. After a positive result, it is about 16.7%. The test provides information, but not certainty.

Try a different population

What if the same test is used in a population with 10% prevalence?

Among 10,000 people, about 1,000 have the disease and 990 test positive. Of the 9,000 without it, about 450 are falsely positive. There are 1,440 positive results, and the proportion with disease is:

990 ÷ 1,440 ≈ 68.8%

The test's performance has not changed, but the meaning of a positive result has. The starting population is different. The original proportion is often called the base rate.

So “this test is accurate” is not enough information. At minimum, ask which population it is used on, how many genuine cases it detects, and how many people without the disease it flags incorrectly.

Bayes' formula is already in the table

Recalculating a probability using new evidence is called Bayesian updating. Here the formula is:

P(disease | positive) = P(positive | disease) × P(disease) ÷ P(positive)

The known condition appears after the vertical bar. The numerator is the probability of both disease and a positive result. The denominator is the probability of any positive result. Dividing gives the proportion of positive results that are genuine cases.

If the formula is difficult to remember, remember the table. Separate genuine positives from false positives, then calculate only within the positive group.

Real test results need interpretation alongside symptoms, medical history, and testing conditions. This example teaches you to distinguish two probabilities; its invented percentages cannot tell you whether you are ill.

Try it

Hover over or tap the links below for hints and answers. Tap again to close.

  1. A hypothetical group of 1,000 people has 2% disease prevalence. The test detects every case. How many people with the disease are expected to test positive?
    Hint
    Question 1 · Hint

    Multiply the total population by prevalence. The question says every genuine case is detected.

    Show answer
    Question 1 · Answer

    1,000×2%=20 people.

  2. If the false-positive rate is 10%, how many of the other 980 people are expected to be flagged incorrectly?
    Hint
    Question 2 · Hint

    The false-positive rate applies to people without the disease, not to the whole group of one thousand.

    Show answer
    Question 2 · Answer

    980×10%=98 people.

  3. Combine the first two answers. What proportion of positive results are genuine cases? Why does detecting every case not mean every positive result is a genuine case?
    Hint
    Question 3 · Hint

    All positives include genuine cases and false alarms. Divide the genuine cases by that combined total.

    Show answer
    Question 3 · Answer

    About 118 people test positive; genuine cases account for 20/118≈16.9%. Detecting every case means none are missed, but people without the condition can still be falsely flagged.

Why shared birthdays are not so rare

Which two people are we comparing?

A class has 23 students. Assume birthdays are independent and equally likely on each of 365 days. What is the probability that at least two students share a birthday?

The answer is slightly more than one half.

This surprises us because we often imagine asking whether someone shares our own birthday. But the question concerns any pair. Two other people matching counts too; you need not be involved.

You can compare yourself with 22 others. Comparing every possible pair in the class gives 23 × 22 ÷ 2 = 253 pairs. There are many more opportunities for a match.

Still, you cannot simply add the probability for each pair. Several pairs can match at once, causing double-counting. Use Chapter 5's method and calculate the opposite event.

First calculate that everyone is different

The first birthday can be any day, without restriction.

The second must avoid that one day, leaving 364 of 365 days.

Given that the first two birthdays differ, the third must avoid two days, leaving 363.

Continue this way. The 23rd person must avoid 22 distinct birthdays, leaving 343 days. Thus:

P(all different) = 365/365 × 364/365 × 363/365 × … × 343/365 ≈ 0.493

“At least one matching pair” and “all different” cover all possibilities, so:

P(at least one shared birthday) = 1 − P(all different) ≈ 0.507

That is about 50.7%.

Start with just three people

If the long product feels cumbersome, try three people:

P(all different) = 365/365 × 364/365 × 363/365

Subtract from 1 to get a matching probability of about 0.82%. With few people, the chance really is small. As the group grows, there are more birthdays to avoid, making complete avoidance less likely.

Real birthdays are not perfectly uniform, and leap days add another complication. We set those details aside to make the method clear. The 50.7% belongs to this simplified model; it does not guarantee a match in every class.

Try it

Hover over or tap the links below for hints and answers. Tap again to close.

  1. Write the probability that four people all have different birthdays as a product. Include all four factors.
    Hint
    Question 1 · Hint

    The first birthday is unrestricted. How many days must each of the next three avoid? Include the fourth person.

    Show answer
    Question 1 · Answer

    365/365×364/365×363/365×362/365. The fourth person must avoid the preceding three distinct birthdays.

  2. Express the probability of at least one shared birthday among ten people as “1 minus ...”. You need not calculate the decimal by hand.
    Hint
    Question 2 · Hint

    First write the probability that all ten birthdays differ. The last person must avoid nine previous birthdays.

    Show answer
    Question 2 · Answer

    1−(365/365×364/365×…×356/365). The ten factors in parentheses give the probability that all ten birthdays differ.

  3. A class has 30 people, including you. How many people do you compare yourself with when asking whether someone shares your birthday? How is that different from asking about any pair in the class?
    Hint
    Question 3 · Hint

    For a match with you, exclude yourself. For any matching pair, which comparisons do not involve you at all?

    Show answer
    Question 3 · Answer

    You compare yourself with 29 others, making 29 pairs. All pairs in the class number 30×29÷2=435. The latter includes pairs that do not involve you.

Three doors: should you switch?

State the rules first

There are three doors. One hides a car; the other two hide goats. Each door is equally likely to hide the car. You want the car, so you choose a door but leave it closed.

The host knows where the car is. Every time, the host opens an unchosen door showing a goat, then offers you a choice: keep your original door or switch to the other unopened door.

If both unchosen doors hide goats, the host may open either. To calculate the overall success rate of always switching, we do not need the host to choose between those goats with equal probability. What matters is that the host always opens a goat door and always offers the switch.

Under these rules, switching wins with probability 2/3; staying wins with probability 1/3.

When does switching win?

Suppose you initially choose door A. The car has three equally likely locations:

Car locationDoor the host opensResult of switching
AEither B or C, both goatsLose
BCWin
CBWin

Switching wins in two of the three equally likely cases.

Another way to see it: if your first choice is right, switching loses. If your first choice is wrong, the host removes the other goat, so switching wins. Your first choice is wrong with probability 2/3, hence switching wins with probability 2/3.

Why not one half each?

Two doors remain, but they remain for different reasons.

Your door remains because you chose it and the host must leave it alone. The other door survives the host's selection: the host knows the answer and must avoid revealing the car. Reducing the number of doors does not erase how they were selected.

Probability depends not just on the options left in front of you, but on how they got there.

Imagine playing 300 games under the same rules. On average, about 100 initial choices are correct and about 200 are wrong. Always switching wins the latter group. The actual number of wins will fluctuate; it need not be exactly 200.

Change the host's behavior, and recalculate

If the host does not know the car's location and randomly opens an unchosen door, the car might be revealed immediately. Even if a goat happens to appear this time, this is a different game.

Or suppose the host offers a switch only when your initial choice is correct. An invitation then tells you that you already have the car; switching loses for certain.

The same visible action—opening a goat door—can arise from different rules. Establish the rules before calculating.

Try it

Hover over or tap the links below for hints and answers. Tap again to close.

  1. Under the standard rules, what is the probability of choosing the car initially? Why does “switching wins” correspond exactly to “the first choice was wrong”?
    Hint
    Question 1 · Hint

    Separate initially right and initially wrong choices. What happens when you switch in each case?

    Show answer
    Question 1 · Answer

    The initial chance is 1/3. If your first choice is wrong, the host must open the other goat door, leaving the car as the only unopened alternative. Switching then wins.

  2. If the host randomly opens any unchosen door, can you still apply the 2/3 switching result directly? Which guarantee has disappeared?
    Hint
    Question 2 · Hint

    The original host must avoid the prize. Could a random opening reveal it immediately?

    Show answer
    Question 2 · Answer

    No. Random opening no longer guarantees avoiding the car. If you consider only games where a goat happens to appear, recalculate the conditional probability under the new rule.

  3. Invent a rule under which the host opens a door and offers a switch only in some cases, so that the invitation reveals whether your first choice was right or wrong. State the rule in one sentence.
    Hint
    Question 3 · Hint

    Could the host make an offer only when you are initially right, or only when you are initially wrong?

    Show answer
    Question 3 · Answer

    For example, the host opens a goat door and offers a switch only when your initial choice contains the car. An invitation then tells you that your original choice was right.

An average gain is not a promised gain

Start with a small game

A game gives you a one-half chance of winning 10 yuan and a one-half chance of losing 10 yuan. Write gains as positive numbers and losses as negative numbers. The average payoff is:

1/2 × 10 + 1/2 × (−10) = 0 yuan

We call this an expected payoff of zero. It does not mean you will break even this time. One play produces either a 10-yuan gain or a 10-yuan loss; zero is not even an available result.

In probability, “expectation” does not mean a wish or a promise. Multiply each result by its probability, then add. More likely results receive more weight in this average.

A die has no face marked 3.5

Roll a fair die and receive that many yuan, with no entry fee. Each of the six outcomes has probability 1/6, so the expected amount is:

1/6 × 1 + 1/6 × 2 + … + 1/6 × 6 = 3.5 yuan

One roll pays 1, 2, 3, 4, 5, or 6 yuan, never 3.5. The 3.5 describes an average. Over many independent plays under unchanged rules, the average amount received approaches it.

If each play costs 4 yuan, the expected net payoff is 3.5 − 4 = −0.5 yuan. Remember entry fees and other costs when calculating a return.

Frequent wins can still mean an overall loss

Compare two games. These payoffs are already net of participation costs:

GameWinLossExpected net payoff
A95% chance of gaining 1 yuan5% chance of losing 10 yuan0.95 × 1 − 0.05 × 10 = 0.45 yuan
B99% chance of gaining 1 yuan1% chance of losing 150 yuan0.99 × 1 − 0.01 × 150 = −0.51 yuan

Game B usually wins more often, yet its expected loss is 0.51 yuan per play. One large loss can cancel many small wins.

Game A has positive expectation, but that does not make it suitable for everyone. If losing 10 yuan once is more than you can afford, a positive average does not solve the problem.

Here we draw on Taleb's approach: consider not only the chance of success, but how much success gains and failure loses. He discusses this distinction in his Edge essay.

What does one average leave out?

Option A pays 5 yuan for sure. Option B gives a one-half chance of gaining 100 yuan and a one-half chance of losing 90 yuan. Both have an expected net payoff of 5 yuan, but the risks clearly differ.

Expected value combines many possibilities into one useful comparison. In doing so, it leaves out detail. Before deciding, return to the outcome table: what is the worst result, can you afford it, and do you need the money immediately?

Try it

Hover over or tap the links below for hints and answers. Tap again to close.

  1. A game gives a one-half chance of gaining 6 yuan and a one-half chance of losing 2 yuan. What is its expected net payoff?
    Hint
    Question 1 · Hint

    Write gains as positive and losses as negative. Multiply each by its probability and add.

    Show answer
    Question 1 · Answer

    0.5×6−0.5×2=2 yuan.

  2. A game gives a 90% chance of gaining 2 yuan and a 10% chance of losing 20 yuan. What is its expected net payoff?
    Hint
    Question 2 · Hint

    Do not look only at the chance of winning. Multiply the possible loss by its own probability too.

    Show answer
    Question 2 · Answer

    0.9×2−0.1×20=−0.2 yuan: an expected loss of 0.2 yuan per play.

  3. A forecast often gets the direction right, but a wrong forecast causes a large loss. Why can the fraction of correct forecasts alone not tell you whether the strategy makes money?
    Hint
    Question 3 · Hint

    Compare the total from many small gains with the amount of one large loss. What costs might also be missing?

    Show answer
    Question 3 · Answer

    The fraction correct says nothing about the size of gains and losses. One large loss can cancel many small gains, and fees or other costs may also matter.

Five heads in a row: what next?

Earlier excess does not require later compensation

A fair coin lands heads five times in a row. Is tails now more likely, to bring the counts closer together?

If the tosses are independent, the sixth still has an equal chance of heads or tails. The first five results have changed neither the coin nor the procedure for the next toss.

Thinking that an outcome is “due” because it has appeared too rarely is the gambler's fallacy. It confuses long-run stability of proportions with a requirement to balance short-run counts.

Keep two questions separate

Before any tosses, the probability that the next five are all heads is:

(1/2)⁵ = 1/32

After those five heads have already occurred, the probability that the next toss is heads is 1/2.

The first question requires five events to happen; the second asks about just one future event. Counting the observed past again would answer a different question.

How can the proportion approach one half without compensation?

Suppose the first ten tosses are all heads. The proportion of heads is then 100%.

Now suppose the next 1,000 tosses happen to contain 500 heads and 500 tails. Together, there are 510 heads and 500 tails. Heads account for about 50.5%.

Heads still lead by ten; tails have not “caught up.” But the total has grown, so those ten extra heads have less effect on the proportion.

The law of large numbers says that, in the model of independent fair tosses under unchanged rules, the proportion of heads almost surely tends to 1/2 as the number of tosses grows. It does not require the difference in counts to go to zero, nor does it make the next toss compensate for past results. Chapter 14 explains “almost surely.”

Has the real process stayed the same?

The coin model has clear assumptions. Real machines wear out; replacing a factory's material can change why failures occur. Applying ten years of old failure records directly to modified equipment needs justification.

A Taleb-attributed excerpt in the Heron Notes warns that an enormous sample can make an inference from experience feel like a rigorous proof. Yet “this happened many times before” and “this must happen in the future” remain different claims.

Past data can be useful. But ask whether old and new conditions match, whether important cases were missed, and whether an event absent from ten thousand observations simply has not occurred yet.

Try it

Hover over or tap the links below for hints and answers. Tap again to close.

  1. Seven independent tosses of a fair coin have all been heads. What is the probability of tails on the eighth?
    Hint
    Question 1 · Hint

    Do independent tosses change their probabilities to balance earlier counts?

    Show answer
    Question 1 · Answer

    1/2. Independent tosses do not compensate for earlier outcomes.

  2. Each independent trial has success probability 1/4. Across 40 trials, what is the expected number of successes? Does that guarantee exactly that many successes?
    Hint
    Question 2 · Hint

    Multiply the number of trials by the success probability. Then ask whether an expectation guarantees the actual count.

    Show answer
    Question 2 · Answer

    40×1/4=10 successes in expectation. The actual count can be above or below 10.

  3. A factory switches to a new material. Can it directly use the preceding ten years' failure rate as the new failure probability? Name one condition that needs checking.
    Hint
    Question 3 · Hint

    The historical record concerns the old material. Could changing the material alter either causes or chances of failure?

    Show answer
    Question 3 · Answer

    Not without checking. At minimum, investigate whether the new material changes the causes of failure. Operating conditions and maintenance can also affect the probability.

Fair rounds can still leave you broke

Include your starting capital

You enter a game with 10 chips. Each round uses a fair coin: heads wins 1 chip, tails loses 1. Rounds are independent and there are no fees.

Agree on two stopping rules: leave if your chips reach 0, and leave if they reach 20.

Which end do you reach first?

The starting point, 10, lies halfway between 0 and 20, and each step is equally likely to go up or down. By symmetry, reaching 20 first and reaching 0 first each have probability 1/2.

Every round can be fair while you still have a one-half chance of going broke first.

Picture a number line

Write the numbers 0 through 20 on paper and place a counter at 10. Move one step right for heads and one step left for tails.

This is a random walk: a random outcome determines each step's direction. The stopping points, 0 and 20, are called boundaries.

If you start at 5, you are five steps from ruin but fifteen from the target. Each step is still fair, but reaching the two ends first is no longer equally likely.

In this model, starting with i chips and aiming for N chips:

P(reach N first) = i/N

P(go broke first) = 1 − i/N

Starting at 5 with a target of 20 gives a 1/4 chance of reaching the target and a 3/4 chance of ruin. Starting at 15 reverses those probabilities.

Where does the formula come from?

Let q(i) be the probability of reaching N first when starting at i. At 0, you have already lost, so q(0)=0. At N, you have already reached the target, so q(N)=1.

From an interior point, the next step takes you to i+1 or i−1 with equal probability. Therefore:

q(i) = [q(i+1) + q(i−1)] ÷ 2

The success probability at each interior point equals the average of its neighbors. This means the increase in success probability is the same at every step to the right.

There are N steps from 0 to N, across which success probability rises from 0 to 1. Each step adds 1/N. At i, the probability is i/N.

This reasoning uses the fact that a fair independent walk between these finite boundaries hits one of them with probability 1. A game can last a long time, but remaining inside forever has probability 0.

What should you ask in real life?

Starting capital and stopping rules affect the final outcome. Looking only at the chance of winning each round misses the fact that, once broke, you cannot continue.

Businesses and investments are more complicated: there may be fees, losses can exceed one step, and the probabilities may be unknown. You cannot simply use i/N to estimate a company's survival chances. But you can ask related questions: how little money would force you to stop? Could one bad result exhaust the available funds?

Try it

Hover over or tap the links below for hints and answers. Tap again to close.

  1. Under this chapter's rules, the boundaries are 0 and 20. Starting with 12 chips, what is the probability of reaching 20 first?
    Hint
    Question 1 · Hint

    In this chapter's formula, i is your initial number of chips and N is the target. Identify each.

    Show answer
    Question 1 · Answer

    12/20=0.6, or 60%.

  2. With the same boundaries, start with 2 chips. What is the probability of going broke first?
    Hint
    Question 2 · Hint

    First find the chance of reaching the target, then subtract from 1. Do not confuse success probability with ruin probability.

    Show answer
    Question 2 · Answer

    Reaching 20 first has probability 2/20=0.1, so ruin first has probability 1−0.1=0.9, or 90%.

  3. Suppose the game introduces a fee per round, or the win probability drops below one half. Can you still use this chapter's formula directly? Which assumption changed?
    Hint
    Question 3 · Hint

    Check the original assumptions: does each step still add or subtract one chip with equal probability?

    Show answer
    Question 3 · Answer

    No. A fee changes the actual gains and losses per step; a lower win probability removes the equal-chance assumption. Recalculate using the new rules.

Probability 1 can allow exceptions

A point has a location but no length

Imagine choosing a real number uniformly between 0 and 1. Include all real numbers, without restricting the number of decimal places.

“Uniformly” means an interval's probability equals its length. Landing between 0 and 0.5 has probability 1/2; landing between 0.49 and 0.51 has probability 0.02.

What is the probability of selecting exactly the number 0.5, specified in advance?

It is 0. A single point has no length. You can also enclose it in shorter and shorter intervals: the probability of the point cannot exceed that of any enclosing interval. Those intervals can be arbitrarily short, so the point's probability must be 0.

Yet 0.5 is still one of the available numbers. It has not been ruled out.

“Not choosing 0.5” has probability 1

If choosing exactly 0.5 has probability 0, not choosing it has probability:

1 − 0 = 1

But “not choosing 0.5” does not include every possible outcome. The exception, 0.5, remains.

This is what almost surely means: an event has probability 1 but may omit outcomes whose combined probability is 0. An event is sure, in the sense used here, if it has no exceptions and includes the entire sample space.

Probability here measures the size of a set, not merely whether the set contains anything. The mathematical way to assign such sizes is called a measure. For this uniform example, you can initially think of measure as length.

Be precise: probability 0 does not mean “extremely small but positive.” It is exactly zero. A zero-probability event can be nonempty, meaning the model includes such outcomes. That does not promise that a finite number of experiments will produce one.

If every point has probability zero, why not the whole interval?

Whatever number you specify in advance has probability 0. Yet the number selected must be some particular number. Is that contradictory?

No. “Selecting some number” includes all outcomes; “selecting this number specified in advance” includes only one. They are different events.

Another question follows: if an interval consists of points, why does adding all those zeros not give zero?

Probability's addition rule allows us to sum finitely many disjoint events, or infinitely many that can be listed in a sequence. The latter are countably infinite, like the positive integers 1, 2, 3, and so on. All the real numbers between 0 and 1 cannot be listed this way; they are uncountably infinite. A rule for countably many events cannot simply be applied to all those individual real-number points.

On a first reading, remember where the difficulty lies: not every kind of “infinitely many” allows the same summation rule.

Three more examples

Rational and irrational numbers

A rational number can be written as a ratio of two integers, such as 1/2, 1/3, or 3/4. There are infinitely many rationals between 0 and 1, but they can be listed: order fractions by increasing denominator, then numerator, and skip duplicates.

In the uniform real-number model, each rational point has probability 0. A countable collection of such points still has total probability 0. Therefore selecting an irrational number has probability 1.

The rational numbers have not disappeared. You can find one in every interval, however short. Being present everywhere in this sense is not the same as occupying positive total length.

Two random numbers are exactly equal

Choose X and Y independently and uniformly between 0 and 1. Treat (X,Y) as a point in a square. All points satisfying X=Y lie on its diagonal.

The square has area, but the diagonal has none, so the probability of equality is 0. The diagonal is nevertheless nonempty: (0.5,0.5), for example, lies on it.

A computer works with finite precision. Two computer-generated random numbers can therefore have a positive probability of being equal. A finite-precision display is not the ideal continuous model itself.

Infinitely many coin tosses

For independent fair tosses, the probability that the first n tosses are all tails is (1/2)ⁿ. As n grows, this approaches 0. The infinite sequence of tails forever therefore has probability 0, yet it remains one sequence in the sample space.

Likewise, getting only tails from any specified toss onward has probability 0. Possible starting positions can be listed as 1, 2, 3, and so on. Combining these cases still gives probability 0.

Thus “infinitely many heads occur” has probability 1, even though not every sequence contains infinitely many heads. Like the law of large numbers in Chapter 12, the statement needs the qualification “almost surely.”

When does probability 1 mean there are no exceptions?

For a fair die, every face has positive probability. An event that leaves out any face has probability below 1. In this model, a probability-1 event must include all six faces.

More generally, in a finite or countable sample space, if every individual outcome has positive probability, no nonempty event can have probability 0.

The condition is that every individual outcome has positive probability, not simply that there are finitely many. Even with two outcomes, mathematics allows assigning probability 0 to one and 1 to the other. The number of outcomes alone does not determine whether a zero-probability event must be empty.

An extension: does this explain election in Calvinism?

Someone might ask: could an elect person's probability of salvation be 1 while final loss remains possible?

The preceding mathematics does not establish that claim. Talking about probability 1 requires a sample space and a probability assignment. Classical Reformed teaching on election and perseverance makes a different kind of claim.

Chapter 17 of the Westminster Confession of Faith states that those truly accepted, effectually called, and sanctified will not totally and finally fall from grace, but will persevere to the end. It also acknowledges that true believers can temporarily fall into serious sin. The fifth head of the Canons of Dort likewise distinguishes human weakness and falling from God's final preservation.

Within that teaching, “truly elect but finally lost” is not a retained zero-probability exception. Adding such an exception changes the doctrine.

Also distinguish whether someone truly is elect from whether another person can judge this accurately. The second concerns the limits of our knowledge. It does not establish a random chance of failure in the first.

StatementWhat it concerns
Exceptions exist but have total probability 0How a probability model assigns probabilities to outcomes
True believers may fall temporarilyHow a doctrine understands human weakness and sin
The truly elect will not finally be lostClassical Reformed teaching about the final outcome
I do not know whether someone's faith is genuineLimits to an observer's information

Mathematics can help distinguish these statements. A shared word cannot by itself settle the theology.

Try it

Hover over or tap the links below for hints and answers. Tap again to close.

  1. If P(A)=1, what is the probability that A does not occur? Must there be no outcomes at all in which A fails?
    Hint
    Question 1 · Hint

    Use the complement formula, then recall a single specified point in a uniform real-number model.

    Show answer
    Question 1 · Answer

    The probability that A fails is 0, but the set of failure outcomes need not be empty. Choosing exactly 0.5 uniformly from the real interval is a nonempty zero-probability event.

  2. In the continuous uniform model, each number specified in advance has probability 0. Does this mean the final result is not a particular number? Distinguish the two events involved.
    Hint
    Question 2 · Hint

    Do ‘selecting some number’ and ‘selecting this number specified in advance’ contain the same outcomes?

    Show answer
    Question 2 · Answer

    No. The selected result is a particular number. “Selecting some number” includes all outcomes; “selecting this number specified in advance” includes only one.

  3. Why does “probability 1 can allow exceptions” not directly prove that “a truly elect person could finally be lost”?
    Hint
    Question 3 · Hint

    The mathematical result requires a sample space and probability measure. Is the confession making that kind of probability statement?

    Show answer
    Question 3 · Answer

    The first is a statement about the measure of sets in a probability model. Classical Reformed teaching about final preservation is a different kind of claim. Without an appropriate probability model, the mathematical conclusion cannot simply be transferred; retaining an exception of final loss would also change the doctrine.

Rare, large losses: understanding fat tails

How much can one result change?

Suppose you record a business's daily net returns. It earns 1 yuan on each of the first 99 days, then loses 150 yuan on day 100. The first 99 days look consistently good; including day 100 leaves a total loss of 51 yuan.

This simple example shows why counting profitable days misses the size of losses. It does not establish that the business has a fat-tailed distribution, but it helps explain why rare, large outcomes matter.

What is the “tail”?

Arrange possible losses from small to large and plot their distribution. Farther to the right means a larger loss. The part covering very large losses is the tail.

Some models make losses ten or a hundred times larger rapidly become extraordinarily unlikely. In other distributions, the probabilities decline much more slowly, leaving large losses with substantial weight. Such distributions are often described as fat-tailed.

“Fat” means relatively more probability in the tail. It does not mean all outcomes are large or every observation is likely to be a disaster. A precise discussion must specify the distribution and the comparison. Here we begin with its implications for judgment.

Taleb's Statistical Consequences of Fat Tails examines these distributions and why limited data can understate the contribution of extreme outcomes.

One new number changes the average

Suppose ten repairs each cost 100 yuan. The average cost is 100 yuan. An eleventh repair involves major damage and costs 10,000 yuan. The new average becomes:

(10 × 100 + 10,000) ÷ 11 = 1,000 yuan

One extra observation raises the average from 100 to 1,000 yuan.

Again, eleven observations do not establish that a distribution is fat-tailed. They illustrate that if some outcomes can greatly exceed ordinary ones, the recorded average may be highly sensitive to a few large values.

Some fat-tailed models have a mean but require very large samples to estimate it reliably. Others do not have a finite mean at all. Calculating an average from a spreadsheet does not establish that a stable, well-estimated population mean lies behind it.

Unseen does not mean impossible

A platform operates for three years without a major accident. You can say the record contains no major accident. That alone does not establish that a major accident is impossible.

Ask whether the record is complete, whether unusual conditions or heavy loads occurred, whether failures can happen together, and how large the resulting damage could be.

The largest historical loss is not automatically a ceiling on future losses. A genuine limit needs support from contracts, physical constraints, or other reliable conditions—not merely “we have never lost more than this.”

Distinguish two kinds of problem

The first is a fully specified simple game: a 99% chance of losing 1 yuan and a 1% chance of losing 1,000 yuan. The probabilities and amounts are given, so expected loss can be calculated directly.

The second is an incompletely understood real risk. You do not know either how often large losses occur or how large they might be. An estimate with many decimal places does not remove that uncertainty.

The first kind teaches calculation. The second requires examining the evidence behind the calculation. This distinction is part of what Taleb asks readers to notice.

Try it

Hover over or tap the links below for hints and answers. Tap again to close.

  1. Option A has a 10% chance of losing 100 yuan and no loss otherwise. What is the expected loss?
    Hint
    Question 1 · Hint

    Multiply each loss by its probability. No loss is an amount of zero.

    Show answer
    Question 1 · Answer

    0.1×100=10 yuan.

  2. Option B has a 99% chance of losing 1 yuan and a 1% chance of losing 1,000 yuan. What is the expected loss? Compared with A, which has the larger expected loss and the larger maximum possible loss? Can these bounded two-outcome examples alone establish a fat-tailed distribution?
    Hint
    Question 2 · Hint

    Calculate expected loss separately from maximum possible loss. Do both examples impose a fixed ceiling?

    Show answer
    Question 2 · Answer

    B has expected loss 0.99×1+0.01×1,000=10.99 yuan, exceeding A's 10 yuan. Its maximum loss is 1,000 yuan, exceeding A's 100 yuan. Both examples have bounded losses; this comparison does not establish that B has a mathematically fat-tailed distribution.

  3. A risk estimate uses 20 years of data, but future losses could exceed the historical maximum. Write two questions you would ask.
    Hint
    Question 3 · Hint

    Did the past include conditions capable of causing a large loss? Is the historical maximum necessarily a true upper limit?

    Show answer
    Question 3 · Answer

    For example: did these 20 years include conditions that could cause a major loss? Is there a reliable upper bound on loss? Could multiple failures occur together?

Before deciding, ask what you can afford to lose

An 80% success rate is not enough to decide

Two proposals each claim an 80% chance of success. If A fails, you waste one day. If B fails, the team loses its main client.

Even if both probabilities are perfectly accurate, the proposals are not equally attractive. The chances of failure match; the situations left afterward differ greatly.

Probability supplies part of what a decision needs. The size of gains and losses, their timing, and who bears them also matter.

In his Edge essay and a paper on binary forecasts and real-world payoffs, Taleb discusses the difference between predicting whether an event occurs and evaluating the resulting gains or losses. This chapter turns that approach into questions you can answer for yourself.

Make the decision concrete

Suppose you plan a paid lecture. The venue and preparation cost 3,000 yuan upfront. Each ticket leaves 100 yuan after other costs.

Selling 30 tickets covers the upfront cost. Selling 20 loses 1,000 yuan; selling 50 earns 2,000 yuan. Writing these numbers down is more useful than saying “I think it will probably work.”

You cannot yet calculate an expected return, because you do not know the probability of each sales outcome. Do not invent percentages just to fill a formula.

You could examine comparable past events or collect advance reservations. Use the new information to decide whether to book a venue and how large it should be.

Answer six questions

  1. What decision am I making? For example, whether to pay a nonrefundable venue fee now.
  2. What outcomes are possible, and how much would each gain or lose? Go beyond the labels “success” and “failure.”
  3. What supports my probability estimate? Are past events comparable in audience, price, and timing?
  4. How large could the worst loss be, and can I bear it? Are there responsibilities missing from the budget?
  5. If the judgment is wrong, who pays? Is that the same person who makes the decision?
  6. What new information would change my decision? For example, too few reservations by a deadline might mean choosing a smaller venue or cancelling.

These questions do not always need precise numbers. Identifying what you do not know is often more useful than filling the gap with an apparently precise answer.

Change the proposal, not just the yes-or-no answer

If risking 3,000 yuan at once is too much, could you begin with a smaller event? Obtain a refundable booking? Wait for enough reservations before paying?

These steps may not improve your ability to predict attendance. They can nevertheless reduce the loss if your prediction is wrong. Deciding involves arranging things so that mistakes are affordable, as well as forecasting what will happen.

In Chapter 13, a player with no chips cannot continue. In real life, preserving money, time, and room to adjust can matter as much as earning a little more now.

Start with one small decision

Choose a decision you will make soon. Write down possible results, their costs, and the information still missing. Then calculate the probabilities and expectations that can reasonably be calculated.

Do not force a calculation where the information is absent. State your assumptions, separate what is known from what is not, and consider whether a small, affordable trial could teach you more.

We began with coins and dice. The habit to take away is simple: when you see a probability, ask both how it was obtained and what follows if the event happens.

Try it

Hover over or tap the links below for hints and answers. Tap again to close.

  1. Choose a familiar small decision. Write at least two possible outcomes and their consequences, without assigning probabilities yet.
    Hint
    Question 1 · Hint

    Try a travel plan, an event, or a small purchase. Describe the time or money spent when things go well and when they do not.

    Show answer
    Question 1 · Answer

    For example, deciding whether to cycle to an appointment: good weather means arriving on time with no fare; heavy rain could cause delay and additional transport costs. Add other outcomes as needed.

  2. Someone says, “The risk is only 0.1%.” Write at least three questions you need answered.
    Hint
    Question 2 · Hint

    Ask whether this is a risk per attempt, per day, or per year; then ask about the evidence and the consequences.

    Show answer
    Question 2 · Answer

    For example: is 0.1% a risk per attempt, per day, or per year? What data support it? How large would the loss be? Does the new situation resemble the original sample?

  3. For your chosen decision, state a concrete rule: what would make you pause or reassess?
    Hint
    Question 3 · Hint

    Use ‘If ... then ...’. Make the condition observable and the response something you can actually do.

    Show answer
    Question 3 · Answer

    For example: if fewer than 30 paid registrations have arrived seven days before the event, stop new spending and reconsider using a smaller venue. State an observable condition and the action it triggers.

Answers and further reading

A short glossary

TermMeaningExample or note
Sample spaceAll possible outcomes of a trialThe six faces of a die
EventA selection of outcomes“Even” contains 2, 4, and 6
Equally likelyOutcomes have the same probabilityThe six faces of a fair die
Relative frequencyOccurrences divided by trialsSeven heads in ten tosses gives 0.7
ComplementAll outcomes where the event does not happenThe complement of “even” is “odd” on a die
Mutually exclusiveTwo events cannot occur togetherOne roll cannot show both 1 and 6
IndependentLearning whether one event occurred does not change the other's probabilityTwo independent coin tosses
Conditional probabilityProbability given a known conditionProbability of hearts given a red card
Base rateA starting proportion before new evidenceDisease prevalence before testing
SensitivityThe proportion of genuine cases testing positiveNot the proportion of positives that are genuine cases
False-positive rateThe proportion without the condition who test positiveThe denominator includes only those without it
Expected valueSum of each result multiplied by its probabilityNot a guarantee for one trial
Random walkMovement according to random stepsGaining or losing one chip at a time
Almost surelyWith probability 1Zero-probability exceptions may remain
Null setA set of measure zero under the chosen measureA single point in a continuous uniform model
Fat tailsA distributional feature in which probabilities of large outcomes decline relatively slowlyA few extreme outcomes may substantially affect the total

Formula reference

  • For finitely many equally likely outcomes: P(event) = favorable outcomes ÷ total outcomes.
  • Complement: P(A) = 1 − P(not A).
  • General addition: P(A or B) = P(A) + P(B) − P(A and B).
  • For mutually exclusive events: P(A or B) = P(A) + P(B).
  • Conditional probability: P(A|B) = P(A and B) ÷ P(B), where P(B)>0.
  • General multiplication: P(A and B) = P(B) × P(A|B), again with P(B)>0.
  • For independent events: P(A and B) = P(A) × P(B).
  • In n independent trials with the same success probability p: P(at least one success) = 1 − (1−p)ⁿ.
  • Bayes' formula: P(A|B) = P(B|A) × P(A) ÷ P(B). In the form used here, require P(A)>0 and P(B)>0.
  • With finitely many numerical outcomes: expected value is the sum of each outcome multiplied by its probability.
  • For independent fair steps of +1 or −1, stopped at 0 or N, starting at i: P(reach N first) = i/N; P(reach 0 first) = 1−i/N.

Exercise answers

Chapter 1

  1. 1/6. The six faces are equally likely, and one of them is 6.
  2. It is 1, not one half. Under the stated conditions, either side landing upward shows a heads design.
  3. Specify the location, the time interval, and what counts as rain. Also identify the source of the estimate, such as the forecasting agency.

Chapter 2

  1. The outcomes less than 3 are 1 and 2, so 2/6=1/3.
  2. No tails means HH, with probability 1/4. Therefore at least one tail has probability 1−1/4=3/4.
  3. 3/4. The four balls, not the two colors, are equally likely. Red includes three elementary outcomes; blue includes one.

Chapter 3

  1. 7/20=0.35, or 35%.
  2. 1/2. Under the fairness and independence assumptions, the first four results do not change the fifth probability.
  3. In a five-person survey, one changed answer moves the result by 20 percentage points. A 5,000-person survey is usually less sensitive to an individual response. But size cannot repair obvious selection bias, such as asking only supporters.

Chapter 4

  1. 1/6+1/6=1/3. A 1 and a 6 are mutually exclusive on one roll.
  2. Odd numbers are 1, 3, and 5; numbers less than 4 are 1, 2, and 3. Simple addition counts 1 and 3 twice. The combined outcomes are 1, 2, 3, and 5, giving 4/6=2/3.
  3. Independence gives 1/2×1/2=1/4.

Chapter 5

  1. The probability of no six on either roll is (5/6)², so the answer is 1−(5/6)²=11/36.
  2. All three failures have probability 0.8³=0.512. The answer is 1−0.512=0.488, or 48.8%.
  3. 1−(19/20)²⁰≈0.642, or about 64.2%. A small probability on each trial can produce a much larger probability of at least one occurrence over many trials.

Chapter 6

  1. Of the 26 black cards, 13 are spades, giving 13/26=1/2.
  2. Among the odd results 1, 3, and 5, only 5 exceeds 3. The answer is 1/3.
  3. 1/2. In the independent, equal-probability model, learning that the younger child is a girl does not change the chance that the older child is a girl.

Chapter 7

  1. Mutually exclusive, but not independent. One toss cannot be both heads and tails. Knowing it is heads changes the probability of tails from 1/2 to 0.
  2. Not mutually exclusive, but independent. TH satisfies both events, and the first toss does not change the second toss's probability.
  3. They can occur together, sharing outcomes 4 and 6, so they are not mutually exclusive. The joint probability is 1/3, while the product is 1/2×1/2=1/4. These differ, so the events are not independent either.

Chapter 8

  1. 1,000×2%=20 people.
  2. 980×10%=98 people.
  3. About 118 people test positive; genuine cases account for 20/118≈16.9%. Detecting every case means none are missed, but people without the condition can still be falsely flagged.

Chapter 9

  1. 365/365×364/365×363/365×362/365. The fourth person must avoid the preceding three distinct birthdays.
  2. 1−(365/365×364/365×…×356/365). The ten factors in parentheses give the probability that all ten birthdays differ.
  3. You compare yourself with 29 others, making 29 pairs. All pairs in the class number 30×29÷2=435. The latter includes pairs that do not involve you.

Chapter 10

  1. The initial chance is 1/3. If your first choice is wrong, the host must open the other goat door, leaving the car as the only unopened alternative. Switching then wins.
  2. No. Random opening no longer guarantees avoiding the car. If you consider only games where a goat happens to appear, recalculate the conditional probability under the new rule.
  3. For example, the host opens a goat door and offers a switch only when your initial choice contains the car. An invitation then tells you that your original choice was right.

Chapter 11

  1. 0.5×6−0.5×2=2 yuan.
  2. 0.9×2−0.1×20=−0.2 yuan: an expected loss of 0.2 yuan per play.
  3. The fraction correct says nothing about the size of gains and losses. One large loss can cancel many small gains, and fees or other costs may also matter.

Chapter 12

  1. 1/2. Independent tosses do not compensate for earlier outcomes.
  2. 40×1/4=10 successes in expectation. The actual count can be above or below 10.
  3. Not without checking. At minimum, investigate whether the new material changes the causes of failure. Operating conditions and maintenance can also affect the probability.

Chapter 13

  1. 12/20=0.6, or 60%.
  2. Reaching 20 first has probability 2/20=0.1, so ruin first has probability 1−0.1=0.9, or 90%.
  3. No. A fee changes the actual gains and losses per step; a lower win probability removes the equal-chance assumption. Recalculate using the new rules.

Chapter 14

  1. The probability that A fails is 0, but the set of failure outcomes need not be empty. Choosing exactly 0.5 uniformly from the real interval is a nonempty zero-probability event.
  2. No. The selected result is a particular number. “Selecting some number” includes all outcomes; “selecting this number specified in advance” includes only one.
  3. The first is a statement about the measure of sets in a probability model. Classical Reformed teaching about final preservation is a different kind of claim. Without an appropriate probability model, the mathematical conclusion cannot simply be transferred; retaining an exception of final loss would also change the doctrine.

Chapter 15

  1. 0.1×100=10 yuan.
  2. B has expected loss 0.99×1+0.01×1,000=10.99 yuan, exceeding A's 10 yuan. Its maximum loss is 1,000 yuan, exceeding A's 100 yuan. Both examples have bounded losses; this comparison does not establish that B has a mathematically fat-tailed distribution.
  3. For example: did these 20 years include conditions that could cause a major loss? Is there a reliable upper bound on loss? Could multiple failures occur together?

Chapter 16

  1. For example, deciding whether to cycle to an appointment: good weather means arriving on time with no fare; heavy rain could cause delay and additional transport costs. Add other outcomes as needed.
  2. For example: is 0.1% a risk per attempt, per day, or per year? What data support it? How large would the loss be? Does the new situation resemble the original sample?
  3. For example: if fewer than 30 paid registrations have arrived seven days before the event, stop new spending and reconsider using a smaller venue. State an observable condition and the action it triggers.

What this book simplifies

TopicAssumption used hereWhat to study next
Counting probabilitiesFinitely many equally likely elementary outcomesAssigning unequal probabilities individually
Repeated trialsUsually independent with unchanged probabilitiesModels with dependence or changing conditions
Test examplesHypothetical prevalence and test performanceRelevant populations and estimation error in real data
Birthdays365 equally likely days; independent peopleLeap days, seasonal patterns, and population differences
Three doorsThe host always reveals a goat and offers a switchConditional probabilities under other host rules
Expected valueBegin with finitely many outcomesWhether a mean exists for infinitely many outcomes
Gambler's ruinIndependent fair steps of one chip, stopped at boundariesUnequal probabilities, fees, and variable stakes
Almost surelyExplanations using length, area, and coin sequencesFull definitions of events and probability spaces in measure theory
Fat tailsThe effect of extreme results on totalsSpecific distributions, tail decay, and existence of moments such as the mean

Where to read next

For a systematic introduction, continue with Volume I of Feller's An Introduction to Probability Theory and Its Applications. It contains classic problems involving coins, counting, random walks, and ruin. It asks more of your algebra, reasoning, and practice than this short introduction. You need not read it straight through; begin with a familiar problem.

To explore how extremes affect statistics, study more probability and statistics before reading Taleb's Statistical Consequences of Fat Tails. For options and hedging, turn to Dynamic Hedging. Both are technical books, not first books for a complete beginner.

The website also keeps a subject-based reading list, with descriptions and ratings out of five. The ratings express this site's recommendations in relation to its subject matter, not Taleb's own ratings.

What can “No Hacking.” tell us?

In the supplied screenshot, Baibanbao mentions Gemini's recommendations of books by Feller and Ian Hacking. Taleb replies: “No Hacking.”

In context, the most natural reading is that he does not want Hacking included in that reading list. Hacking is both the author's surname and a word that gives the reply the flavor of an English warning sign.

But those two words do not explain his reasons. Nor does his silence about Feller, by itself, establish an explicit recommendation of Feller. The screenshot does not tell us whether he objects to probability's intellectual history, or to any particular argument by Hacking.

This book therefore retains the screenshot and reading suggestions without presenting guessed motives as Taleb's stated views.

Taleb's ideas and the sources

This book draws on the site's original lessons and Taleb-attributed excerpts collected in the Heron Notes, a supplied Chinese notes document. That document also contains translations, summaries, and discussion; not every sentence in it is Taleb's own. The main ideas used here are to distinguish probability from outcome size, attend to extreme losses, and examine assumptions when extrapolating from past data.

Conditional probability, expectation, random walks, and almost-sure events are mathematical concepts in probability theory. Taleb's questions about risk help frame the lessons; these mathematical results are not presented as his inventions.

  1. Taleb's 2008 Edge essay, discussing event probabilities, outcome sizes, and their different roles in decisions.
  2. Statistical Consequences of Fat Tails, Taleb's technical monograph, first posted in 2020; the linked page includes revised versions.
  3. On the Statistical Differences between Binary Forecasts and Real World Payoffs, Taleb's paper, first posted as a preprint in 2019.
  4. A March 1, 2025 post attributed to Taleb in the Heron Notes: a very large sample can make inference from experience seem like rigorous proof. Chapter 12 discusses this point. The notes do not preserve a direct link to the post.
  5. A March 24, 2025 post attributed to Taleb in the Heron Notes: he jokes that academics saying “almost always” mean “always,” while ordinary people saying “always” mean “almost always.” This is a joke about usage, not the mathematical definition of almost surely. The notes do not preserve a direct link.
  6. The Westminster Confession of Faith, Chapter 17, and the Canons of Dort, fifth head, support the discussion of classical Reformed teaching in Chapter 14.
  7. The two screenshots and the list labeled “Taleb's library” were supplied for this project. The shelf list has not been checked book by book and does not establish that Taleb recommends every title.