Base Rate Neglect
A common cognitive error where individuals judge probability based solely on how well an event matches a stereotype (representativeness heuristic), ignoring the underlying frequency of the categories involved (base rates).
3Blue1Brown · YouTube · 15:11
This segment introduces Bayes' theorem through the classic 'librarian vs. farmer' problem studied by Kahneman and Tversky. It demonstrates how humans often ignore base rates (priors) in favor of representativeness (likelihoods). By constructing a representative sample grid, it calculates the posterior probability geometrically, showing that even strong evidence must be weighed against prior odds. Finally, it derives the algebraic formula from this geometric setup. This segment explores how geometric representations and natural frequencies can make Bayesian probability more intuitive. It begins by visualizing the joint probability (E|H) as the area of a rectangle within a larger sample space. The discussion then shifts to cognitive psychology, contrasting abstract percentages with concrete 'out of 100' counts to explain why people often fall prey to the conjunction fallacy, as demonstrated by the classic Linda problem. By reframing probabilities as proportions of people or areas, the video demystifies Bayes' theorem, showing it simply asks for the fraction of evidence-supporting cases where the hypothesis is also true. Finally, it addresses criticisms regarding ambiguous priors in personality-based problems, concluding that while context influences initial beliefs, evidence should serve to update rather than rigidly determine our convictions.
Use the learning inspector for key ideas and moments, or open the reading tabs for the complete notes.
Generated from the video's visuals and explanation; not verbatim speech.
Bayes' theorem is central to scientific discovery, machine learning, and even treasure hunting. Understanding it requires grasping three levels: what the symbols mean, why the formula works, and when to apply it. We begin with a famous cognitive bias experiment involving a character named Steve.
Steve’s description may suggest a librarian. The illustrative setup considers only librarians and farmers, with twenty farmers per librarian. This is the example’s background assumption, not a current occupational statistic. Matching a description is not the same as a posterior probability.
To resolve this, imagine a population of 210 people: 10 librarians and 200 farmers. Suppose 40% of librarians fit Steve's description, while only 10% of farmers do. This yields 4 matching librarians and 20 matching farmers. Out of the 24 total people who fit the description, only 4 are librarians. Thus, the probability Steve is a librarian given the evidence is approximately 16.7%, far lower than intuition suggests.
We can formalize this calculation. Let H be the hypothesis 'Steve is a librarian' and E be the evidence 'fits the description'. The term is the prior probability. The term P(E|H) = 0.4 is the likelihood—how probable the evidence is if the hypothesis is true. Similarly, P(¬H)P(E|¬H) accounts for the alternative case where the hypothesis is false. The prior refers to this two-category example population, and conditioning requires .
Combining these terms gives the full expansion of Bayes' theorem: P(H|E) equals the product of the prior and likelihood, divided by the sum of such products across all mutually exclusive hypotheses. The denominator represents the total probability of observing the evidence regardless of which hypothesis is true. This result, P(H|E), is called the posterior—the updated belief after incorporating new data.
We begin by looking at the geometry behind Bayes' theorem. Consider the term in the numerator. Geometrically, this represents the intersection of two events: the hypothesis being true and the evidence occurring. If we view the total sample space as a large square, the prior probability defines the width of a vertical strip, and the likelihood defines the height of a shaded region within that strip. Multiplying these dimensions gives us the exact area of the top-left rectangle. This area corresponds precisely to the joint probability—the chance that both the hypothesis and the evidence happen together. Visualizing probability as physical area helps ground abstract formulas in spatial reality.
To understand human intuition better, let's step back from pure algebra. A powerful technique involves imagining a representative sample—a specific group of people—rather than dealing with abstract ratios. For instance, instead of thinking about percentages, imagine a population of 210 individuals divided into librarians and farmers. This approach aligns with findings from psychologists Daniel Kahneman and Amos Tversky. They discovered that framing statistical problems using concrete numbers of people significantly improves accuracy compared to using isolated probabilities. When we anchor our reasoning to a fixed sample size, complex conditional relationships become tangible counts, reducing cognitive load and minimizing common errors in probabilistic judgment.
The Linda problem compares being a bank teller with being both a bank teller and active in the feminist movement. The second event is a subset of the first, so . A compelling narrative can distract from inclusion; strict inequality is not required in every case.
Rephrasing the question as counts out of a hundred makes the subset relationship easier to see. The video reports a reduction in errors in the described experiment; it does not establish a zero-error guarantee for every person or problem. A subset’s count is at most the parent set’s count.
While counting people works well for discrete events, real-world phenomena often involve continuous variables, such as weather patterns or measurement tolerances. In these cases, representing probabilities as grids of dots becomes cumbersome. Here, geometry offers a superior alternative. Imagine a bar chart where the width represents time or intensity, and the colored portion indicates the likelihood of rain. As conditions change continuously, the boundary between sunny and rainy outcomes shifts smoothly across the canvas. Unlike static icons, geometric areas naturally accommodate infinitesimal changes. Furthermore, sketching these regions on paper provides an immediate visual feedback loop, making it easier to manipulate equations and spot inconsistencies during manual calculation without relying solely on symbolic manipulation.
Returning to Bayes' theorem itself, viewing it through the lens of proportions renders its meaning almost self-evident. The formula states that the posterior probability equals the joint probability divided by the marginal probability . Translated into plain language: First, restrict your attention only to those cases where the evidence has occurred. Then, among that filtered subgroup, calculate what fraction also satisfies the hypothesis . Whether you measure this fraction by counting heads in a crowd or measuring pixels on a screen, the underlying arithmetic remains identical. Recognizing this structural simplicity transforms a daunting equation into a straightforward instruction for filtering data based on observed outcomes. Conditioning requires a positive evidence-region probability.
Critics sometimes argue that applying base rates to individual profiles, like Steve the librarian/farmer puzzle, ignores contextual ambiguity. Who exactly is Steve? Is he drawn randomly from national census data, or is he someone personally known to the observer? These questions highlight that priors are subjective starting points influenced by background knowledge. Changing assumptions about the source population alters the baseline widths in our diagrams. Similarly, tweaking stereotypes affects the relative heights of likelihood rectangles. Yet regardless of where one places the dividing lines initially, the mechanism of updating stays constant. Evidence does not dictate absolute truth instantly; instead, it nudges existing beliefs along a spectrum. Mastering this dynamic interplay between flexible priors and hard data constitutes the essence of rational inference.
A common cognitive error where individuals judge probability based solely on how well an event matches a stereotype (representativeness heuristic), ignoring the underlying frequency of the categories involved (base rates).
The hypothesis probability before incorporating this evidence. Here it is in the stipulated two-occupation population. Different sampling populations and background information can change the prior.
The conditional probability of observing the specific evidence assuming the hypothesis is true. It measures how compatible the data is with the model.
The revised probability of the hypothesis after taking the new evidence into account. It combines the prior belief with the strength of the evidence.
Used here as the denominator in Bayes' rule. It sums the joint probabilities of the evidence occurring under each possible hypothesis (true or false) to normalize the distribution.
In geometric interpretations of Bayes' theorem, the product of the prior and the likelihood corresponds to the area of a rectangular sub-region. This area physically represents the joint probability that both the hypothesis and the evidence occur simultaneously within the total sample space.
A logical error where individuals judge a combined event (e.g., 'bank teller AND feminist') as more probable than one of its constituent parts ('bank teller'). This violates the axioms of probability because the intersection of sets can never be larger than either original set.
Counts within a common reference population can make conditioning and inclusion easier to understand. This is a useful presentation technique, not a theorem that eliminates every reasoning error.
Bayes' theorem calculates the posterior probability by taking the proportion of cases supporting the hypothesis () strictly within the universe of cases already confirmed by the evidence (). It answers: Given that happened, how much of that territory overlaps with ?
Initial beliefs (priors) depend heavily on context definition—for example, whether a subject is assumed to be a random stranger or a personal acquaintance. While debatable contexts shift the baseline distribution, proper use of evidence updates these views systematically rather than replacing them arbitrarily.
The posterior is the fraction of evidence-supporting cases in which the hypothesis also holds: for . The geometric and natural-frequency examples combine prior and likelihood rather than ignore base rates. Priors depend on the stipulated sampling population and background information.
The conjunction fallacy manifests when individuals judge the probability of a combined event (being a bank teller AND active in the feminist movement) as higher than the probability of one of its constituent parts (being a bank teller). Mathematically, the second event is a subset of the first, so .
Conditions: Comparing the probability of a subset event against its superset event; Events are defined such that one is contained within the other
The calculation assumes a population of 210 people: 10 librarians and 200 farmers. With 40% of librarians fitting the description (yielding 4 matching librarians) and 10% of farmers fitting it (yielding 20 matching farmers), there are 24 total matches.
Conditions: Population consists only of librarians and farmers with a 1:20 ratio; Likelihoods are stipulated as 40% for librarians and 10% for farmers; Conditioning requires
In this two-category example, the prior is the probability of the hypothesis 'Steve is a librarian' before seeing evidence, calculated as based on the stipulated population. The likelihood is the probability of the evidence 'fits the description' given the hypothesis is true, stipulated as 0.4.
Conditions: Hypothesis H: 'Steve is a librarian'; Evidence E: 'Fits the description'; Population assumption: 10 librarians, 200 farmers; Likelihood assumption: 40% for librarians, 10% for farmers
In the geometric model, the total sample space is represented as a large square. The prior probability determines the width of a vertical strip representing the hypothesis, and the likelihood determines the height of the shaded region within that strip where the evidence occurs.
Conditions: Total sample space is normalized to unit area; Events H and E are treated as measurable subsets
For continuous variables like weather patterns or measurement tolerances, representing probabilities as grids of dots becomes cumbersome. Geometry offers a superior alternative by using areas (e.g., bar charts where width represents time/intensity and color indicates likelihood).
Conditions: Variables are continuous rather than discrete; Probabilities need to be visualized for intuitive understanding
Critics point out that priors are subjective starting points influenced by background knowledge. The definition of the source population matters significantly: whether Steve is drawn randomly from national census data or is someone personally known to the observer changes the baseline widths in probability diagrams.
Conditions: Application of base rates to specific individual profiles; Ambiguity exists regarding the sampling frame (random vs. personal acquaintance)
Framing problems using natural frequencies (e.g., counts out of a fixed population like 210 people) aligns with findings from psychologists Kahneman and Tversky. Concrete numbers of people make complex conditional relationships tangible, reducing cognitive load.
Conditions: Problems involve conditional probabilities or Bayesian inference; Comparison is between natural frequency framing and percentage/probability framing
Structurally, Bayes' theorem asks you to first restrict attention to only those cases where the evidence has occurred. Then, among that filtered subgroup, it calculates what fraction also satisfies the hypothesis .
Conditions: Evidence E has occurred (conditioning); Marginal probability