Probability is not merely a list of fixed numbers attached to events. A probability describes uncertainty relative to the information currently available. When new information rules out some outcomes, the relevant sample space changes, so the probabilities within it must be recalculated. Conditional probability is the mathematical language for performing that update without losing track of the original experiment. The central habit of this lesson is therefore simple: identify the new universe of possible outcomes before doing arithmetic.
Learning objectives and a guiding question
By the end of this lesson, you will be able to interpret and calculate a conditional probability from words, tables, trees, and formulas. You will derive and use the multiplication rule, distinguish independence from disjointness, and combine cases with the law of total probability. You will also use Bayes’ theorem to reverse the direction of a condition while respecting prior probabilities. Every formula will be connected to a restricted sample space so that the notation records reasoning rather than replacing it. These skills prepare you for probability distributions, statistical inference, reliability analysis, and evidence-based decision making.
The guiding question is, “What remains possible after I learn this information?” Suppose a card is selected from a standard deck and you are told that the card is a face card. The original sample space contained equally likely cards, but only face cards remain compatible with the information. If the question asks for the probability that the card is a king, the relevant count is four kings out of those twelve remaining cards. The answer is therefore the horizontal fraction , not . The new information has changed the denominator from fifty-two to twelve.
This example contains the entire conceptual structure of conditioning. The phrase “given that the card is a face card” identifies the restricted universe. The desired event, drawing a king, is then intersected with that universe because an outcome must satisfy both descriptions. The denominator measures the size or probability of the given event, while the numerator measures the part that also satisfies the target event. Before memorizing any rule, practice saying those two roles aloud.
From a sample space to a restricted sample space
Let denote the original sample space, and let and be events within it. Learning that occurred removes every outcome outside from consideration. The event therefore acts as the new sample space, even though the original probability model was built on . The portion of that survives this restriction is the intersection . The symbol means “and,” so contains outcomes belonging to both events.
The diagram below shows this change of viewpoint. In the original space, event occupies one region and event occupies another, with an overlap between them. After conditioning on , the region outside becomes irrelevant rather than impossible in the original experiment. Within the restricted region, only the overlap counts as success for event . This visual distinction prevents a common error: retaining the old denominator after the information has changed.
The restriction idea works whether outcomes are equally likely or not. Counting is convenient for cards, dice, and finite tables, but probability mass can also be unequal or continuous. In every case, the denominator must represent all probability still under consideration after the condition is known. The numerator must represent the probability that satisfies both the target and the condition. That invariant meaning is more reliable than any particular computational shortcut.
Derive and interpret the conditional-probability formula
For an event with positive probability, conditional probability is defined by . The expression is read as the probability of given . The vertical bar is not division and does not mean “such that” in this context. It announces that is being treated as the reference space. The requirement matters because dividing by zero cannot produce a probability distribution. Every symbol in the formula therefore records a specific part of the restriction process.
The denominator rescales the surviving probability mass so that the new sample space has total probability one. The numerator selects the part of that surviving mass that also belongs to . Because is a subset of , its probability cannot exceed . Consequently, the quotient always lies between zero and one when the probability model is valid. This range check is a quick way to detect an inverted fraction or an incorrect denominator.
Consider a class of students containing students who take chemistry, who take physics, and who take both. If one student is chosen and you learn that the student takes physics, then is the group of physics students. The target is taking chemistry, and the overlap contains students. Thus . The factors of cancel because conditioning replaces the original class with the physics subgroup.
Read the order of the notation carefully
Conditional probability is directional, so and usually describe different questions. In the class example, because twelve students satisfy the condition. Reversing the order gives because eighteen students now satisfy the condition. The same overlap appears in both numerators, but the reference groups differ. Reading the notation from left to right as “target given condition” makes that difference explicit.
Natural language can hide this order, especially in medical, legal, and engineering contexts. “The probability of a defect given alarm” asks what fraction of alarms correspond to defects. “The probability of an alarm given defect” asks how frequently the system detects an actual defect. The first quantity concerns the credibility of an alarm, while the second concerns detection sensitivity. Treating them as interchangeable is sometimes called the inverse fallacy.
A useful translation routine has three steps. First, underline the information introduced by words such as “given,” “among,” “of those,” or “provided that.” Second, place that event to the right of the vertical bar because it defines the denominator. Third, place the requested event to the left and form the intersection in the numerator. This routine turns a linguistic problem into a probability structure before numbers can distract you. It should become a deliberate habit whenever conditional language appears.
Build the multiplication rule from the definition
Starting with , multiply both sides by . The result is . This multiplication rule says that the probability of following the branch and then reaching equals the probability of entering times the conditional probability of within . The rule does not assume that and are independent. In fact, the conditional factor is precisely what allows the rule to represent dependence.
The intersection is symmetric, meaning , so the same probability can be factored in the reverse order. Therefore as well. Equating the two factorizations gives . This equality is the algebraic core of Bayes’ theorem. It also shows why reversing a conditional probability requires both events’ base probabilities.
The rule extends to a sequence through repeated conditioning. For three events, . Each factor describes the next event under all information accumulated earlier along the path. Probability trees encode exactly this sequence, with multiplication along a single path. Adding across mutually exclusive completed paths then combines different ways the target event can occur.
Use tables and trees as reasoning tools
A two-way table makes conditional denominators visible by organizing counts across two categorical variables. To condition on a row category, divide the desired cell by that row total. To condition on a column category, divide the desired cell by that column total. The grand total is used only for an unconditional probability. Writing the relevant total beside the vertical bar helps keep the reference group fixed.
A probability tree emphasizes temporal or logical stages instead of rectangular counts. Branches leaving the same node must sum to one because they represent all possibilities under the information at that node. Multiply branch probabilities along a path to obtain the probability of the corresponding intersection. Add path probabilities only when those paths are mutually exclusive ways to reach the event of interest. The diagram below shows how the table and tree express the same joint structure.
Choose the representation that makes the hidden denominator easiest to see. Tables are especially effective when actual frequencies are available or when two categories are crossed. Trees are especially effective for sequential events, repeated trials with changing probabilities, and diagnostic pathways. A Venn-style area diagram emphasizes intersections and restricted spaces but may not preserve exact proportions. Moving among representations is a form of verification because each one exposes different mistakes.
Distinguish independence from disjointness
Events and are independent when learning that one occurred does not change the probability of the other. In symbols, independence means whenever . Substituting the conditional definition gives the equivalent multiplication test . The word “equivalent” means either equation can be used to establish the same relationship under the stated conditions. Independence is therefore a statement about information, not simply about whether two events look unrelated.
Disjoint events have no common outcomes, so and . If both disjoint events have positive probability, then learning that occurred forces the probability of to zero. Because that updated value differs from the original positive value, the events are dependent. Thus disjointness and independence are nearly opposites for nontrivial events. Confusing them often comes from using the informal word “separate” without specifying its mathematical meaning.
Suppose a fair six-sided die is rolled, with meaning “the result is even” and meaning “the result exceeds three.” Here , , and because four and six satisfy both. The product does not equal , so the events are dependent. Equivalently, differs from . A verbal impression cannot replace this explicit comparison. Both valid tests lead to the same conclusion.
Combine cases with the law of total probability
A collection of events forms a partition when the events are mutually exclusive and together cover the entire sample space. Every outcome then belongs to exactly one partition category. Event can be divided into the disjoint pieces . Adding their probabilities and applying the multiplication rule gives . The sigma symbol means to add one term for every index from one through .
Suppose a factory obtains of its parts from line 1 and from line 2. The defect rates are and , respectively. The overall defect probability is . The result is equivalent to . It is a weighted average because each conditional defect rate is weighted by its line’s share of production.
The weights must refer to the same population and must sum to one. Simply averaging and would incorrectly treat the two production lines as equally common. A total-probability calculation preserves both within-group behavior and group prevalence. This separation is essential whenever rates differ across hospitals, schools, machines, or demographic groups. Ask for the rate inside each case and the prevalence of each case. Both questions are necessary for the overall rate.
Reverse a condition with Bayes’ theorem
Bayes’ theorem follows by solving the equality for the desired conditional probability. The result is , provided . The numerator is the joint probability that both and occur. The denominator is the total probability of the observed evidence . The quotient therefore asks what fraction of all evidence-producing cases came from source .
When partition the possible sources, substitute the law of total probability into the denominator. This produces . The symbol marks the particular source being evaluated, while the index runs across every possible source in the denominator. The numerator supports one explanation of the evidence, and the denominator represents all explanations included in the model. Bayes’ theorem is consequently normalized comparison, not a mysterious reversal trick.
The diagram below organizes the update into prior, likelihood, evidence, and posterior. The prior represents belief or prevalence before observing . The likelihood measures how compatible the evidence is with source . The evidence normalizes all source-weighted likelihoods. The posterior is the updated probability after the evidence is included.
Understand base rates through natural frequencies
Consider a condition present in of a population, a test with sensitivity, and a false-positive rate. Sensitivity is , where denotes having the condition. The false-positive rate is , where the superscript denotes the complement, or not having the condition. The desired probability after a positive result is , which reverses the stated sensitivity. A low prevalence can make this posterior much smaller than intuition expects.
Natural frequencies make the calculation concrete by imagining representative people. About have the condition, and of them test positive. About do not have the condition, and of them test positive. There are therefore positive tests in total, of which are true positives. The conditional probability is , or approximately .
This result does not mean the test is useless or inaccurate. It means that false positives drawn from a large unaffected group can outnumber true positives drawn from a small affected group. The prevalence is called a base rate because it describes how common the condition was before testing. Ignoring that rate creates the base-rate fallacy. Decisions should also consider consequences, follow-up tests, and uncertainty in the stated rates, which probability alone does not determine.
Check units, complements, and numerical meaning
Probabilities are dimensionless ratios, so they do not carry physical units such as meters or seconds. Counts used to form a probability must nevertheless refer to compatible units of observation, such as students out of students or defective parts out of inspected parts. Dividing eight chemistry-and-physics students by twelve physics students cancels the student count and leaves a pure ratio. Percentages are another notation for such ratios, with . Keeping counts labeled until the final division makes the reference population clear.
Complements provide useful checks because . Within any fixed condition , the target event and its complement exhaust the restricted sample space. Similarly, conditional branches leaving the same tree node must sum to one. If a sensitivity is , then the false-negative rate under the condition is . This complement is different from the false-positive rate because the two rates condition on different populations.
A defensible final statement should name the condition, target, denominator, and interpretation. Instead of writing only , say that among outcomes in , one quarter also belong to . Then confirm that the result lies between zero and one and is reasonable relative to the unrestricted probability. If the conditional value rises, the new information favors the target; if it falls, the information weighs against it. These verbal checks turn a calculation into an explanation. They also make an incorrect reference population easier to notice.
Diagnose common mistakes deliberately
The most common mistake is reversing and . Prevent it by identifying the restricted group before inserting any numbers. A second mistake is using the grand total as the denominator even after the phrase “given that” has narrowed the population. A third is multiplying without either establishing independence or using a conditional factor. Each error changes the probability model, not merely the arithmetic.
Another mistake is dividing by an event with zero probability. In an elementary discrete model, conditioning on an impossible event is undefined because there is no restricted population to normalize. More advanced probability theory can define conditional structures for continuous variables, but it does not repair ordinary division by zero. State the condition when using the elementary definition. Domain restrictions are part of the theorem and should not be hidden.
Finally, avoid treating a model’s categories and rates as unquestionable facts. A Bayes calculation is only as appropriate as its partition, likelihoods, and prior information. Selection bias can make observed frequencies unrepresentative, and dependence between repeated observations can invalidate a simple product. Report assumptions alongside results, especially in health, safety, and policy applications. Mathematical precision includes being explicit about what the model leaves out.
Practice with retrieval, representation, and explanation
First, let , , and . Calculate and decide whether and are independent. Then calculate and explain why it differs from the first conditional probability. Your response should identify both denominators rather than giving numbers alone. Check independence using both the conditional and product criteria.
Second, return to the factory with production shares and and defect rates and . Draw a two-stage tree, label every branch, and verify that all terminal path probabilities sum to one. Find the overall defect probability using total probability. Then find the probability that a defective part came from line 2 by applying Bayes’ theorem. Interpret the posterior as a fraction of all defective parts, not of all manufactured parts.
Third, design a small two-way table about two school activities with at least students. Choose counts so the two events are dependent but not disjoint. Compute both conditional directions, the intersection, and the product of the marginal probabilities. Explain how the table reveals dependence and how changing one cell could produce independence. Creating a valid example requires deeper control than merely recognizing one.
Solutions and checks
For the first problem, , which differs from , so the events are not independent. Also, because the denominator changes when the condition changes. The product test agrees because , not . The intersection is too small for independence. Both conditional directions should be reported distinctly.
For the factory, the defective paths have probabilities and . Adding them gives , or . Bayes’ theorem then gives , or . Line 2 supplies fewer parts but more defective parts. The posterior therefore exceeds its production share.
The third problem has many correct constructions. Verify that all four interior counts are nonnegative, row and column totals agree with the grand total, and neither event is empty. Dependence is shown when or, equivalently, when . A complete explanation must specify which table total serves as each denominator. Recalculate all marginal totals after changing a cell.
Connect conditional probability to what comes next
Conditional probability becomes more powerful when outcomes are assigned numerical values through random variables. A conditional distribution describes how the entire distribution of a variable changes after information is observed. Conditional expectation then summarizes an updated distribution with a probability-weighted average. These ideas support regression, Markov chains, Bayesian inference, and stochastic processes. The restricted-sample-space principle remains unchanged even when the notation becomes more advanced.
Statistical inference also distinguishes probabilities about data under a model from probabilities assigned to hypotheses or parameters. Bayes’ theorem can connect these quantities only after a prior model and likelihood have been specified. Frequentist procedures use conditional reasoning too, although they interpret probability differently. Learning to name the random event and condition now will prevent serious confusion later. Notation is most useful when every symbol has a clear referent in the underlying experiment.
The durable lesson is that information changes the denominator before it changes the answer. Begin with the new reference group, locate the target within it, and only then compute. Use tables, trees, and natural frequencies to make that structure visible. Verify the result through range, complement, and representation checks. Conditional probability then becomes a disciplined method for learning from evidence rather than a collection of disconnected formulas.