A random variable turns outcomes of an uncertain process into numbers that can be analyzed. The mapping is fixed even though the outcome is not known in advance. A probability distribution then describes how probability is allocated among possible numerical values or intervals. Expected value summarizes center, while variance and standard deviation summarize spread. This lesson connects outcome spaces, functions, graphs, calculations, assumptions, and interpretation.
Begin with outcomes and a fixed mapping
An experiment is any repeatable or conceptually repeatable process with an uncertain outcome. Its sample space is the set of possible outcomes. The Greek capital omega, , is a conventional symbol for that space. An event is a subset of outcomes. Probability is assigned to events before a random variable converts outcomes into numerical values.
A random variable is a function . The arrow means that maps each outcome in the sample space to one real number. The capital letter names the random variable. A lowercase commonly represents one possible numerical value. Randomness lies in which outcome occurs, not in whether the mapping changes its rule.
For two coin tosses, the sample space is . Define as the number of heads. Then , , , and . Different outcomes can map to the same numerical value. The distribution combines their probabilities after the mapping.
Distinguish random variables from their realized values
Before an outcome is observed, represents a numerical quantity whose value is uncertain. After observation, a realized value such as can be recorded. The statement describes the event containing every outcome mapped to one. For the two-toss example, that event is . Probability can therefore be written .
The random variable is not itself a probability. It is a numerical function to which probabilities are transferred from outcomes. A value such as can have a probability, but the number two is not a probability. Keeping function, value, event, and probability separate prevents notation from collapsing. Each object plays a different role.
Random variables can carry physical units. If is waiting time, it may be measured in seconds. If is mass, it may be measured in kilograms. Probabilities are dimensionless, while density units depend on the variable. Expected values inherit the variable’s units.
Classify discrete and continuous variables
A discrete random variable has a finite or countably infinite set of possible values. Counts such as number of defects, number of arrivals, or number of successes are usually discrete. The values may be listed in principle even if the list is infinite. Gaps can exist between possible values. Probability may be assigned directly to individual values.
A continuous random variable can take values across intervals. Idealized measurement quantities such as time, length, temperature, or concentration are often modeled continuously. Any one exact value has probability zero under a continuous density model. Intervals can still have positive probability. Probability comes from area rather than from density height alone.
The distinction belongs to the model as well as the measurement. A digital instrument may record temperature to the nearest , producing discrete recorded values. A continuous model may remain a useful approximation to the underlying temperature. Modeling choices simplify reality for a purpose. Their consequences should be acknowledged.
Construct a probability mass function
For a discrete random variable , the probability mass function is . The subscript reminds the reader which random variable the function describes. Every mass satisfies . The masses sum to one, written . The sigma symbol instructs us to add over every possible value.
For the number of heads in two fair tosses, , , and . The middle value receives probability from two outcomes. Adding gives . No probability belongs to values outside zero, one, and two. A bar graph can display the separate masses.
The mass function provides probabilities of events by addition. For example, . The symbol means greater than or equal to. The event includes one or two heads. Discreteness makes endpoint inclusion visible because individual points can carry probability.
Use the cumulative distribution function
The cumulative distribution function is . It records all probability at or below the input value. The definition works for discrete, continuous, and mixed variables. A CDF never decreases as increases. Its values lie between zero and one.
For a discrete variable, the CDF is a step function. It jumps at each value with positive probability. The jump size at equals . Between possible values, the CDF stays constant. Right-continuity determines how each jump endpoint is included.
Interval probabilities can be recovered through subtraction. For , . The left endpoint is excluded because removes probability at or below . Different endpoint choices matter for discrete variables. For a continuous distribution, individual endpoints have zero probability and the choices agree.
Define a continuous probability density
A continuous random variable can be described by a probability density function . The density must satisfy . Its total area is one, written . The integral sign represents continuous accumulation. Infinite limits mean the entire real line is included.
Probability over an interval is . The probability is area under the density curve between the endpoints. Density height is probability per unit of , not probability at one point. A density can exceed one if it remains narrow enough that total area is one. Only accumulated area must lie between zero and one.
If is measured in seconds, density has units of inverse seconds. Multiplying density by the differential width produces a dimensionless probability contribution. This unit analysis explains why density height is not itself probability. Changing measurement units changes numerical density height. The same interval probability remains invariant after proper transformation.
Understand zero probability at a point
For a continuous variable, for every exact value . This does not mean that the variable can never take a value. It means an individual point has zero width and therefore zero density area. Uncountably many zero-probability points collectively form intervals with positive probability. The logic differs from adding a countable list of zeros.
Measurement language often rounds continuous values. An instrument reading may represent an interval such as . That interval can have positive probability. The displayed number is not an infinitely precise event. Resolution links continuous models with discrete records.
For continuous , endpoint inclusion does not affect interval probability. Thus under the model. For discrete , adding or removing an endpoint may change probability. Recognizing the variable type determines whether inequality symbols matter numerically. This is a conceptual check before integration or summation.
Calculate expected value for discrete variables
The expected value of a discrete random variable is when the sum exists. Each possible value is multiplied by its probability and the products are added. The operator means expectation. Expected value is a probability-weighted center. It is not necessarily one of the possible outcomes.
For the two-toss head count, . In this case, the expectation is a possible value. For a fair six-sided die, expectation is even though no face displays . The value describes long-run average behavior across repeated independent trials. It does not predict the next outcome.
Expectation is linear. For constants and , . This rule does not require independence because only one variable appears. More generally, even when and are dependent. Linearity makes totals easier to analyze than their full distributions.
Calculate expectation for continuous variables
For a continuous random variable, expectation is when the integral exists. The value is weighted by density and accumulated over the real line. The formula parallels the discrete probability-weighted sum. Integrals replace sums because values vary continuously. The result inherits the units of .
Suppose is uniform on . Its density is over that interval and zero elsewhere. The expectation is . Symmetry also places the center halfway between the endpoints. Units reduce correctly during integration.
An expectation need not exist even when a valid probability distribution exists. Heavy tails can make the positive and negative weighted contributions fail to converge. A distribution also can have a mean but no finite variance. Formulas require existence conditions. Software output should not replace checking those mathematical assumptions.
Interpret expected value in decisions
Consider a game that pays with probability and loses with probability . Let net payoff be . Its expected value is . Over many independent plays, average payoff tends toward a loss of about twenty-five cents per play. One play still produces either gain or loss.
Expected monetary value is not the only decision criterion. A rare catastrophic loss can matter more than its average contribution suggests. Risk tolerance, utility, legal constraints, and resource limits may alter a rational choice. Two distributions with the same expectation can have very different spreads and tails. Center must be interpreted with variability.
Expectation can also describe conservation across random outcomes. If repeated fair transfers redistribute money without external cost, total expected money may remain fixed. Individual outcomes can still vary substantially. Expected value summarizes an ensemble or long-run average. It does not make uncertainty disappear.
Define variance and standard deviation
Let denote the population mean. Variance is . The difference measures deviation from the mean. Squaring prevents positive and negative deviations from canceling. Averaging squared deviations quantifies spread.
Variance has squared units. If is measured in meters, variance is measured in square meters. Standard deviation is . The Greek letter sigma, , commonly denotes population standard deviation. Taking the square root returns to the original units.
The computational identity is when the required expectations exist. It follows by expanding and using linearity of expectation. The identity can simplify hand calculations. In floating-point computation with very large nearly equal terms, a numerically stable algorithm may be preferable. Algebraic equivalence does not guarantee identical numerical behavior under finite precision.
Compute a discrete variance
Let , , and . The mean is . Next compute . Therefore variance is . The standard deviation is .
If counts events, the mean and standard deviation are expressed in count units, while variance is in squared count units. The calculation uses population probabilities rather than a sample denominator correction. Every probability contributes to both moments. A probability table helps prevent omitted values. The probabilities should sum to one before moments are trusted.
The mean is not required to be an attainable count. It locates the distribution’s balance point. The standard deviation describes a typical scale of variation but is not a strict boundary. Individual values can lie more than one standard deviation away. Distribution shape determines how probability is arranged around the center.
Transform location and scale
Let , where and are constants. The expectation is . Adding shifts every value and therefore shifts the mean by . Multiplying by scales all values and the mean. Units must remain compatible when quantities are added.
Variance transforms as . The shift does not affect spread because it moves every value equally. The factor is squared because deviations are squared. Standard deviation becomes . Absolute value keeps spread nonnegative when is negative.
If and , let . Then . The variance is . The standard deviation changes from to . Subtracting shifts the center but contributes nothing to the spread.
Recognize common discrete models
A Bernoulli random variable records one success or failure. It takes value one with probability and zero with probability . Its mean is , and its variance is . The word success labels the event of interest without implying desirability. The model represents one binary trial.
A binomial variable counts successes in independent Bernoulli trials with constant success probability . Its mean is , and its variance is . Independence and constant probability are substantive assumptions. Sampling without replacement from a small population generally violates independence. Changing trial conditions can invalidate constant .
A geometric variable models the number of trials until the first success under repeated independent equal-probability trials. A Poisson variable often models counts of events in a fixed exposure when a constant-rate and independence structure is plausible. Each distribution is a model with assumptions, not merely a formula shape. Context should be checked before parameters are estimated. Similar-looking count data can arise from different processes.
Recognize common continuous models
The uniform distribution assigns constant density across a finite interval. Equal-length subintervals then have equal probability. The normal distribution has a symmetric bell shape determined by mean and standard deviation . It extends over the entire real line. Many measurement and aggregation processes are approximately normal in central regions.
An exponential distribution models nonnegative waiting times under a constant-rate memoryless process. Its density decreases as time increases. A lognormal distribution models positive quantities whose logarithms are approximately normal. It often has a long right tail. Choosing among these models requires process knowledge and diagnostic checking.
Normality should not be assumed merely because a histogram looks vaguely bell shaped. Sample size, bin choices, tails, truncation, and mixtures can mislead. Quantile plots and subject knowledge provide additional evidence. Even an approximate model may be useful for one purpose and poor for another. The relevant features depend on the intended calculation.
Connect distributions to graphs and simulation
A PMF graph uses separated bars or points because discrete values carry masses. A density graph uses a continuous curve whose areas represent probability. A CDF graph always rises from near zero toward one. These graphs encode different objects. Their vertical axes should not be labeled interchangeably.
Simulation generates outcomes according to a specified distribution. Repeated simulated values form empirical frequencies and averages. As the number of trials grows, these summaries often approach theoretical probabilities and expectations under appropriate conditions. Simulation can build intuition and approximate difficult calculations. It does not validate whether the chosen distribution describes reality.
Empirical distributions summarize observed or simulated data without asserting a smooth theoretical family. The empirical CDF increases by observation weights at recorded values. Comparing it with a fitted model can reveal discrepancies. Graphical comparison is strongest when paired with context and uncertainty. A model is a proposed data-generating description, not the data themselves.
Diagnose common probability mistakes
One mistake is treating density height as probability. Only an integral over an interval gives continuous probability. Another is expecting the mean to be a possible outcome. The mean is a weighted center and can lie between attainable values. A third is forgetting that standard deviation, not variance, shares the variable’s units.
Another mistake is failing to square a scale factor in variance. If values double, deviations double and squared deviations quadruple. A shift changes the mean but not spread. Unit analysis makes these transformations easier to remember. Variance units expose a missing square.
Model assumptions also cause errors. A binomial model requires a fixed number of trials, binary outcomes, independence, and constant success probability. A normal model can assign impossible negative values to inherently positive quantities if variability is large. A distribution name should be accompanied by an argument for why its structure applies. Calculation cannot repair a mismatched model.
Practice a complete distribution routine
First define the experiment, sample space, and random-variable mapping. Second classify the numerical variable as discrete or continuous under the chosen model. Third specify a PMF, density, or CDF with domain and units. Fourth check nonnegativity and total probability one. Fifth calculate the requested probability or moment and interpret it in context.
Suppose delivery time is uniform from to . The density is over that interval. The probability of delivery between and is interval width divided by total width, or . The mean is the midpoint . Units cancel in probability but remain in expectation.
As a discrete check, let count successes in ten independent trials with . A binomial model gives . Its variance is , and standard deviation is . These summaries do not imply that every outcome lies between roughly one and four successes. The complete distribution is needed for exact event probabilities.
Consolidate distribution reasoning
A random variable maps outcomes to numbers. A PMF assigns probability mass to discrete values, while a density assigns probability per unit to continuous regions. A CDF accumulates probability at or below each input and works for either type. Expected value locates a weighted center. Variance and standard deviation quantify spread on squared and original scales.
Transformations change center and spread in predictable ways. Common distribution families package particular assumptions about support, independence, rate, symmetry, and tail behavior. Their formulas become meaningful only when those assumptions fit the process. Graphs and simulation help compare representations. Units keep probabilities, densities, values, and moments distinct.
The strongest solution begins with the outcome process and mapping rather than with a memorized distribution name. It checks normalization, support, endpoint meaning, and dimensional consistency. It interprets averages as long-run or ensemble centers rather than guarantees for individuals. Probability distributions organize uncertainty without eliminating it. Their value lies in making assumptions and consequences explicit.