Short answer

The population is the group you want to learn about; the sample is the group drawn to supply evidence about it. A population parameter describes the whole group, while a sample statistic describes the sampled group. The sampling frame is a separate link: the source from which units are selected. 2 1

On this page

At a glance

QuestionStatistical populationStatistical sample
What does it identify?Group the inquiry concernsGroup drawn for the study
What summarizes it?ParameterStatistic
College GPA exampleAll US college students in the example100 randomly selected students
Main roleDefines the intended inferenceSupplies evidence for that inference

These roles follow Penn State’s terminology and worked example. 2

What each thing is

The target population answers “Who or what is this study about?” The actual sampling frame answers “What source supplies the units we can select?” SIPP makes this distinction concrete: its universe concerns civilian, noninstitutionalized US residents, but its frame is the current Master Address File, from which housing units are sampled. 1

Key differences

Population and sample summaries have different referents even when both are averages or percentages. An average calculated for sampled students is a sample statistic; the average for the entire target student population is a parameter. Using the former to estimate the latter does not make them the same quantity. 2

How to tell them apart

Look for the statement defining everyone the conclusions concern, then look for the selection procedure and the observations analyzed. Label those separately rather than calling every listed group “the sample.” The limit: a selection description alone does not establish exactly which observations were ultimately collected; SIPP’s supplied text primarily describes selection. 1

Where they overlap

The concepts concern related units, not unrelated groups: Penn State describes samples as drawn from populations. Both can support summaries of the same characteristic, such as GPA. Their connection is what makes inference possible, but a shared characteristic does not erase the difference between the group studied and the group being described. 2

Edge cases

Selection can operate at a different level from the target population. SIPP selects geographic primary sampling units and then addresses, while its universe is defined in terms of people. An address list therefore should not be mistaken for the population of residents, nor should selected geographic units be mistaken for individual observations. 1

Why the distinction exists

The distinction separates what researchers want to know from the evidence available to estimate it. It also keeps sampling design visible: SIPP uses a complex design, and treating it as a simple random sample can underestimate standard errors. Knowing the sample’s size alone does not capture that design. 2 1

Common misconceptions

“Sample” does not mean “simple random sample”: SIPP uses staged selection and oversamples low-income households. Also, Penn State’s definition emphasizes representation, but that wording should not be read as proof that any selected group successfully represents its target. The supplied sources do not establish representativeness for an arbitrary study. 1 2

Examples

First, in Penn State’s GPA example, the target is all US college students, the sample contains 100 randomly selected students, and their mean GPA of 2.9 is a statistic—not the population mean. Second, its birth example uses 32 sampled births to investigate the population relationship between birth weight and gestation length. The observed relationship supplies evidence about the broader relationship. 2

Sources

  1. US Census Bureau: Sampling - U.S. Census Bureau
  2. Penn State STAT Online: terminology

Research and drafting are AI-assisted, with citations beside the claims they support. The founder reviews each article before it is selected. This is editorial review, not specialist certification. About WhatDiffers

Report an error