Aimathic
Login | English | Deutsch

Free math worksheets

Build your own math worksheets from 30,000+ problems for grades 3 to 12, from fractions to AP Calculus. Every problem comes with step-by-step solutions.

Sampling distribution of a difference in proportions

Click problems to add them to your worksheet.

55075712
In two independent populations, the true proportions are \(p_1=0.64\) and \(p_2=0.52\). For repeated independent random samples, what is the mean of the sampling distribution of \(\hat p_1-\hat p_2\)?

Hints

- Focus on the population proportions, not on any one pair of samples. - Keep the subtraction in the stated order: population \(1\) minus population \(2\).

Solution

1. The sampling distribution of \(\hat p_1-\hat p_2\) is centered at \(p_1-p_2\). 2. \(0.64-0.52=0.12\).

Answer

The mean is \(0.12\).
55075812
The true proportions of households using a certain service are \(p_A=0.38\) in population A and \(p_B=0.51\) in population B. What is the mean of the sampling distribution of \(\hat p_B-\hat p_A\)?

Hints

- Match the order of the sample statistics to the order of the population proportions. - A positive result means the first named population has the larger proportion.

Solution

1. The requested order is B minus A, so the mean is \(p_B-p_A\). 2. \(0.51-0.38=0.13\).

Answer

The mean is \(0.13\).
55075912
A random sample is taken from each of two stores. Let \(\hat p_1\) be the sample proportion of customers who are satisfied at store \(1\), and let \(\hat p_2\) be the corresponding sample proportion at store \(2\). In one pair of samples, \(\hat p_1-\hat p_2=0.08\). Interpret this value in context.

Hints

- Interpret the sign as well as the size of the difference. - Convert the decimal difference to percentage points if that makes the comparison clearer.

Solution

1. The statistic compares the two sample proportions in the order store \(1\) minus store \(2\). 2. A difference of \(0.08\) means the observed satisfaction proportion at store \(1\) is \(0.08\), or \(8\) percentage points, higher than at store \(2\).

Answer

The sample satisfaction proportion at store \(1\) is \(8\) percentage points higher than the sample satisfaction proportion at store \(2\).
55076012
Two independent populations both have true success proportion \(0.45\). Repeated random samples are taken from each population. What is the mean of the sampling distribution of \(\hat p_1-\hat p_2\)?

Hints

- Separate the center of the sampling distribution from the value produced by a particular pair of samples. - Ask what happens to the population difference when the two population proportions are equal.

Solution

1. The mean is the difference in the two population proportions. 2. \(0.45-0.45=0\). 3. Individual sample differences can vary even though the sampling distribution is centered at \(0\).

Answer

The mean is \(0\).
55076112
A study concerns two populations with \(p_1=0.58\) and \(p_2=0.53\). One pair of samples happens to give \(\hat p_1-\hat p_2=0.07\). A student says the sampling distribution of \(\hat p_1-\hat p_2\) should be centered at \(0.07\). Is the student correct? State the correct center.

Hints

- Decide whether a sampling-distribution parameter should come from population values or from one observed statistic. - The observed difference is one outcome from the sampling distribution, not its defining center.

Solution

1. The center of the sampling distribution is determined by the population proportions, not by one observed sample difference. 2. The correct center is \(p_1-p_2=0.58-0.53=0.05\). 3. Therefore, the student's statement is not correct.

Answer

No. The correct center is \(0.05\).
55076212
The histogram shows \(40\) simulated values of \(\hat p_1-\hat p_2\) from repeated samples taken with the same design. Estimate the center of the sampling distribution from the graph.
Figure for problem 550762

Hints

- Look for the horizontal value around which the simulated differences balance. - Do not use the height of a single bar as the center; consider the whole distribution.

Solution

1. The simulated differences cluster symmetrically around about \(0.15\). 2. Therefore, the center of the sampling distribution is approximately \(0.15\).

Answer

Approximately \(0.15\).
55076312
Two sampling designs use the same populations, with \(p_1=0.60\) and \(p_2=0.40\). Design X uses \(n_1=n_2=100\), while design Y uses \(n_1=n_2=400\). Compare the centers and standard deviations of the sampling distributions of \(\hat p_1-\hat p_2\).

Hints

- The mean depends on the population proportions, whereas the spread also depends on sample size. - Compare how multiplying both sample sizes by the same factor affects the terms under the square root. - Distinguish a change in center from a change in variability.

Solution

1. Both distributions have mean \(p_1-p_2=0.20\), so their centers are the same. 2. In design Y, each sample size is \(4\) times as large as in design X. 3. Because standard deviation varies with the reciprocal square root of sample size, design Y has one-half the standard deviation of design X.

Answer

The two distributions have the same center, \(0.20\). Design Y has half the standard deviation of design X.
55076412
Independent random samples of size \(100\) are taken from two large populations with \(p_1=0.60\) and \(p_2=0.50\). Find the mean and standard deviation of the sampling distribution of \(\hat p_1-\hat p_2\).

Hints

- Use the population proportions to determine both the center and spread. - The two variance contributions are added before taking the square root. - Keep the two sample sizes paired with their corresponding population proportions.

Solution

1. The mean is \(0.60-0.50=0.10\). 2. The standard deviation is \(\sqrt{\frac{0.60(0.40)}{100}+\frac{0.50(0.50)}{100}}\). 3. This equals \(\sqrt{0.0049}=0.07\).

Answer

Mean: \(0.10\). Standard deviation: \(0.07\).
55076512
Two independent populations have \(p_1=0.72\) and \(p_2=0.55\). Random samples of sizes \(n_1=160\) and \(n_2=240\) are taken. Find the mean and standard deviation of \(\hat p_1-\hat p_2\). Round the standard deviation to four decimal places.

Hints

- Compute the center and spread separately. - Each population contributes its own proportion and sample size to the variance. - Delay rounding until after the square root is evaluated.

Solution

1. The mean is \(0.72-0.55=0.17\). 2. The standard deviation is \(\sqrt{\frac{0.72(0.28)}{160}+\frac{0.55(0.45)}{240}}\). 3. The value is approximately \(0.0479\).

Answer

Mean: \(0.17\). Standard deviation: \(0.0479\).
55076612
For two independent populations, \(p_1=0.45\), \(p_2=0.30\), \(n_1=200\), and \(n_2=150\). Calculate the standard deviation of the sampling distribution of \(\hat p_1-\hat p_2\). Round to four decimal places.

Hints

- This question asks only for spread, not for the center. - Compute each population's variance contribution before adding them. - Take the square root only after the two contributions have been combined.

Solution

1. Use \(\sqrt{\frac{p_1(1-p_1)}{n_1}+\frac{p_2(1-p_2)}{n_2}}\). 2. Substitution gives \(\sqrt{\frac{0.45(0.55)}{200}+\frac{0.30(0.70)}{150}}\). 3. The standard deviation is approximately \(0.0514\).

Answer

\(0.0514\).
55076712
For fixed population proportions, both sample sizes in a two-sample design are multiplied by \(4\). By what factor does the standard deviation of \(\hat p_1-\hat p_2\) change?

Hints

- First determine how the variance changes when both denominators are multiplied by the same factor. - Remember that standard deviation is the square root of variance. - The population proportions do not change in this comparison.

Solution

1. Each variance term is divided by \(4\) when its sample size is multiplied by \(4\). 2. The total variance is therefore one-fourth as large. 3. Taking the square root makes the standard deviation one-half as large.

Answer

The standard deviation is multiplied by \(\frac{1}{2}\).
55076812
Two designs use populations with \(p_1=p_2=0.50\). Design A uses \(n_1=100\) and \(n_2=100\). Design B uses \(n_1=400\) and \(n_2=100\). Which design has the smaller standard deviation for \(\hat p_1-\hat p_2\), and what are the two standard deviations? Round to four decimal places.

Hints

- Compute the two variance contributions separately for each design. - Increasing only one sample size changes only one term under the square root. - Compare the final standard deviations rather than only the sample sizes.

Solution

1. For design A, \(\sigma_A=\sqrt{\frac{0.25}{100}+\frac{0.25}{100}}\approx0.0707\). 2. For design B, \(\sigma_B=\sqrt{\frac{0.25}{400}+\frac{0.25}{100}}\approx0.0559\). 3. Design B has the smaller spread because only one of the two variance contributions was reduced.

Answer

Design B; \(\sigma_A\approx0.0707\) and \(\sigma_B\approx0.0559\).
55076912
Independent random samples of sizes \(n_1=50\) and \(n_2=80\) are taken without replacement from populations of sizes \(N_1=1000\) and \(N_2=2000\). The population proportions are \(p_1=0.40\) and \(p_2=0.60\). Verify the conditions needed for an approximately normal sampling distribution of \(\hat p_1-\hat p_2\).

Hints

- Check the data-collection condition separately from the numerical conditions. - For sampling without replacement, compare each sample size with \(10\%\) of its own population. - For normality, inspect expected successes and expected failures in both populations.

Solution

1. Randomization is satisfied because the problem states that both samples are random and independent. 2. The \(10\%\) condition holds because \(50\le0.10(1000)=100\) and \(80\le0.10(2000)=200\). 3. The expected successes and failures are \(20\), \(30\), \(48\), and \(32\), all at least \(10\). 4. Therefore, the sampling distribution can be modeled as approximately normal.

Answer

All required conditions are satisfied, so an approximately normal model is appropriate.
55077012
Two independent random samples are taken from very large populations. For population \(1\), \(p_1=0.08\) and \(n_1=100\). For population \(2\), \(p_2=0.50\) and \(n_2=80\). Does the usual normality condition hold for the sampling distribution of \(\hat p_1-\hat p_2\)? Explain.

Hints

- The normality condition must succeed for both successes and failures in both populations. - One failed expected-count check is enough to make the condition fail. - Do not replace the population proportion with a pooled proportion for this sampling-distribution check.

Solution

1. Population \(1\) has expected successes \(n_1p_1=100(0.08)=8\). 2. Because \(8<10\), the normality condition fails even though the other expected counts are at least \(10\). 3. Therefore, the usual approximately normal model is not justified by the condition.

Answer

No. The condition fails because \(n_1p_1=8<10\).
55077112
A repeated-sampling simulation for \(\hat p_1-\hat p_2\) is shown. The design uses independent samples, but for population \(1\), \(p_1=0.04\) and \(n_1=100\). Explain how the graph and the numerical condition agree about whether a normal model is appropriate.
Figure for problem 550771

Hints

- Compare the overall shape of the histogram with a bell-shaped distribution. - Check the expected number of successes for population \(1\). - Connect the visual evidence to the theoretical condition rather than treating them as unrelated checks.

Solution

1. The histogram is noticeably right-skewed rather than bell-shaped. 2. Numerically, \(n_1p_1=100(0.04)=4\), which is below \(10\). 3. The failed expected-success condition is consistent with the visibly nonnormal simulated sampling distribution.

Answer

A normal model is not appropriate; the histogram is skewed and \(n_1p_1=4<10\).
55077212
Two independent simple random samples are taken without replacement. Sample \(1\) has size \(90\) from a population of size \(700\), and sample \(2\) has size \(120\) from a population of size \(2000\). The population proportions are \(p_1=0.50\) and \(p_2=0.40\). The expected-success and expected-failure counts are all at least \(10\). Is the usual approximately normal sampling model for \(\hat p_1-\hat p_2\) justified? Explain.

Hints

- Do not stop after checking expected successes and failures. - Sampling without replacement requires a population-size check for each sample. - A condition must hold for both samples, not just one of them.

Solution

1. The randomization and normality conditions are satisfied. 2. For population \(1\), \(10\%\) of \(700\) is \(70\), but \(90>70\). 3. Therefore, the \(10\%\) condition fails for sample \(1\), so the usual independent-sampling approximation is not justified.

Answer

No. The \(10\%\) condition fails because \(90>0.10(700)=70\).
55077312
Independent random samples of sizes \(n_1=200\) and \(n_2=250\) are taken from large populations with \(p_1=0.58\) and \(p_2=0.50\). Conditions for a normal model are satisfied. Find \(P(\hat p_1-\hat p_2>0.10)\). Round to four decimal places.

Hints

- Find the center and spread of the sampling distribution before standardizing. - The requested event is an upper-tail event. - Use the standard normal distribution only after converting the cutoff to a z-score.

Solution

1. The mean is \(0.58-0.50=0.08\). 2. The standard deviation is \(\sqrt{\frac{0.58(0.42)}{200}+\frac{0.50(0.50)}{250}}\approx0.0471\). 3. For \(0.10\), \(z=\frac{0.10-0.08}{0.0471}\approx0.425\). 4. The upper-tail probability is approximately \(0.3355\).

Answer

\(P(\hat p_1-\hat p_2>0.10)\approx0.3355\).
55077412
For two independent populations, \(p_1=0.64\), \(p_2=0.52\), \(n_1=150\), and \(n_2=180\). Conditions for a normal model are satisfied. Find \(P(0.08<\hat p_1-\hat p_2<0.16)\). Round to four decimal places.

Hints

- The interval is centered on the population difference, which creates useful symmetry. - Standardize both endpoints using the same mean and standard deviation. - The desired probability is the area between the two standardized cutoffs.

Solution

1. The mean is \(0.64-0.52=0.12\). 2. The standard deviation is \(\sqrt{\frac{0.64(0.36)}{150}+\frac{0.52(0.48)}{180}}\approx0.0541\). 3. The z-scores for \(0.08\) and \(0.16\) are approximately \(-0.740\) and \(0.740\). 4. The probability between these z-scores is approximately \(0.5406\).

Answer

\(P(0.08<\hat p_1-\hat p_2<0.16)\approx0.5406\).
55077512
Independent random samples are taken from populations with \(p_1=0.55\) and \(p_2=0.43\), using \(n_1=300\) and \(n_2=250\). Conditions for a normal model are satisfied. Find the value \(c\) such that \(P(\hat p_1-\hat p_2<c)=0.90\). Round \(c\) to four decimal places.

Hints

- This is an inverse normal problem rather than a direct tail-probability calculation. - Determine the z-score that leaves \(0.90\) of the standard normal area to its left. - Convert the standardized percentile back to the scale of \(\hat p_1-\hat p_2\).

Solution

1. The mean is \(0.55-0.43=0.12\). 2. The standard deviation is \(\sqrt{\frac{0.55(0.45)}{300}+\frac{0.43(0.57)}{250}}\approx0.04249\). 3. The \(90\)th percentile of the standard normal distribution is \(z\approx1.2816\). 4. Thus \(c=0.12+1.2816(0.04249)\approx0.1745\).

Answer

\(c\approx0.1745\).
55077612
Two independent populations have \(p_1=0.48\) and \(p_2=0.52\). Random samples of size \(400\) are taken from each population, and the normal-model conditions hold. Find the probability that the first sample proportion exceeds the second sample proportion.

Hints

- Translate “the first sample proportion exceeds the second” into an inequality involving their difference. - The sampling distribution is centered at a negative value in this problem. - Pay attention to whether the desired area is to the left or right of \(0\).

Solution

1. The event that the first sample proportion exceeds the second is \(\hat p_1-\hat p_2>0\). 2. The mean is \(0.48-0.52=-0.04\), and the standard deviation is approximately \(0.03533\). 3. For the cutoff \(0\), \(z=\frac{0-(-0.04)}{0.03533}\approx1.132\). 4. Therefore, \(P(\hat p_1-\hat p_2>0)\approx1-0.8712=0.1288\).

Answer

Approximately \(0.1288\).
55077712
The graph compares simulated sampling distributions of \(\hat p_1-\hat p_2\) for two designs that use the same two population proportions. Which design most likely uses the larger sample sizes? Explain using the centers and spreads shown.
Figure for problem 550777

Hints

- Compare the location of the two distributions before comparing their variability. - Ask which feature of a sampling distribution changes when sample sizes increase while population proportions stay fixed. - A larger sample does not systematically shift the center of an unbiased sampling distribution.

Solution

1. Both simulated distributions are centered near \(0.10\), which is consistent with using the same population difference. 2. Design A is visibly narrower than design B. 3. Larger sample sizes reduce the standard deviation of \(\hat p_1-\hat p_2\) without changing its mean, so design A most likely uses the larger sample sizes.

Answer

Design A, because it has about the same center as design B but a smaller spread.
55077812
The graph shows a normal model for the sampling distribution of \(\hat p_1-\hat p_2\) with mean \(0.08\) and standard deviation \(0.04\). The shaded region begins at \(0\) and extends to the right. State the event represented by the shaded region and find its probability.
Figure for problem 550778

Hints

- Use the shaded side of the cutoff to translate the graph into an inequality. - Standardize the boundary value \(0\) using the displayed sampling-distribution parameters. - A cutoff below the mean should leave more than half of the area to its right.

Solution

1. The shaded region represents \(\hat p_1-\hat p_2>0\), meaning the first sample proportion exceeds the second. 2. Standardizing \(0\) gives \(z=\frac{0-0.08}{0.04}=-2\). 3. The area to the right of \(z=-2\) is approximately \(0.9772\).

Answer

The event is \(\hat p_1-\hat p_2>0\), and its probability is approximately \(0.9772\).
55077912
Two independent populations have \(p_1=0.60\) and \(p_2=0.40\). The first sample size is \(n_1=150\). The variance of the sampling distribution of \(\hat p_1-\hat p_2\) is \(0.0028\). Find the second sample size \(n_2\).

Hints

- Work with variance rather than taking a square root, because the given quantity is already a variance. - Separate the known population's contribution from the unknown one. - After isolating the term containing \(n_2\), solve the resulting reciprocal equation.

Solution

1. The variance equation is \(0.0028=\frac{0.60(0.40)}{150}+\frac{0.40(0.60)}{n_2}\). 2. The first variance contribution is \(\frac{0.24}{150}=0.0016\), leaving \(0.0012\) for the second contribution. 3. Solve \(\frac{0.24}{n_2}=0.0012\), which gives \(n_2=200\).

Answer

\(n_2=200\).
55078012
For two independent populations, \(p_1=0.70\), \(p_2=0.50\), \(n_1=120\), and \(n_2=180\). A student computes the standard deviation of \(\hat p_1-\hat p_2\) by first pooling the population proportions and then using the pooled value in both variance terms. Explain why that method is inappropriate for the ordinary sampling distribution, then calculate the correct standard deviation and the student's pooled standard deviation. Round both to four decimal places.

Hints

- Ask what assumption pooling represents and whether that assumption is part of this sampling-distribution problem. - For an ordinary difference-in-proportions sampling distribution, each population contributes its own variance term. - Compare the two numerical spreads only after establishing which model each formula represents.

Solution

1. The ordinary sampling distribution uses the actual population proportions separately: \(\sigma=\sqrt{\frac{0.70(0.30)}{120}+\frac{0.50(0.50)}{180}}\approx0.0560\). 2. Pooling is tied to a hypothesis-test null model that assumes equal population proportions; no such equality is assumed here. 3. The pooled proportion would be \(\frac{120(0.70)+180(0.50)}{300}=0.58\). 4. The student's value is \(\sqrt{0.58(0.42)(\frac{1}{120}+\frac{1}{180})}\approx0.0582\). 5. The pooled calculation therefore changes the spread and is not the correct sampling-distribution standard deviation.

Answer

Pooling is inappropriate here. The correct standard deviation is approximately \(0.0560\); the pooled calculation gives approximately \(0.0582\).

All problems may be used, copied and printed free of charge for school and tutoring, including paid tutoring. Commercial adaptations as well as publication or redistribution on the internet are not permitted.