Aimathic
Login | English | Deutsch

Free math worksheets

Build your own math worksheets from 28,000 problems for grades 3 to 12, from fractions to calculus. Every problem includes step-by-step solutions.

Idea of a sampling distribution

Click problems to add them to your worksheet.

54754212
A school district has \(1200\) seniors. A simple random sample of \(80\) seniors is selected, and the sample proportion \(\hat p\) who plan to attend a two-year college is calculated. Identify the population, the sample, the parameter, the statistic, and the sampling distribution in this setting.

Hints

- Separate the full group of interest from the individuals actually observed. - Distinguish a fixed population quantity from a value computed from one sample. - Imagine repeating the sampling method many times to identify the distribution being requested.

Solution

1. The population is all \(1200\) seniors in the district. 2. The sample is the \(80\) selected seniors. 3. The parameter is the true proportion \(p\) of all district seniors who plan to attend a two-year college. 4. The statistic is the observed sample proportion \(\hat p\). 5. The sampling distribution is the distribution of \(\hat p\) over all possible simple random samples of \(80\) seniors from the population.

Answer

Population: all \(1200\) district seniors. Sample: the \(80\) selected seniors. Parameter: the population proportion \(p\). Statistic: the sample proportion \(\hat p\). Sampling distribution: the distribution of \(\hat p\) across all possible simple random samples of size \(80\).
54754312
A researcher takes one random sample of \(200\) voters and obtains \(\hat p=0.54\). The researcher calls the single value \(0.54\) “the sampling distribution of \(\hat p\).” Explain the error and describe what the sampling distribution actually represents.

Hints

- Decide whether the quoted number is one outcome or a collection of possible outcomes. - Imagine repeating the entire sampling process. - Identify what varies from sample to sample.

Solution

1. The value \(0.54\) is one realized statistic from one sample. 2. A sampling distribution contains the possible values of \(\hat p\) and their probabilities under repeated use of the same sampling method. 3. It describes sample-to-sample variability, not the result of only one sample.

Answer

The value \(0.54\) is one sample proportion, not a distribution. The sampling distribution is the probability distribution of \(\hat p\) across all possible random samples of size \(200\) drawn by the same method.
54756712
A finite population has mean \(\mu\). Instead of taking a partial sample, a researcher measures every member of the population and computes the mean. Describe the sampling distribution of this mean if the same census procedure is conceptually repeated.

Hints

- Identify whether any members are left unobserved. - Ask what could vary if the same complete measurement procedure were repeated. - A distribution with only one possible value has no spread.

Solution

1. A census includes the entire population every time. 2. The computed mean is always exactly \(\mu\). 3. Therefore, the sampling distribution places probability \(1\) at \(\mu\) and has standard deviation \(0\).

Answer

The sampling distribution is degenerate at \(\mu\): \(P(\bar X=\mu)=1\), with standard deviation \(0\).
54723612
A probability simulation is repeated in \(50\) independent runs of \(200\) trials each. The \(50\) estimated probabilities have mean \(0.413\) and standard deviation \(0.034\). Interpret the standard deviation in this context. What would generally happen to this run-to-run standard deviation if each run used \(2000\) trials?

Hints

- Distinguish variability among estimates from the probability being estimated. - More trials stabilize relative frequencies. - The number of independent runs and trials per run play different roles.

Solution

1. The standard deviation \(0.034\) describes typical run-to-run variation of the estimated probability around its mean. 2. Increasing the number of trials per run reduces sampling variability. 3. The run-to-run standard deviation would generally become smaller with \(2000\) trials.

Answer

The value \(0.034\) measures typical variation among run estimates. With \(2000\) trials per run, that variation should generally decrease.
54754512
A simulation repeatedly takes a random sample of the same size from a population and records the sample mean. The results of \(1000\) repetitions are summarized below. <table> <tr><th>Sample mean</th><th>\(8\)</th><th>\(9\)</th><th>\(10\)</th><th>\(11\)</th><th>\(12\)</th></tr> <tr><th>Frequency</th><td>\(50\)</td><td>\(220\)</td><td>\(430\)</td><td>\(240\)</td><td>\(60\)</td></tr> </table> a) Use the simulation to estimate \(P(\bar X\ge11)\). b) Estimate the mean of the sampling distribution. c) Explain what would change if the simulation were repeated with many more samples.

Hints

- Convert relevant frequencies to relative frequencies. - Treat the table as an empirical probability distribution. - Distinguish increasing the number of simulated samples from increasing the size of each original sample.

Solution

1. The simulated probability of \(\bar X\ge11\) is \(\frac{240+60}{1000}=0.30\). 2. The simulated sampling-distribution mean is \(\frac{8(50)+9(220)+10(430)+11(240)+12(60)}{1000}=10.04\). 3. With many more repetitions, random simulation noise would generally decrease and the empirical distribution would more closely approximate the true sampling distribution.

Answer

a) \(0.30\). b) \(10.04\). c) More repetitions would usually make the simulated distribution more stable and closer to the true sampling distribution.
54754612
A population has mean \(\mu\) and standard deviation \(\sigma\). Compare the sampling distributions of \(\bar X\) for random samples of sizes \(4\) and \(36\). a) Give the mean and standard deviation of each sampling distribution. b) How many times as large is the standard deviation for \(n=4\)? c) Which sampling distribution would generally be more nearly normal if the population is skewed?

Hints

- Separate what happens to the center from what happens to the spread. - Compare the square roots of the two sample sizes. - Think about how much averaging occurs in each statistic.

Solution

1. Both sampling distributions have mean \(\mu\). 2. For \(n=4\), the standard deviation is \(\frac{\sigma}{2}\); for \(n=36\), it is \(\frac{\sigma}{6}\). 3. The first standard deviation is \(3\) times the second. 4. The \(n=36\) sampling distribution would generally be more nearly normal because it averages more observations.

Answer

a) Both means are \(\mu\). The standard deviations are \(\sigma/2\) and \(\sigma/6\). b) The \(n=4\) standard deviation is \(3\) times as large. c) The \(n=36\) distribution would generally be more nearly normal.
54755112
A simulation uses \(500\) random samples of size \(25\) to approximate the sampling distribution of a sample mean. Consider two changes: a) Increase the number of simulated samples from \(500\) to \(10{,}000\), keeping sample size \(25\). b) Increase each sample size from \(25\) to \(100\), keeping \(500\) simulated samples. Explain the main effect of each change.

Hints

- Decide which change alters the statistic being repeatedly computed. - Separate simulation precision from the mathematical variability of the statistic. - Track what remains fixed in each scenario.

Solution

1. In a), the target sampling distribution is unchanged because the sample size remains \(25\). More repetitions make the simulation's picture of that distribution smoother and more stable. 2. In b), the target sampling distribution changes because the statistic is now based on samples of size \(100\). Its spread is smaller, and its shape is generally closer to normal. 3. The number of repetitions controls simulation accuracy; the sample size controls the actual sampling distribution.

Answer

a) More repetitions improve the empirical approximation to the same sampling distribution. b) Larger samples create a different sampling distribution with less variability and generally a more normal shape.
54755712
For samples of size \(16\), the sampling distribution of \(\bar X\) has mean \(12\) and standard deviation \(1.5\). Let \(S\) be the sample sum from the same samples. a) Find the mean and standard deviation of the sampling distribution of \(S\). b) If a particular sample has \(\bar x=14\), what is its sample sum?

Hints

- Write the exact relationship between the sum and mean for a fixed sample size. - Apply that transformation to both the center and spread. - Use the same relationship for the particular observed sample.

Solution

1. For every sample, \(S=16\bar X\). 2. Therefore, \(E(S)=16(12)=192\). 3. The standard deviation is \(16(1.5)=24\). 4. If \(\bar x=14\), then \(S=16(14)=224\).

Answer

a) Mean \(192\); standard deviation \(24\). b) The sample sum is \(224\).
54755812
Two populations are sampled independently using samples of size \(100\). Population A has proportion \(p=0.10\), and Population B has proportion \(p=0.50\). Let \(\hat p_A\) and \(\hat p_B\) be the sample proportions. a) Find the standard deviation of each sampling distribution. b) Which sample proportion is more variable, and why?

Hints

- Use the population proportion and its complement to describe sampling spread. - Keep the sample size fixed while comparing the two populations. - Compare the products that determine the variance.

Solution

1. For Population A, \(\sigma_{\hat p_A}=\sqrt{\frac{0.10(0.90)}{100}}=0.03\). 2. For Population B, \(\sigma_{\hat p_B}=\sqrt{\frac{0.50(0.50)}{100}}=0.05\). 3. The second sampling distribution is more variable because \(p(1-p)\) is larger at \(p=0.50\) than at \(p=0.10\).

Answer

a) \(\sigma_{\hat p_A}=0.03\) and \(\sigma_{\hat p_B}=0.05\). b) \(\hat p_B\) is more variable.
54756212
A population has standard deviation \(15\). A simulation takes many random samples of size \(25\) and records their means. The simulated sample means have standard deviation \(3.2\). Compare the simulated spread with the theoretical standard deviation of the sampling distribution. Explain why the values need not be exactly equal.

Hints

- Compute the theoretical sampling spread from the population spread and sample size. - Compare the two standard deviations numerically. - Recall that a finite simulation has random simulation variability.

Solution

1. The theoretical standard deviation of the sample mean is \(\frac{15}{\sqrt{25}}=3\). 2. The simulated standard deviation \(3.2\) is \(0.2\) larger than the theoretical value. 3. A finite simulation has random simulation variability, so its empirical standard deviation need not equal the theoretical standard deviation exactly.

Answer

The theoretical standard deviation is \(3\). The simulated value \(3.2\) is slightly larger, which can occur because a finite simulation does not reproduce the theoretical sampling distribution exactly.
54756912
The possible values in the sampling distribution of a sample proportion are spaced \(0.02\) apart. The distribution is centered at \(0.64\). a) Find the sample size. b) Find the population proportion.

Hints

- Relate the support spacing to a one-person change in the sample count. - Use the center of the sampling distribution to identify the population quantity. - Treat the sample size and population proportion as separate features.

Solution

1. Consecutive sample proportions differ by \(1/n\). 2. Since \(1/n=0.02\), the sample size is \(n=50\). 3. The mean of the sampling distribution of \(\hat p\) is the population proportion, so \(p=0.64\).

Answer

a) \(n=50\). b) \(p=0.64\).
54757212
A population has mean \(70\) and standard deviation \(12\). Compare the distribution of one randomly selected observation \(X\) with the sampling distribution of \(\bar X\) for samples of size \(36\). a) Compare their means. b) Compare their standard deviations. c) Explain why the sampling distribution is not the distribution of the \(36\) raw observations in one sample.

Hints

- Identify what one value represents in each distribution. - Apply the sample-size effect only to the average. - Separate within-sample variation from sample-to-sample variation.

Solution

1. Both \(X\) and \(\bar X\) have mean \(70\). 2. The standard deviation of \(X\) is \(12\), while the standard deviation of \(\bar X\) is \(12/\sqrt{36}=2\). 3. The sampling distribution contains one mean from each repeated sample, whereas the raw-data distribution contains individual observations within one sample.

Answer

a) Both means are \(70\). b) The standard deviations are \(12\) for an individual observation and \(2\) for the sample mean. c) The sampling distribution describes statistics across repeated samples, not raw values inside one sample.
54758012
A simulation takes \(1000\) random samples of the same size from a population and records each sample mean. The results are grouped below. <table> <tr><th>Sample-mean interval</th><th>Midpoint</th><th>Frequency</th></tr> <tr><td>\(40\le\bar X<44\)</td><td>\(42\)</td><td>\(100\)</td></tr> <tr><td>\(44\le\bar X<48\)</td><td>\(46\)</td><td>\(250\)</td></tr> <tr><td>\(48\le\bar X<52\)</td><td>\(50\)</td><td>\(400\)</td></tr> <tr><td>\(52\le\bar X<56\)</td><td>\(54\)</td><td>\(200\)</td></tr> <tr><td>\(56\le\bar X<60\)</td><td>\(58\)</td><td>\(50\)</td></tr> </table> a) Estimate the probability that a new sample mean is at least \(52\). b) Use the interval midpoints to estimate the mean of the simulated sampling distribution.

Hints

- Treat the simulation frequencies as empirical probabilities. - Identify every interval included in the requested event. - For the center, weight each representative value by how often its interval occurred.

Solution

1. There are \(200+50=250\) simulated means at least \(52\), so the estimated probability is \(\frac{250}{1000}=0.25\). 2. The midpoint-weighted mean is \(\frac{100(42)+250(46)+400(50)+200(54)+50(58)}{1000}=49.4\).

Answer

a) Approximately \(0.25\). b) Approximately \(49.4\).
54720212
A Monte Carlo simulation estimates an event probability using \(400\) independent trials and observes \(265\) successes. a) Find the estimated probability. b) Estimate the Monte Carlo standard error using \(\sqrt{\hat p(1-\hat p)/n}\). c) Give an approximate interval of two standard errors on either side of the estimate, and interpret it as simulation uncertainty rather than uncertainty in the model itself.

Hints

- The simulation estimate is a sample proportion. - Insert that proportion and the number of independent trials into the standard-error expression. - Distinguish numerical simulation error from uncertainty about whether the probability model is appropriate.

Solution

1. \(\hat p=\frac{265}{400}=0.6625\). 2. The estimated standard error is \(\sqrt{\frac{0.6625\cdot0.3375}{400}}\approx0.0236\). 3. Two standard errors give \(0.6625\pm0.0472\), or approximately \((0.615,0.710)\). 4. This interval describes random Monte Carlo variation from using a finite number of simulated trials, assuming the simulation model is fixed.

Answer

a) \(\hat p=0.6625\). b) Estimated Monte Carlo standard error: approximately \(0.0236\). c) Approximate two-standard-error interval: \((0.615,0.710)\). It describes simulation-to-simulation variation from using \(400\) trials, not uncertainty about the fixed probability model.
54721712
A sample of six delivery times, in minutes, is \(18,21,24,25,29,35\). A bootstrap trial samples six values from this list with replacement and calculates the sample mean. Describe how to use many bootstrap trials to estimate the probability that a resampled mean exceeds \(27\) minutes. Out of \(4000\) trials, \(932\) means exceeded \(27\). Give the estimate.

Hints

- A resample should have the same size as the original sample. - What feature distinguishes resampling with replacement from selecting a new subset? - Use the long-run fraction of resampled means meeting the condition.

Solution

1. For each trial, draw six times independently from the observed list, allowing a value to be selected more than once. 2. Compute the mean of the six resampled values and record whether it exceeds \(27\). 3. The estimated probability is \(\frac{932}{4000}=0.233\).

Answer

The bootstrap estimate is \(0.233\), or \(23.3\%\).
54754412
A population consists of \(2,5,9\). Two values are sampled independently with replacement. Let \(R\) be the range of the two sampled values. a) Construct the sampling distribution of \(R\). b) Find \(E(R)\).

Hints

- Treat ordered samples as equally likely because sampling is with replacement. - Group samples according to the statistic rather than the raw pair. - Use the completed sampling distribution for the long-run average range.

Solution

1. There are nine equally likely ordered samples. The three samples with matching values have range \(0\). 2. The pairs using \(2\) and \(5\) give range \(3\) in two orders; those using \(5\) and \(9\) give range \(4\) in two orders; those using \(2\) and \(9\) give range \(7\) in two orders. 3. The expected range is \(0(3/9)+3(2/9)+4(2/9)+7(2/9)=\frac{28}{9}\approx3.11\).

Answer

a) <table> <tr><th>\(r\)</th><th>\(0\)</th><th>\(3\)</th><th>\(4\)</th><th>\(7\)</th></tr> <tr><th>\(P(R=r)\)</th><td>\(\frac{1}{3}\)</td><td>\(\frac{2}{9}\)</td><td>\(\frac{2}{9}\)</td><td>\(\frac{2}{9}\)</td></tr> </table> b) \(E(R)=\frac{28}{9}\approx3.11\).
54754712
In a population, the proportion with a certain characteristic is \(p=0.40\). A random sample of \(3\) independent individuals is selected, and \(\hat p\) is the sample proportion with the characteristic. a) Construct the sampling distribution of \(\hat p\). b) Find its mean and standard deviation.

Hints

- List every possible count of individuals with the characteristic. - Convert each count to a sample proportion. - Use the resulting distribution to summarize its center and spread.

Solution

1. The number with the characteristic can be \(0,1,2,3\), so \(\hat p\) can be \(0,\frac{1}{3},\frac{2}{3},1\). 2. The corresponding probabilities are \((0.60)^3=0.216\), \(3(0.40)(0.60)^2=0.432\), \(3(0.40)^2(0.60)=0.288\), and \((0.40)^3=0.064\). 3. The sampling-distribution mean is \(E(\hat p)=0.40\). 4. Its standard deviation is \(\sqrt{\frac{0.40(0.60)}{3}}\approx0.2828\).

Answer

a) <table> <tr><th>\(\hat p\)</th><th>\(0\)</th><th>\(\frac{1}{3}\)</th><th>\(\frac{2}{3}\)</th><th>\(1\)</th></tr> <tr><th>Probability</th><td>\(0.216\)</td><td>\(0.432\)</td><td>\(0.288\)</td><td>\(0.064\)</td></tr> </table> b) Mean \(0.40\); standard deviation approximately \(0.2828\).
54754812
A finite population consists of \(10,20,30,40\). A simple random sample of size \(2\) is selected without replacement. Let \(M\) be the smaller sampled value. a) Construct the sampling distribution of \(M\). b) Find \(E(M)\).

Hints

- List the possible samples under the stated sampling method. - Compute the requested statistic for each sample. - Group identical statistic values before finding the expected value.

Solution

1. The six equally likely samples are \(\{10,20\},\{10,30\},\{10,40\},\{20,30\},\{20,40\},\{30,40\}\). 2. Their minimum values are \(10,10,10,20,20,30\). 3. Thus, the probabilities of \(10,20,30\) are \(3/6,2/6,1/6\). 4. The expected minimum is \(10(1/2)+20(1/3)+30(1/6)=\frac{50}{3}\approx16.67\).

Answer

a) <table> <tr><th>\(m\)</th><th>\(10\)</th><th>\(20\)</th><th>\(30\)</th></tr> <tr><th>\(P(M=m)\)</th><td>\(\frac{1}{2}\)</td><td>\(\frac{1}{3}\)</td><td>\(\frac{1}{6}\)</td></tr> </table> b) \(E(M)=\frac{50}{3}\approx16.67\).
54754912
A population consists of \(0,10,20\). A sample of size \(2\) is selected in two different ways: method A samples with replacement, and method B samples without replacement. Let \(\bar X_A\) and \(\bar X_B\) be the sample means. a) Construct both sampling distributions. b) Compare their means and variances.

Hints

- Keep the two sampling mechanisms separate when listing possible samples. - Group repeated sample means and assign probabilities from the appropriate sample space. - Compare both center and spread after constructing the distributions.

Solution

1. With replacement, the nine ordered samples produce means \(0,5,10,15,20\) with probabilities \(1/9,2/9,3/9,2/9,1/9\). 2. Without replacement, the three equally likely samples produce means \(5,10,15\), each with probability \(1/3\). 3. Both sampling distributions have mean \(10\). 4. The with-replacement variance is \(\frac{100}{3}\), while the without-replacement variance is \(\frac{50}{3}\). Sampling without replacement produces less variability here.

Answer

a) Method A: <table> <tr><th>\(\bar x_A\)</th><th>\(0\)</th><th>\(5\)</th><th>\(10\)</th><th>\(15\)</th><th>\(20\)</th></tr> <tr><th>Probability</th><td>\(\frac{1}{9}\)</td><td>\(\frac{2}{9}\)</td><td>\(\frac{1}{3}\)</td><td>\(\frac{2}{9}\)</td><td>\(\frac{1}{9}\)</td></tr> </table> Method B: <table> <tr><th>\(\bar x_B\)</th><th>\(5\)</th><th>\(10\)</th><th>\(15\)</th></tr> <tr><th>Probability</th><td>\(\frac{1}{3}\)</td><td>\(\frac{1}{3}\)</td><td>\(\frac{1}{3}\)</td></tr> </table> b) Both means are \(10\). The variances are \(\frac{100}{3}\) with replacement and \(\frac{50}{3}\) without replacement.
54755012
The three panels come from a study of commute times. One panel shows the population distribution, one shows the data from one random sample of \(50\) commuters, and one shows the sample means from \(5000\) independent random samples of size \(50\). a) Identify which panel represents each distribution. b) Describe how the observational unit and the spread differ among the three distributions. c) Explain why these are three different distributions even though they come from the same population study.
Figure for problem 547550

Hints

- Compare the regularity and spread of the three panels. - Decide whether one plotted value represents one commuter or an average from an entire sample. - A distribution based on many repeated averages should generally be narrower than a distribution of individual values.

Solution

1. Panel a) is the population distribution. It shows the right-skewed distribution of individual commute times. 2. Panel b) is the distribution of the \(50\) observations in one sample. It is also right-skewed but is more irregular because it contains only one finite sample. 3. Panel c) is the sampling distribution of the sample mean. It is much narrower and more nearly symmetric because every plotted value is an average of \(50\) commute times. 4. The observational units are different: individual commuters in the population, individual commuters in one sample, and sample means across repeated samples.

Answer

a) Panel a): population distribution. Panel b): one-sample data distribution. Panel c): sampling distribution of the sample mean. b) The first two distributions use individual commute times, while the third uses sample means. The sampling distribution is much less spread out because each value averages \(50\) observations. c) They describe different sets of values: all population observations, the observations from one sample, and statistics from many repeated samples.
54755212
A population consists of \(1,2,5,8\), with population mean \(4\). Two values are sampled independently with replacement, and \(\bar X\) is the sample mean. a) Find \(P(\bar X=4)\). b) The mean of the sampling distribution of \(\bar X\) is \(4\). Explain why this does not contradict part a).

Hints

- Translate the required sample mean into a required sum of the two observations. - Check all possible population-value pairs. - Distinguish a distribution's center from the values in its support.

Solution

1. For \(\bar X=4\), the two sampled values would need to sum to \(8\). 2. No ordered pair from \(\{1,2,5,8\}\) has sum \(8\), so \(P(\bar X=4)=0\). 3. The mean of a distribution is a weighted balance point and does not have to be one of the distribution's possible values. 4. Averaging all \(16\) equally likely sample means gives \(4\), even though no individual sample mean equals \(4\).

Answer

a) \(P(\bar X=4)=0\). b) A distribution's mean need not be an attainable value; it is the weighted average of all possible sample-mean values.
54755312
A population consists of \(1,4,6\). Two values are sampled independently with replacement. Let \(M\) be the larger sampled value. a) Construct the sampling distribution of \(M\). b) Find \(E(M)\).

Hints

- List or classify the equally likely ordered samples. - Count how many samples produce each possible maximum. - Use the resulting distribution to calculate its expected value.

Solution

1. There are nine equally likely ordered samples. Only \((1,1)\) has maximum \(1\), so its probability is \(1/9\). 2. The samples \((1,4),(4,1),(4,4)\) have maximum \(4\), so the probability is \(3/9\). 3. The remaining five samples have maximum \(6\), so the probability is \(5/9\). 4. The expected maximum is \(1(1/9)+4(3/9)+6(5/9)=\frac{43}{9}\approx4.78\).

Answer

a) <table> <tr><th>\(m\)</th><th>\(1\)</th><th>\(4\)</th><th>\(6\)</th></tr> <tr><th>\(P(M=m)\)</th><td>\(\frac{1}{9}\)</td><td>\(\frac{1}{3}\)</td><td>\(\frac{5}{9}\)</td></tr> </table> b) \(E(M)=\frac{43}{9}\approx4.78\).
54755412
Two independent random samples are taken from the same population with mean \(\mu\) and standard deviation \(\sigma\). The first sample has size \(25\), and the second has size \(100\). Let \(D=\bar X_1-\bar X_2\). a) Find the mean and standard deviation of the sampling distribution of \(D\). b) Which sample contributes more to the variability of \(D\)?

Hints

- Treat each sample mean as a random statistic with its own sampling spread. - Use the independence of the two samples when combining their variability. - Compare the two variance contributions rather than only the sample sizes.

Solution

1. The mean is \(E(D)=\mu-\mu=0\). 2. Independence gives \(\sigma_D=\sqrt{\frac{\sigma^2}{25}+\frac{\sigma^2}{100}}=\frac{\sigma}{\sqrt{20}}\). 3. The first sample contributes variance \(\sigma^2/25\), which is four times the second sample's contribution \(\sigma^2/100\).

Answer

a) Mean \(0\); standard deviation \(\sigma/\sqrt{20}\). b) The sample of size \(25\) contributes more variability.
54755512
A sample mean based on \(50\) observations has standard error \(s\). The population and sampling method remain unchanged. Find the smallest new sample size that makes the standard error at most \(70\%\) of \(s\).

Hints

- Compare the new and old sampling spreads as a ratio. - Use how standard error depends on the square root of sample size. - Round upward because the final spread must not exceed the target.

Solution

1. Standard error is proportional to \(1/\sqrt{n}\). 2. The ratio of new to old standard error is \(\sqrt{\frac{50}{n}}\). 3. Require \(\sqrt{\frac{50}{n}}\le0.70\), so \(n\ge\frac{50}{0.70^2}\approx102.04\). 4. The smallest integer is \(103\).

Answer

\(n=103\).
54755612
A population consists of \(-1\) and \(1\), each equally likely. Two observations are sampled independently with replacement. Let \(S^2\) be the sample variance calculated with denominator \(n-1\). a) Construct the sampling distribution of \(S^2\). b) Find \(E(S^2)\) and compare it with the population variance.

Hints

- List all ordered samples and compute the statistic for each one. - Separate samples with matching values from samples with different values. - Use the sampling probabilities to find the statistic's expected value.

Solution

1. The four ordered samples are equally likely. The samples \((-1,-1)\) and \((1,1)\) have sample variance \(0\). 2. The samples \((-1,1)\) and \((1,-1)\) have mean \(0\) and sample variance \(2\). 3. Thus, \(P(S^2=0)=1/2\) and \(P(S^2=2)=1/2\). 4. The expected sample variance is \(0(1/2)+2(1/2)=1\), equal to the population variance \(1\).

Answer

a) <table> <tr><th>\(s^2\)</th><th>\(0\)</th><th>\(2\)</th></tr> <tr><th>\(P(S^2=s^2)\)</th><td>\(\frac{1}{2}\)</td><td>\(\frac{1}{2}\)</td></tr> </table> b) \(E(S^2)=1\), which equals the population variance.
54755912
A finite population of \(300\) values has mean \(50\) and standard deviation \(12\). A simple random sample of \(90\) values is selected without replacement, and \(\bar X\) is calculated. a) Find the mean of the sampling distribution of \(\bar X\). b) Use the finite-population correction to find its standard deviation. c) Compare it with \(12/\sqrt{90}\) and explain the difference.

Hints

- The sampling mechanism does not shift the center of the sample mean. - Account for the fraction of the finite population that is observed. - Compare how replacement versus no replacement affects uncertainty.

Solution

1. The sampling-distribution mean is the population mean, \(50\). 2. The finite-population standard deviation is \(\frac{12}{\sqrt{90}}\sqrt{\frac{300-90}{300-1}}\approx1.060\). 3. Ignoring the finite population gives \(\frac{12}{\sqrt{90}}\approx1.265\). 4. Sampling \(30\%\) of the population without replacement removes substantial uncertainty, so the corrected spread is smaller.

Answer

a) Mean \(50\). b) Standard deviation approximately \(1.060\). c) The uncorrected value is approximately \(1.265\); it is larger because it ignores the reduced variability from sampling a large fraction without replacement.
54756012
A city wants to estimate average household water use. Method A takes a simple random sample of \(100\) households. Method B randomly selects \(20\) blocks and surveys \(5\) neighboring households on each selected block. Both methods collect \(100\) observations. Households on the same block tend to have similar water use. Which method's sample mean would generally have the more variable sampling distribution? Explain.

Hints

- Compare the dependence among observations under the two designs. - Ask whether every additional household contributes equally new information. - Sampling-distribution spread depends on the sampling method as well as the number of observations.

Solution

1. Method A spreads observations across the city and produces more nearly independent information. 2. In Method B, households within the same block tend to be similar, so several observations may repeat much of the same information. 3. The effective amount of independent information is therefore smaller for Method B. 4. Method B's sample mean would generally have the more variable sampling distribution.

Answer

Method B would generally be more variable because within-block similarity reduces the effective independent information despite the same nominal sample size.
54756312
For random samples of size \(49\), the sampling distribution of \(\bar X\) has standard deviation \(4\). a) Find the population standard deviation. b) Find the sample size needed to reduce the sampling-distribution standard deviation to \(2\).

Hints

- Use the first sampling distribution to recover the population spread. - Keep that population parameter fixed when changing the sample size. - Solve the second standard-error relationship for the new sample size.

Solution

1. Since \(4=\frac{\sigma}{\sqrt{49}}\), the population standard deviation is \(\sigma=4(7)=28\). 2. For a target standard deviation of \(2\), require \(2=\frac{28}{\sqrt{n}}\). 3. Thus, \(\sqrt{n}=14\), so \(n=196\).

Answer

a) \(\sigma=28\). b) \(n=196\).
54756412
The sampling distribution of a sample proportion \(\hat p\) has mean \(0.35\) and standard deviation \(\sqrt{0.002275}\). Assume independent observations. Find the population proportion \(p\) and the sample size \(n\).

Hints

- Use the center of the sampling distribution to identify the population proportion. - Square the stated sampling standard deviation to obtain the variance. - Solve the sampling-variance relationship for the sample size.

Solution

1. The mean of the sampling distribution is the population proportion, so \(p=0.35\). 2. The variance is \(0.002275\), and \(\operatorname{Var}(\hat p)=\frac{p(1-p)}{n}\). 3. Thus, \(0.002275=\frac{0.35(0.65)}{n}\). 4. Since \(0.35(0.65)=0.2275\), \(n=\frac{0.2275}{0.002275}=100\).

Answer

\(p=0.35\) and \(n=100\).
54756612
Independent random samples are taken from two populations. Population 1 has proportion \(p_1=0.40\) with sample size \(100\), and Population 2 has proportion \(p_2=0.30\) with sample size \(150\). Let \(D=\hat p_1-\hat p_2\). Find the mean and standard deviation of the sampling distribution of \(D\).

Hints

- Find the center of each sample proportion before taking their difference. - Use the independence of the samples when combining variability. - Add variance contributions even though the statistic subtracts the proportions.

Solution

1. The mean is \(E(D)=p_1-p_2=0.40-0.30=0.10\). 2. Independence gives \(\operatorname{Var}(D)=\frac{0.40(0.60)}{100}+\frac{0.30(0.70)}{150}=0.0038\). 3. Therefore, \(\sigma_D=\sqrt{0.0038}\approx0.0616\).

Answer

Mean \(0.10\); standard deviation approximately \(0.0616\).
54756812
A website estimates a population proportion by posting an open poll and using the proportion among volunteers. The site repeats this procedure on many days and calls the resulting distribution a sampling distribution. Explain why the repeated values can form a distribution but why standard simple-random-sample formulas may not describe its center or spread.

Hints

- Separate repeated outcomes from the method that generates them. - Check whether every population member has a controlled chance to enter the sample. - Sampling-distribution formulas depend on the sampling design, not only the sample size.

Solution

1. Repeating the volunteer procedure produces a distribution of the resulting proportions, so repeated-sample variation can be observed. 2. The respondents are self-selected rather than randomly sampled from the population. 3. The respondent composition can systematically differ from the population and can vary according to factors not represented in simple random sampling. 4. Therefore, the usual simple-random-sample center and standard-error formulas are not justified for this procedure.

Answer

The repeated volunteer proportions do have a distribution, but it is generated by a self-selection mechanism. It need not be centered at the population proportion or have the spread predicted for a simple random sample.
54757112
For sample proportions based on independent samples of size \(64\), the population proportion \(p\) is unknown. a) What value of \(p\) gives the greatest possible standard deviation of \(\hat p\)? b) What is that maximum standard deviation?

Hints

- Focus on the part of the sampling variance that changes with the population proportion. - Consider where a proportion and its complement are most balanced. - Use the fixed sample size after identifying the maximizing proportion.

Solution

1. The sampling variance is \(\frac{p(1-p)}{64}\). 2. The product \(p(1-p)\) is largest at \(p=0.50\), where it equals \(0.25\). 3. The maximum standard deviation is \(\sqrt{\frac{0.25}{64}}=0.0625\).

Answer

a) \(p=0.50\). b) The maximum standard deviation is \(0.0625\).
54757312
A company has two equally large departments with very different average commute times. Method A takes a simple random sample of \(100\) employees from the whole company. Method B randomly samples \(50\) employees from each department and combines the results using equal weights. Commute times are relatively similar within each department. Which method should generally produce a less variable sampling distribution for the estimated company mean? Explain.

Hints

- Identify the major source of differences in the population. - Compare whether each design fixes or randomizes the representation of the two groups. - Sampling designs can reduce variation without changing the total sample size.

Solution

1. Method B guarantees balanced representation of both equally large departments. 2. Because commute times are relatively homogeneous within each department but differ between departments, stratifying removes department-composition variation from sample to sample. 3. Method A can randomly overrepresent one department, adding variability. 4. Therefore, Method B should generally produce the less variable sampling distribution.

Answer

Method B should generally be less variable because stratification controls the sample's department composition and uses the within-department similarity.
54757412
A finite population has \(5\) members, \(3\) of whom have a certain characteristic. A simple random sample of \(2\) members is selected without replacement, and \(\hat p\) is the sample proportion with the characteristic. a) Construct the sampling distribution of \(\hat p\). b) Find its mean and compare it with the population proportion.

Hints

- List the possible counts of the characteristic in the sample. - Count samples using the available members of each type. - Convert counts to proportions before finding the center.

Solution

1. The sample can contain \(0,1,\) or \(2\) members with the characteristic, so \(\hat p\) can be \(0,0.5,1\). 2. The probabilities are \(\frac{\binom{3}{0}\binom{2}{2}}{\binom{5}{2}}=0.10\), \(\frac{\binom{3}{1}\binom{2}{1}}{\binom{5}{2}}=0.60\), and \(\frac{\binom{3}{2}\binom{2}{0}}{\binom{5}{2}}=0.30\). 3. The sampling-distribution mean is \(0(0.10)+0.5(0.60)+1(0.30)=0.60\), equal to the population proportion \(3/5\).

Answer

a) <table> <tr><th>\(\hat p\)</th><th>\(0\)</th><th>\(0.5\)</th><th>\(1\)</th></tr> <tr><th>Probability</th><td>\(0.10\)</td><td>\(0.60\)</td><td>\(0.30\)</td></tr> </table> b) The mean is \(0.60\), equal to the population proportion.
54757512
A finite population of \(500\) people has proportion \(p=0.40\) with a characteristic. A simple random sample of \(100\) people is selected without replacement. a) Find the standard deviation of \(\hat p\) using the finite-population correction. b) Compare it with the independent-sampling value that ignores the correction.

Hints

- Start with the usual sample-proportion spread. - Account for the substantial fraction of the population being sampled. - Interpret why sampling without replacement reduces remaining uncertainty.

Solution

1. Ignoring the finite population gives \(\sqrt{\frac{0.40(0.60)}{100}}\approx0.0490\). 2. The correction factor is \(\sqrt{\frac{500-100}{500-1}}=\sqrt{\frac{400}{499}}\). 3. The corrected standard deviation is \(0.0490\sqrt{\frac{400}{499}}\approx0.0439\). 4. The corrected value is smaller because sampling without replacement removes one-fifth of the population from further selection.

Answer

a) Approximately \(0.0439\). b) The uncorrected value is approximately \(0.0490\), which is larger.
54757712
In a large population, \(2\%\) of items have a rare label. Independent random samples of \(100\) items are taken, and \(\hat p\) is the sample proportion with the label. a) State the mean and standard deviation of the sampling distribution of \(\hat p\). b) List the first four possible values of \(\hat p\). c) Is a normal model appropriate for this sampling distribution? Explain.

Hints

- Connect the sample proportion to the number of labeled items in a sample. - Determine how much the proportion changes when the count changes by one. - Consider whether both possible outcome counts are expected to occur often enough for symmetry.

Solution

1. The center is \(\mu_{\hat p}=p=0.02\). 2. The standard deviation is \(\sigma_{\hat p}=\sqrt{\frac{0.02(0.98)}{100}}=0.014\). 3. Because the sample count can be \(0,1,2,3,\ldots\), the first four proportions are \(0,0.01,0.02,0.03\). 4. The expected number with the label is only \(100(0.02)=2\), so the distribution is strongly right-skewed rather than approximately normal.

Answer

a) Mean \(0.02\); standard deviation \(0.014\). b) \(0,0.01,0.02,0.03\). c) No. The expected labeled count is only \(2\), so the sampling distribution is strongly right-skewed.
54757812
Two simulation studies repeatedly sample from the same population. Study A uses samples of size \(25\), and the simulated sampling distribution of \(\bar X\) has standard deviation about \(6\). Study B uses a larger fixed sample size, and its simulated sampling distribution has standard deviation about \(3\). a) Estimate the sample size used in Study B. b) How should the centers of the two simulated sampling distributions compare?

Hints

- Compare the two observed spreads as a ratio. - Think about how sampling spread changes when sample size is multiplied. - Changing sample size affects precision, not the population quantity being estimated.

Solution

1. For samples from the same population, the standard deviation of \(\bar X\) is inversely proportional to \(\sqrt n\). 2. Reducing the standard deviation from \(6\) to \(3\) is a factor of \(2\), so the sample size must increase by a factor of \(2^2=4\). 3. Study B therefore uses approximately \(25\cdot4=100\) observations per sample. 4. Both sampling distributions are centered at the same population mean.

Answer

a) Approximately \(100\) observations. b) Their centers should be approximately equal because both estimate the same population mean.
54757912
A population has three equally common categories: A, B, and C. Two independent observations are selected with replacement. Let \(D\) be the number of different categories represented in the sample. Construct the sampling distribution of \(D\), and find its mean and standard deviation.

Hints

- List the ordered pairs of category labels. - Separate pairs with matching labels from pairs with different labels. - Summarize the resulting two-value distribution.

Solution

1. There are \(9\) equally likely ordered samples. 2. The \(3\) samples AA, BB, and CC give \(D=1\), so \(P(D=1)=\frac13\). 3. The other \(6\) samples give \(D=2\), so \(P(D=2)=\frac23\). 4. The mean is \(1\cdot\frac13+2\cdot\frac23=\frac53\). 5. The variance is \(\frac29\), so the standard deviation is \(\frac{\sqrt2}{3}\approx0.471\).

Answer

<table> <tr><th>\(D\)</th><th>\(1\)</th><th>\(2\)</th></tr> <tr><th>Probability</th><td>\(\frac13\)</td><td>\(\frac23\)</td></tr> </table> The mean is \(\frac53\), and the standard deviation is \(\frac{\sqrt2}{3}\approx0.471\).
54758112
The finite population is \(\{2,5,9,12,17\}\). A simple random sample of \(4\) values is selected without replacement, and \(\bar X\) is calculated. Construct the sampling distribution of \(\bar X\), then find its mean and standard deviation.

Hints

- Describe each sample by the value left out. - Use the population total to obtain each sample mean efficiently. - Treat the complete list of simple random samples as equally likely.

Solution

1. Each sample is determined by the one population value omitted, and the population total is \(45\). 2. Omitting \(2,5,9,12,17\) gives sample means \(10.75,10,9,8.25,7\), respectively. 3. The five samples are equally likely, so each mean has probability \(0.20\). 4. The sampling-distribution mean is \(9\), equal to the population mean. 5. The variance is \(1.725\), so the standard deviation is \(\sqrt{1.725}\approx1.313\).

Answer

<table> <tr><th>\(\bar X\)</th><th>\(7\)</th><th>\(8.25\)</th><th>\(9\)</th><th>\(10\)</th><th>\(10.75\)</th></tr> <tr><th>Probability</th><td>\(0.20\)</td><td>\(0.20\)</td><td>\(0.20\)</td><td>\(0.20\)</td><td>\(0.20\)</td></tr> </table> The mean is \(9\), and the standard deviation is approximately \(1.313\).
54758212
In a large population, the proportion who prefer option A is \(p=0.35\). Independent random samples of \(200\) people are taken. Let \(\hat p\) be the sample proportion preferring A, and let \(\hat q\) be the sample proportion not preferring A. a) Find the mean and standard deviation of the sampling distribution of \(\hat q\). b) Describe the relationship between \(\hat p\) and \(\hat q\) across repeated samples.

Hints

- Express the second sample proportion using the first one. - A complement changes the center but not the amount of sample-to-sample variation. - Think about what happens to one statistic whenever the other increases.

Solution

1. The complementary population proportion is \(q=1-0.35=0.65\). 2. The sampling-distribution mean is \(\mu_{\hat q}=0.65\). 3. Its standard deviation is \(\sqrt{\frac{0.65(0.35)}{200}}\approx0.0337\). 4. In every sample, \(\hat q=1-\hat p\), so the two statistics move in exactly opposite directions and have the same standard deviation.

Answer

a) Mean \(0.65\); standard deviation approximately \(0.0337\). b) \(\hat q=1-\hat p\) in every sample, so their sampling distributions are mirror images and the statistics are perfectly negatively related.
54758512
Two independent observations \(X_1\) and \(X_2\) are selected with replacement from the population \(\{0,1,2\}\), with each value equally likely. Let \(D=X_2-X_1\). Construct the sampling distribution of \(D\), and find its mean and standard deviation.

Hints

- Preserve the order of the two observations when listing samples. - Group pairs that yield the same signed difference. - Look for symmetry before carrying out all summary calculations.

Solution

1. The \(9\) ordered pairs are equally likely. 2. Differences \(-2,-1,0,1,2\) occur in \(1,2,3,2,1\) pairs, respectively. 3. The corresponding probabilities are \(\frac19,\frac29,\frac39,\frac29,\frac19\). 4. Symmetry gives mean \(0\). 5. The variance is \(\frac43\), so the standard deviation is \(\frac{2}{\sqrt3}\approx1.155\).

Answer

<table> <tr><th>\(D\)</th><th>\(-2\)</th><th>\(-1\)</th><th>\(0\)</th><th>\(1\)</th><th>\(2\)</th></tr> <tr><th>Probability</th><td>\(\frac19\)</td><td>\(\frac29\)</td><td>\(\frac39\)</td><td>\(\frac29\)</td><td>\(\frac19\)</td></tr> </table> The mean is \(0\), and the standard deviation is \(\frac{2}{\sqrt3}\approx1.155\).
54758612
A population has mean \(12\) and standard deviation \(5\). A sampling plan selects a first observation \(X_1\), then deliberately pairs it with \(X_2=24-X_1\). Let \(\bar X=\frac{X_1+X_2}{2}\). a) Describe the sampling distribution of \(\bar X\). b) Compare its standard deviation with the standard deviation \(\frac{5}{\sqrt2}\) that would result from two independent observations.

Hints

- Substitute the pairing rule into the statistic before using a general sampling formula. - Determine whether the sample mean can vary at all. - Check the independence condition behind the usual standard-error expression.

Solution

1. Substituting the pairing rule gives \(\bar X=\frac{X_1+(24-X_1)}{2}=12\) for every sample. 2. The sampling distribution therefore places probability \(1\) at \(12\). 3. Its standard deviation is \(0\). 4. The independent-sampling value would be \(\frac{5}{\sqrt2}\approx3.536\), which does not apply because the observations are perfectly dependent.

Answer

a) \(\bar X=12\) with probability \(1\), so its sampling distribution is concentrated at \(12\). b) Its standard deviation is \(0\), compared with approximately \(3.536\) for two independent observations.
54758712
A large population has proportion \(p=0.30\). Compare the sampling distributions of \(\hat p\) for independent random samples of sizes \(50\) and \(200\). For each distribution, state its mean, standard deviation, and spacing between consecutive possible values.

Hints

- The target population proportion determines both centers. - Compare the sample sizes inside the square-root expression for spread. - One additional success changes a sample proportion by the reciprocal of the sample size.

Solution

1. Both sampling distributions have mean \(0.30\). 2. For \(n=50\), the standard deviation is \(\sqrt{\frac{0.30(0.70)}{50}}\approx0.0648\), and possible proportions are spaced by \(\frac1{50}=0.02\). 3. For \(n=200\), the standard deviation is \(\sqrt{\frac{0.30(0.70)}{200}}\approx0.0324\), and possible proportions are spaced by \(\frac1{200}=0.005\). 4. Quadrupling the sample size halves the spread and makes the support four times finer.

Answer

For \(n=50\): mean \(0.30\), standard deviation approximately \(0.0648\), spacing \(0.02\). For \(n=200\): mean \(0.30\), standard deviation approximately \(0.0324\), spacing \(0.005\).
54758812
A population value is either \(0\) with probability \(0.40\) or \(1\) with probability \(0.60\). Three independent observations are selected. Let \(M\) be the sample mode. Because the sample size is odd, there is no tie. Construct the sampling distribution of \(M\), and find its mean and standard deviation.

Hints

- Determine how many occurrences are needed for a value to be the mode. - Group samples according to the number of \(1\) values. - Once the two probabilities are known, treat the statistic as a two-value random variable.

Solution

1. The mode is \(1\) when at least two of the three observations equal \(1\). 2. The probability is \(3(0.60)^2(0.40)+(0.60)^3=0.648\). 3. Therefore \(P(M=0)=1-0.648=0.352\). 4. Since \(M\) takes values \(0\) and \(1\), its mean is \(0.648\). 5. Its standard deviation is \(\sqrt{0.648(0.352)}\approx0.478\).

Answer

<table> <tr><th>\(M\)</th><th>\(0\)</th><th>\(1\)</th></tr> <tr><th>Probability</th><td>\(0.352\)</td><td>\(0.648\)</td></tr> </table> The mean is \(0.648\), and the standard deviation is approximately \(0.478\).
54722112
Two teaching methods were used with separate groups, and the observed difference in mean scores is \(4.2\) points. a) Describe a randomization simulation for testing the null claim that the method labels have no effect. b) Describe a bootstrap simulation for estimating the sampling variability of the difference in group means. c) Explain why the two resampling procedures answer different questions.

Hints

- Identify which feature must be destroyed to represent the null hypothesis. - Identify which empirical group distributions must be preserved to estimate uncertainty. - Compare the reference distribution produced by each resampling scheme.

Solution

1. For a randomization test, pool all observed scores, repeatedly shuffle the method labels while preserving the original group sizes, and recompute the difference in means. The simulated distribution represents differences expected under exchangeable labels. 2. For a bootstrap, resample with replacement separately within each observed group, preserving group sizes, and recompute the difference. This approximates sampling variation around the observed group distributions. 3. Label shuffling imposes the null relationship being tested. Within-group bootstrap resampling preserves the observed group difference and estimates uncertainty rather than a null distribution.

Answer

a) Shuffle method labels across all scores while keeping the two group sizes fixed. b) Resample scores with replacement within each group. c) Randomization simulates a null model; bootstrap resampling estimates sampling variability around the observed data.
54756112
A population rating is equally likely to be any integer from \(1\) through \(6\). Two ratings are sampled independently with replacement, and \(\bar X\) is their mean. a) Construct the sampling distribution of \(\bar X\). b) Find its mean and standard deviation. c) Find \(P(\bar X\ge5)\).

Hints

- Organize ordered pairs by their sum. - Convert each possible sum into a sample mean. - Use the frequency pattern to calculate the requested summaries and tail probability.

Solution

1. The \(36\) ordered samples are equally likely. The possible means \(1,1.5,2,2.5,3,3.5,4,4.5,5,5.5,6\) occur with frequencies \(1,2,3,4,5,6,5,4,3,2,1\). 2. The sampling-distribution mean is \(3.5\). 3. The variance is \(\frac{35}{24}\), so the standard deviation is \(\sqrt{\frac{35}{24}}\approx1.2076\). 4. Means at least \(5\) occur in \(3+2+1=6\) of the \(36\) samples, so the probability is \(1/6\).

Answer

a) <table> <tr><th>\(\bar x\)</th><th>\(1\)</th><th>\(1.5\)</th><th>\(2\)</th><th>\(2.5\)</th><th>\(3\)</th><th>\(3.5\)</th><th>\(4\)</th><th>\(4.5\)</th><th>\(5\)</th><th>\(5.5\)</th><th>\(6\)</th></tr> <tr><th>Probability</th><td>\(\frac{1}{36}\)</td><td>\(\frac{2}{36}\)</td><td>\(\frac{3}{36}\)</td><td>\(\frac{4}{36}\)</td><td>\(\frac{5}{36}\)</td><td>\(\frac{6}{36}\)</td><td>\(\frac{5}{36}\)</td><td>\(\frac{4}{36}\)</td><td>\(\frac{3}{36}\)</td><td>\(\frac{2}{36}\)</td><td>\(\frac{1}{36}\)</td></tr> </table> b) Mean \(3.5\); standard deviation \(\sqrt{35/24}\approx1.2076\). c) \(P(\bar X\ge5)=\frac{1}{6}\).
54756512
A population consists of \(0,1,2\), each equally likely. Three observations are sampled independently with replacement. Let \(M\) be the sample median. a) Construct the sampling distribution of \(M\). b) Find \(E(M)\).

Hints

- Classify samples by how many observations fall at each extreme. - Use symmetry to avoid repeating equivalent counting. - Assign the remaining samples to the middle median value.

Solution

1. There are \(27\) equally likely ordered samples. The median is \(0\) when at least two observations are \(0\), which occurs in \(7\) samples. 2. By symmetry, the median is \(2\) in \(7\) samples. 3. The remaining \(13\) samples have median \(1\). 4. Thus, \(E(M)=0(7/27)+1(13/27)+2(7/27)=1\).

Answer

a) <table> <tr><th>\(m\)</th><th>\(0\)</th><th>\(1\)</th><th>\(2\)</th></tr> <tr><th>\(P(M=m)\)</th><td>\(\frac{7}{27}\)</td><td>\(\frac{13}{27}\)</td><td>\(\frac{7}{27}\)</td></tr> </table> b) \(E(M)=1\).
54757012
A finite population consists of \(0,2,4,10\). A simple random sample of size \(3\) is selected without replacement. Let \(D\) be the sample median minus the sample mean. Construct the sampling distribution of \(D\), and find its mean and standard deviation.

Hints

- List samples by the one population value omitted. - For each sample, calculate both center measures before taking their difference. - Treat the resulting statistic values as an equally likely distribution.

Solution

1. The four equally likely samples are obtained by omitting one population value. 2. Omitting \(0,2,4,10\) gives \(D=-\frac43,-\frac23,-2,0\), respectively. 3. Each value has probability \(\frac14\). 4. The mean is \(-1\). 5. The variance is \(\frac59\), so the standard deviation is \(\frac{\sqrt5}{3}\approx0.745\).

Answer

<table> <tr><th>\(D\)</th><th>\(-2\)</th><th>\(-\frac43\)</th><th>\(-\frac23\)</th><th>\(0\)</th></tr> <tr><th>Probability</th><td>\(\frac14\)</td><td>\(\frac14\)</td><td>\(\frac14\)</td><td>\(\frac14\)</td></tr> </table> The mean is \(-1\), and the standard deviation is \(\frac{\sqrt5}{3}\approx0.745\).
54757612
A population has mean \(\mu\) and standard deviation \(\sigma\). Before taking a sample, a fair coin is flipped. If it lands heads, one observation is selected; if it lands tails, four independent observations are selected. Let \(M\) be the mean of the observations selected. Find the mean and standard deviation of the sampling distribution of \(M\).

Hints

- Consider the two possible sample sizes separately. - Compare the centers of the two conditional sampling distributions. - Combine variances only after accounting for how often each sampling plan is used.

Solution

1. Conditional on a sample size of \(1\), the sample mean has mean \(\mu\) and variance \(\sigma^2\). 2. Conditional on a sample size of \(4\), the sample mean has mean \(\mu\) and variance \(\frac{\sigma^2}{4}\). 3. Both conditional means equal \(\mu\), so the unconditional mean is \(\mu\). 4. The unconditional variance is \(\frac12\sigma^2+\frac12\cdot\frac{\sigma^2}{4}=\frac{5\sigma^2}{8}\). 5. The standard deviation is \(\sigma\sqrt{\frac58}\).

Answer

The sampling distribution has mean \(\mu\) and standard deviation \(\sigma\sqrt{\frac58}\).
54758312
A systematic sample selects every second item from the production sequence shown. The starting position is chosen uniformly at random from the first two positions. Let \(\bar X\) be the sample mean. a) Construct the sampling distribution of \(\bar X\). b) Compare its center and shape with what would be expected from simple random samples of \(5\) items.
Figure for problem 547583

Hints

- Follow each possible random starting position through the positions shown. - Do not assume that a fairly large sample fraction guarantees a small spread when the sampling interval matches a population pattern. - Compare whether each method mixes the two item values within a sample.

Solution

1. Starting at position \(1\) selects positions \(1,3,5,7,9\), whose values are all \(0\), so \(\bar X=0\). 2. Starting at position \(2\) selects positions \(2,4,6,8,10\), whose values are all \(10\), so \(\bar X=10\). 3. Each start has probability \(0.5\), so the sampling distribution places probability \(0.5\) at each of \(0\) and \(10\). 4. Its mean is \(5\), equal to the population mean, but it is extremely spread out and two-point rather than concentrated near \(5\). 5. Simple random samples of \(5\) items would usually contain a mix of both values and have means closer to \(5\).

Answer

a) \(P(\bar X=0)=0.5\) and \(P(\bar X=10)=0.5\). b) The center is \(5\), but this systematic design produces a two-point, highly variable distribution. Simple random sampling would produce means concentrated more closely around \(5\).
54758412
A population consists of \(0,2,\) and \(6\), each equally likely. Two observations are selected independently with replacement. Define the sample midrange \(R\) as the average of the smaller and larger observations. Construct the sampling distribution of \(R\), and find its mean and standard deviation.

Hints

- List ordered samples because the two selections are independent. - Several different samples can produce the same midpoint of the extremes. - Combine equal statistic values before finding the numerical summaries.

Solution

1. The \(9\) ordered samples are equally likely. 2. The midrange values \(0,1,2,3,4,6\) occur in \(1,2,1,2,2,1\) ordered samples, respectively. 3. Their probabilities are therefore \(\frac19,\frac29,\frac19,\frac29,\frac29,\frac19\). 4. The mean is \(\frac83\). 5. The variance is \(\frac{28}{9}\), so the standard deviation is \(\frac{2\sqrt7}{3}\approx1.764\).

Answer

<table> <tr><th>\(R\)</th><th>\(0\)</th><th>\(1\)</th><th>\(2\)</th><th>\(3\)</th><th>\(4\)</th><th>\(6\)</th></tr> <tr><th>Probability</th><td>\(\frac19\)</td><td>\(\frac29\)</td><td>\(\frac19\)</td><td>\(\frac29\)</td><td>\(\frac29\)</td><td>\(\frac19\)</td></tr> </table> The mean is \(\frac83\), and the standard deviation is \(\frac{2\sqrt7}{3}\approx1.764\).

All problems may be used, copied and printed free of charge for school and tutoring, including paid tutoring. Commercial adaptations as well as publication or redistribution on the internet are not permitted.