Aimathic
Login | English | Deutsch

Free math worksheets

Build your own math worksheets from 30,000+ problems for grades 3 to 12, from fractions to AP Calculus. Every problem comes with step-by-step solutions.

Hypothesis test for a mean or mean difference

Click problems to add them to your worksheet.

54664412
A package-delivery company claims that its mean delivery time for local shipments during July 2026 is less than \(2\,\text{days}\). Write an investigative question for evaluating that claim.

Hints

- Identify the population mean named in the claim. - Preserve the direction of the comparison. - State the type of shipments and time period represented.

Solution

1. The variable is delivery time in days for a local shipment. 2. The parameter is the population mean delivery time for all local shipments during July 2026. 3. The directional claim is that the mean is less than \(2\,\text{days}\). 4. The question should state the population and benchmark.

Answer

One valid question is: “For all local shipments during July 2026, is the mean delivery time less than \(2\,\text{days}\)?”
54772312
Professor Daniel Kim wants to test whether a population mean is less than \(40\). Lena Ortiz writes \(H_0:\mu<40\) and \(H_a:\mu=40\). Correct the hypotheses and explain the role of equality in the null hypothesis.

Hints

- Put the directional research question in the alternative hypothesis. - Ask which hypothesis supplies the fixed population value used when calculating the test statistic. - The null and alternative should be written about the population mean, not the sample mean.

Solution

1. The research claim “less than \(40\)” belongs in the alternative hypothesis. 2. The null hypothesis specifies the reference value used to form the test distribution, so it contains equality. 3. The correct hypotheses are \(H_0:\mu=40\) and \(H_a:\mu<40\).

Answer

Use \(H_0:\mu=40\) and \(H_a:\mu<40\). The equality belongs in the null hypothesis because the test evaluates sample evidence relative to that fixed reference value.
54781912
Dr. Farah Malik has only one observation from a population and wants to perform a one-sample \(t\)-test for the population mean. Explain why the standard one-sample \(t\)-test cannot be carried out with \(n=1\).

Hints

- Identify the sample quantity used to estimate the standard error in a \(t\)-test. - Ask whether variability can be estimated from a single observed value. - Check the degrees-of-freedom rule when \(n=1\).

Solution

1. A one-sample \(t\)-test estimates the standard error using the sample standard deviation \(s\). 2. With only one observation, there is no within-sample variation from which to calculate a sample standard deviation. 3. The degrees of freedom would be \(n-1=0\), so there is no corresponding \(t\)-reference distribution for the standard procedure. 4. Therefore, a standard one-sample \(t\)-test cannot be performed from a single observation.

Answer

The test cannot be performed with \(n=1\) because the sample standard deviation is undefined and the procedure would have \(0\) degrees of freedom.
54791512
A valid one-sample \(t\)-test is carried out with significance level \(\alpha=0.05\), and the exact p-value is \(0.050000\), not a rounded display. Using the decision rule “reject \(H_0\) when \(p\le\alpha\),” what is the correct decision? Explain why the fact that equality occurs matters.

Hints

- Read the inequality in the decision rule carefully. - Compare the exact p-value with the significance level, including the equality case. - Distinguish an exact value from a rounded software display.

Solution

1. The stated decision rule rejects whenever the p-value is less than or equal to the significance level. 2. Here, \(p=0.050000\) and \(\alpha=0.05\), so \(p=\alpha\). 3. Because equality is included in the rejection rule, reject \(H_0\). 4. This conclusion depends on the p-value being given as exact; a displayed rounded value of \(0.050\) could hide a value slightly above or below \(0.05\).

Answer

Reject \(H_0\) under the stated rule because the exact p-value equals \(\alpha\), and the rule uses \(p\le\alpha\).
54765112
A cereal manufacturer wants to test whether the population mean fill weight is less than the labeled \(52\,\text{g}\). Miles Turner writes the hypotheses as \(H_0:\bar{x}=52\) and \(H_a:\bar{x}<52\). Identify the error, write the hypotheses correctly, and explain why the corrected hypotheses use a different symbol from Miles's version.

Hints

- Ask whether a hypothesis should describe a sample already observed or a population quantity that is unknown. - Distinguish the symbol for a sample mean from the symbol for a population mean. - Make sure the direction of the alternative matches the manufacturer's question about being below the label value.

Solution

1. Hypotheses must be statements about the population parameter, not the sample statistic. 2. Let \(\mu\) be the population mean fill weight for the cereal boxes. 3. The correct hypotheses are \(H_0:\mu=52\) and \(H_a:\mu<52\). 4. The sample mean \(\bar{x}\) is observed from the sample and is used as evidence about the unknown population mean \(\mu\); it is not the parameter being tested.

Answer

The error is using \(\bar{x}\), a sample statistic, in the hypotheses. The correct hypotheses are \(H_0:\mu=52\) and \(H_a:\mu<52\), where \(\mu\) is the population mean fill weight.
54765712
A coffee roaster tests whether the population mean roast loss is different from \(15\%\). A valid one-sample \(t\)-test gives a p-value of \(0.080\). Operations manager Ana Costa says, “There is an \(8\%\) chance that the true population mean roast loss is \(15\%\).” Explain why Ana's statement is incorrect. Give a correct interpretation of the p-value and state the conclusion at \(\alpha=0.05\).

Hints

- Ask what assumption is made before a p-value is calculated. - Separate a probability about possible sample results from a probability about a fixed population parameter. - Base the formal decision on the stated significance level, then translate it back to the original claim.

Solution

1. A p-value is calculated assuming the null hypothesis is true; it is not the probability that the null hypothesis is true. 2. The p-value means that if the population mean roast loss were \(15\%\), the probability of obtaining a sample result at least as far from \(15\%\) as the observed result, in either direction, would be about \(0.080\). 3. Since \(0.080>0.05\), fail to reject \(H_0\). 4. The data do not provide convincing evidence that the population mean roast loss differs from \(15\%\).

Answer

The p-value is not the probability that \(H_0\) is true. Assuming \(\mu=15\%\), there is about an \(8\%\) chance of getting a sample result at least as extreme as the observed one, in either direction. Because \(0.080>0.05\), fail to reject \(H_0\); there is not convincing evidence that the population mean roast loss differs from \(15\%\).
54767512
A valid two-sided one-sample \(t\)-test for a population mean produces \(t=2.12\) with \(17\) degrees of freedom. The significance level is \(\alpha=0.05\). Find the p-value to four decimal places, make the test decision, and explain why this result should be described carefully rather than as overwhelming evidence.

Hints

- A two-sided test counts extreme results on both sides of the reference distribution. - Compare the resulting probability with the significance level only after accounting for both tails. - The distance of a p-value from the cutoff can help communicate the strength of evidence without changing the formal decision.

Solution

1. For a two-sided test, the p-value is twice the upper-tail probability beyond \(|t|=2.12\) for \(T_{17}\). 2. The p-value is approximately \(0.0490\). 3. Since \(0.0490<0.05\), reject \(H_0\) at the \(5\%\) significance level. 4. The p-value is only slightly below \(0.05\), so the evidence is just sufficient under this decision rule; it should not be characterized as overwhelmingly strong evidence against \(H_0\).

Answer

The two-sided p-value is approximately \(0.0490\). Reject \(H_0\) at \(\alpha=0.05\), but describe the evidence as relatively modest because the p-value is very close to the significance cutoff.
54768712
A valid one-sample \(t\)-test for a population mean gives a p-value of \(0.12\). Dr. Marco Rossi says, “Since we failed to reject the null hypothesis, we have shown that the null value is the true population mean.” Explain why this conclusion is too strong and state what a failure to reject the null hypothesis actually means.

Hints

- A test decision describes the strength of evidence in the sample, not certainty about a fixed parameter. - Think about the difference between “not enough evidence against” and “proved true.” - Consider whether a larger or more precise study could potentially lead to a different evidence assessment.

Solution

1. A hypothesis test evaluates whether the sample provides sufficiently strong evidence against the null hypothesis under a chosen significance rule. 2. A p-value of \(0.12\) may be too large to reject \(H_0\) at common significance levels such as \(0.05\). 3. Failing to reject \(H_0\) means the data do not provide convincing evidence for the alternative hypothesis at the chosen significance level. 4. It does not prove that the null hypothesis is true or that the population mean equals the null value exactly.

Answer

A failure to reject \(H_0\) is not proof that \(H_0\) is true. It means only that the sample does not provide sufficiently strong evidence for the alternative hypothesis at the chosen significance level.
54769312
A random sample of \(80\) observations is used to test a claim about a population mean. The population standard deviation is unknown, but the sample standard deviation is available. Zoe Williams says, “Since the sample is large, use a one-sample \(z\)-test instead of a one-sample \(t\)-test.” Evaluate Zoe's recommendation.

Hints

- Decide which measure of population spread is actually known before selecting the reference distribution. - A large sample makes the \(t\)- and standard normal distributions similar, but it does not make an unknown population standard deviation known. - Separate the shape condition from the question of whether population variability is known.

Solution

1. The population standard deviation \(\sigma\) is unknown, so sampling uncertainty must be estimated using the sample standard deviation \(s\). 2. The appropriate procedure for a population mean with unknown \(\sigma\) is a one-sample \(t\)-test. 3. A large sample makes the \(t\)-distribution close to the standard normal distribution, but the reference distribution remains \(t\) when \(\sigma\) is estimated by \(s\). 4. Therefore, use a one-sample \(t\)-test with \(df=79\), assuming the other conditions are met.

Answer

Use a one-sample \(t\)-test with \(df=79\), not a \(z\)-test. The population standard deviation is unknown, so the sample standard deviation is used and the reference distribution is \(t\).
54769912
Dr. Noura El-Sayed measures reaction time for a random sample of participants before and after a training program. The paired difference is defined as \(d=\text{before}-\text{after}\). The research question is whether the population mean reaction time is higher after training than before it. Caleb Lee writes \(H_a:\mu_d>0\). Determine whether Caleb's alternative hypothesis has the correct direction. Write the correct hypotheses and explain the sign.

Hints

- Test the sign convention using a simple example in which the after value is larger than the before value. - Translate the verbal claim through the exact definition of the paired difference. - Keep the null centered at the no-change value.

Solution

1. Higher reaction time after training means \(\text{after}>\text{before}\). 2. Since \(d=\text{before}-\text{after}\), a higher after value corresponds to a negative difference. 3. Therefore, the correct hypotheses are \(H_0:\mu_d=0\) and \(H_a:\mu_d<0\). 4. Caleb's \(H_a:\mu_d>0\) would instead represent lower reaction times after training. 5. These hypotheses test a before/after population mean difference; without randomized treatment assignment, they do not by themselves establish that the training program caused any change.

Answer

Caleb's direction is incorrect. With \(d=\text{before}-\text{after}\), higher reaction time after training gives negative differences, so use \(H_0:\mu_d=0\) and \(H_a:\mu_d<0\). The before/after comparison alone does not establish causation.
54770512
In a one-sample \(t\)-test, the test statistic is \(t=-2.4\). Yuna Park says this means the sample mean is \(2.4\) units below the null value. Explain the correct meaning of \(t=-2.4\) and identify the scale represented by the number \(2.4\).

Hints

- A test statistic compares an observed difference with the amount of sampling variability expected under the null. - Ask what units remain after dividing a difference by its standard error. - The sign gives direction, while the magnitude gives standardized distance from the null value.

Solution

1. A \(t\) statistic standardizes the difference between the sample mean and the null hypothesized population mean. 2. The value \(t=-2.4\) means the sample mean is \(2.4\) estimated standard errors below the null value. 3. It does not mean the original measurement difference is \(2.4\) units; the actual difference depends on the standard error.

Answer

\(t=-2.4\) means the sample mean is \(2.4\) estimated standard errors below the null hypothesized mean. The number \(2.4\) is on a standardized scale, not necessarily the original measurement scale.
54771712
A valid \(95\%\) one-sample \(t\)-confidence interval for a population mean is \((10.2,14.6)\). Dr. Anika Bose wants to test \(H_0:\mu=15\) versus \(H_a:\mu\ne15\) at \(\alpha=0.05\). Use the confidence interval to determine the test decision and state the conclusion.

Hints

- Identify the parameter value that the null hypothesis claims. - A corresponding confidence interval shows which parameter values remain plausible at that confidence level. - The interval's location also indicates the direction of the observed departure from the null value.

Solution

1. A two-sided test at \(\alpha=0.05\) corresponds to checking whether the null value is contained in the \(95\%\) confidence interval for the same population mean. 2. The null value \(15\) is not contained in \((10.2,14.6)\). 3. Therefore, reject \(H_0\). 4. The data provide convincing evidence that the population mean differs from \(15\); the interval suggests it is lower than \(15\).

Answer

Reject \(H_0\). Since \(15\) is outside the \(95\%\) confidence interval, there is convincing evidence at \(\alpha=0.05\) that the population mean differs from \(15\), with the observed direction below \(15\).
54774712
A delivery company tests \(H_0:\mu=2.0\) days against \(H_a:\mu>2.0\) days, where \(\mu\) is the population mean delivery time for a certain service. Describe a Type I error and a Type II error in this context.

Hints

- Start by identifying what decision corresponds to rejecting the null hypothesis. - For one error, imagine rejecting a true null; for the other, imagine not rejecting when the alternative is true. - Translate each logical error back into the delivery-time claim.

Solution

1. A Type I error occurs when the null hypothesis is true but is rejected. 2. In context, that means concluding that the population mean delivery time is greater than \(2.0\) days when the population mean is actually \(2.0\) days. 3. A Type II error occurs when the alternative is true but the null hypothesis is not rejected. 4. In context, that means failing to conclude that the population mean delivery time is greater than \(2.0\) days when it actually is greater than \(2.0\) days.

Answer

Type I error: conclude that \(\mu>2.0\) days when \(\mu=2.0\) days. Type II error: fail to conclude that \(\mu>2.0\) days when \(\mu>2.0\) days is actually true.
54775912
Dr. Priya Shah measures typing speed for the same \(20\) participants before and after a training session and plans a paired \(t\)-test for the population mean change. One participant has a recorded “before” speed but no “after” speed. How should that incomplete observation be handled for the paired test, and how many degrees of freedom will the test use if the remaining \(19\) pairs satisfy the conditions for \(t\) inference? Also explain what concern remains if the missing “after” value is related to that participant's outcome.

Hints

- Identify the actual observations analyzed in a paired test. - A participant contributes only if a difference can be calculated for that participant. - Separate the mechanical degrees-of-freedom calculation from the question of whether the remaining complete cases are representative.

Solution

1. A paired analysis uses one difference for each participant with both measurements available. 2. The participant with no “after” value has no paired difference, so the lone “before” value cannot be included in the paired \(t\)-test. 3. The analysis therefore uses \(19\) complete differences. 4. A one-sample \(t\)-test on \(19\) differences uses \(df=19-1=18\). 5. If the missing “after” measurement is related to the participant's outcome, the \(19\) complete pairs may be systematically different from the target population, so population inference can be biased even though the degrees-of-freedom calculation is correct.

Answer

Exclude the incomplete pair from the paired analysis. Use the \(19\) complete differences, giving \(18\) degrees of freedom. If the missing value is outcome-related, the complete pairs may not be representative, so the population conclusion may be biased.
54777712
Dr. Irene Kwon uses significance level \(\alpha=0.01\) for a valid one-sample \(t\)-test of a population mean. Mason Reed says, “This means there is a \(1\%\) chance that the null hypothesis is true.” Explain the correct interpretation of \(\alpha=0.01\).

Hints

- Interpret the significance level under the assumption that the null hypothesis is true. - Think about repeated use of the same decision rule. - Keep a procedure's error rate separate from a probability assigned to a hypothesis.

Solution

1. The significance level is a property of the testing procedure under the assumption that the null hypothesis is true. 2. If the null hypothesis is true and the same testing procedure is repeated many times under the same conditions, about \(1\%\) of those tests would reject the null hypothesis. 3. The significance level is not the probability that the null hypothesis is true.

Answer

\(\alpha=0.01\) means the testing procedure has a \(1\%\) Type I error rate: when \(H_0\) is true, it rejects \(H_0\) about \(1\%\) of the time in repeated use. It does not mean there is a \(1\%\) chance that \(H_0\) is true.
54778312
Dr. Luis Ferreira is planning a one-sample test for a population mean. The sample size, population variability, true alternative mean, and test direction will stay fixed. Dr. Ferreira considers lowering the significance level from \(\alpha=0.05\) to \(\alpha=0.01\). How will this change affect the Type I error rate and the power of the test?

Hints

- Think of the significance level as controlling how easy it is to enter the rejection region under the null hypothesis. - A stricter rejection rule affects both false rejections and successful detection of a real effect. - Hold the sample size and true alternative value fixed while comparing the two rules.

Solution

1. Lowering \(\alpha\) makes the rejection region more difficult to reach when the null hypothesis is true. 2. Therefore, the Type I error rate decreases from \(5\%\) to \(1\%\). 3. With all other features fixed, the more restrictive rejection rule also makes it harder to reject the null when the alternative is true. 4. Therefore, the power decreases and the Type II error probability increases.

Answer

Lowering \(\alpha\) from \(0.05\) to \(0.01\) decreases the Type I error rate but also decreases power, so the Type II error probability increases.
54778912
In Study A, a one-sample \(t\)-test has \(n=25\), sample standard deviation \(s=10\), and a sample mean that is \(4\) units above the null value. In Study B, the sample standard deviation is also \(10\), but \(n=100\). How far above the null value must Study B's sample mean be to produce the same test statistic as Study A?

Hints

- Find the standardized distance produced by Study A first. - Compare the standard errors for the two sample sizes. - The same test statistic requires the raw distance from the null to scale with the standard error.

Solution

1. Study A has test statistic \(t=\frac{4}{10/\sqrt{25}}=\frac{4}{2}=2\). 2. In Study B, the standard error is \(\frac{10}{\sqrt{100}}=1\). 3. To produce \(t=2\), the sample mean must be \(2\cdot1=2\) units above the null value. 4. The larger sample needs only half the raw difference because its standard error is half as large.

Answer

Study B's sample mean must be \(2\) units above the null value.
54779512
In a valid one-sample \(t\)-test, the sample mean is exactly equal to the null-hypothesis mean, so the test statistic is \(t=0\). What is the p-value for a two-sided alternative? What is the p-value for a one-sided alternative such as \(H_a:\mu>\mu_0\)? Explain.

Hints

- Locate the observed statistic at the center of the reference distribution. - For a one-sided test, consider the area from the observed statistic into the relevant tail. - For a two-sided test, ask what values are at least as far from \(0\) as \(0\) itself.

Solution

1. A \(t\)-distribution is symmetric about \(0\), so exactly half of its probability lies above \(0\) and half below \(0\). 2. For a one-sided upper-tail alternative, the p-value is \(P(T\ge0)=0.5\). 3. For a two-sided alternative, outcomes at least as far from \(0\) as an observed statistic of \(0\) include the entire distribution, so the p-value is \(1\).

Answer

The two-sided p-value is \(1\). For an upper-tailed one-sided alternative, the p-value is \(0.5\).
54780112
A valid two-sided one-sample \(t\)-test of \(H_0:\mu=50\) gives p-value \(0.030\). Without constructing either interval, determine whether \(50\) is contained in the corresponding \(95\%\) confidence interval and whether it is contained in the corresponding \(99\%\) confidence interval.

Hints

- Match a two-sided test at significance level \(\alpha\) with a confidence interval having confidence level \(1-\alpha\). - Compare the same p-value with \(0.05\) and \(0.01\). - Higher-confidence intervals are wider, so inclusion can change as the confidence level rises.

Solution

1. Since \(0.030<0.05\), the two-sided test rejects \(H_0:\mu=50\) at \(\alpha=0.05\). Therefore, \(50\) is outside the corresponding \(95\%\) confidence interval. 2. Since \(0.030>0.01\), the test fails to reject \(H_0\) at \(\alpha=0.01\). Therefore, \(50\) is inside the corresponding \(99\%\) confidence interval. 3. The wider \(99\%\) interval can include the null value even when the narrower \(95\%\) interval excludes it.

Answer

\(50\) is not in the \(95\%\) confidence interval, but it is in the \(99\%\) confidence interval.
54780712
A random sample of \(16\) observations has sample mean \(50.8\), sample median \(52.0\), and sample standard deviation \(3.2\). The sample distribution is roughly symmetric with no outliers. Amara Okafor wants to test \(H_0:\mu=50\) versus \(H_a:\mu\ne50\) and uses the sample median in the numerator of the \(t\)-statistic. Explain the error, then calculate the correct test statistic and the p-value to three decimal places.

Hints

- Match the sample statistic to the population parameter named in the hypotheses. - The null hypothesis is about a population mean. - After choosing the correct point estimate, standardize its distance from the null value.

Solution

1. A one-sample \(t\)-test for a population mean uses the sample mean, not the sample median, as the point estimate. The random sample and stated shape support the small-sample \(t\)-procedure. 2. The correct test statistic is \(t=\frac{50.8-50}{3.2/\sqrt{16}}=1.00\). 3. With \(df=15\), the two-sided p-value is approximately \(0.333\). 4. The sample median is a different statistic and does not belong in this test statistic for \(\mu\).

Answer

Amara should use \(\bar{x}\), not the sample median. The correct statistic is \(t=1.00\), with two-sided p-value approximately \(0.333\).
54781312
Two one-sample tests use the same sample size, population variability, significance level, and one-sided alternative \(H_a:\mu>\mu_0\). In Scenario A, the true population mean is only slightly greater than \(\mu_0\). In Scenario B, the true population mean is much greater than \(\mu_0\). Which scenario has greater power, and why?

Hints

- Hold the rejection rule and sampling variability fixed. - Compare where the sampling distribution is centered under the two true alternative means. - A center farther into the direction of the alternative makes rejection more likely.

Solution

1. Power is the probability of rejecting the null hypothesis when the alternative is true. 2. With sample size and variability fixed, sample means are more likely to fall far into the upper rejection region when the true mean is farther above \(\mu_0\). 3. Therefore, Scenario B has greater power because the true effect is larger relative to the sampling variability.

Answer

Scenario B has greater power. A true mean farther above the null value is easier for the test to distinguish from \(H_0\) when the other test conditions are unchanged.
54782512
A valid two-sided one-sample \(t\)-test has \(19\) degrees of freedom and observed test statistic \(t=-2.20\). At \(\alpha=0.05\), the critical values are \(-2.093\) and \(2.093\). Make the test decision using the critical-value approach and explain what it implies about the p-value.

Hints

- Locate the observed statistic relative to both critical boundaries. - A two-sided rejection region has one part in each tail. - Connect membership in the rejection region with the corresponding p-value comparison to \(\alpha\).

Solution

1. The two-sided rejection region consists of test statistics less than \(-2.093\) or greater than \(2.093\). 2. The observed statistic \(-2.20\) lies beyond the lower critical value \(-2.093\). 3. Therefore, reject \(H_0\) at \(\alpha=0.05\). 4. Because the critical-value and p-value approaches give the same decision, the two-sided p-value must be less than \(0.05\).

Answer

Reject \(H_0\). Since \(t=-2.20\) lies in the rejection region, the two-sided p-value is less than \(0.05\).
54783712
Quality engineer Nikhil Rao chooses a random starting day from a long period of stable production, then samples one randomly selected bottle from each of \(50\) consecutive days. Assume this \(50\)-day block is representative of the production period of interest. During the block, machine settings and ambient conditions often persist from one day to the next. Nikhil plans a one-sample \(t\)-test for the population mean fill amount. Explain why the sample size alone does not guarantee that the usual one-sample \(t\)-test is valid.

Hints

- The sampling setup is intended to remove selection bias as the main issue; focus on whether successive observations are independent. - Ask whether knowing one day's fill amount could provide information about the next day's fill amount. - Increasing the number of observations does not automatically repair a dependence structure.

Solution

1. A one-sample \(t\)-test relies on observations that are independent enough for the standard-error calculation and reference distribution to be appropriate. 2. Even though the block was selected to represent the production period and one bottle was randomly sampled each day, consecutive daily observations can still be associated because operating conditions persist over time. 3. A sample size of \(50\) can help with the shape of the sampling distribution, but it does not remove dependence created by serial persistence. 4. The dependence must be addressed, or a method appropriate for serially dependent data must be used, before applying the usual one-sample \(t\)-test.

Answer

The usual one-sample \(t\)-test is not justified merely because \(n=50\). Consecutive daily observations may be serially dependent, and a large sample does not repair dependence in the data-collection process.
54784912
Before collecting data, Dr. Jae-min Lee knows that only an increase above the benchmark mean would matter scientifically. Dr. Camille Laurent suggests using a two-sided one-sample \(t\)-test “to be safer.” Both tests would use \(\alpha=0.05\). If the true population mean is above the benchmark, compare the power of the pre-specified upper-tailed test with the power of the two-sided test. Explain why the direction must be chosen before seeing the data.

Hints

- Compare where each test places its allowed rejection probability. - Think about how far into the upper tail the statistic must fall under each design. - The alternative hypothesis is part of the study plan, not a choice made after inspecting results.

Solution

1. An upper-tailed test places the full \(0.05\) rejection probability in the upper tail, while a two-sided test splits the rejection probability between two tails. 2. When the true mean is above the benchmark, the upper-tailed test therefore has a less extreme upper rejection cutoff and greater power to detect that increase. 3. The direction must be specified before examining the data; choosing the tail after seeing the sample would change the error rate and invalidate the planned significance level.

Answer

For detecting a true increase, the pre-specified upper-tailed test has greater power than the two-sided test at the same \(\alpha=0.05\). The direction must be chosen in advance, not selected after observing the sample.
54786112
A random sample of \(36\) observations has mean \(78\) and sample standard deviation \(12\). Ethan Brooks tests \(H_0:\mu=75\) against \(H_a:\mu\ne75\) and obtains \(t=1.50\). Ethan reports \(p\approx0.071\) from the upper tail of the \(t\) distribution with \(35\) degrees of freedom. Correct the p-value to four decimal places and make the decision at \(\alpha=0.05\).

Hints

- Match the tail calculation to the direction stated in the alternative hypothesis. - A two-sided alternative counts extreme results on both sides of the null value. - Compare the corrected p-value with the significance level only after using the proper tails.

Solution

1. The alternative is two-sided, so outcomes at least as far from \(0\) in either direction count as evidence against \(H_0\). 2. The upper-tail probability for \(t=1.50\) with \(35\) degrees of freedom is approximately \(0.0713\). 3. By symmetry, the two-sided p-value is \(2(0.0713)\approx0.1426\). 4. Since \(0.1426>0.05\), fail to reject \(H_0\).

Answer

The correct two-sided p-value is approximately \(0.1426\). Fail to reject \(H_0\) at \(\alpha=0.05\).
54787312
Dr. Gabriela Souza wants to test a claim about a population mean. The response variable is annual household income, which is strongly right-skewed, and Dr. Souza has a random sample of \(120\) households. Marcus Lee argues that a test for a median should replace the one-sample \(t\)-test because the raw data are not normal. Explain why changing to a procedure for the median would answer a different question, and assess whether the large sample can support inference about the mean.

Hints

- Start by identifying the population parameter named in the research question. - A different statistical procedure may target a different parameter even if it seems more robust. - For a large sample, distinguish the shape of the raw population from the shape of the sampling distribution of the mean.

Solution

1. The research question concerns the population mean, so replacing the procedure with one that tests a median changes the parameter being studied. 2. The one-sample \(t\)-procedure is designed for inference about the population mean. 3. With a random sample of \(120\), the sampling distribution of the sample mean is typically approximately normal by the central limit theorem unless the population has exceptionally severe features. 4. Thus, strong skewness in the raw incomes does not by itself justify switching to a test of a different parameter.

Answer

A test for a median would answer a different question. With a random sample of \(120\), mean-based \(t\) inference can be reasonable despite a strongly skewed population, provided there are no extraordinary features that defeat the large-sample approximation.
54787912
A one-sample \(t\)-test gives \(p=0.20\). Lena Kovács says, “That means there is an \(80\%\) chance the null hypothesis is false, so the Type II error probability is \(0.80\).” Explain both errors in Lena's reasoning.

Hints

- Ask what assumption is made when a p-value is computed. - Distinguish a probability about data from a probability about a hypothesis. - Type II error is a property of a test under a particular alternative, not the complement of an observed p-value.

Solution

1. A p-value is calculated under the assumption that the null hypothesis is true; it is not the probability that the null hypothesis itself is true or false. 2. Therefore, \(p=0.20\) does not imply an \(80\%\) probability that \(H_0\) is false. 3. A Type II error probability depends on a specific alternative population mean, the sample size, variability, significance level, and test direction. 4. It cannot be obtained as \(1-p\) from the observed test result.

Answer

The p-value is not a probability assigned to \(H_0\), and \(1-p\) is not the Type II error probability. The Type II error rate must be evaluated for a specified alternative and test design.
54788512
Two independent studies use the same valid one-sample \(t\)-test at \(\alpha=0.05\). Study 1 reports \(p=0.049\), and Study 2 reports \(p=0.051\). Luca Bianchi says Study 1 found strong evidence while Study 2 found essentially no evidence because only the first result is “significant.” Evaluate Luca's interpretation.

Hints

- Separate a binary decision rule from the continuous information in a p-value. - Compare how far apart the two p-values actually are. - A significance threshold is a decision convention, not a sudden boundary between evidence and no evidence.

Solution

1. Under the stated decision rule, Study 1 rejects \(H_0\) because \(0.049<0.05\), while Study 2 fails to reject because \(0.051>0.05\). 2. The two p-values are nevertheless extremely close and represent nearly the same strength of evidence against their respective null hypotheses, assuming comparable designs. 3. Crossing the \(0.05\) threshold changes the formal decision but does not create a large discontinuity in the underlying evidence.

Answer

The formal decisions differ at \(\alpha=0.05\), but it is misleading to describe the evidence as strong in one study and absent in the other. \(p=0.049\) and \(p=0.051\) provide very similar evidence.
54789112
A valid one-sample \(t\)-test reports \(p=0.03\). Dr. Fatima Zahra says, “That means there is a \(97\%\) chance that an identical replication will also reject the null hypothesis.” Explain why the p-value does not support that statement.

Hints

- Recall what probability a p-value actually describes. - A future study will have a new random sample and a new test statistic. - The probability of rejecting under a specified alternative is a design property of the test, not the complement of one observed p-value.

Solution

1. The p-value describes how unusual the observed test statistic would be under the null hypothesis; it is not a probability of replication success. 2. Whether a future study rejects depends on the true population mean, population variability, sample size, significance level, and random sampling variation. 3. The long-run probability of rejection under a particular alternative is the test's power, which cannot be found as \(1-p\) from one observed study.

Answer

The statement is incorrect. \(1-0.03=0.97\) is not the probability that a replication will reject. Replication probability depends on the test's power under the true population conditions.
54790312
A random sample of \(100\) observations is drawn from a population of more than \(1000\) observations. The sample ranges from \(30\) to \(70\), with sample mean \(55\) and sample standard deviation \(10\). Arjun Mehta says \(H_0:\mu=50\) cannot be rejected because \(50\) lies inside the range of observed values. Test \(H_0:\mu=50\) against \(H_a:\mu\ne50\) at \(\alpha=0.05\), giving the p-value to seven decimal places, and explain why the sample range is not the relevant criterion.

Hints

- A hypothesis about a mean is evaluated using the sampling variability of the sample mean. - The range describes individual observations, not uncertainty in the population mean estimate. - Compare the sample mean with the null mean in standard-error units.

Solution

1. The test concerns the population mean, so the relevant statistic is the sample mean and its standard error, not whether the null value falls between the sample minimum and maximum. 2. The standard error is \(10/\sqrt{100}=1\). 3. The test statistic is \(t=(55-50)/1=5.00\) with \(99\) degrees of freedom. 4. The two-sided p-value is approximately \(0.0000025\). 5. Since the p-value is far below \(0.05\), reject \(H_0\).

Answer

\(t=5.00\) and \(p\approx0.0000025\), so reject \(H_0\). The fact that \(50\) lies within the range of individual observations does not determine whether the population mean equals \(50\).
54790912
A one-sample test is designed to compare a population mean temperature with \(20\,^\circ\text{C}\). The sample data are converted to Fahrenheit before analysis, but Lina Chen leaves the null value as \(20\) and tests \(H_0:\mu=20\,^\circ\text{F}\). Explain the error and state the correct null value on the Fahrenheit scale.

Hints

- A test statistic compares quantities measured on the same scale. - Convert the benchmark using the same transformation applied to the observations. - Leaving the numerical null value unchanged after a unit conversion changes the hypothesis itself.

Solution

1. Converting the sample data changes the measurement scale, so the hypothesized mean must be converted to the same scale. 2. The Fahrenheit equivalent of \(20\,^\circ\text{C}\) is \(20\cdot\frac{9}{5}+32=68\,^\circ\text{F}\). 3. Testing against \(20\,^\circ\text{F}\) would test a completely different population mean. 4. The correct Fahrenheit-scale null hypothesis is \(H_0:\mu=68\,^\circ\text{F}\).

Answer

The null value must be converted with the data. The correct null hypothesis is \(H_0:\mu=68\,^\circ\text{F}\), not \(20\,^\circ\text{F}\).
54792712
A simple random sample of \(150\) adults is selected for a one-sample \(t\)-test about mean weekly exercise time. Only \(70\) people respond, and people who exercise very little are believed to be less likely to respond. Explain why the large original sample does not eliminate the inferential problem created by nonresponse.

Hints

- Distinguish the people selected from the people whose responses are analyzed. - Ask whether the chance of responding may depend on the variable being studied. - Larger samples reduce random error, not systematic selection bias.

Solution

1. The original sample was random, but the analyzed data come only from the respondents. 2. If response probability is related to exercise time, respondents can systematically differ from nonrespondents. 3. The resulting sample mean can therefore be biased for the target population mean. 4. Increasing the original sample size does not remove a systematic nonresponse mechanism, so a standard one-sample \(t\)-test on respondents alone may not support population inference.

Answer

Nonresponse can bias the analyzed sample because response is related to the outcome. A large original random sample does not repair systematic missingness among the people who actually provide data.
55623012
Two studies are described below. Study I: A random sample of \(16\) batteries is used to test whether the population mean lifetime differs from the manufacturer's claim of \(10\,\text{h}\). Study II: A random sample of \(16\) laptops is tested before and after a software update, and the goal is to test whether the population mean change in battery life differs from \(0\). For each study, identify the population parameter and the appropriate \(t\)-test structure. Explain the design feature that makes the two studies require different analyses.

Hints

- Count how many measurements each sampled unit contributes in each study. - A repeated measurement on the same unit creates a natural difference. - Match the test to the population parameter actually named by the research question.

Solution

1. Study I has one quantitative measurement per sampled battery, so the parameter is the population mean lifetime \(\mu\). The appropriate structure is a one-sample \(t\)-test of \(H_0:\mu=10\). 2. Study II measures the same laptop twice, so each laptop contributes one paired difference \(d=\text{after}-\text{before}\). The parameter is the population mean paired difference \(\mu_d\), and the appropriate structure is a one-sample \(t\)-test on the differences, such as \(H_0:\mu_d=0\). 3. The key distinction is that Study II contains matched repeated measurements, while Study I contains one measurement per independent sampled unit.

Answer

Study I: parameter \(\mu\), the population mean battery lifetime; use a one-sample \(t\)-test against \(10\,\text{h}\). Study II: parameter \(\mu_d\), the population mean within-laptop change; use a paired \(t\)-test, equivalently a one-sample \(t\)-test on the paired differences.
55623112
A random sample of \(20\) runners is timed on the same course before and after a training program. The report gives the before mean and standard deviation, the after mean and standard deviation, and \(n=20\), but it does not give the standard deviation of the \(20\) within-runner differences. The goal is to test \(H_0:\mu_d=0\), where \(d=\text{after}-\text{before}\). Can the paired \(t\)-test statistic be determined from the reported summaries alone? If not, identify the missing quantity and explain why the two separate standard deviations are insufficient.

Hints

- Write the standard error for a paired mean difference. - Ask which sample standard deviation appears in that formula. - Two measurements on the same runner can move together, so their relationship matters for the variability of the difference.

Solution

1. A paired \(t\)-test is based on the sample of within-runner differences, not on the two sets of times treated separately. 2. Its standard error is \(s_d/\sqrt{20}\), where \(s_d\) is the sample standard deviation of the paired differences. 3. The separate before and after standard deviations do not determine \(s_d\) because the variability of the differences also depends on how each runner's two times are related. 4. Therefore, the test statistic cannot be determined without \(s_d\), the raw paired data, or equivalent information such as the within-pair correlation.

Answer

No. The missing quantity is \(s_d\), the sample standard deviation of the \(20\) paired differences. Separate before and after standard deviations do not determine the variability of within-runner changes.
54763912
A ceramics studio states that the population mean firing time for a certain glaze is \(74\,\text{min}\). A random sample of \(25\) firing cycles has mean \(76.3\,\text{min}\) and standard deviation \(5.0\,\text{min}\). The sample distribution is roughly symmetric with no outliers. At the \(\alpha=0.05\) significance level, test whether the population mean firing time is greater than \(74\,\text{min}\). State the hypotheses, calculate the test statistic to two decimal places and the p-value to four decimal places, and give a conclusion in context.

Hints

- Identify the population quantity named in the studio's claim and the direction of the question. - Compare the observed sample mean with the claimed value in units of estimated sampling variability. - The conclusion should reflect both the p-value and the direction stated in the alternative hypothesis.

Solution

1. Test \(H_0:\mu=74\) against \(H_a:\mu>74\), where \(\mu\) is the population mean firing time. 2. The random sample and sample-shape conditions support a one-sample \(t\)-test with \(df=24\). 3. The test statistic is \(t=\frac{76.3-74}{5.0/\sqrt{25}}=2.30\). 4. The one-sided p-value is \(P(T_{24}\ge2.30)\approx0.0152\). 5. Because \(0.0152<0.05\), reject \(H_0\). The sample provides convincing evidence that the population mean firing time is greater than \(74\,\text{min}\).

Answer

\(H_0:\mu=74\), \(H_a:\mu>74\); \(t=2.30\) with \(df=24\); p-value \(\approx0.0152\). Reject \(H_0\). There is convincing evidence that the population mean firing time exceeds \(74\,\text{min}\).
54764512
An audio lab randomly selects \(18\) studio monitors from a large inventory and measures each monitor before and after a firmware update. For each monitor, the response is the time needed to enter standby mode. Mei Tanaka wants to know whether the population mean standby time is lower after the update than before it. Define an appropriate parameter, write the null and alternative hypotheses, identify the appropriate \(t\)-test, and explain why treating the before and after measurements as two independent samples would be inappropriate.

Hints

- Decide what one numerical comparison can be made within each monitor. - The direction of the alternative must match what “lower after” means for the difference you define. - Think about whether observations coming from the same physical unit can reasonably be treated as independent.

Solution

1. Let \(d=\text{after time}-\text{before time}\), and let \(\mu_d\) be the population mean paired difference in standby time for monitors represented by the random sample. 2. The hypotheses are \(H_0:\mu_d=0\) and \(H_a:\mu_d<0\). 3. The appropriate procedure is a one-sample \(t\)-test for the population mean difference using the \(18\) paired differences. 4. The before and after values from the same monitor are dependent, so an independent two-sample procedure would ignore the pairing built into the design. 5. This test can assess whether mean standby time is lower after the update than before it; the before/after design alone does not establish that the firmware update caused any change.

Answer

Let \(d=\text{after}-\text{before}\). Test \(H_0:\mu_d=0\) versus \(H_a:\mu_d<0\) with a one-sample \(t\)-test on the paired differences. A two-sample test is inappropriate because each after measurement is linked to the before measurement from the same monitor. The comparison does not by itself establish causation.
54766312
Ten performance halls are randomly selected from a regional network. Each hall's echo-decay time is measured before and after an acoustic treatment. Let \(d=\text{after}-\text{before}\), in seconds. The observed differences are \(-0.12,-0.05,-0.18,-0.09,-0.14,-0.02,-0.11,-0.08,-0.16,-0.07\), and the differences show no strong skewness or outliers. At \(\alpha=0.01\), test whether the population mean echo-decay time is lower after the treatment than before it. Give the test statistic to two decimal places and the p-value to six decimal places.

Hints

- Turn each before-and-after pair into one change with a consistent subtraction order. - Check the direction of the alternative against what “lower after” means for that change. - Separate evidence of a before/after mean difference from evidence that the treatment caused that difference.

Solution

1. Test \(H_0:\mu_d=0\) versus \(H_a:\mu_d<0\), where \(d=\text{after}-\text{before}\). 2. The sample has \(\bar{d}=-0.102\,\text{s}\) and \(s_d\approx0.0498\,\text{s}\). The random sample and difference distribution support a one-sample paired-difference \(t\)-test with \(df=9\). 3. The test statistic is \(t=\frac{-0.102}{0.0498/\sqrt{10}}\approx-6.47\). 4. The one-sided p-value is approximately \(0.000058\). 5. Since the p-value is less than \(0.01\), reject \(H_0\). There is convincing evidence that the population mean echo-decay time is lower after the acoustic treatment than before it. 6. Because every sampled hall was observed before and after treatment without a randomized untreated comparison, this result does not by itself establish that the treatment caused the decrease.

Answer

\(H_0:\mu_d=0\), \(H_a:\mu_d<0\); \(t\approx-6.47\), p-value \(\approx0.000058\). Reject \(H_0\). The data provide convincing evidence that population mean echo-decay time is lower after treatment than before treatment, but this before/after design alone does not establish causation.
54766912
A concert venue claims that the population mean time required for a sound check is \(35\,\text{min}\). A random sample of \(40\) sound checks has mean \(33.2\,\text{min}\) and standard deviation \(6.4\,\text{min}\). At \(\alpha=0.05\), test whether the population mean sound-check time is less than \(35\,\text{min}\). Include the hypotheses, test statistic to two decimal places, p-value to four decimal places, and contextual conclusion.

Hints

- The alternative should reflect the direction “less than.” - Measure how far the sample mean falls below the claimed mean relative to estimated sampling variability. - Match the p-value tail to the direction of the research question.

Solution

1. Test \(H_0:\mu=35\) against \(H_a:\mu<35\). 2. The random sample and sample size \(n=40\) support a one-sample \(t\)-test with \(df=39\). 3. The test statistic is \(t=\frac{33.2-35}{6.4/\sqrt{40}}\approx-1.78\). 4. The lower-tail p-value is approximately \(0.0415\). 5. Because \(0.0415<0.05\), reject \(H_0\). The sample provides convincing evidence that the population mean sound-check time is less than \(35\,\text{min}\).

Answer

\(H_0:\mu=35\), \(H_a:\mu<35\); \(t\approx-1.78\), p-value \(\approx0.0415\). Reject \(H_0\). There is convincing evidence that the population mean sound-check time is less than \(35\,\text{min}\).
54768112
A before-and-after study records a quantitative measurement for the same \(22\) devices under two settings. The distribution of the “before” measurements is strongly right-skewed, and the distribution of the “after” measurements is also strongly right-skewed. However, the \(22\) paired differences are roughly symmetric with no outliers. Dr. Aisha Rahman wants to use a one-sample \(t\)-test for the population mean paired difference. Which distribution's shape is relevant for checking the sample-data condition, and does the stated shape information support the test?

Hints

- Identify the actual observations that enter the one-sample analysis after pairing is taken into account. - The shapes of the original two measurement sets can differ from the shape of their within-device differences. - Check conditions on the data object the test statistic is actually built from.

Solution

1. A matched-pairs analysis turns each device's two measurements into one difference. 2. The one-sample \(t\)-test is applied to the sample of paired differences, so the shape condition concerns the distribution of those differences, not the two original measurement distributions separately. 3. The \(22\) differences are roughly symmetric with no outliers, so the stated shape information supports the sample-data condition for the paired \(t\)-test. 4. Other conditions must still be checked, including whether the paired differences can be treated as independent and whether the sampling design supports inference to the intended population.

Answer

The relevant shape is the distribution of the \(22\) paired differences. Because those differences are roughly symmetric with no outliers, the stated shape information supports use of the paired one-sample \(t\)-test, assuming the independence and sampling conditions also hold.
54771112
A museum randomly selects \(12\) days and records the average relative humidity in a storage room on each selected day. The percentages are \(46.2,45.9,47.1,44.8,46.5,45.7,46.8,47.4,45.1,46.0,46.9,45.5\). The sample shows no strong skewness or outliers. At \(\alpha=0.01\), test whether the population mean daily relative humidity is greater than \(45\%\). Give the test statistic to two decimal places and the p-value to six decimal places.

Hints

- Summarize the raw sample before comparing it with the hypothesized population mean. - The small sample makes the stated shape information important for using the reference distribution. - The alternative hypothesis tells you which tail of the reference distribution matters.

Solution

1. Test \(H_0:\mu=45\%\) versus \(H_a:\mu>45\%\). 2. The sample has \(\bar{x}\approx46.158\%\) and \(s\approx0.810\%\). The random sample and shape support a one-sample \(t\)-test with \(df=11\). 3. The standard error is approximately \(0.810/\sqrt{12}=0.234\%\), giving \(t\approx4.96\). 4. The one-sided p-value is approximately \(0.000216\). 5. Since the p-value is below \(0.01\), reject \(H_0\). There is convincing evidence that the population mean daily relative humidity exceeds \(45\%\).

Answer

\(H_0:\mu=45\%\), \(H_a:\mu>45\%\); \(t\approx4.96\), p-value \(\approx0.000216\). Reject \(H_0\). There is convincing evidence that the population mean daily relative humidity is greater than \(45\%\).
54773512
A random sample of \(36\) automated warehouse runs has mean throughput \(72\) packages per minute and sample standard deviation \(9\) packages per minute. The population of runs is much larger than the sample. Using the same data, compare these two one-sided tests at \(\alpha=0.05\): a) \(H_0:\mu=70\) versus \(H_a:\mu>70\) b) \(H_0:\mu=71.5\) versus \(H_a:\mu>71.5\) Calculate each test statistic to two decimal places and each p-value to three decimal places, then explain why the p-values differ even though the sample data are identical.

Hints

- The sample size and sample standard deviation stay fixed, so the estimated sampling variability is the same in both tests. - Focus on how far the sample mean lies above each proposed null value. - A larger standardized distance in the direction of the alternative produces stronger evidence against the null hypothesis.

Solution

1. The standard error for both tests is \(\frac{9}{\sqrt{36}}=1.5\) packages per minute, with \(df=35\). 2. For part a, \(t=\frac{72-70}{1.5}\approx1.33\), giving an upper-tail p-value of approximately \(0.096\). Fail to reject \(H_0\) at \(\alpha=0.05\). 3. For part b, \(t=\frac{72-71.5}{1.5}\approx0.33\), giving an upper-tail p-value of approximately \(0.370\). Fail to reject \(H_0\) at \(\alpha=0.05\). 4. The second null value is closer to the observed sample mean, so its test statistic is closer to \(0\) and its p-value is larger.

Answer

a) \(t\approx1.33\), p-value \(\approx0.096\); fail to reject \(H_0\). b) \(t\approx0.33\), p-value \(\approx0.370\); fail to reject \(H_0\). Both tests fail to reject \(H_0\), but the p-values differ because changing the null value changes how far the observed sample mean lies from the value assumed by \(H_0\).
54774112
Before collecting data, a food laboratory plans a two-sided test of \(H_0:\mu=400\,\text{mg}\) versus \(H_a:\mu\ne400\,\text{mg}\) for the population mean sodium content of a product. The resulting one-sample \(t\)-statistic is \(t=-1.80\) with \(39\) degrees of freedom. The two-sided p-value is approximately \(0.0796\). After seeing that the sample mean is below \(400\,\text{mg}\), analyst Sofia Vega changes the alternative to \(H_a:\mu<400\,\text{mg}\), reports the one-sided p-value \(0.0398\), and rejects \(H_0\) at \(\alpha=0.05\). Evaluate this reasoning and give the appropriate conclusion for the planned study.

Hints

- Consider when the direction of an alternative hypothesis should be decided. - Use the p-value that matches the question the study was designed to test. - Compare that p-value with the stated significance level before writing the conclusion.

Solution

1. The direction of a one-sided alternative should be chosen from the research question before examining the sample result, not selected afterward because the observed statistic points in that direction. 2. The planned test was two-sided, so the relevant p-value is \(0.0796\). 3. Because \(0.0796>0.05\), fail to reject \(H_0\). 4. The data do not provide convincing evidence at the \(0.05\) level that the population mean sodium content differs from \(400\,\text{mg}\).

Answer

Sofia's switch to a one-sided alternative after seeing the data is not valid for the planned analysis. Using the planned two-sided p-value \(0.0796\), fail to reject \(H_0\); there is not convincing evidence that \(\mu\ne400\,\text{mg}\).
54775312
Two valid two-sided one-sample \(t\)-tests produce the same test statistic, \(t=2.10\), but Test A has \(5\) degrees of freedom and Test B has \(50\) degrees of freedom. Find each p-value to four decimal places and compare the decisions at \(\alpha=0.05\). Explain why the p-values are different even though the test statistics are equal.

Hints

- A \(t\)-statistic alone does not determine a p-value unless the degrees of freedom are also known. - Compare the tail behavior of a \(t\)-distribution with few degrees of freedom to one with many degrees of freedom. - Use the same two-sided tail rule for both tests before comparing with \(0.05\).

Solution

1. For Test A, the two-sided p-value is \(2P(T_5\ge2.10)\approx0.0898\). 2. For Test B, the two-sided p-value is \(2P(T_{50}\ge2.10)\approx0.0408\). 3. Test A fails to reject \(H_0\) at \(\alpha=0.05\), while Test B rejects \(H_0\). 4. A \(t\)-distribution with fewer degrees of freedom has heavier tails, so the same standardized statistic corresponds to a larger tail probability.

Answer

Test A: p-value \(\approx0.0898\), so fail to reject \(H_0\). Test B: p-value \(\approx0.0408\), so reject \(H_0\). The p-values differ because the two tests use different \(t\)-distributions.
54776512
A random sample of \(25\) observations has mean \(18.4\). A valid two-sided one-sample \(t\)-test of \(H_0:\mu=20\) produces test statistic \(t=-2.00\). Determine the sample standard deviation used in the test, then find the two-sided p-value to four decimal places.

Hints

- Work backward from the definition of the one-sample \(t\)-statistic. - The sample size determines the denominator's square-root factor and the degrees of freedom. - For a two-sided test, use both tails beyond the magnitude of the observed statistic.

Solution

1. The test statistic satisfies \(-2.00=\frac{18.4-20}{s/\sqrt{25}}\). 2. Since \(\sqrt{25}=5\), solving \(-2.00=\frac{-1.6}{s/5}\) gives \(s=4.0\). 3. The test has \(df=24\). 4. The two-sided p-value is \(2P(T_{24}\ge2.00)\approx0.0569\).

Answer

The sample standard deviation is \(4.0\), and the two-sided p-value is approximately \(0.0569\).
54777112
A performance lab randomly selects \(20\) devices from a large population and measures each device before and after a software optimization. For each device, let \(d=\text{after score}-\text{before score}\). The sample of differences has mean \(\bar{d}=3.1\) points and standard deviation \(s_d=2.4\) points, with no strong skewness or outliers. At \(\alpha=0.05\), test whether the population mean after-before score difference is greater than \(2\) points. Give the test statistic to two decimal places and the p-value to four decimal places.

Hints

- Reduce the matched measurements to one difference per device. - The null value is not zero, so compare the sample mean difference with the specific value named in the claim. - Use the direction of the claim to choose the appropriate tail for the p-value.

Solution

1. Test \(H_0:\mu_d=2\) against \(H_a:\mu_d>2\). 2. Because the data are paired, use a one-sample \(t\)-test on the \(20\) differences, with \(df=19\). The random sample and stated shape support population inference for the mean paired difference. 3. The test statistic is \(t=\frac{3.1-2}{2.4/\sqrt{20}}\approx2.05\). 4. The upper-tail p-value is approximately \(0.0272\). 5. Because \(0.0272<0.05\), reject \(H_0\). There is convincing evidence that the population mean after-before score difference exceeds \(2\) points.

Answer

\(t\approx2.05\) with \(df=19\) and p-value \(\approx0.0272\). Reject \(H_0\); there is convincing evidence that \(\mu_d>2\) points.
54783112
A manufacturer claims that the mean compression strength of a material is \(100\,\text{MPa}\). A random sample of \(10\) specimens from a large production population has mean \(104\,\text{MPa}\) and sample standard deviation \(5\,\text{MPa}\). The boxplot shows the sample is roughly symmetric with no apparent outliers. Test \(H_0:\mu=100\) against \(H_a:\mu>100\) at \(\alpha=0.05\). Give the test statistic to three decimal places and the p-value to four decimal places, and include the role of the boxplot in deciding whether the procedure is reasonable.
Figure for problem 547831

Hints

- With a small sample, inspect the sample shape before relying on a \(t\)-procedure. - Compare the observed sample mean with the null value in standard-error units. - The direction of the alternative determines which tail contributes to the p-value.

Solution

1. Because the sample is small, the roughly symmetric shape with no apparent outliers supports using a one-sample \(t\)-procedure. The random sample from a large production population supplies the sampling basis. 2. The test statistic is \(t=\frac{104-100}{5/\sqrt{10}}\approx2.530\) with \(9\) degrees of freedom. 3. The upper-tail p-value is approximately \(0.0161\). 4. Since \(0.0161<0.05\), reject \(H_0\). There is statistically significant evidence that the population mean compression strength exceeds \(100\,\text{MPa}\).

Answer

\(t\approx2.530\), \(df=9\), and \(p\approx0.0161\). Reject \(H_0\); the data provide significant evidence that the population mean compression strength is greater than \(100\,\text{MPa}\). The boxplot supports the small-sample \(t\)-procedure because it shows no strong skewness or outliers.
54784312
A school district with more than \(800\) students studies the number of absences per student during a school year. In a random sample of \(80\) students, the mean is \(12.6\) absences and the sample standard deviation is \(8.0\) absences. The district wants to test whether the population mean exceeds \(10\) absences. Explain why a one-sample \(t\)-procedure can be appropriate even though the response is a count, then carry out the test at \(\alpha=0.05\). Give the test statistic to three decimal places and the p-value to four decimal places.

Hints

- Decide whether the response is quantitative before choosing an inference procedure. - For a large random sample, focus on the sampling distribution of the mean rather than requiring the raw data themselves to be normal. - Use the direction of the claim to choose the appropriate tail.

Solution

1. The response is quantitative, the random sample is less than \(10\%\) of the population, and the large sample makes the sampling distribution of the sample mean approximately normal unless the population is extraordinarily skewed. 2. The hypotheses are \(H_0:\mu=10\) and \(H_a:\mu>10\). 3. The test statistic is \(t=\frac{12.6-10}{8/\sqrt{80}}\approx2.907\) with \(79\) degrees of freedom. 4. The upper-tail p-value is approximately \(0.0024\). 5. Since \(0.0024<0.05\), reject \(H_0\). There is significant evidence that the population mean number of absences exceeds \(10\).

Answer

A one-sample \(t\)-procedure is reasonable because the variable is quantitative and the random sample is large relative to the shape issue while remaining a small fraction of the population. \(t\approx2.907\), \(df=79\), and \(p\approx0.0024\), so reject \(H_0\).
54785512
A random sample of \(10{,}000\) manufactured parts is taken from a production population much larger than \(100{,}000\) parts. The sample has mean length \(50.2\,\text{mm}\) and sample standard deviation \(5.0\,\text{mm}\). Test \(H_0:\mu=50\,\text{mm}\) against \(H_a:\mu\ne50\,\text{mm}\) at \(\alpha=0.05\). Give the test statistic to two decimal places and the p-value to six decimal places. After making the statistical decision, explain why statistical significance alone does not establish that the difference is practically important.

Hints

- A very large sample can make the standard error quite small. - Separate evidence that a difference exists from the size of the observed difference. - Practical importance requires a contextual benchmark beyond statistical significance.

Solution

1. The standard error is \(5/\sqrt{10{,}000}=0.05\,\text{mm}\). 2. The test statistic is \(t=\frac{50.2-50}{0.05}=4.00\) with \(9999\) degrees of freedom. 3. The two-sided p-value is approximately \(0.000064\), so reject \(H_0\) at \(\alpha=0.05\). 4. The observed difference from the benchmark is only \(0.2\,\text{mm}\). Whether that size matters in manufacturing depends on tolerances, costs, and consequences, which the p-value does not measure.

Answer

\(t=4.00\) and \(p\approx0.000064\), so reject \(H_0\). The result is statistically significant, but the practical importance of a \(0.2\,\text{mm}\) difference must be judged using engineering context rather than the p-value alone.
54786712
A random sample of \(12\) observations has mean \(105\) and sample standard deviation \(9\). The sample distribution is approximately symmetric with no outliers. To test \(H_0:\mu=100\) against \(H_a:\mu>100\), Nadia El-Hassan uses the standard normal distribution because the test statistic is about \(1.92\). Compare Nadia's normal-tail p-value with the correct one-sample \(t\) p-value, giving both p-values to four decimal places, and explain why the \(t\) distribution is appropriate.

Hints

- For a small sample, check the sample shape as well as the random-sampling condition. - Determine whether the population standard deviation is known or estimated from the sample. - Compare the tail behavior of the normal and \(t\) reference distributions.

Solution

1. The random sample and stated shape support a small-sample one-sample \(t\)-procedure. The test statistic is \(t=\frac{105-100}{9/\sqrt{12}}\approx1.925\). 2. A standard normal upper-tail probability at \(1.925\) is approximately \(0.0271\). 3. Because the population standard deviation is unknown and estimated by \(s=9\), the correct reference distribution is \(t\) with \(11\) degrees of freedom. 4. The corresponding upper-tail \(t\) p-value is approximately \(0.0403\), which is larger because the \(t\) distribution has heavier tails.

Answer

Nadia's normal-tail value is approximately \(0.0271\), but the correct one-sample \(t\) p-value is approximately \(0.0403\) with \(11\) degrees of freedom.
54789712
A large population is known from long-term process data to be approximately normal. A random sample of only \(6\) observations has mean \(23\) and sample standard deviation \(2.5\). Test \(H_0:\mu=20\) against \(H_a:\mu>20\) at \(\alpha=0.05\). Give the test statistic to three decimal places and the p-value to four decimal places. Explain why the small sample size does not automatically prevent use of the one-sample \(t\)-procedure here.

Hints

- For a small sample, look for information about the population shape. - Use the sample standard deviation to estimate the standard error of the mean. - The direction of the alternative determines the relevant tail.

Solution

1. Because the population is stated to be approximately normal and large relative to the sample, a one-sample \(t\)-procedure can be used even with a very small random sample. 2. The test statistic is \(t=\frac{23-20}{2.5/\sqrt{6}}\approx2.939\) with \(5\) degrees of freedom. 3. The upper-tail p-value is approximately \(0.0161\). 4. Since \(0.0161<0.05\), reject \(H_0\). There is significant evidence that the population mean exceeds \(20\).

Answer

\(t\approx2.939\), \(df=5\), and \(p\approx0.0161\). Reject \(H_0\). The small sample is acceptable because the population is stated to be approximately normal.
54793312
Dr. Owen Price tests \(H_0:\mu=50\) at \(\alpha=0.05\). Dr. Price begins with \(20\) observations and, whenever the result is not significant, collects \(10\) more observations and repeats the ordinary one-sample \(t\)-test. Sampling stops as soon as a test gives \(p<0.05\). Explain why treating the final test as an ordinary fixed-sample \(5\%\) test is problematic.

Hints

- Count how many opportunities the analysis may have to declare significance. - The usual significance level assumes a fixed decision rule, not repeated attempts until the threshold is crossed. - Consider what happens under the null hypothesis when many chances to reject are allowed.

Solution

1. An ordinary \(\alpha=0.05\) test controls the Type I error rate for a pre-specified analysis of one sample size. 2. Repeatedly testing after additional data creates multiple opportunities to cross the \(0.05\) threshold even when \(H_0\) is true. 3. Stopping specifically when significance appears makes the overall probability of at least one false rejection larger than the nominal \(5\%\) level. 4. A valid sequential design requires a procedure that accounts for the repeated looks at the data rather than applying an unchanged fixed-sample rule each time.

Answer

The final ordinary \(t\)-test does not retain a \(5\%\) overall Type I error rate under this optional-stopping rule. Repeated looks at the data increase the chance of a false rejection unless the sequential procedure is adjusted appropriately.
54772912
A marine laboratory is evaluating a new sensor that is intended to have a population mean operating time greater than \(50\) hours before recalibration. The lab plans to test \(H_0:\mu=50\) against \(H_a:\mu>50\) using a random sample of \(16\) sensors. Assume the sample distribution is roughly symmetric with no outliers, and suppose the sample standard deviation is \(3.2\) hours. At significance level \(\alpha=0.05\), how large must the sample mean be for the one-sample \(t\)-test to reject \(H_0\)? Give the cutoff sample mean to two decimal places.

Hints

- Think about the test statistic value that marks the start of the rejection region for this one-sided test. - Express that boundary in terms of the unknown sample mean rather than starting with a p-value. - The sample standard deviation and sample size determine how far the sample mean must be from the null value.

Solution

1. With \(n=16\), the test uses \(df=15\). For an upper-tailed test at \(\alpha=0.05\), the critical value is \(t\approx1.753\). 2. The rejection boundary satisfies \(\frac{\bar{x}-50}{3.2/\sqrt{16}}=1.753\). 3. Since \(\frac{3.2}{\sqrt{16}}=0.8\), the boundary is \(\bar{x}=50+1.753\cdot0.8\approx51.402\) hours. 4. Therefore, the test rejects \(H_0\) when \(\bar{x}\) is greater than approximately \(51.40\) hours.

Answer

The sample mean must be greater than approximately \(51.40\) hours for the test to reject \(H_0\) at \(\alpha=0.05\).
54792112
A company has \(300\) employees. A simple random sample of \(40\) employees selected without replacement has mean commute time \(32.3\) minutes and sample standard deviation \(9.0\) minutes. The sample has no extreme outliers. Test \(H_0:\mu=30\) against \(H_a:\mu>30\) at \(\alpha=0.05\). First use the usual standard error \(s/\sqrt{n}\). Then use the finite-population-adjusted standard error \(\sqrt{\frac{N-n}{N-1}}\frac{s}{\sqrt{n}}\). Use a \(t\) distribution with \(39\) degrees of freedom for both calculations. Give each standard error and test statistic to three decimal places and each p-value to four decimal places, then compare the decisions.

Hints

- Compute the usual standard error and decision before applying the finite-population factor. - Sampling without replacement from a substantial fraction of the population reduces the standard error. - Recompute the test statistic and p-value using the adjusted standard error, then compare the two decisions.

Solution

1. The usual unadjusted standard error is \(9/\sqrt{40}\approx1.423\) minutes. 2. The unadjusted statistic is \(t=\frac{32.3-30}{1.423}\approx1.616\), giving an upper-tail p-value of approximately \(0.0570\). This calculation fails to reject \(H_0\) at \(\alpha=0.05\). 3. The finite-population factor is \(\sqrt{\frac{300-40}{300-1}}\approx0.933\). 4. The adjusted standard error is \(0.933(1.423)\approx1.327\) minutes. 5. The adjusted statistic is \(t=\frac{32.3-30}{1.327}\approx1.733\), giving an upper-tail p-value of approximately \(0.0455\). 6. Using the finite-population adjustment, reject \(H_0\). There is significant evidence at the \(5\%\) level that the company's mean commute time exceeds \(30\) minutes. 7. The adjustment matters because the sample contains \(40/300\approx13.3\%\) of the finite population, so sampling without replacement reduces the sampling variability.

Answer

Unadjusted: \(SE\approx1.423\), \(t\approx1.616\), and \(p\approx0.0570\), so fail to reject \(H_0\). Adjusted: \(SE\approx1.327\), \(t\approx1.733\), and \(p\approx0.0455\), so reject \(H_0\) at \(\alpha=0.05\).

All problems may be used, copied and printed free of charge for school and tutoring, including paid tutoring. Commercial adaptations as well as publication or redistribution on the internet are not permitted.