Aimathic
Login | English | Deutsch

Free math worksheets

Build your own math worksheets from 28,000 problems for grades 3 to 12, from fractions to calculus. Every problem includes step-by-step solutions.

Hypothesis test for a difference in means

Click problems to add them to your worksheet.

54767812
Two independent groups of students use different note-taking systems for the same type of course. Let \(\mu_D\) be the population mean exam score for students using the digital system and \(\mu_P\) the population mean exam score for students using the paper system. Researchers want to know whether the digital system leads to a higher population mean score. Write the null and alternative hypotheses using the difference \(\mu_D-\mu_P\), and identify the appropriate testing procedure when the population standard deviations are unknown.

Hints

- The subtraction order is already fixed, so translate “higher” directly into a sign for that difference. - The null statement should represent no population mean difference. - Use the study design to decide whether the observations form one paired sample or two independent samples.

Solution

1. The null hypothesis represents no population mean difference: \(H_0:\mu_D-\mu_P=0\). 2. A higher population mean score for the digital system corresponds to \(H_a:\mu_D-\mu_P>0\). 3. With two independent samples and unknown population standard deviations, the appropriate procedure is a two-sample \(t\)-test for the difference between population means, provided its conditions are satisfied.

Answer

\(H_0:\mu_D-\mu_P=0\) and \(H_a:\mu_D-\mu_P>0\). Use a two-sample \(t\)-test for the difference between population means, assuming the required conditions hold.
54772012
A company compares two independent production methods and asks, “Do the methods have different population mean defect-repair times?” No direction is favored before the data are collected. Using \(\mu_1-\mu_2\), write the null and alternative hypotheses and explain why the alternative should be two-sided.

Hints

- Translate “different” without adding a direction that the question does not state. - The no-difference value for a subtraction of two population means is straightforward. - Choose the alternative before looking at which sample mean happens to be larger.

Solution

1. No population mean difference is represented by \(\mu_1-\mu_2=0\). 2. The null hypothesis is \(H_0:\mu_1-\mu_2=0\). 3. The research question asks about any difference, whether Method 1 is higher or lower. 4. Therefore, the alternative is \(H_a:\mu_1-\mu_2\ne0\).

Answer

\(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\). The alternative is two-sided because the question asks whether the means differ in either direction.
54788212
A two-sample test compares mean completion times. The observed difference is \(6\) minutes and the estimated standard error of the difference is \(2.5\) minutes. A student reports the test statistic as “\(t=2.4\) minutes.” Correct the statement and interpret the test statistic.

Hints

- Track the units in both the difference and its standard error. - A standardized statistic compares a distance with a spread measured in the same units. - Interpret the result in standard-error units rather than original measurement units.

Solution

1. The test statistic is \(t=6/2.5=2.4\). 2. The minutes in the numerator and denominator cancel, so a \(t\)-statistic has no units. 3. The value \(2.4\) means the observed sample-mean difference is \(2.4\) estimated standard errors above the null difference of \(0\).

Answer

The correct statement is \(t=2.4\), with no units. The observed difference is \(2.4\) estimated standard errors above the null value.
54791212
Software reports a two-sided two-sample \(t\)-test p-value as \(0.000\) because it displays only three decimal places. A student writes, “The p-value is exactly zero.” Explain why that statement is not justified and give an appropriate way to report the result from the displayed output.

Hints

- Distinguish a rounded display from an exact mathematical value. - A finite test statistic leaves some tail area in a continuous reference distribution. - Report only the precision supported by the software output.

Solution

1. A displayed value of \(0.000\) means the p-value rounded to three decimal places is \(0.000\); it does not mean the exact tail probability is zero. 2. A continuous \(t\) reference distribution gives a positive tail probability for any finite observed test statistic. 3. From three-decimal rounding alone, an appropriate report is that the p-value is very small; for example, \(p<0.001\) is justified by a display of \(0.000\).

Answer

Do not report \(p=0\). The software display indicates an extremely small positive p-value; reporting \(p<0.001\) is appropriate from the three-decimal output.
54764812
A wildlife rehabilitation center compares mean recovery time for two feeding protocols using independent groups of animals randomly assigned to the protocols. A two-sample \(t\)-test of \(H_0:\mu_1-\mu_2=0\) versus \(H_a:\mu_1-\mu_2\ne0\) gives a p-value of \(0.031\). a) Interpret the p-value in context. b) State the test decision at \(\alpha=0.05\) and at \(\alpha=0.01\). c) For each significance level, give the corresponding conclusion about the population mean recovery times.

Hints

- A p-value is calculated under the assumption made by the null hypothesis. - For a two-sided alternative, “as extreme” includes departures in either direction. - Compare the same p-value separately with each stated significance level before writing conclusions.

Solution

1. If the two protocols truly have equal population mean recovery times, the probability of obtaining a sample mean difference at least as extreme as the observed difference, in either direction, is about \(0.031\). 2. Since \(0.031<0.05\), reject \(H_0\) at \(\alpha=0.05\). There is convincing evidence that the population mean recovery times differ. 3. Since \(0.031>0.01\), fail to reject \(H_0\) at \(\alpha=0.01\). There is not convincing evidence at the \(1\%\) level that the population mean recovery times differ.

Answer

a) Assuming equal population mean recovery times, the chance of obtaining a sample mean difference at least as extreme as the one observed, in either direction, is about \(0.031\). b) At \(\alpha=0.05\), reject \(H_0\). At \(\alpha=0.01\), fail to reject \(H_0\). c) The data provide convincing evidence of a difference at the \(5\%\) level, but not at the \(1\%\) level.
54766012
A valid \(95\%\) confidence interval for \(\mu_1-\mu_2\), the difference in population mean response times for two independent systems, is \((1.2,4.8)\,\text{s}\). Without recomputing a test statistic, determine the decision for the two-sided test \(H_0:\mu_1-\mu_2=0\) versus \(H_a:\mu_1-\mu_2\ne0\) at \(\alpha=0.05\). State the conclusion and explain how the interval supports it.

Hints

- Identify the parameter value that represents “no difference.” - A confidence interval lists parameter values that remain plausible at the corresponding confidence level. - The signs of the endpoints also indicate the direction supported by the data.

Solution

1. A two-sided test at \(\alpha=0.05\) corresponds to checking whether the null value is contained in a \(95\%\) confidence interval for the same parameter. 2. The interval \((1.2,4.8)\) does not contain \(0\). 3. Therefore, reject \(H_0\) at \(\alpha=0.05\). 4. Because the whole interval is positive, the data provide convincing evidence that \(\mu_1>\mu_2\), and therefore that the population mean response times differ.

Answer

Reject \(H_0\). Since \(0\) is not in the \(95\%\) confidence interval \((1.2,4.8)\,\text{s}\), there is convincing evidence at \(\alpha=0.05\) that the two population mean response times differ; the positive interval indicates \(\mu_1>\mu_2\).
54766612
A researcher takes independent random samples from two large populations to compare a quantitative outcome. Sample A has \(n_A=14\) and shows strong right skewness with one clear outlier. Sample B has \(n_B=18\) and is roughly symmetric with no outliers. Both samples satisfy the \(10\%\) condition. Is a two-sample \(t\)-test for the difference in population means justified by the stated conditions? Explain.

Hints

- Meeting independence-related conditions does not automatically settle the shape requirement. - Pay special attention to distribution shape when either sample is small. - One problematic sample can prevent the two-sample procedure from meeting its conditions.

Solution

1. The random sampling and \(10\%\) conditions are satisfied. 2. Both sample sizes are below \(30\), so the observed sample distributions must be free from strong skewness and outliers unless population normality is otherwise established. 3. Sample A has strong skewness and an outlier, so the sample-data condition is not satisfied. 4. Therefore, the usual two-sample \(t\)-test is not justified from the information given.

Answer

No. Although the randomization and \(10\%\) conditions are met, Sample A is small and has strong skewness plus an outlier. The sample-data condition fails, so the usual two-sample \(t\)-test is not justified.
54768412
A two-sample \(t\)-test compares population mean processing times using the subtraction \(\mu_A-\mu_B\). Software reports \(t=-2.45\) and a two-sided p-value of \(0.018\). A student concludes, “System A has a higher population mean processing time because the result is statistically significant.” Correct the student's conclusion.

Hints

- Statistical significance tells you whether a null value is inconsistent with the data at a chosen significance level, not which group is larger by itself. - Use the sign of the statistic together with the stated subtraction order. - Translate a negative value of \(A-B\) into a comparison between A and B.

Solution

1. The p-value indicates evidence against the null hypothesis of equal population mean processing times at significance levels greater than \(0.018\). 2. The negative test statistic means the observed difference \(\bar{x}_A-\bar{x}_B\) is negative relative to the null value \(0\). 3. Thus, the data support a difference in population means, with the observed direction \(\mu_A<\mu_B\), not \(\mu_A>\mu_B\). 4. Statistical significance does not reverse the meaning of the subtraction order.

Answer

The p-value supports a difference in the population mean processing times at significance levels greater than \(0.018\), but the negative test statistic for \(\mu_A-\mu_B\) indicates that System A has the lower observed mean. The evidence points toward \(\mu_A<\mu_B\), not \(\mu_A>\mu_B\).
54770812
A valid two-sample \(t\)-test for \(H_0:\mu_1-\mu_2=0\) reports an estimated difference of \(-3.6\), a standard error of \(1.2\), and about \(40\) degrees of freedom. The alternative is two-sided. Calculate the test statistic and p-value, and state the test decision at \(\alpha=0.01\).

Hints

- Standardize the observed mean difference relative to the null difference. - A two-sided alternative uses extreme results in both directions. - Compare the resulting p-value with the stated significance level, not with a confidence percentage.

Solution

1. The test statistic is \(t=\frac{-3.6-0}{1.2}=-3.00\). 2. With about \(40\) degrees of freedom, the two-sided p-value for \(|t|=3.00\) is approximately \(0.00463\). 3. Since \(0.00463<0.01\), reject \(H_0\). 4. The data provide convincing evidence that the two population means differ.

Answer

\(t=-3.00\), p-value \(\approx0.00463\). Reject \(H_0\) at \(\alpha=0.01\); there is convincing evidence that the population means differ.
54771412
Two independent random samples are taken from two populations to compare their mean weekly practice time. A student combines all observations into one list, finds one overall mean, and proposes a one-sample \(t\)-test against \(0\). Explain why this procedure does not answer the research question and identify the appropriate parameter and test.

Hints

- Ask what information is lost when group labels are removed. - The research question concerns a contrast between two populations, not whether one combined mean equals zero. - Match the procedure to the number of independent populations represented in the data.

Solution

1. Combining both groups into one overall mean removes the group distinction that the research question is about. 2. The parameter of interest is the difference in population means, \(\mu_1-\mu_2\), not a single overall population mean. 3. The appropriate null hypothesis is \(H_0:\mu_1-\mu_2=0\), with an alternative determined by the research question. 4. For two independent samples with unknown population standard deviations, use a two-sample \(t\)-test for the difference between population means, assuming conditions are satisfied.

Answer

Pooling the observations into one mean destroys the comparison between the groups. The parameter is \(\mu_1-\mu_2\), and the appropriate procedure is a two-sample \(t\)-test for the difference between population means.
54773812
Two independent random samples of \(8\) measurements are taken from two populations that are each known to be approximately normal. A researcher wants to test whether the population means differ. A student says a two-sample \(t\)-test is invalid because neither sample size is at least \(30\). Evaluate the student's claim.

Hints

- The sample-size guideline is connected to the shape of the sampling distribution. - Look for information given about the shapes of the two population distributions. - Separate the sample-size issue from the independence and randomization requirements.

Solution

1. A sample size of at least \(30\) is one way to justify an approximately normal sampling distribution when population shape is not known to be normal. 2. Here, both population distributions are stated to be approximately normal, so small sample sizes do not by themselves prevent two-sample \(t\) inference. 3. Because the samples are independent random samples, a two-sample \(t\)-test is appropriate if the remaining sampling conditions, including the \(10\%\) condition when sampling without replacement, are satisfied.

Answer

The student's claim is incorrect. Since both populations are approximately normal, samples of size \(8\) can be used for a two-sample \(t\)-test, provided the other conditions are satisfied.
54775012
Two studies test the same hypotheses about the difference between two population means at the same significance level. The populations have the same variability in both studies, and the true population mean difference is the same. Study 1 uses \(20\) observations in each group. Study 2 uses \(80\) observations in each group. Which study has greater power to detect the true difference, and why?

Hints

- Hold the significance level, true difference, and population variability fixed. - Compare how the sample sizes affect the variability of the estimated difference. - Power is the chance of rejecting the null hypothesis when the alternative is true.

Solution

1. Increasing both sample sizes decreases the standard error of the difference in sample means. 2. With the same true population mean difference, a smaller standard error tends to produce a test statistic farther from the null value. 3. This makes rejection of a false null hypothesis more likely. 4. Therefore, Study 2 has greater power.

Answer

Study 2 has greater power because its larger samples produce a smaller standard error, making a true difference easier to detect.
54775612
A valid two-sample \(t\)-test comparing population mean completion times is performed with all times measured in minutes. The test statistic is \(t=-2.35\), and the two-sided p-value is \(0.024\). Suppose every observation is converted from minutes to seconds and the same test of equal population means is repeated. What happens to the test statistic and p-value? Explain.

Hints

- Track what a common multiplication does to both the sample-mean difference and the sample standard deviations. - The test statistic is a ratio of an estimated difference to its estimated standard error. - If that ratio is unchanged, the reference distribution gives the same p-value.

Solution

1. Converting minutes to seconds multiplies every observation, both sample means, both sample standard deviations, and the standard error by \(60\). 2. The sample-mean difference in the numerator of the test statistic is multiplied by \(60\), and the standard error in the denominator is also multiplied by \(60\). 3. The scale factor cancels, so the test statistic remains \(t=-2.35\). 4. The degrees of freedom and the tail direction are unchanged, so the p-value remains \(0.024\).

Answer

The test statistic remains \(t=-2.35\), and the p-value remains \(0.024\). A common positive unit conversion scales the estimated difference and its standard error by the same factor.
54776812
An experiment uses \(24\) matched pairs of participants. Within each pair, one participant is randomly assigned to Method A and the other to Method B. The response is a quantitative score. A student ignores the matching and proposes an ordinary two-sample \(t\)-test using all \(24\) scores from Method A and all \(24\) scores from Method B as independent groups. Explain why a paired \(t\)-test is the appropriate analysis instead.

Hints

- Use the study design, not just the equal group sizes, to decide whether observations are independent. - Ask whether each observation in one method has a designated partner in the other method. - The matched-pairs analysis reduces each pair to one quantitative difference.

Solution

1. The study design deliberately links each Method A observation with a specific Method B observation through the matched pair. 2. The within-pair responses may be related, so treating all \(48\) scores as two independent samples ignores the dependence created by matching. 3. Form one difference between the A and B responses within each of the \(24\) pairs. 4. Test the population mean of those paired differences with a one-sample \(t\)-test, provided the conditions for paired \(t\) inference are satisfied.

Answer

Use a paired \(t\)-test on the \(24\) within-pair differences. The ordinary two-sample test is inappropriate because the matching creates dependence between observations within each pair.
54777412
A two-sided two-sample \(t\)-test is carried out using the subtraction \(\bar{x}_A-\bar{x}_B\). The test statistic is \(t=2.60\), and the p-value is \(0.013\). Suppose the analyst instead defines the difference as \(\bar{x}_B-\bar{x}_A\) and repeats the same two-sided test. What are the new test statistic and p-value? Explain.

Hints

- Track the effect of reversing the subtraction on the numerator of the test statistic. - The estimated standard error does not depend on which group is listed first. - For a two-sided test, compare the two statistics by magnitude.

Solution

1. Reversing the subtraction changes the observed difference to its negative while leaving the standard error unchanged. 2. Therefore, the test statistic changes from \(2.60\) to \(-2.60\). 3. A two-sided p-value depends on the magnitude of the test statistic, not its sign. 4. Therefore, the p-value remains \(0.013\).

Answer

The new test statistic is \(t=-2.60\), and the two-sided p-value remains \(0.013\).
54778012
Software reports a valid two-sample \(t\)-test with test statistic \(t=1.84\) and \(27.6\) degrees of freedom. A student says the output must be wrong because degrees of freedom have to be whole numbers. Explain why the reported degrees of freedom can be valid.

Hints

- Distinguish a one-sample degrees-of-freedom rule from the approximation used for two independent samples. - The two groups may have different sample sizes and different estimated variances. - Software can use the resulting noninteger value directly when finding the p-value.

Solution

1. A two-sample \(t\)-procedure does not need to assume equal population variances. 2. When the two sample variances and sample sizes are combined using a Welch-type approximation, the resulting degrees of freedom can be noninteger. 3. The value \(27.6\) is therefore a valid approximate degrees-of-freedom value for the reference \(t\)-distribution used by the software.

Answer

The output can be correct. Welch two-sample \(t\) procedures use an approximate degrees-of-freedom calculation that can produce a noninteger value such as \(27.6\).
54778612
Two studies use the same sample sizes, significance level, and two-sided hypotheses to test a difference between two population means. The true population mean difference is the same in both studies. Study A uses a noisy measurement method, while Study B uses a more precise method that produces substantially smaller within-group standard deviations. Which study should have greater power to detect the true difference? Explain.

Hints

- Hold the sample sizes, significance level, and true difference fixed. - Compare how measurement variability affects the standard error of the estimated difference. - Greater separation from the null in standard-error units leads to a greater chance of rejection when the alternative is true.

Solution

1. Smaller within-group standard deviations produce a smaller standard error for the difference in sample means when sample sizes are fixed. 2. For the same true mean difference, a smaller standard error makes the observed standardized difference tend to lie farther from the null value. 3. This increases the probability of rejecting the false null hypothesis. 4. Therefore, Study B should have greater power.

Answer

Study B should have greater power because its lower within-group variability produces a smaller standard error, making the same true mean difference easier to detect.
54781012
A researcher wants to compare the population medians of two independent groups. A student proposes using a standard two-sample \(t\)-test and says that a significant result would show the population medians differ. Explain why this procedure does not directly test the researcher's stated parameter.

Hints

- Identify the population parameter named in the research question. - Identify the sample statistics used by the proposed test. - An inference procedure must match the parameter it is designed to estimate or test.

Solution

1. A standard two-sample \(t\)-test is built to test a hypothesis about the difference between two population means. 2. Its test statistic uses sample means and their estimated sampling variability. 3. A hypothesis about population medians is a different inferential target and is not directly tested by the standard two-sample \(t\)-procedure. 4. Therefore, a significant two-sample \(t\)-test would provide evidence about population means, not automatically about population medians.

Answer

The proposed test does not match the parameter. A standard two-sample \(t\)-test tests a difference in population means, not a difference in population medians.
54782812
A transit agency tests whether the mean weekday ridership on routes using a new schedule is greater than the mean weekday ridership on comparable routes using the old schedule. The hypotheses are \(H_0:\mu_N-\mu_O=0\) and \(H_a:\mu_N-\mu_O>0\). Describe a Type I error and a Type II error in this context.

Hints

- Translate each error type by first deciding whether the null hypothesis is actually true or false. - Then distinguish between rejecting the null hypothesis and failing to reject it. - State the error using the population means, not just the observed sample means.

Solution

1. A Type I error means rejecting \(H_0\) when it is true. 2. In context, that means concluding that the new schedule has greater population mean weekday ridership when the two schedules actually have equal population means. 3. A Type II error means failing to reject \(H_0\) when the stated alternative is true. 4. In context, that means failing to conclude that the new schedule has greater population mean weekday ridership when its population mean actually is greater.

Answer

Type I error: conclude that the new schedule has greater population mean weekday ridership when the two population means are actually equal. Type II error: fail to conclude that the new schedule has greater population mean weekday ridership when it actually does.
54783412
A two-sample \(t\)-test of \(H_0:\mu_A-\mu_B=0\) versus \(H_a:\mu_A-\mu_B>0\) gives \(p=0.04\). A manager concludes, “This proves the mean for A is more than \(5\) units higher than the mean for B.” Explain why that conclusion does not follow from the test, and state the hypotheses that would address the manager’s \(5\)-unit claim.

Hints

- Compare the benchmark in the original null hypothesis with the benchmark in the manager’s claim. - Statistical significance relative to one value does not automatically imply significance relative to a larger value. - Rewrite the claim directly as a statement about the population mean difference.

Solution

1. The reported test evaluates whether the population mean difference is greater than \(0\), not whether it is greater than \(5\). 2. A p-value of \(0.04\) can provide evidence for a positive difference at common significance levels, but it does not establish that the difference exceeds \(5\). 3. To assess the stronger claim, test \(H_0:\mu_A-\mu_B=5\) against \(H_a:\mu_A-\mu_B>5\).

Answer

The conclusion is not justified because the original test compares the mean difference with \(0\), not \(5\). The relevant hypotheses are \(H_0:\mu_A-\mu_B=5\) and \(H_a:\mu_A-\mu_B>5\).
54784012
A two-sample \(t\)-test comparing \(\mu_A-\mu_B\) gives a positive test statistic. Statistical software reports a two-sided p-value of \(0.032\). What would the p-value be for the alternative \(H_a:\mu_A-\mu_B>0\)? What would it be for \(H_a:\mu_A-\mu_B<0\)? Explain why the two one-sided p-values are very different.

Hints

- Use the sign of the observed test statistic to identify which direction the data favor. - A two-sided p-value combines equally extreme outcomes from both tails of a symmetric reference distribution. - The opposite one-sided alternative uses the tail on the other side of the distribution.

Solution

1. Because the test statistic is positive, the observed result lies in the direction of \(H_a:\mu_A-\mu_B>0\). 2. For a symmetric \(t\) distribution, the one-sided p-value in that direction is half the two-sided value: \(0.032\div2=0.016\). 3. The opposite one-sided p-value is the probability in the lower tail up to the positive statistic, which is \(1-0.016=0.984\). 4. The alternatives use opposite tails, so the same positive statistic is evidence for one direction and evidence against the other.

Answer

For \(H_a:\mu_A-\mu_B>0\), \(p=0.016\). For \(H_a:\mu_A-\mu_B<0\), \(p=0.984\).
54786412
A researcher wants to test \(H_a:\mu_A-\mu_B>0\). Software instead defines the difference as \(\mu_B-\mu_A\) and reports \(t=-2.10\) with a lower-tail p-value of \(0.020\). Does this software output provide evidence in the researcher’s intended direction? Explain how the subtraction order and tail correspond.

Hints

- Rewrite the researcher’s claim using the subtraction order chosen by the software. - Reversing a difference reverses its sign and the direction of the alternative. - Match the tail of the p-value to the rewritten alternative.

Solution

1. The researcher’s claim \(\mu_A-\mu_B>0\) is equivalent to \(\mu_B-\mu_A<0\). 2. Because the software uses \(B-A\), evidence for the researcher’s claim should appear as a negative test statistic and a small lower-tail probability. 3. The reported \(t=-2.10\) and lower-tail p-value \(0.020\) are therefore aligned with the intended alternative. 4. At significance levels above \(0.020\), such as \(0.05\), the result provides statistically significant evidence in the researcher’s intended direction.

Answer

Yes. With the software’s \(B-A\) order, the intended claim becomes \(\mu_B-\mu_A<0\). Thus \(t=-2.10\) with lower-tail \(p=0.020\) is evidence that \(\mu_A>\mu_B\).
54787612
Before collecting data, a study team sets \(\alpha=0.05\) for a two-sample \(t\)-test of a difference in population means. The resulting p-value is \(0.073\). After seeing the result, a team member proposes changing the significance level to \(0.10\) so the result can be called statistically significant. Evaluate this proposal.

Hints

- Treat the significance level as part of the study design rather than a label chosen after the result is known. - Compare the reported p-value with the level that was actually planned. - Consider what happens to the Type I error rule if the threshold is changed in response to the data.

Solution

1. With the pre-specified level \(\alpha=0.05\), the result is not statistically significant because \(0.073>0.05\). 2. A significance level represents a planned tolerance for Type I error and should be chosen before examining the test result. 3. Raising \(\alpha\) after seeing the p-value makes the decision rule data-dependent and no longer preserves the originally planned error rate. 4. The study should report the p-value and the decision under the pre-specified \(0.05\) level rather than changing the rule after the fact.

Answer

The proposal is inappropriate. Under the pre-specified \(\alpha=0.05\), fail to reject \(H_0\). Changing \(\alpha\) to \(0.10\) only after seeing \(p=0.073\) is a post hoc change to the decision rule.
54788812
Two independent samples of customer ratings use a five-point scale from \(1\) to \(5\). Group A has \(n_A=8\), with seven ratings of \(1\) and one rating of \(5\). Group B has \(n_B=9\), with eight ratings of \(1\) and one rating of \(5\). A researcher proposes a standard two-sample \(t\)-test for the population mean ratings. Assess whether the usual small-sample \(t\) procedure is well justified.

Hints

- Consider both the sample sizes and the shapes of the observed responses. - Small-sample \(t\) methods need stronger shape conditions than large-sample methods. - A bounded discrete scale can produce severe nonnormality when responses pile up at an endpoint.

Solution

1. Both samples are very small. 2. Each sample is extremely concentrated at the lower endpoint with one value at the upper endpoint, producing a highly nonnormal shape. 3. With such small samples, there is not enough large-sample protection for the mean-based \(t\) approximation. 4. The standard two-sample \(t\)-test is therefore poorly justified for these data without a stronger modeling argument or a different analysis suited to the response scale.

Answer

The usual two-sample \(t\)-test is not well justified. Both groups are very small and have extremely nonnormal five-point rating distributions.
54790612
A researcher records productivity scores for \(100\) workers, sorts the workers by those scores, labels the top \(50\) “Group A” and the bottom \(50\) “Group B,” and then proposes a two-sample \(t\)-test to determine whether the groups have different population mean productivity. Explain why this test does not provide a meaningful inferential comparison of two preexisting populations or treatments.

Hints

- Ask how group membership was determined. - A valid comparison requires group definitions that are not created from the response being tested. - Statistical inference cannot treat a guaranteed sorting difference as if it arose from independent group populations or treatment assignment.

Solution

1. The groups were created directly from the response variable being compared. 2. A difference in their sample means is guaranteed by the sorting rule rather than arising from independent samples or assigned treatments. 3. The usual two-sample \(t\)-test framework assumes groups defined independently of the observed response values used in the test. 4. A small p-value would therefore reflect the outcome-based group construction, not evidence about a genuine population or treatment difference.

Answer

The proposed test is not meaningful because Group A and Group B were defined by sorting the very productivity scores being compared. The resulting mean difference is built into the group-definition rule.
54791812
A two-sided Welch two-sample \(t\)-test has test statistic \(t=1.96\) and \(8\) degrees of freedom. A student says the p-value must be about \(0.05\) because \(1.96\) is the familiar \(95\%\) standard normal cutoff. The correct two-sided p-value is approximately \(0.0857\). Explain the discrepancy and make the decision at \(\alpha=0.05\).

Hints

- A numerical test statistic has meaning only relative to its reference distribution. - Compare the tail thickness of the \(t\) and normal distributions at low degrees of freedom. - Use the stated p-value for the decision rather than a memorized normal cutoff.

Solution

1. The reference distribution is a \(t\) distribution with \(8\) degrees of freedom, not the standard normal distribution. 2. A \(t\) distribution with only \(8\) degrees of freedom has heavier tails than the normal distribution. 3. Therefore, \(|t|=1.96\) is less unusual under this \(t\) distribution and gives the larger two-sided p-value \(0.0857\). 4. Since \(0.0857>0.05\), fail to reject \(H_0\).

Answer

The p-value is approximately \(0.0857\), not \(0.05\), because the reference distribution has only \(8\) degrees of freedom and heavier tails than the standard normal distribution. Fail to reject \(H_0\) at \(\alpha=0.05\).
54792412
A researcher tests \(H_0:\mu_A-\mu_B=0\) against \(H_a:\mu_A-\mu_B>0\). Software reports \(t=-2.40\) and an upper-tail p-value of \(0.008\). Explain why these two reported values are inconsistent with each other for the stated alternative.

Hints

- Locate the observed test statistic relative to zero. - Identify which tail the alternative hypothesis uses. - A statistic pointing opposite the alternative should not produce a tiny one-sided p-value in that alternative’s tail.

Solution

1. The stated alternative uses the upper tail because it looks for positive values of \(\mu_A-\mu_B\). 2. A negative test statistic such as \(-2.40\) lies well to the left of \(0\), opposite the direction of the alternative. 3. The upper-tail area to the right of a negative statistic must be greater than \(0.5\), not as small as \(0.008\). 4. Therefore, either the sign of the statistic, the tail used for the p-value, or the stated subtraction order is wrong in the report.

Answer

The output is inconsistent. For an upper-tailed test, \(t=-2.40\) would produce a large p-value greater than \(0.5\), not \(0.008\).
54793012
Two independent large random samples have heavily overlapping ranges of individual observations. A valid two-sample \(t\)-test for equal population means reports \(t=-3.05\) and a two-sided p-value of \(0.003\). A student says the null hypothesis cannot be rejected because the raw data ranges overlap. Evaluate the student’s reasoning.

Hints

- Distinguish variation among individuals from uncertainty in sample means. - A test of means does not require the two groups’ raw values to be separated. - Use the reported p-value for the formal inference.

Solution

1. A two-sample \(t\)-test compares population means using the observed mean difference relative to its standard error. 2. Overlap among individual observations or raw ranges is compatible with a statistically detectable difference in population means. 3. The reported p-value \(0.003\) is below \(0.05\), so the data provide significant evidence that the population means differ. 4. The overlapping ranges do not override the test result because they describe individual variation, not the sampling uncertainty of the mean difference.

Answer

The student is incorrect. The overlapping raw ranges do not prevent a significant difference in means. With \(p=0.003\), reject equality of the population means at \(\alpha=0.05\).
54793612
For the same two independent samples, software reports a \(95\%\) confidence interval for \(\mu_A-\mu_B\) of \((0.8,3.2)\). The same report also gives a two-sided two-sample \(t\)-test of \(H_0:\mu_A-\mu_B=0\) with p-value \(0.20\). Explain why these two results cannot both come from the same standard \(t\)-inference procedure.

Hints

- Match a \(95\%\) two-sided confidence interval with a two-sided test at \(\alpha=0.05\). - Check whether the null difference is inside or outside the interval. - The confidence-interval and hypothesis-test decisions must agree when they use the same \(t\) procedure.

Solution

1. A \(95\%\) two-sample \(t\)-confidence interval and a two-sided \(t\)-test at \(\alpha=0.05\) are dual procedures when based on the same estimate, standard error, and degrees of freedom. 2. The interval \((0.8,3.2)\) excludes the null value \(0\). 3. Therefore, the corresponding two-sided test must reject \(H_0\) at \(\alpha=0.05\), which requires a p-value below \(0.05\). 4. A reported p-value of \(0.20\) would fail to reject, so at least one result is incorrect or the two outputs were not produced by the same inferential setup.

Answer

The results are inconsistent. Because the \(95\%\) interval excludes \(0\), the corresponding two-sided \(t\)-test must have \(p<0.05\), not \(p=0.20\).
54721512
A school compares two tutoring schedules. Eight students’ score gains are \(2,3,4,5,7,8,9,12\). Four students used Schedule A and four used Schedule B, and the observed difference in mean gain was \(3.5\) points. To estimate how unusual a difference of at least \(3.5\) would be if the schedule labels made no difference, describe one randomization trial. In \(5000\) shuffled trials, \(146\) produced a difference whose absolute value was at least \(3.5\). Estimate the probability of such an extreme result under the random-label model.

Hints

- What parts of the original data should remain fixed if the schedule labels are assumed irrelevant? - How can one shuffled assignment preserve the original group sizes? - Which simulated differences should count as at least as extreme as the observed one?

Solution

1. Keep the eight score gains fixed and randomly assign four of them to label A and the other four to label B. 2. Compute the difference in the two shuffled group means and record whether its absolute value is at least \(3.5\). 3. The estimated probability is \(\frac{146}{5000}=0.0292\).

Answer

One trial randomly reallocates four of the fixed gains to each schedule, computes the difference in means, and checks whether its absolute value is at least \(3.5\). The estimated probability is \(0.0292\), or \(2.92\%\).
54764212
A theater company is comparing two wireless cue systems. During a randomized trial, \(24\) cue events use the new system and \(27\) use the current system. Cue latency, in milliseconds, is measured for each event. The new system has \(\bar{x}_N=42.1\,\text{ms}\) and \(s_N=5.4\,\text{ms}\); the current system has \(\bar{x}_C=45.6\,\text{ms}\) and \(s_C=6.1\,\text{ms}\). Both sample distributions are roughly symmetric with no outliers. Test \(H_0:\mu_N-\mu_C=0\) versus \(H_a:\mu_N-\mu_C<0\) at \(\alpha=0.05\). Calculate the test statistic and p-value, then state the conclusion in context.

Hints

- Keep the subtraction order in the hypotheses, sample difference, and conclusion the same. - Ask how far the observed difference is from the null value relative to its estimated sampling variability. - Use the direction of the alternative hypothesis when deciding which tail contributes to the p-value.

Solution

1. Random assignment and the sample shapes support a two-sample \(t\)-test for \(\mu_N-\mu_C\). 2. The estimated difference is \(42.1-45.6=-3.5\,\text{ms}\), and the standard error is \(\sqrt{\frac{5.4^2}{24}+\frac{6.1^2}{27}}\approx1.610\,\text{ms}\). 3. The test statistic is \(t=\frac{-3.5}{1.610}\approx-2.17\). Technology gives approximately \(49\) degrees of freedom. 4. The one-sided p-value is approximately \(0.0173\). 5. Because \(0.0173<0.05\), reject \(H_0\). The trial provides convincing evidence that the new system has a lower population mean cue latency than the current system.

Answer

\(t\approx-2.17\), p-value \(\approx0.0173\). Reject \(H_0\). There is convincing evidence that the new cue system has a lower population mean latency.
54765412
A materials lab compares the population mean wear resistance of two coatings. For the subtraction \(\mu_N-\mu_S\), where \(N\) is the new coating and \(S\) is the standard coating, a two-sample test produces \(t=1.91\) with about \(40\) degrees of freedom. The research question is whether the new coating has greater mean wear resistance. A student reports a two-sided p-value of about \(0.0633\) and concludes there is no evidence at \(\alpha=0.05\). Explain the student's error, find the appropriate p-value, and give the correct conclusion.

Hints

- Translate the phrase “greater mean” into a direction for the population mean difference. - Decide whether evidence in the opposite direction should count toward this particular research question. - The sign of the test statistic should agree with the subtraction order and the stated direction.

Solution

1. The research question is directional, so the alternative is \(H_a:\mu_N-\mu_S>0\), not a two-sided alternative. 2. With \(t=1.91\) and about \(40\) degrees of freedom, the upper-tail p-value is approximately \(0.0317\). 3. The reported two-sided p-value is approximately twice the appropriate one-sided p-value because the observed statistic is in the direction specified by the alternative. 4. Since \(0.0317<0.05\), reject \(H_0\). There is convincing evidence that the new coating has greater population mean wear resistance than the standard coating.

Answer

The student used a two-sided p-value even though the research question gives \(H_a:\mu_N-\mu_S>0\). The appropriate p-value is about \(0.0317\). Reject \(H_0\) at \(\alpha=0.05\); the data provide convincing evidence that the new coating has greater population mean wear resistance.
54767212
A large randomized experiment compares two manufacturing processes. A valid two-sample \(t\)-test finds that the new process has a population mean cycle time lower than the current process, with p-value less than \(0.001\). The estimated mean reduction is only \(0.7\,\text{s}\). Explain what the very small p-value establishes and what it does not establish about the practical importance of the new process.

Hints

- Separate the strength of evidence from the size of an estimated effect. - A p-value answers a question about compatibility with a null model, not about business or engineering importance. - Consider what additional standard would be needed to call a time reduction practically meaningful.

Solution

1. The very small p-value provides strong evidence against the null hypothesis of equal population mean cycle times in favor of the stated alternative. 2. The randomized experiment supports a causal conclusion that the process change affects mean cycle time for units represented by the experiment. 3. Statistical significance does not measure the size or practical importance of the effect. 4. The estimated reduction of \(0.7\,\text{s}\) must be evaluated using engineering, cost, or operational criteria to decide whether the effect is practically meaningful.

Answer

The p-value below \(0.001\) gives strong statistical evidence that the population mean cycle times differ in the stated direction. It does not show that the difference is practically important. The estimated \(0.7\,\text{s}\) reduction must be judged against a meaningful operational standard.
54769012
Two independent random samples are drawn with replacement from two finite populations. A researcher plans a two-sample \(t\)-test for the difference in population means and worries that each sample contains more than \(10\%\) of its population. Does the usual \(10\%\) condition for sampling without replacement create a problem here? Explain the role of sampling with replacement.

Hints

- Ask why a finite-population percentage condition is needed in the first place. - Compare what happens to the population after a draw with and without replacement. - Do not confuse this issue with the separate conditions involving randomization and sample shape.

Solution

1. The \(10\%\) condition is used when sampling without replacement to make dependence among sampled observations small enough for the usual independence approximation. 2. With sampling with replacement, each draw returns the selected unit before the next draw, so the finite-population depletion that motivates the \(10\%\) condition is absent. 3. Therefore, exceeding \(10\%\) of the finite population is not itself a violation in this with-replacement design. 4. The researcher still needs to check the other requirements for the two-sample \(t\)-test.

Answer

No. The usual \(10\%\) condition applies to sampling without replacement. With replacement, the sampling mechanism does not create the same finite-population dependence, so exceeding \(10\%\) is not itself a problem. Other test conditions must still be checked.
54769612
Two training methods are compared using independent groups. The sampling conditions for a two-sample \(t\)-test are satisfied. Method A has \(n_A=31\), \(\bar{x}_A=18.2\,\text{min}\), and \(s_A=3.1\,\text{min}\). Method B has \(n_B=29\), \(\bar{x}_B=20.0\,\text{min}\), and \(s_B=3.6\,\text{min}\). Test \(H_0:\mu_A-\mu_B=0\) versus \(H_a:\mu_A-\mu_B<0\) at \(\alpha=0.05\). Calculate the test statistic and p-value and state the conclusion.

Hints

- Preserve the subtraction order \(A-B\) when finding the observed difference. - Combine the two independent estimates of sampling variability before standardizing the difference. - The alternative points to the lower tail of the reference distribution.

Solution

1. The estimated difference is \(18.2-20.0=-1.8\,\text{min}\). 2. The standard error is \(\sqrt{\frac{3.1^2}{31}+\frac{3.6^2}{29}}\approx0.870\,\text{min}\). 3. The test statistic is \(t=\frac{-1.8}{0.870}\approx-2.07\). Technology gives about \(55\) degrees of freedom. 4. The lower-tail p-value is approximately \(0.0216\). 5. Since \(0.0216<0.05\), reject \(H_0\). There is convincing evidence that Method A has a lower population mean completion time than Method B.

Answer

\(t\approx-2.07\), p-value \(\approx0.0216\). Reject \(H_0\). The data provide convincing evidence that Method A has a lower population mean completion time than Method B.
54770212
Independent random samples of seniors are taken from two high schools to compare population mean commute times. A valid two-sample \(t\)-test gives a p-value of \(0.002\), and the observed mean commute time is higher at School A. Explain what conclusion is supported about the two school populations and why the study does not establish that attending School A causes longer commutes.

Hints

- Separate the strength of statistical evidence from the type of conclusion allowed by the study design. - Random sampling and random assignment answer different questions. - Ask whether the explanatory group membership was assigned by researchers or merely observed.

Solution

1. The small p-value provides strong evidence that the two school populations do not have equal mean commute times, in the direction indicated by the observed difference. 2. Because the students were randomly sampled, the result can support population inference to the school populations represented by the sampling design. 3. School attendance was not randomly assigned as a treatment in an experiment. 4. Therefore, the test does not establish that attending School A causes longer commutes; other differences between the school populations may explain the association.

Answer

The data provide strong evidence that the population mean commute times differ, with School A higher in the observed direction. Random sampling supports inference to the sampled school populations, but the observational design does not establish that school attendance causes the difference.
54772612
A two-sample \(t\)-test uses \(H_a:\mu_1-\mu_2>0\). The observed test statistic is \(t=-0.80\) with \(30\) degrees of freedom. Find the one-sided p-value and explain why it is large even though \(|t|=0.80\) is not extremely small.

Hints

- A one-sided p-value depends on direction as well as distance from zero. - Locate the observed statistic relative to the tail named by the alternative hypothesis. - Do not automatically halve a two-sided p-value without checking whether the observed statistic points in the alternative's direction.

Solution

1. The alternative looks for positive values of \(\mu_1-\mu_2\), so the p-value is the upper-tail probability \(P(T_{30}\ge-0.80)\). 2. This probability is approximately \(0.785\). 3. The observed statistic is negative, which is opposite the direction specified by the alternative. 4. Most of the reference distribution lies at or above \(-0.80\), so the one-sided p-value is large.

Answer

The p-value is approximately \(0.785\). It is large because the observed test statistic points opposite the positive direction specified by \(H_a\).
54773212
A regional testing center compares the mean score on a certification assessment for candidates trained with two independent programs. Program A has \(n_A=34\), \(\bar{x}_A=81.2\), and \(s_A=7.4\). Program B has \(n_B=31\), \(\bar{x}_B=76.5\), and \(s_B=6.8\). The conditions for two-sample \(t\) inference are satisfied. Test \(H_0:\mu_A-\mu_B=2\) against \(H_a:\mu_A-\mu_B>2\) at \(\alpha=0.05\). Calculate the test statistic and p-value and state the conclusion in context.

Hints

- The null hypothesis does not say the population means are equal, so identify the difference specified by the null. - Compare the observed sample-mean difference with that null difference in units of estimated sampling variability. - Match the tail of the p-value to the direction stated in the alternative hypothesis.

Solution

1. The observed difference is \(81.2-76.5=4.7\). 2. The standard error is \(\sqrt{\frac{7.4^2}{34}+\frac{6.8^2}{31}}\approx1.761\). 3. The test statistic compares the observed difference with the null difference of \(2\): \(t=\frac{4.7-2}{1.761}\approx1.53\). 4. Using about \(63\) degrees of freedom, the upper-tail p-value is approximately \(0.065\). 5. Because \(0.065>0.05\), fail to reject \(H_0\). The data do not provide convincing evidence that Program A's population mean score exceeds Program B's by more than \(2\) points.

Answer

\(t\approx1.53\) with p-value \(\approx0.065\). Fail to reject \(H_0\). There is not convincing evidence that \(\mu_A-\mu_B>2\).
54776212
Two independent groups are compared with a valid two-sample \(t\)-test. Separate \(95\%\) confidence intervals for the two population means overlap slightly, but the two-sample test of \(H_0:\mu_1-\mu_2=0\) gives p-value \(0.032\). A student argues that the test result must be wrong because the two individual confidence intervals overlap. Evaluate the argument.

Hints

- Identify the parameter tested by the two-sample procedure. - Separate intervals for two means are not the same object as an interval for their difference. - Base the test decision on the p-value from the stated valid test.

Solution

1. Overlap of separate confidence intervals for \(\mu_1\) and \(\mu_2\) is not the criterion for a test about \(\mu_1-\mu_2\). 2. The two-sample test uses the sampling variability of the estimated difference directly. 3. A p-value of \(0.032\) is below \(0.05\), so the valid two-sided test rejects \(H_0\) at the \(5\%\) significance level. 4. Slight overlap of the separate mean intervals can occur even when a confidence interval for the difference excludes \(0\).

Answer

The student's argument is incorrect. Separate confidence-interval overlap does not determine the two-sample test decision. With p-value \(0.032\), reject \(H_0\) at \(\alpha=0.05\).
54779212
Two hypothetical studies observe the same sample-mean difference of \(3\) units and the same sample standard deviation of \(6\) units in both groups. Study A has \(25\) observations per group, while Study B has \(100\) observations per group. Assume the conditions for two-sample \(t\) inference are satisfied. Compare the two-sided test statistics and p-values for testing \(H_0:\mu_1-\mu_2=0\). Explain why the same observed difference gives much stronger evidence in Study B.

Hints

- The observed difference and sample standard deviations are fixed, so compare only the standard errors created by the two sample sizes. - A test statistic measures the observed difference in standard-error units. - A larger magnitude test statistic produces a smaller two-sided p-value.

Solution

1. For Study A, the standard error is \(\sqrt{\frac{6^2}{25}+\frac{6^2}{25}}\approx1.697\), so \(t\approx\frac{3}{1.697}=1.77\). The two-sided p-value is approximately \(0.0835\). 2. For Study B, the standard error is \(\sqrt{\frac{6^2}{100}+\frac{6^2}{100}}\approx0.849\), so \(t\approx\frac{3}{0.849}=3.54\). The two-sided p-value is approximately \(0.00051\). 3. The larger samples in Study B produce a much smaller standard error, so the same raw difference is farther from \(0\) in standard-error units.

Answer

Study A: \(t\approx1.77\), p-value \(\approx0.0835\). Study B: \(t\approx3.54\), p-value \(\approx0.00051\). Study B gives much stronger evidence because its larger samples make the estimated difference much more precise.
54779812
A valid two-sample test uses \(H_0:\mu_1-\mu_2=4\) versus \(H_a:\mu_1-\mu_2>4\). The observed sample-mean difference is \(7\), the estimated standard error is \(1.5\), and the test uses \(30\) degrees of freedom. A student calculates \(t=7/1.5\) because “the null distribution is centered at zero.” Correct the test statistic, find the p-value, and explain the student's error.

Hints

- Identify the parameter value assumed by the null hypothesis before standardizing. - A test statistic measures the observed estimate's distance from the null value, not automatically from zero. - Use the direction of the alternative to select the p-value tail.

Solution

1. Under the null hypothesis, the sampling distribution of the estimated difference is centered at the null difference \(4\), not at \(0\). 2. The test statistic is \(t=\frac{7-4}{1.5}=2.00\). 3. With \(30\) degrees of freedom, the upper-tail p-value is approximately \(0.0273\). 4. The student's calculation incorrectly measures the observed difference from \(0\) instead of from the value specified by \(H_0\).

Answer

The correct test statistic is \(t=2.00\), and the one-sided p-value is approximately \(0.0273\). The null reference point is \(4\), not \(0\).
54780412
Independent random samples satisfy the conditions for a two-sample \(t\)-test of \(H_0:\mu_1-\mu_2=0\). The observed sample-mean difference is \(2\). Group 1 has \(n_1=25\) and \(s_1=4\); Group 2 has \(n_2=36\) and \(s_2=6\). A student calculates the standard error by averaging the two group standard errors. Correct the standard error, then find the two-sided test statistic and p-value.

Hints

- Use the stated independent-sample conditions, then work with the two variance contributions rather than averaging group standard errors. - The null value for the mean difference is \(0\). - After standardizing the observed difference, use both tails for the stated alternative.

Solution

1. Independent variance contributions add, so the standard error is \(\sqrt{\frac{4^2}{25}+\frac{6^2}{36}}=\sqrt{1.64}\approx1.281\). 2. The test statistic is \(t=\frac{2-0}{1.281}\approx1.56\). 3. A Welch calculation gives about \(59\) degrees of freedom. 4. The two-sided p-value is approximately \(0.124\). 5. Averaging the group standard errors is incorrect because the variability of an independent difference is obtained by adding variances, not averaging standard deviations.

Answer

The correct standard error is approximately \(1.281\), the test statistic is \(t\approx1.56\), and the two-sided p-value is approximately \(0.124\).
54781612
Two independent groups are being compared. A one-sample test for Group A against a benchmark value gives p-value \(0.03\), while the corresponding one-sample test for Group B against the same benchmark gives p-value \(0.20\). A student concludes that the two population means must differ because one result is statistically significant and the other is not. Explain why that conclusion does not follow and identify the appropriate test for comparing the groups.

Hints

- Ask whether either reported test has the difference \(\mu_A-\mu_B\) as its parameter. - “Significant” versus “not significant” is not itself a test of a difference between effects. - Use one procedure whose null hypothesis directly represents equality of the two population means.

Solution

1. A significant result in one separate test and a nonsignificant result in another does not directly test the difference between the two population means. 2. The two one-sample tests can have different standard errors and different distances from the benchmark, so their significance labels are not a valid comparison statistic. 3. To compare the groups directly, test \(H_0:\mu_A-\mu_B=0\) with a two-sample \(t\)-test, provided its conditions are satisfied.

Answer

The conclusion is not justified from the two separate p-values. The groups should be compared directly with a two-sample \(t\)-test for \(\mu_A-\mu_B\).
54782212
Ten stores are assigned to Promotion A and ten stores to Promotion B. Each store reports daily sales for \(30\) days. A student treats the \(300\) daily sales values from each promotion as \(300\) independent observations and proposes a two-sample \(t\)-test for mean sales. Explain the problem with this analysis.

Hints

- Identify what was independently assigned to each promotion. - Repeated measurements from the same unit can be correlated. - The number of recorded rows is not always the number of independent observations for inference.

Solution

1. Daily sales measurements from the same store are repeated observations on one experimental unit and are likely related across days. 2. Treating all \(300\) daily values in a promotion as independent ignores the clustering within stores and greatly overstates the effective sample size. 3. The independent units assigned to promotions are the stores, not the individual store-days. 4. A standard two-sample \(t\)-test on all \(600\) daily observations as if they were independent is therefore inappropriate.

Answer

The proposed test uses pseudoreplication. The \(30\) daily values from one store are not \(30\) independent experimental units; the store is the unit assigned to a promotion.
54784612
Independent random samples from two populations give \(n_A=15\), \(\bar{x}_A=25\), \(s_A=4\) and \(n_B=20\), \(\bar{x}_B=20\), \(s_B=10\). Assume both population distributions are approximately normal with no outliers. A student says a two-sample \(t\)-test cannot be used because the sample standard deviations are too different. Use the unequal-variance two-sample \(t\) procedure to test \(H_0:\mu_A-\mu_B=0\) against \(H_a:\mu_A-\mu_B\ne0\). Use \(26.34\) degrees of freedom.

Hints

- Check the random-sampling and population-shape information before using a small-sample procedure. - Let each group contribute its own estimated variability to the uncertainty of the difference; equal variances are not required. - Compare the resulting two-sided p-value with the stated significance level.

Solution

1. The independent random samples and the stated population shapes support a two-sample \(t\) procedure. The unequal-variance procedure does not require the population variances to be equal. 2. The estimated standard error is \(\sqrt{\frac{4^2}{15}+\frac{10^2}{20}}\approx2.463\). 3. The test statistic is \(t=\frac{25-20}{2.463}\approx2.030\). 4. With about \(26.34\) degrees of freedom, the two-sided p-value is approximately \(0.0526\). 5. At \(\alpha=0.05\), fail to reject \(H_0\). The data do not provide significant evidence of a difference in the population means at the \(5\%\) level.

Answer

The unequal-variance procedure is appropriate despite the different sample standard deviations. \(t\approx2.030\), \(df\approx26.34\), and \(p\approx0.0526\), so fail to reject \(H_0\) at \(\alpha=0.05\).
54785212
In a randomized experiment, \(100\) participants are assigned to Treatment A and \(100\) to Treatment B. Before the outcome is measured, \(30\) participants assigned to A drop out, compared with \(5\) assigned to B. A researcher proposes a two-sample \(t\)-test using only the participants who completed the study. Explain why a small p-value from that test would not, by itself, preserve the original randomized comparison.

Hints

- Ask what random assignment protected against at the start of the experiment. - Then consider whether the same kinds of participants remained in both groups. - A significance calculation cannot repair a systematic change in who is being compared.

Solution

1. Random assignment initially creates comparable treatment groups and supports a causal comparison. 2. Large, unequal dropout after assignment can destroy that comparability if dropout is related to treatment or outcome. 3. A two-sample \(t\)-test on completers treats the remaining observations as the groups being compared, but the groups may now differ systematically because of attrition. 4. Therefore, statistical significance among completers does not by itself restore the causal protection provided by the original random assignment.

Answer

Differential attrition can undermine the randomized comparison. A small p-value from a completers-only two-sample \(t\)-test does not by itself establish that the observed difference is a causal treatment effect.
54787012
A \(90\%\) two-sample \(t\)-confidence interval for \(\mu_A-\mu_B\) is \((0.2,4.8)\), based on \(30\) degrees of freedom. For these degrees of freedom, use \(t^*=1.697\) for a \(90\%\) interval and \(t^*=2.042\) for a \(95\%\) interval. A student says the result automatically means a two-sided test of \(H_0:\mu_A-\mu_B=0\) would reject at both \(\alpha=0.10\) and \(\alpha=0.05\). Evaluate the claim.

Hints

- Use the midpoint and half-width of the reported interval to recover its point estimate and margin of error. - Divide the \(90\%\) margin of error by its critical value to find the standard error. - Construct the \(95\%\) interval and check whether it contains the null value.

Solution

1. The \(90\%\) interval has center \(\frac{0.2+4.8}{2}=2.5\) and margin of error \(2.3\). 2. Because \(0\) is outside the \(90\%\) interval, the corresponding two-sided test rejects at \(\alpha=0.10\). 3. The estimated standard error is \(2.3/1.697\approx1.355\). 4. The \(95\%\) margin of error is \(2.042\cdot1.355\approx2.768\), giving the interval \(2.5\pm2.768\), or approximately \((-0.27,5.27)\). 5. Because the \(95\%\) interval contains \(0\), the corresponding two-sided test fails to reject at \(\alpha=0.05\).

Answer

The claim is false. The test rejects at \(\alpha=0.10\), but the corresponding \(95\%\) interval is approximately \((-0.27,5.27)\), so the test fails to reject at \(\alpha=0.05\).
54789412
Software reports a two-sided Welch two-sample \(t\)-test with test statistic \(t=5.00\) but only \(1\) degree of freedom. The two-sided p-value is approximately \(0.126\). Explain why the p-value can be this large even though \(|t|=5\) sounds extreme, and make the decision at \(\alpha=0.05\).

Hints

- A test statistic alone does not determine a p-value without a reference distribution. - Compare the tail behavior of a \(t\) distribution with very few degrees of freedom to one with many. - Use the reported p-value for the formal decision.

Solution

1. The p-value is determined by both the magnitude of the test statistic and the reference distribution’s degrees of freedom. 2. With only \(1\) degree of freedom, the \(t\) distribution has extremely heavy tails, so values far from \(0\) are much less unusual than they would be with many degrees of freedom. 3. Thus, \(|t|=5\) corresponds to a two-sided p-value of about \(0.126\) for \(df=1\). 4. Since \(0.126>0.05\), fail to reject \(H_0\).

Answer

Fail to reject \(H_0\). With only \(1\) degree of freedom, the \(t\) distribution is so heavy-tailed that \(|t|=5\) still gives \(p\approx0.126\).
54790012
A study is designed to test \(H_0:\mu_A-\mu_B=0\) against \(H_a:\mu_A-\mu_B>0\). Suppose the true population difference is actually substantially negative. Compared with the case \(\mu_A-\mu_B=0\), what happens to the probability of rejecting \(H_0\) in favor of the stated upper-tailed alternative? Explain.

Hints

- Locate the rejection region implied by the alternative hypothesis. - Then imagine shifting the true mean difference in the opposite direction. - Power depends on where the true parameter lies relative to the alternative and rejection region.

Solution

1. The rejection region for the stated alternative is in the upper tail of the test-statistic distribution. 2. If the true population difference is negative, the sampling distribution shifts in the direction opposite the rejection region. 3. As a result, rejection in favor of \(\mu_A-\mu_B>0\) becomes less likely than it is at the null boundary. 4. For a sufficiently negative true difference, that rejection probability can be far below the significance level.

Answer

The probability of rejecting in favor of \(\mu_A-\mu_B>0\) decreases, potentially far below \(\alpha\), because the true difference shifts the test statistic away from the upper-tail rejection region.
54774412
A repair shop is comparing the time required to complete the same diagnostic procedure with two independent software systems. The procedure times, in minutes, for random samples are: <table><tr><th>System A</th><td>14.2</td><td>15.1</td><td>13.8</td><td>14.7</td><td>16.0</td><td>15.4</td><td>14.9</td><td>13.9</td></tr><tr><th>System B</th><td>16.1</td><td>15.8</td><td>17.0</td><td>16.5</td><td>15.2</td><td>16.8</td><td>17.3</td><td>15.9</td><td>16.4</td></tr></table> The sample distributions show no strong skewness or outliers. At \(\alpha=0.01\), test whether System A has a lower population mean procedure time than System B.

Hints

- Summarize each sample before comparing the two population means. - Keep the subtraction order consistent with the direction named in the alternative hypothesis. - Use the observed difference relative to its estimated sampling variability to judge the strength of evidence.

Solution

1. Test \(H_0:\mu_A-\mu_B=0\) against \(H_a:\mu_A-\mu_B<0\). 2. From the data, \(\bar{x}_A=14.75\), \(s_A\approx0.762\), \(\bar{x}_B\approx16.333\), and \(s_B\approx0.656\). 3. The standard error is \(\sqrt{\frac{0.762^2}{8}+\frac{0.656^2}{9}}\approx0.347\). 4. The test statistic is \(t=\frac{14.75-16.333}{0.347}\approx-4.57\). A Welch calculation gives about \(14\) degrees of freedom. 5. The lower-tail p-value is approximately \(0.00022\). 6. Because \(0.00022<0.01\), reject \(H_0\). There is convincing evidence that System A has a lower population mean procedure time than System B.

Answer

\(t\approx-4.57\) with about \(14\) degrees of freedom and p-value \(\approx0.00022\). Reject \(H_0\); the data provide convincing evidence that \(\mu_A<\mu_B\).
54785812
A city has \(500\) registered food trucks. A simple random sample of \(80\) has mean daily revenue \(\$1240\) and sample standard deviation \(\$300\). Another city has \(1200\) registered food trucks. A simple random sample of \(60\) has mean daily revenue \(\$1090\) and sample standard deviation \(\$280\). Because the samples were taken without replacement, estimate each sampling variance with \(\frac{N-n}{N-1}\frac{s^2}{n}\). Test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2\ne0\) at \(\alpha=0.05\). Use \(128\) degrees of freedom, and compare the adjusted standard error with the usual unadjusted standard error.

Hints

- Apply the finite-population factor separately to each city’s estimated sampling variance. - Add the two adjusted variance contributions before taking the square root. - Use the adjusted standard error in the test statistic and compare the two-sided p-value with \(0.05\).

Solution

1. The estimated difference is \(1240-1090=150\) dollars. 2. The usual unadjusted standard error is \(\sqrt{\frac{300^2}{80}+\frac{280^2}{60}}\approx49.312\) dollars. 3. Applying the finite-population correction to each sample gives \(SE=\sqrt{\frac{500-80}{500-1}\frac{300^2}{80}+\frac{1200-60}{1200-1}\frac{280^2}{60}}\approx46.790\) dollars. 4. The adjusted test statistic is \(t=\frac{150}{46.790}\approx3.206\). 5. With \(128\) degrees of freedom, the two-sided p-value is approximately \(0.0017\). 6. Since \(0.0017<0.05\), reject \(H_0\). There is significant evidence that the cities differ in mean daily food-truck revenue. 7. The adjusted standard error is smaller because sampling a substantial fraction of a finite population reduces sampling variability.

Answer

The unadjusted standard error is approximately \(\$49.312\), while the finite-population-adjusted standard error is approximately \(\$46.790\). The adjusted test gives \(t\approx3.206\), \(df=128\), and \(p\approx0.0017\), so reject \(H_0\) at \(\alpha=0.05\).

All problems may be used, copied and printed free of charge for school and tutoring, including paid tutoring. Commercial adaptations as well as publication or redistribution on the internet are not permitted.