Aimathic
Login | English | Deutsch

Free math worksheets

Build your own math worksheets from 30,000+ problems for grades 3 to 12, from fractions to AP Calculus. Every problem comes with step-by-step solutions.

Evaluate statistical models and inferences

Click problems to add them to your worksheet.

55558411
Researchers record weekly exercise time and resting heart rate for a group of adults. Adults who exercise more tend to have lower resting heart rates. No one is assigned an exercise amount. Does this study by itself establish that more exercise causes a lower resting heart rate?

Hints

- Check whether the explanatory variable was assigned by the researchers. - Distinguish observing a relationship from creating comparable treatment groups.

Solution

1. The study observes existing behavior rather than assigning exercise amounts. 2. It can show an association between exercise time and resting heart rate, but other variables may help explain the relationship. 3. Therefore, the study by itself does not establish causation.

Answer

No. It shows an association, not a cause-and-effect conclusion.
55558511
A school district randomly selects \(150\) students from all high school students in the district and asks whether they use the public library each week. No treatment is assigned. What does the random selection primarily support: generalizing the survey result to high school students in the district, or making a causal claim about library use?

Hints

- Identify whether the random step chooses who enters the study or determines a treatment they receive. - Match that random step to the kind of conclusion it can support.

Solution

1. Random selection helps make the sample representative of the district's high school student population. 2. No treatment is assigned, so the study does not create the conditions needed for a causal comparison. 3. Therefore, random selection supports generalization, not causation.

Answer

It supports generalizing the survey result to high school students in the district, not making a causal claim.
55558611
The histogram shows results from a no-effect randomization model. The dashed line marks the observed difference from the actual study. Does the observed difference look typical or unusual under the no-effect model? What does that suggest about the no-effect explanation?
Figure for problem 555586

Hints

- Compare the observed line with the region where most simulated outcomes appear. - Results deep in a tail occur less often under the model used to generate the simulation.

Solution

1. Most simulated differences are concentrated near \(0\). 2. The observed difference is marked far in the right tail, where very few simulated results occur. 3. Therefore, the observed difference is unusual under the no-effect model and provides evidence against that explanation.

Answer

Unusual. The observed difference is far in the tail, so the simulation provides evidence against the no-effect model.
54934511
A news article reports that months with higher ice-cream sales also have more drowning incidents and concludes that buying ice cream increases drowning risk. Explain why the evidence does not justify that causal conclusion, and identify one plausible confounding variable.

Hints

- Ask whether anyone was randomly assigned to the proposed cause. - Look for a third factor that changes both variables during the same months. - An observational association can have alternative explanations.

Solution

1. The evidence is observational: no one was randomly assigned to buy ice cream, so the association does not isolate a causal effect. 2. Temperature or season is a plausible confounder because warmer months can increase both ice-cream sales and swimming exposure, which can increase drowning incidents.

Answer

The study shows an observational association, not causation. Temperature or season is a plausible confounder affecting both ice-cream sales and swimming exposure.
54934711
A simple random sample of \(800\) adults records weekly exercise time and nightly sleep duration. Adults who exercise more also tend to report more sleep. What can this study reasonably generalize to the population, and why can it not establish that increasing exercise causes more sleep?

Hints

- Use the sampling method to decide the scope of generalization. - Use the assignment method to decide whether causation is supported. - Random sampling and random assignment answer different inference questions.

Solution

1. Because the adults were selected with a simple random sample, the observed association can be generalized to the adult population represented by the sampling frame, subject to ordinary sampling and measurement limitations. 2. Exercise was observed rather than randomly assigned, so confounding or reverse causation can explain the association. The design therefore does not establish a causal effect.

Answer

The random sample supports a population-level association between exercise and sleep, but not a causal claim because exercise was not randomly assigned.
54934811
Two trials evaluate the same headache treatment. Trial A is double-blind and placebo-controlled. Trial B tells every participant and evaluator which treatment was received, and Trial B reports a much larger effect. Explain why Trial A gives the more credible estimate and why blinding still does not guarantee an unbiased study.

Hints

- Consider how treatment knowledge can influence both participants and evaluators. - Identify exactly which information blinding hides. - No single design feature controls every source of bias.

Solution

1. In Trial B, participants' expectations can change reported symptoms, and evaluators' expectations can influence measurement or interaction. 2. Double blinding in Trial A reduces both expectation effects tied to treatment knowledge and evaluator measurement bias. 3. Blinding does not eliminate other problems such as attrition, unrepresentative recruitment, poor adherence, or faulty outcome measures.

Answer

Trial A is more credible because double blinding reduces participant-expectation and evaluator-measurement effects. Blinding does not guarantee an unbiased result because attrition, recruitment, adherence, and measurement problems can remain.
54934911
A delivery company analyzes \(50{,}000\) shipments. A new routing rule lowers mean delivery time by \(0.2\) minute, and the result has a simulated tail probability below \(0.001\). Before seeing the data, the company decided that a reduction smaller than \(2\) minutes would not justify implementation costs. Should the company call the result practically significant under its stated rule? Explain why a result can be statistically strong yet practically unimportant.

Hints

- Compare the observed effect size with the decision threshold chosen before the analysis. - Keep evidence that an effect is nonzero separate from evidence that it is large enough to matter. - Sample size affects statistical precision, not the practical threshold itself.

Solution

1. The result is not practically significant because the observed reduction of \(0.2\) minute is far below the prespecified \(2\)-minute threshold. 2. With a very large sample, a small effect can be estimated precisely enough to produce strong statistical evidence against a no-effect model even when the effect is too small to matter operationally.

Answer

No. The \(0.2\)-minute reduction is below the \(2\)-minute practical threshold. A very large sample can make a tiny effect statistically detectable without making it practically important.
54935211
A linear model relating outdoor temperature \(x\) to electricity demand \(y\) was fitted using days with temperatures from \(40^\circ\text{F}\) to \(80^\circ\text{F}\). A report uses the line to predict demand at \(105^\circ\text{F}\). Explain why this prediction is not well supported by the fitted data.

Hints

- Compare the requested input with the range of inputs used to fit the model. - A good fit inside one range does not determine behavior far outside it. - Ask what evidence the data actually provide near \(105^\circ\text{F}\).

Solution

1. The prediction at \(105^\circ\text{F}\) is a far extrapolation beyond the observed range of \(40^\circ\text{F}\) to \(80^\circ\text{F}\). 2. A relationship that is approximately linear within the fitted range can bend, level off, or change under more extreme conditions, so the data do not establish that the fitted line remains valid at \(105^\circ\text{F}\).

Answer

The prediction is an unsupported extrapolation far beyond the data range; a linear pattern from \(40^\circ\text{F}\) to \(80^\circ\text{F}\) need not continue to \(105^\circ\text{F}\).
54935711
Across counties, those with more public-library visits per resident tend to have higher average reading scores. A columnist concludes that any individual who visits a library more often will have a higher reading score. Explain the statistical error in moving from the county-level association to that individual claim.

Hints

- Compare the unit represented by each data point with the unit named in the conclusion. - Group averages can behave differently from relationships within groups. - Do not infer an individual relationship merely from an aggregate association.

Solution

1. The evidence compares county-level averages, while the conclusion concerns individual people. 2. Group averages do not show that the same relationship holds within counties or for individual residents. County characteristics such as income or school funding can also affect both averages. 3. Therefore, the aggregate association cannot by itself justify the individual-level claim.

Answer

The conclusion commits an aggregate-to-individual inference error: a relationship between county averages does not establish the same relationship for individual residents.
54936011
A cross-sectional survey finds that teenagers with more daily screen time report higher anxiety. A headline says, “Screen time causes anxiety.” Give two distinct noncausal explanations for the association: one involving reverse direction and one involving a third variable.

Hints

- Ask whether the proposed outcome could also influence the proposed cause. - Then look for a third factor connected to both measured variables. - The two explanations should be statistically different mechanisms.

Solution

1. Reverse causation is plausible: teenagers experiencing anxiety may turn to screens more often, so anxiety could partly influence screen time. 2. A confounder such as sleep disruption, social isolation, family stress, or school workload could affect both screen use and anxiety. 3. Because the survey is observational and cross-sectional, these alternatives prevent the association from establishing a causal direction.

Answer

For example, anxiety could lead to more screen use, and a third variable such as sleep disruption could increase both screen time and anxiety. The survey therefore does not establish that screen time causes anxiety.
54936111
A company uses emails sent per day as its only measure of employee productivity. A model predicts email count well from hours online, so management claims it predicts productivity well. Explain the measurement flaw and describe a better evaluation approach.

Hints

- Separate what the model predicts from what management actually wants to measure. - Think of cases where the proxy and the desired outcome move in opposite directions. - Broad constructs usually require several valid indicators.

Solution

1. The model predicts email count, but email count is an unvalidated proxy for the broader construct of productivity. 2. High counts can reflect fragmented or unnecessary communication, while valuable focused work may generate few emails. 3. A better evaluation uses several role-appropriate outcomes, such as completed-work quality, timeliness, customer results, and peer or supervisor assessment.

Answer

Email count is not a validated measure of productivity: many emails can reflect inefficiency, while productive work may require few messages. Use multiple job-relevant quality and outcome measures instead.
54936211
A degree-9 polynomial has zero training error on \(10\) data points but average error \(18\) units on \(5\) new test points from the same process. A quadratic model has training error \(3\) units and test error \(4\) units. Which model should be preferred for prediction, and what modeling problem does the degree-9 polynomial demonstrate?

Hints

- Judge predictive performance using data that were not used to fit the models. - Compare training performance with test performance for the highly flexible model. - Perfect fit to training observations can include accidental noise.

Solution

1. The quadratic model should be preferred because its test error of \(4\) units is much smaller than the degree-9 model's \(18\) units. 2. The degree-9 polynomial fits the training data perfectly but generalizes poorly to new data, which is overfitting.

Answer

Prefer the quadratic model. The degree-9 polynomial demonstrates overfitting: perfect training fit but much worse performance on new data.
54937011
In a randomized experiment with only \(20\) participants, the treatment group's baseline score mean is \(70\), while the control group's baseline mean is \(62\). A critic says randomization must have failed because the baseline means are not equal. Explain why the unequal means do not by themselves show that the random assignment was invalid.

Hints

- Randomization creates a probability mechanism, not guaranteed equality of every group summary. - Consider how much chance imbalance is possible with a small sample. - Separate an unusual realized imbalance from evidence that the assignment process itself was invalid.

Solution

1. Random assignment balances groups in expectation over repeated experiments; it does not guarantee identical baseline summaries in every realized sample. 2. With only \(20\) participants, chance imbalance can be substantial. 3. The difference should be reported and considered in analysis, but the imbalance alone does not prove manipulation or failure of the randomization procedure.

Answer

Valid random assignment can still produce unequal baseline means by chance, especially in a small sample. The imbalance alone does not prove that randomization failed.
54937411
A travel-time model is trained on trips from January through March. A separate held-out test set from April contains only trips made on clear weekdays. The model will be used every day, including weekends and periods of rain or snow. Explain why the held-out accuracy can still be a poor estimate of performance in actual use.

Hints

- The test set is separate from training, so focus on a different requirement for useful evaluation. - Compare the conditions represented in the test set with the conditions expected in deployment. - A held-out set can still be biased for the intended use population.

Solution

1. Although the April observations are genuinely held out from training, the test set does not represent the full conditions under which the model will be deployed. 2. Clear weekdays may have more regular traffic and fewer unusual delays than weekends or periods of rain or snow. 3. Therefore, the held-out accuracy can be systematically optimistic for the intended deployment population.

Answer

The test set is independent of training but unrepresentative of deployment. Restricting evaluation to clear weekdays can make accuracy look better than it will be on weekends or in adverse weather.
54937511
Five teams study the same treatment. Two favorable studies with tail probabilities below \(0.05\) are published, while three studies with no clear effect remain unpublished. A review summarizes only the published studies. Identify the bias and explain what evidence a stronger review should seek.

Hints

- Ask what determined whether each completed study became visible. - Selection can operate on whole studies, not only on participants. - Search for eligible evidence before filtering by its result.

Solution

1. The review has publication bias because study visibility depends on whether results are favorable or statistically significant. 2. Omitting null studies makes the evidence appear more consistently positive than the full set of completed studies. 3. A stronger review should seek registrations, unpublished results, completed protocols, and all prespecified outcomes from every eligible study.

Answer

The review has publication bias. It should include registered, unpublished, and null studies, along with all prespecified outcomes, rather than selecting studies according to their results.
54938111
A charity's total annual donations rise from \(\$20{,}000\) to \(\$40{,}000\). The number of donors rises from \(100\) to \(400\). A headline says, “Donor generosity doubled.” a) Find the mean donation in each year. b) Evaluate the headline. c) Give a more accurate summary of what changed.

Hints

- Separate a total from a per-person average. - Use the relevant denominator for the word “generosity.” - Several components can change in opposite directions while the total rises.

Solution

1. The first-year mean is \(\frac{20000}{100}=\$200\) per donor. 2. The second-year mean is \(\frac{40000}{400}=\$100\) per donor. 3. Total donations doubled, but the average donation per donor fell by \(\$100\), or \(50\%\), so the headline's claim about individual generosity is unsupported. 4. A more accurate summary is that many more donors produced a higher total despite a lower average gift.

Answer

a) \(\$200\) and \(\$100\) per donor. b) The headline is misleading; average giving fell by \(50\%\). c) Total donations doubled because the donor count quadrupled, while the mean gift was cut in half.
54938411
A model estimates that the mean daily demand next month will be between \(490\) and \(510\) units with \(95\%\) confidence. Historical day-to-day demand varies widely, and a prediction interval for one future day is \((420,580)\). Which interval should a manager use when planning staffing for one specific future day, and why?

Hints

- Identify whether the operational decision concerns an average or one future observation. - Individual outcomes vary around the mean even if the mean were known precisely. - Match the interval's target to the quantity the manager needs to plan for.

Solution

1. The confidence interval \((490,510)\) estimates uncertainty in the mean daily demand, not the variability of an individual future day. 2. The prediction interval \((420,580)\) includes both uncertainty about the mean and natural day-to-day variation. 3. Staffing one specific future day therefore calls for the prediction interval.

Answer

Use the prediction interval \((420,580)\), because the decision concerns one future day's demand rather than the population or period mean.
54939311
A forecasting model has mean absolute error \(12\) units on a test period. A simple seasonal baseline that predicts the value from the same month one year earlier has mean absolute error \(8\) units. The model's developers emphasize that its \(R^2\) is \(0.90\) and call it highly accurate. Which method is better for the stated prediction goal, and why does the high \(R^2\) not overturn that comparison?

Hints

- Use the evaluation measure tied directly to the stated prediction goal. - Compare both methods on the same held-out period. - A goodness-of-fit summary and out-of-sample prediction error do not answer the same question.

Solution

1. The seasonal baseline is better on the stated prediction criterion because its mean absolute error is \(8\) units rather than \(12\). 2. Test-period prediction error directly measures performance on the goal being evaluated, while \(R^2\) is a different fit summary. 3. A high \(R^2\) therefore does not compensate for worse out-of-sample prediction error than a simple baseline.

Answer

The seasonal baseline predicts better because its test mean absolute error is \(8\) rather than \(12\) units. The model's high \(R^2\) is a different fit measure and does not override worse test error.
55015411
Study A selects a simple random sample of students and records their sleep time and grades without assigning a treatment. Study B recruits volunteers and randomly assigns them to two study schedules. a) Which study can support a population-wide association for the population represented by its sampling frame? b) Which study can support a causal comparison for its participants? c) Explain why neither study automatically supports both conclusions.

Hints

- Use the method of selecting participants to judge generalization. - Use the method of assigning conditions to judge causation. - Do not treat random sampling and random assignment as interchangeable.

Solution

1. Study A can support a population-wide association because it uses a simple random sample, but it cannot establish causation because sleep time was not assigned. 2. Study B can support a causal comparison for its participants because the schedules were randomly assigned, but volunteers may not represent a larger student population. 3. Random sampling supports generalization, while random assignment supports causal comparison. Each study includes only one of these design features.

Answer

a) Study A. b) Study B. c) Random sampling supports generalization; random assignment supports causation. Neither study has both features.
55015511
The residual plot comes from a fitted linear model. Does the plot show clear evidence that the linear model has missed a curved relationship? Explain using the residual pattern.
Figure for problem 550155

Hints

- Look for a systematic shape in the residuals rather than expecting every point to lie on zero. - Compare the display with an arch or other repeated curved pattern. - Random-looking scatter above and below zero does not by itself indicate missed curvature.

Solution

1. The residuals are scattered above and below zero without a clear arch, curve, or systematic trend as \(x\) increases. 2. Therefore, this residual plot does not show clear evidence of missing curvature and is consistent with a linear relationship.

Answer

No. The residuals show no clear systematic curved pattern around zero, so this plot is consistent with a linear model.
55015611
The following steps describe a randomization test, but they are out of order. A) Count how many simulated differences are at least as extreme as the observed difference. B) Randomly shuffle the treatment labels and calculate a simulated difference. C) Calculate the observed difference between the treatment groups. D) Repeat the shuffle-and-calculate process many times. a) Put the steps in a logical order. b) State how the simulated tail proportion is calculated.

Hints

- The observed statistic must be known before simulated results can be compared with it. - One shuffled assignment produces one value in the no-effect distribution. - The final proportion uses extreme simulated trials as the numerator.

Solution

1. First calculate the observed difference, so C comes first. 2. Shuffle the treatment labels and calculate one simulated difference, so B comes next. 3. Repeat that process many times, so D follows. 4. Finally count the simulated differences at least as extreme as the observed difference, so A is last. 5. The tail proportion is the number of simulated differences at least as extreme as the observed difference divided by the total number of shuffled trials.

Answer

a) C, B, D, A. b) Divide the number of simulated results at least as extreme as the observed result by the total number of randomization trials.
54934611
A school recruits \(120\) volunteers for a study of a new study-planning app. The volunteers are randomly assigned to use the app or not use it. The app group scores an average of \(3.2\) points higher, and a randomization simulation gives a tail proportion of \(0.010\). a) What does random assignment allow the researchers to conclude? b) What does the volunteer sample prevent them from concluding? c) Interpret the simulated tail proportion.

Hints

- Treat random assignment and random sampling as serving different inferential purposes. - Use the tail proportion to judge compatibility with a no-effect model. - Match the scope of the conclusion to how participants entered the study.

Solution

1. Random assignment makes the treatment groups comparable on average, so the small tail proportion supports a causal conclusion about the app for the participating students in this experiment. 2. Because participants volunteered rather than being randomly sampled, the result cannot automatically be generalized to all students at the school. 3. Under a no-effect randomization model, only about \(1.0\%\) of simulated assignments produced a difference at least as favorable as \(3.2\) points.

Answer

a) The study supports a causal effect of the app for the participating students in this experiment. b) It does not justify generalizing the effect to every student at the school. c) A difference at least this large occurred in about \(1\%\) of no-effect simulations.
54935111
Researchers test a treatment on \(20\) unrelated outcomes and highlight the only outcome with a simulated tail probability below \(0.05\), which is \(0.030\). Under a model where the treatment affects none of the outcomes, assume each test has an independent \(0.05\) false-flag probability. Compute the probability of at least one false flag among the \(20\) tests and use it to explain why the highlighted \(0.030\) should be treated cautiously.

Hints

- It is easier to compute the complement event in which none of the tests produces a false flag. - Then compare the resulting family-level probability with the single-test rate. - The analysts had twenty opportunities to find an apparently unusual result.

Solution

1. The probability of no false flags is \((0.95)^{20}\approx0.3585\). 2. Therefore, the probability of at least one false flag is \(1-0.3585\approx0.6415\). 3. With about a \(64\%\) chance of at least one false flag under the no-effect model, selecting the one favorable result after twenty searches provides much weaker evidence than a single prespecified test would.

Answer

\(1-(0.95)^{20}\approx0.6415\). Because at least one false flag is common across twenty independent tests, the selected tail probability \(0.030\) is weak evidence unless the multiple search is accounted for.
54935311
The residual plot shows residuals from a fitted linear model, ordered by increasing \(x\). a) Describe the residual pattern. b) Is a linear model appropriate? Explain. c) What kind of model feature should be considered instead?
Figure for problem 549353

Hints

- Look at the signs and sizes in sequence rather than averaging them. - A useful residual plot should not leave a visible structure unexplained. - Match the shape of the residual pattern to a missing feature in the model.

Solution

1. The residuals begin negative, become positive in the middle, and return to negative, forming a systematic curved pattern. 2. A linear model is not appropriate because residuals should fluctuate without a clear pattern around zero; the model underpredicts in the middle and overpredicts at both ends. 3. A curved relationship, such as a quadratic model, should be considered and checked against the data.

Answer

a) The residuals form an arch-shaped pattern. b) No. Their systematic pattern shows that the line misses curvature. c) Consider a curved model, such as a quadratic, and recheck residuals.
54935511
The bar chart compares support for two options. A commentator says Option B has “five times as much support” because its visible bar is five times taller above the baseline shown. a) Read the actual percentages. b) Find the difference in percentage points and the relative increase from A to B. c) Explain why the commentator's comparison is misleading.
Figure for problem 549355

Hints

- Use the axis values, not the apparent bar-height ratio. - Calculate both an absolute percentage-point difference and a relative change. - Compare the visible bar heights with the actual percentages they represent.

Solution

1. Option A has \(48\%\) support and Option B has \(52\%\). 2. The difference is \(52-48=4\) percentage points. The relative increase is \(\frac{52-48}{48}\approx0.0833=8.33\%\). 3. The y-axis begins at \(47\%\) rather than \(0\%\), so the visible bar heights are \(48-47=1\) and \(52-47=5\). The second visible height is five times the first, but that comparison uses distances above an arbitrary truncated baseline, not the support percentages.

Answer

a) A: \(48\%\); B: \(52\%\). b) Difference: \(4\) percentage points; relative increase: about \(8.33\%\). c) The visible heights are \(1\) and \(5\) percentage points above a truncated \(47\%\) baseline. The fivefold visual comparison is not a comparison of the actual support percentages.
54936311
A coach selects the \(20\) players with the lowest free-throw percentages after one unusually poor week and gives them a new warm-up routine. Their average rises the next week, but there is no comparison group. Explain how regression toward the mean could produce the increase and describe a stronger design.

Hints

- Notice that selection used an extreme first measurement. - Ask what may happen on a second measurement without any treatment. - A comparison group should experience the same timing and selection process.

Solution

1. The players were selected for extreme low results that partly reflect temporary random variation. 2. On a later week, the unusually poor variation is unlikely to repeat as strongly, so percentages can move closer to typical levels even if the routine has no effect. 3. A stronger design randomly assigns similarly low-performing players to the new routine or a comparison routine and compares later changes.

Answer

The selected extreme lows can be followed by less extreme results simply because temporary poor variation does not repeat; this is regression toward the mean. Randomly assign comparable low-performing players to treatment and control routines to test the effect.
54936411
Among \(1000\) transactions, \(50\) are fraudulent and \(950\) are legitimate. A model labels every transaction “legitimate” and is advertised as \(95\%\) accurate. a) Verify the accuracy. b) Find the percentage of fraudulent transactions the model detects. c) Explain why accuracy alone is a misleading performance measure here. d) Name one additional performance quantity that should be reported.

Hints

- Count correct predictions separately for each actual class. - Examine performance on the rare outcome rather than only the total. - A useful metric should reflect the kind of error that matters for the task.

Solution

1. The model correctly labels all \(950\) legitimate transactions, so accuracy is \(\frac{950}{1000}=95\%\). 2. It detects \(0\) of \(50\) fraudulent transactions, so its fraud-detection rate is \(0\%\). 3. The classes are highly imbalanced, so a model can obtain high overall accuracy by always predicting the common class while failing completely on the important rare class. 4. The report should include the fraud-detection rate, false-alarm rate, or a full confusion matrix.

Answer

a) \(95\%\). b) \(0\%\). c) The common legitimate class dominates the accuracy calculation and hides total failure on fraud. d) Report fraud-detection rate and false-alarm rate, or the full confusion matrix.
54936611
A screening model is tested on \(100\) positive cases and \(900\) negative cases. <table><thead><tr><th>Threshold</th><th>Positive cases flagged</th><th>Negative cases flagged</th></tr></thead><tbody><tr><td>Low</td><td>\(90\)</td><td>\(180\)</td></tr><tr><td>High</td><td>\(65\)</td><td>\(45\)</td></tr></tbody></table> a) Find the detection rate and false-alarm rate for each threshold. b) Describe the tradeoff when the threshold is raised. c) Explain why the “best” threshold depends on context.

Hints

- Use actual positives as the denominator for detection and actual negatives for false alarms. - Compare both rates before judging a threshold. - A threshold decision reflects the costs of different mistakes, not only mathematical accuracy.

Solution

1. At the low threshold, the detection rate is \(\frac{90}{100}=90\%\), and the false-alarm rate is \(\frac{180}{900}=20\%\). 2. At the high threshold, the detection rate is \(\frac{65}{100}=65\%\), and the false-alarm rate is \(\frac{45}{900}=5\%\). 3. Raising the threshold reduces false alarms but misses more positive cases. 4. The best threshold depends on the relative costs of missed positives, false alarms, follow-up resources, and the intended use.

Answer

a) Low: \(90\%\) detection and \(20\%\) false alarms. High: \(65\%\) detection and \(5\%\) false alarms. b) The higher threshold lowers both detections and false alarms. c) The choice depends on the consequences and costs of the two error types.
54936711
An employee salary survey receives answers from \(800\) of \(1000\) selected employees. The respondents' mean salary is \(\$62{,}000\). Later payroll records show that the \(200\) nonrespondents have a mean salary of \(\$120{,}000\). a) Find the mean salary for all \(1000\) selected employees. b) Quantify the bias in the respondent-only mean relative to the full selected sample. c) Explain what the missing-data pattern suggests.

Hints

- Combine group totals rather than averaging the two group means equally. - Weight each mean by its group size. - Ask whether response status is associated with the measured variable.

Solution

1. The combined salary total is \(800\cdot62{,}000+200\cdot120{,}000=73{,}600{,}000\) dollars. 2. The full-sample mean is \(\frac{73600000}{1000}=\$73{,}600\). 3. The respondent-only mean is \(73{,}600-62{,}000=\$11{,}600\) too low. 4. Nonresponse is related to salary, so treating the observed salaries as representative of all selected employees is inappropriate.

Answer

a) \(\$73{,}600\). b) The respondent-only mean underestimates by \(\$11{,}600\). c) The missingness is related to the outcome, producing nonresponse bias.
54936811
A randomized trial finds no clear overall treatment effect. After seeing the data, analysts examine \(12\) demographic subgroups and report that the treatment appears effective in one subgroup with a tail probability of \(0.04\). Explain why this subgroup result should be treated as exploratory rather than as strong confirmatory evidence.

Hints

- Count how many subgroup opportunities were examined before one was highlighted. - Ask whether the subgroup was chosen before or after seeing the outcomes. - Post hoc findings require stronger confirmation than prespecified tests.

Solution

1. The subgroup was selected after the analysts inspected many possible subgroup results, creating multiple opportunities for a chance finding. 2. A tail probability of \(0.04\) is therefore less persuasive than it would be for one prespecified subgroup analysis. 3. The result is better treated as a hypothesis to be tested in a planned, independent replication.

Answer

Because the subgroup was found after searching \(12\) groups, the result may be a chance selection from multiple comparisons. It should be treated as exploratory and confirmed with a prespecified independent analysis.
54936911
The boxplots compare delivery times for two services. Both have a mean of \(30\) minutes. A report concludes that the services are equally reliable because their means match. a) Compare the medians and interquartile ranges. b) Which service is more consistent? c) Explain why equal means do not imply equal reliability.
Figure for problem 549369

Hints

- Compare both centers and spreads shown by the boxes. - Consistency is reflected by how tightly outcomes cluster. - One summary statistic cannot determine an entire distribution.

Solution

1. Service A has median \(30\) and interquartile range \(32-28=4\) minutes. Service B has median \(25\) and interquartile range \(45-15=30\) minutes. 2. Service A is more consistent because its central half and full whisker range are much narrower. 3. Equal means describe only one measure of center; Service B's broad, skewed distribution can have the same mean while producing far less predictable delivery times.

Answer

a) A: median \(30\), IQR \(4\). B: median \(25\), IQR \(30\). b) Service A. c) Reliability depends on variability and distribution shape, not only the mean.
54937211
A classifier is evaluated on two groups. <table><thead><tr><th>Group</th><th>Actual positives</th><th>Positives correctly flagged</th><th>Actual negatives</th><th>Negatives incorrectly flagged</th></tr></thead><tbody><tr><td>A</td><td>\(100\)</td><td>\(90\)</td><td>\(100\)</td><td>\(10\)</td></tr><tr><td>B</td><td>\(20\)</td><td>\(10\)</td><td>\(180\)</td><td>\(10\)</td></tr></tbody></table> a) Find the overall accuracy for each group. b) Find the missed-positive rate for each group. c) Explain why equal accuracy does not imply equal performance across groups.

Hints

- First recover correctly classified negatives from the false-positive counts. - Use actual positives as the denominator for missed-positive rate. - A shared overall percentage can be built from very different kinds of errors.

Solution

1. Group A has \(90\) correct positives and \(90\) correct negatives, so accuracy is \(\frac{180}{200}=90\%\). 2. Group B has \(10\) correct positives and \(170\) correct negatives, so accuracy is also \(\frac{180}{200}=90\%\). 3. Group A misses \(10\) of \(100\) positives, a \(10\%\) missed-positive rate. Group B misses \(10\) of \(20\), a \(50\%\) rate. 4. Different class proportions allow the same total accuracy to hide a much worse error rate on positive cases in Group B.

Answer

a) Both groups have \(90\%\) accuracy. b) Group A: \(10\%\); Group B: \(50\%\). c) Accuracy combines error types and is affected by group class proportions, so it can conceal unequal missed-positive rates.
54937611
A treatment has different estimated effects in two population groups. Group 1 makes up \(80\%\) of the population and has an estimated effect of \(+10\) units. Group 2 makes up \(20\%\) and has an estimated effect of \(-5\) units. a) Find the population-weighted average effect. b) Explain why reporting only the average effect can be misleading. c) What additional result should accompany the average?

Hints

- Weight each group effect by its share of the population. - An average can combine effects with opposite signs. - Decision-making may require knowing who benefits and who may be harmed.

Solution

1. The weighted average effect is \(0.80\cdot10+0.20\cdot(-5)=8-1=7\) units. 2. The positive average hides that the treatment may be harmful for Group 2, so the same conclusion does not apply uniformly to every subgroup. 3. The report should include group-specific effect estimates, uncertainty, group sizes, and evidence about whether the difference between subgroup effects is reliable.

Answer

a) \(+7\) units. b) The average conceals a negative estimated effect in Group 2. c) Report subgroup effects with uncertainty and evaluate the evidence for effect differences.
54937911
A model predicts whether a student will pass a final exam. One predictor is “course completed,” a status entered only after final grades are posted. The model has \(99\%\) test accuracy on a randomly split historical data set. Explain why the \(99\%\) test accuracy is not credible evidence of real prediction performance.

Hints

- Ask when each predictor becomes available relative to the prediction decision. - A held-out set is useful only when it contains the same information that would be available in real use. - Random splitting does not remove a predictor that already reveals future information.

Solution

1. “Course completed” is recorded only after the final outcome is known, so it leaks future or outcome information into the predictor set. 2. A random train-test split does not solve the problem because the leaked variable appears in both sets. 3. In real use before the exam, that information would be unavailable, so the reported test setting does not reproduce the intended prediction task.

Answer

The model uses future-information leakage. Because the leaked predictor appears in both random splits, the \(99\%\) test accuracy measures an unrealistic task rather than true pre-exam prediction performance.
54938011
A loan-default model had \(92\%\) accuracy when developed in \(2018\). Using the same threshold on \(2026\) applications, accuracy is \(76\%\), and applicants now have different income patterns and loan types. What statistical problem does this change suggest, and why does it make the \(2018\) validation result unreliable for current use?

Hints

- Compare the population on which performance was validated with the population now being scored. - Look for changes in predictor distributions and operating conditions over time. - A performance estimate is tied to the distribution on which it was evaluated.

Solution

1. The applicant distribution and possibly the relationship between predictors and default outcomes have shifted over time. 2. The \(2018\) validation data therefore no longer represent the \(2026\) deployment population. 3. This distribution shift explains why the old accuracy estimate cannot be assumed to describe current performance.

Answer

The evidence suggests distribution shift. The \(2018\) validation population no longer matches \(2026\) applicants, so the old performance estimate is not reliable for current deployment.
54938311
A treatment reduces the event rate from \(2\%\) to \(1\%\). An advertisement says, “The treatment cuts risk by \(50\%\).” a) Verify the relative-risk reduction. b) Find the absolute risk reduction in percentage points. c) For \(10{,}000\) similar people, estimate the difference in event counts. d) Evaluate whether the advertisement gives a complete picture.

Hints

- Calculate change once relative to the original risk and once directly on the percentage scale. - Translate both rates into counts for a common population size. - A relative statement can sound large when the baseline rate is small.

Solution

1. The relative reduction is \(\frac{2\%-1\%}{2\%}=\frac{1}{2}=50\%\). 2. The absolute reduction is \(2\%-1\%=1\) percentage point. 3. Without treatment, the expected count is \(10{,}000\cdot0.02=200\); with treatment, it is \(10{,}000\cdot0.01=100\), a difference of \(100\) events. 4. The relative claim is mathematically correct but incomplete without the starting risk and absolute change.

Answer

a) \(50\%\) relative reduction. b) \(1\) percentage point absolute reduction. c) About \(100\) fewer events per \(10{,}000\) people. d) The claim is correct but should also report the baseline risk and absolute reduction.
54938611
An app reports that monthly active users rose \(40\%\), from \(600\) in February to \(840\) in March. January had \(1000\) users, February was affected by a two-day outage, and March of the previous year had \(830\) users. a) Verify the advertised February-to-March increase. b) Find the January-to-March change and the year-over-year March change. c) Evaluate the choice of comparison window.

Hints

- Calculate each change relative to its own starting value. - Compare periods that differ in whether an unusual disruption occurred. - A claim can be numerically correct yet misleading because of endpoint selection.

Solution

1. The February-to-March increase is \(\frac{840-600}{600}=0.40=40\%\). 2. The January-to-March change is \(\frac{840-1000}{1000}=-0.16=-16\%\). 3. The year-over-year March change is \(\frac{840-830}{830}\approx0.0120=1.20\%\). 4. The advertised comparison starts from an unusually low outage month and therefore exaggerates ordinary growth; broader and comparable periods give a different picture.

Answer

a) \(40\%\). b) January to March: \(-16\%\); year-over-year March: about \(+1.20\%\). c) The selected outage month creates a misleadingly low baseline and overstates growth.
54938711
The residual plot comes from a fitted linear model. a) What happens to the residual spread as \(x\) increases, and what model feature is questionable? b) How should prediction uncertainty differ across the displayed \(x\)-ranges? c) Why is one constant “plus or minus \(3\)” prediction rule inappropriate?
Figure for problem 549387

Hints

- Compare the center and spread of residuals separately. - Prediction uncertainty should reflect the local size of typical errors. - A single error margin assumes similar variability everywhere.

Solution

1. The residuals remain centered near zero, but their spread increases as \(x\) increases. Constant residual variability is therefore questionable. 2. Predictions should be most precise at low \(x\), less precise in the middle, and least precise at high \(x\). 3. A constant \(\pm3\) rule is wider than needed for many low-\(x\) cases but far too narrow for high-\(x\) cases, so it misrepresents changing uncertainty.

Answer

a) Residual variability increases with \(x\), so constant spread is questionable. b) Prediction intervals should widen as \(x\) increases. c) One fixed margin ignores the strong change in residual spread across the domain.
54938811
A crop experiment estimates that a fertilizer increases yield by \(5\) bushels per acre in dry conditions and by \(20\) bushels per acre in wet conditions. A report averages the results and says the fertilizer “adds \(12.5\) bushels per acre in all conditions.” a) Identify the problem with the report. b) What do the two effects show about how the fertilizer result depends on moisture? c) State a more accurate conclusion.

Hints

- Compare the effect across levels of the environmental condition. - A single effect is questionable when the change depends strongly on another condition. - Context-specific estimates can be more informative than an unweighted average.

Solution

1. The report treats a context-dependent effect as one constant effect and therefore misrepresents both dry and wet conditions. 2. The fertilizer effect changes with moisture conditions rather than remaining constant. 3. A more accurate conclusion is that the estimated increase is about \(5\) bushels per acre when dry and \(20\) when wet; any overall average must be weighted by the frequency of those conditions and should not replace the separate effects.

Answer

a) The average is incorrectly applied uniformly to two conditions with very different effects. b) The fertilizer effect depends on whether conditions are dry or wet. c) Report the separate \(5\)- and \(20\)-bushel effects and, if needed, a properly weighted overall average.
54938911
A modeling team fits \(50\) candidate models and repeatedly checks performance on the same test set. It selects the model with the highest test accuracy and reports that accuracy as an unbiased estimate of future performance. Explain why the reported test accuracy is likely optimistic.

Hints

- Track whether information from the supposed test set changed any modeling decision. - Selecting the largest value among many noisy estimates favors positive random fluctuation. - An independent test set must remain untouched until model choices are finished.

Solution

1. Because test-set results influenced which model was selected, the test set became part of the model-development process rather than remaining an independent final evaluation. 2. Choosing the maximum among many noisy test accuracies favors models that benefited from positive random fluctuation on that particular test set. 3. The selected accuracy is therefore likely biased upward relative to performance on genuinely new data.

Answer

Repeated test-set reuse makes the test data part of model selection. Choosing the best of \(50\) noisy test results favors an unusually high value, so the reported accuracy is likely optimistic.
54935011
The histogram shows a randomization distribution of mean differences under a no-effect model. The observed difference is \(8.0\) units. Of \(1000\) simulated differences, \(12\) were at least \(8.0\). a) Estimate the one-sided tail probability. b) State what the tail probability means in this study. c) Does it equal the probability that the no-effect model is true? Explain.
Figure for problem 549350

Hints

- Use extreme simulated outcomes divided by all simulated outcomes. - Keep the conditioning direction clear: the model is assumed when the probability is calculated. - Distinguish probability of data under a model from probability of a model given data.

Solution

1. The estimated tail probability is \(\frac{12}{1000}=0.012\). 2. If the no-effect model and randomization process were appropriate, about \(1.2\%\) of assignments would produce a mean difference at least as large as \(8.0\). 3. The tail probability is calculated assuming the no-effect model; it is not the probability that the model itself is true.

Answer

a) \(0.012\). b) Under the no-effect model, a result at least this large occurs in about \(1.2\%\) of randomizations. c) No. It is a probability of simulated data under the model, not a probability assigned to the model.
54935411
A university reports admission results for two departments. <table><thead><tr><th>Department</th><th>Women admitted</th><th>Women applicants</th><th>Men admitted</th><th>Men applicants</th></tr></thead><tbody><tr><td>A</td><td>\(80\)</td><td>\(100\)</td><td>\(18\)</td><td>\(20\)</td></tr><tr><td>B</td><td>\(2\)</td><td>\(20\)</td><td>\(60\)</td><td>\(100\)</td></tr></tbody></table> a) Compare admission rates within each department. b) Compare the overall admission rates. c) Explain how the overall comparison reverses the within-department comparisons.

Hints

- Compute rates using applicants within the same department as denominators. - Then recompute after pooling each gender across departments. - Compare how each group is weighted toward the high-rate and low-rate departments.

Solution

1. In Department A, the rates are \(\frac{80}{100}=80\%\) for women and \(\frac{18}{20}=90\%\) for men. In Department B, they are \(\frac{2}{20}=10\%\) for women and \(\frac{60}{100}=60\%\) for men. 2. Overall, women have \(\frac{82}{120}\approx68.3\%\), while men have \(\frac{78}{120}=65\%\). 3. More women applied to high-admission Department A, while more men applied to low-admission Department B. Different group weights reverse the aggregated comparison.

Answer

a) Men have higher admission rates in both departments: \(90\%\) versus \(80\%\) in A and \(60\%\) versus \(10\%\) in B. b) Overall, women have about \(68.3\%\) admitted and men \(65\%\). c) The applicant groups are distributed very differently across departments with very different base admission rates.
54935611
Participants are randomly assigned to an exercise program or a control group. By the end, \(30\%\) of the exercise group and \(5\%\) of the control group have dropped out. Among completers, the exercise group has a lower mean blood pressure. a) Why does random assignment at the start not completely protect the completer-only comparison? b) Give a mechanism that could bias the observed effect. c) State one analysis or reporting step that would improve evaluation of the result.

Hints

- Random assignment balances groups before later events occur. - Ask who is missing from each final comparison and why. - A trustworthy report should show how different reasonable assumptions about missing outcomes affect the conclusion.

Solution

1. Differential dropout can destroy the comparability created by random assignment among the participants who remain. 2. If exercise participants with little improvement or difficulty following the program are more likely to leave, the completers can make the program look more effective than it was for everyone assigned. 3. The report should compare dropout reasons and baseline traits, include outcomes for all assigned participants when possible, and show how conclusions change under reasonable assumptions about the missing outcomes.

Answer

a) Unequal attrition can make the remaining groups systematically different. b) For example, exercise participants who do not improve may be more likely to drop out. c) Analyze all assigned participants when possible and report how the conclusion changes under reasonable assumptions about missing outcomes.
54935811
From \(2005\) to \(2025\), streaming subscriptions and average college-textbook prices both increased. A regression between the annual series has \(R^2=0.96\), and a report claims streaming subscriptions caused textbook prices to rise. Explain why the claim is unsupported and identify a more meaningful modeling step.

Hints

- Separate goodness of fit from evidence about cause. - Ask what third variable changes systematically with both annual series. - Account for the shared time trend before interpreting the relationship.

Solution

1. \(R^2\) describes the strength of a linear association; it does not establish a causal mechanism or rule out confounding. 2. Both variables have strong upward time trends, so later years are high in both series even if the variables are unrelated. 3. A better analysis would account for time trend and inflation, compare year-to-year changes after adjusting for inflation and use predictors connected to a plausible textbook-price mechanism.

Answer

The high \(R^2\) can be caused by a shared upward time trend and does not establish causation. A more meaningful analysis should account for the time trend and inflation, then evaluate variables tied to a plausible pricing mechanism.
54935911
Study 1 estimates an effect of \(5\) units with a \(95\%\) confidence interval of \((2, 8)\). A larger replication estimates an effect of \(-1\) unit with interval \((-4, 2)\). A commentator says the second study proves the first study was fraudulent. a) Explain why that conclusion is not justified by the intervals. b) Compare the ranges of plausible effects. c) State two reasonable next steps for evaluating the disagreement.

Hints

- An estimate from one sample is not expected to match another exactly. - Compare both the centers and the full plausible ranges. - Investigate methodological differences before making claims about conduct.

Solution

1. Different random samples and study conditions can produce different estimates; disagreement does not by itself establish misconduct. 2. Study 1 supports positive effects from \(2\) to \(8\), while Study 2 includes negative, zero, and small positive effects from \(-4\) to \(2\). They share the boundary value \(2\) but give substantially different centers. 3. Researchers should compare protocols, populations, measurements, and attrition, and combine or repeat well-designed studies when appropriate.

Answer

a) Sampling variation or design differences can explain disagreement; the intervals alone do not show fraud. b) The first supports a positive effect, while the replication supports values from moderately negative through small positive. c) Audit design differences and conduct or combine additional carefully designed replications.
54936511
A risk model assigns a \(30\%\) event probability to each of \(200\) cases. The event actually occurs in \(90\) of those cases. a) Find the observed event rate. b) Evaluate the model's calibration for this risk group. c) Explain why this calculation alone does not show whether the model ranks high-risk and low-risk cases well.

Hints

- Compare the predicted percentage with the observed relative frequency. - State both the direction and size of the discrepancy. - Probability accuracy and ranking performance are different model properties.

Solution

1. The observed event rate is \(\frac{90}{200}=0.45=45\%\). 2. The observed rate is \(45\%-30\%=15\) percentage points higher than predicted, so the model underpredicts risk for this group and is poorly calibrated here. 3. Calibration compares predicted probabilities with observed frequencies; ranking ability requires comparing whether cases with higher assigned risks actually tend to experience more events than cases with lower assigned risks.

Answer

a) \(45\%\). b) The model underpredicts by \(15\) percentage points for this group. c) Calibration within one risk group does not measure how well the model orders cases across different risk levels.
54937111
An analyst checks an experiment's result every day and stops as soon as the simulated tail probability falls below \(0.05\). Under a true no-effect model, \(142\) of \(1000\) simulated experiments using this rule eventually stop with a “significant” result. a) Estimate the false-positive rate of the stopping rule. b) Compare it with the intended \(5\%\) rate. c) Explain why repeated unplanned checking changes the inference.

Hints

- Use falsely stopped simulations divided by all no-effect simulations. - Compare both the absolute and multiplicative increase over the target rate. - A decision rule includes when the data are examined and when collection stops.

Solution

1. The simulated false-positive rate is \(\frac{142}{1000}=0.142=14.2\%\). 2. This is \(14.2\%-5.0\%=9.2\) percentage points above the intended rate and is about \(\frac{14.2}{5}=2.84\) times as large. 3. Each additional look creates another opportunity for random fluctuation to cross the threshold; the original tail-probability interpretation assumes a fixed analysis plan.

Answer

a) \(14.2\%\). b) It is \(9.2\) percentage points higher, about \(2.84\) times the intended rate. c) Repeated optional checking adds chances to stop on a random extreme result.
54937311
A city introduces a road-safety law. Crashes fall from \(100\) to \(70\) in the city. During the same period, crashes in a similar comparison city without the law fall from \(110\) to \(88\). a) Find the before-after change in each city. b) Use the comparison city's change to form an adjusted difference in changes. c) Explain why the raw decline of \(30\) crashes may overstate the law's effect.

Hints

- Compute each location's change using the same subtraction order. - Subtract the comparison change from the treated change. - Use the untreated location to estimate what might have happened anyway.

Solution

1. The treated city changes by \(70-100=-30\) crashes. The comparison city changes by \(88-110=-22\) crashes. 2. The adjusted difference in changes is \((-30)-(-22)=-8\) crashes. 3. A broad trend reduced crashes even without the law, as shown by the comparison city. The raw decline attributes that shared trend to the law, while the adjusted comparison suggests an additional decline of about \(8\) crashes.

Answer

a) Treated city: \(-30\); comparison city: \(-22\). b) \(-8\) crashes. c) Most of the decline may reflect a trend affecting both cities rather than the law alone.
54937711
A study records \(10\) daily observations from each of \(30\) students and analyzes the \(300\) observations as if they came from \(300\) unrelated students. a) Identify the violated model assumption. b) Explain how the mistake affects estimated uncertainty. c) Give two valid ways to restructure the analysis or design.

Hints

- Identify which observations share the same source. - Repeated measurements do not add as much new information as measurements from new independent units. - Either change the unit of analysis or use a model that represents the grouping.

Solution

1. Observations from the same student are likely correlated, violating the assumption that all \(300\) observations are independent. 2. Treating repeated observations as independent overstates the amount of separate information and usually makes standard errors and intervals too small. 3. The analysis can summarize each student first or use a method that treats observations from the same student as related; the design can also sample more students rather than only more days per student.

Answer

a) Independence is violated within each student's repeated measurements. b) Uncertainty is understated because the effective sample size is less than \(300\). c) Analyze student-level summaries or use a method that accounts for repeated observations from each student, and increase the number of students when possible.
54937811
A data set has \(100\) missing scores. An analyst replaces every missing value with the observed mean of \(75\) and then fits a model as if all values were observed. a) Explain how this replacement affects the data's mean. b) Explain how it affects the data's variability. c) Why can confidence intervals from the completed data be too narrow?

Hints

- Consider what happens when many new values are placed exactly at the center. - A filled-in value is not the same as an observed value. - Valid uncertainty should include uncertainty about what was missing.

Solution

1. Replacing missing values by the observed mean leaves the completed-data mean at \(75\) if the observed mean was \(75\). 2. Every imputed value lies exactly at the center, so the completed data have artificially reduced variability. 3. The method treats guessed values as known and ignores uncertainty about missing outcomes, causing model uncertainty and confidence-interval width to be understated.

Answer

a) It preserves the observed mean of \(75\). b) It artificially reduces variability by adding many central values. c) The analysis ignores uncertainty in the imputed values and therefore appears too precise.
54938211
In a randomized trial, \(100\) people are assigned to a program and \(100\) to control. Only \(60\) assigned to the program actually participate. The mean outcome is \(74\) for everyone assigned to the program and \(70\) for everyone assigned to control. Actual participants have mean \(78\). a) Find the effect estimate based on original random assignment. b) Why is comparing actual participants with the control group potentially biased? c) Which comparison best preserves the benefit of randomization?

Hints

- Keep the groups defined by the random mechanism. - A choice made after assignment can be related to the outcome. - Distinguish the effect of being offered a program from the effect among self-selected users.

Solution

1. The assignment-based effect estimate is \(74-70=4\) units. 2. People who choose to participate may differ in motivation, availability, or baseline condition from nonparticipants and controls, so the \(78-70=8\) comparison mixes treatment with self-selection. 3. Comparing all people according to original assignment preserves randomized group comparability and estimates the effect of offering or assigning the program.

Answer

a) \(4\) units. b) Actual participation is self-selected after assignment, so the \(8\)-unit comparison is confounded. c) Compare all \(200\) people by their original randomized assignment.
54938511
A screening test has a \(90\%\) positive rate among people with a condition and a \(5\%\) positive rate among people without it. In the screened population, \(2\%\) have the condition. a) For \(10{,}000\) people, find the expected numbers of true positives and false positives. b) Among positive tests, estimate the percentage who actually have the condition. c) Explain why “\(90\%\) sensitivity” does not mean a person with a positive result has a \(90\%\) chance of having the condition.

Hints

- Begin by splitting the population according to the base rate. - Apply each conditional positive rate within the appropriate group. - Check which event appears after the conditioning bar in each probability.

Solution

1. Of \(10{,}000\) people, \(200\) have the condition and \(9800\) do not. 2. The expected true positives are \(200\cdot0.90=180\). The expected false positives are \(9800\cdot0.05=490\). 3. The percentage with the condition among all positive tests is \(\frac{180}{180+490}\approx0.2687=26.87\%\). 4. Sensitivity conditions on actually having the condition, while the requested probability conditions on receiving a positive test; the low base rate and false positives matter.

Answer

a) \(180\) true positives and \(490\) false positives. b) About \(26.87\%\). c) Sensitivity is \(P(+\mid C)\), not \(P(C\mid +)\); the population base rate changes the latter probability.
54939011
Students are randomly assigned by classroom to receive a new tutoring program or the usual program. Students in neighboring classrooms share tutoring materials and study together after school. a) Identify the problem this behavior creates for keeping the treatment and control conditions separate. b) Explain how the sharing could affect the estimated treatment difference. c) Suggest one design or analysis change.

Hints

- Ask whether one participant's assigned treatment can change another participant's experience. - Consider what happens when the control group receives part of the active program. - Choose assignment units that reduce contact across treatment conditions.

Solution

1. The study's assumption that one group's treatment does not affect outcomes in the other group is threatened by spillover or interference. 2. Control students may receive part of the tutoring benefit, making the groups more similar and often reducing the estimated difference; other social effects could also complicate direction. 3. Researchers could assign more widely separated schools or groups, measure how much sharing occurs, and report how the results change when cross-group exposure is considered.

Answer

a) Spillover makes the treatment and control conditions overlap. b) Shared materials can contaminate the control condition and usually shrink the observed contrast. c) Assign more separated groups and measure or account for cross-group exposure.
54939111
Fifty matched pairs of participants are formed using similar baseline scores. Within every pair, one participant is randomly assigned to Treatment A and the other to Treatment B. The mean of the \(50\) within-pair differences, A minus B, is \(3.2\) points. The standard deviation of the within-pair differences is \(4\) points, while individual outcome scores have a standard deviation of \(15\) points. An analyst discards the pair labels and compares the two treatment groups as unrelated samples. a) Does discarding pair labels remove the random assignment? Explain. b) Why is the unpaired analysis likely less precise? c) What statistic and analysis should use the design most effectively?

Hints

- Separate how treatment was assigned from how outcomes are later analyzed. - Matching is useful when partners share variation that can cancel in a difference. - Use the same observational unit in the analysis that the randomization created.

Solution

1. The original within-pair random assignment still occurred, so discarding labels does not erase the random treatment mechanism or necessarily change the treatment-group mean difference. 2. Matching made partners similar at baseline, and the within-pair differences vary by only \(4\) points compared with \(15\)-point variation among individual outcomes. Ignoring pairs fails to remove that shared baseline variation and therefore loses precision. 3. The analysis should use the \(50\) within-pair differences, centered at the observed mean difference of \(3.2\) points, with a method designed for paired data.

Answer

a) No. Treatment was still randomly assigned within pairs, although the analysis ignores useful design information. b) Ignoring matching leaves large between-participant variation in the error instead of canceling it within pairs. c) Analyze the \(50\) A-minus-B pair differences, whose mean is \(3.2\) points.
54939211
Two models fit the historical data about equally well. Model A forecasts next year's demand as \(120\) units, while Model B forecasts \(180\) units because they make different assumptions about whether recent growth will continue. A report presents only Model A's forecast. a) Identify the omitted source of uncertainty. b) Why is a narrow interval around \(120\) potentially misleading? c) Describe a more responsible report.

Hints

- Separate uncertainty inside one chosen model from uncertainty about choosing the model itself. - Good historical fit does not make two different future assumptions equivalent. - Show how the conclusion changes under other plausible model assumptions.

Solution

1. The report omits model uncertainty: different plausible model structures and assumptions produce very different forecasts. 2. A narrow interval conditional on Model A can describe random error within that model while ignoring the larger uncertainty about which model is appropriate. 3. A responsible report should show forecasts from both justified models, explain their assumptions, evaluate them on relevant validation data, and show how the forecast changes under the different plausible models.

Answer

a) Uncertainty about which model and assumptions are appropriate. b) It can look precise while ignoring another plausible forecast far from \(120\). c) Report alternative models, assumptions, validation performance, and sensitivity of conclusions to model choice.
54939411
A study recruits \(300\) volunteers from one fitness center and randomly assigns them to an energy drink or a placebo. Fifteen people leave the drink group and two leave the placebo group. Among completers, the drink group has a mean heart rate \(4.0\) beats per minute higher. A label-shuffling simulation applied only to the completers gives a tail proportion of \(0.020\), and a \(95\%\) confidence interval for the completer mean difference is \((0.5,7.5)\) beats per minute. A headline says, “The drink raises every adult's heart rate by exactly \(4\) beats per minute.” a) Evaluate the evidence for a causal effect. b) Evaluate generalization to all adults. c) Explain two problems with the word “exactly” and the claim about every adult. d) Explain how differential dropout affects the interpretation of the label-shuffling result and confidence interval. e) Write a conclusion supported by the study.

Hints

- Evaluate random assignment, sampling, missing outcomes, and uncertainty as separate parts of the design. - Ask whether the set of completers can still be treated as if it arose solely from the original random assignment. - Distinguish a sample average from a population effect and from every individual's response. - A defensible conclusion should preserve both the observed completer difference and the study's design limitations.

Solution

1. Random assignment initially creates comparable treatment groups and supports causal inference if outcomes are retained. The much larger dropout in the drink group can destroy that comparability among completers, so the completer-only difference is not a clean randomized estimate of the causal effect. 2. Volunteers from one fitness center are not a random sample of all adults, so broad population generalization is unsupported. 3. The \(4.0\) value is a sample mean difference, not an exact population effect; the confidence interval shows uncertainty about the completer mean difference. Also, an average difference does not imply that every individual changes by the same amount. 4. Shuffling treatment labels only among the observed completers does not reproduce the original random-assignment mechanism if treatment assignment affected who remained in the study. Therefore, the \(0.020\) tail proportion cannot by itself be treated as a clean randomization-test probability for the original experiment. The confidence interval likewise describes uncertainty for the analyzed completers and does not repair attrition bias. 5. A supported conclusion is that drink-group completers had a higher mean heart rate than placebo completers, with an estimated difference of \(4.0\) beats per minute and completer interval \((0.5,7.5)\), but differential attrition weakens causal interpretation and the volunteer sample does not support generalization to all adults.

Answer

a) Random assignment initially supports causation, but differential attrition means the completer-only comparison is not a clean causal estimate. b) The volunteer sample from one fitness center does not support generalization to all adults. c) \(4.0\) is an estimate rather than an exact population effect, and a mean effect is not an identical effect for every individual. d) Because dropout differs after assignment, shuffling labels only among completers need not represent the original randomization; the \(0.020\) tail proportion and completer confidence interval do not remove attrition bias. e) Drink-group completers had a higher mean heart rate, but differential attrition weakens causal interpretation and generalization to all adults is unsupported.

All problems may be used, copied and printed free of charge for school and tutoring, including paid tutoring. Commercial adaptations as well as publication or redistribution on the internet are not permitted.