Aimathic
Login | English | Deutsch

Free math worksheets

Build your own math worksheets from 30,000 problems for grades 3 to 12, from fractions to calculus. Every problem includes step-by-step solutions.

Evaluate statistical models and inferences

Click problems to add them to your worksheet.

54934711
A simple random sample of \(800\) adults records weekly exercise time and nightly sleep duration. Adults who exercise more also tend to report more sleep. a) What population claim can the random sample support? b) Why does the study not establish that increasing exercise causes more sleep? c) Give one plausible confounding variable.

Hints

- Use the sampling method to decide the population scope. - Use the treatment-assignment method to decide whether causation is supported. - Look for a third variable connected to both measured quantities.

Solution

1. The random sample can support generalizing an association between exercise and sleep to the adult population represented by the sampling frame. 2. Exercise was observed rather than randomly assigned, so differences may be caused by other variables or reverse direction of influence. 3. Work schedule, health status, age, or stress could affect both exercise time and sleep duration.

Answer

a) The study can support a population-level association between exercise and sleep. b) It cannot establish causation because there was no random assignment. c) For example, work schedule or overall health may affect both variables.
54934811
Two trials evaluate the same headache treatment. Trial A is double-blind and placebo-controlled. Trial B tells every participant and evaluator which treatment was received, and Trial B reports a much larger effect. Explain why Trial A gives the more credible estimate and why blinding still does not guarantee an unbiased study.

Hints

- Consider how treatment knowledge can influence both participants and evaluators. - Identify exactly which information blinding hides. - No single design feature controls every source of bias.

Solution

1. In Trial B, participants' expectations can change reported symptoms, and evaluators' expectations can influence measurement or interaction. 2. Double blinding in Trial A reduces both expectation effects tied to treatment knowledge and evaluator measurement bias. 3. Blinding does not eliminate other problems such as attrition, unrepresentative recruitment, poor adherence, or faulty outcome measures.

Answer

Trial A is more credible because double blinding reduces participant-expectation and evaluator-measurement effects. Blinding does not guarantee an unbiased result because attrition, recruitment, adherence, and measurement problems can remain.
54935211
A linear model relating outdoor temperature \(x\) to electricity demand \(y\) was fitted using days with temperatures from \(40^\circ\text{F}\) to \(80^\circ\text{F}\). A report uses the line to predict demand at \(105^\circ\text{F}\). a) Identify the model-use problem. b) Why can a line that fits well from \(40^\circ\text{F}\) to \(80^\circ\text{F}\) still fail at \(105^\circ\text{F}\)? c) What evidence would make the \(105^\circ\text{F}\) prediction more credible?

Hints

- Compare the prediction input with the range used to fit the model. - A strong fit within one domain does not determine behavior outside it. - Credible prediction requires evidence near the intended input values.

Solution

1. The prediction is an extrapolation far outside the observed temperature range. 2. The relationship can bend, level off, change slope, or be affected by different operating conditions beyond the data range even if it is nearly linear within the observed range. 3. Data from temperatures near \(105^\circ\text{F}\) or a validated model with justified behavior over that expanded domain would make the prediction more credible.

Answer

a) It is a far extrapolation beyond the model's data domain. b) A local linear pattern need not continue under more extreme conditions. c) Collect relevant high-temperature data or validate a model over the wider range.
54936111
A company uses emails sent per day as its only measure of employee productivity. A model predicts email count well from hours online, so management claims it predicts productivity well. Explain the measurement flaw and describe a better evaluation approach.

Hints

- Separate what the model predicts from what management actually wants to measure. - Think of cases where the proxy and the desired outcome move in opposite directions. - Broad constructs usually require several valid indicators.

Solution

1. The model predicts email count, but email count is an unvalidated proxy for the broader construct of productivity. 2. High counts can reflect fragmented or unnecessary communication, while valuable focused work may generate few emails. 3. A better evaluation uses several role-appropriate outcomes, such as completed-work quality, timeliness, customer results, and peer or supervisor assessment.

Answer

Email count is not a validated measure of productivity: many emails can reflect inefficiency, while productive work may require few messages. Use multiple job-relevant quality and outcome measures instead.
54937511
Five teams study the same treatment. Two favorable studies with tail probabilities below \(0.05\) are published, while three studies with no clear effect remain unpublished. A review summarizes only the published studies. Identify the bias and explain what evidence a stronger review should seek.

Hints

- Ask what determined whether each completed study became visible. - Selection can operate on whole studies, not only on participants. - Search for eligible evidence before filtering by its result.

Solution

1. The review has publication bias because study visibility depends on whether results are favorable or statistically significant. 2. Omitting null studies makes the evidence appear more consistently positive than the full set of completed studies. 3. A stronger review should seek registrations, unpublished results, completed protocols, and all prespecified outcomes from every eligible study.

Answer

The review has publication bias. It should include registered, unpublished, and null studies, along with all prespecified outcomes, rather than selecting studies according to their results.
54938111
A charity's total annual donations rise from \(\$20{,}000\) to \(\$40{,}000\). The number of donors rises from \(100\) to \(400\). A headline says, “Donor generosity doubled.” a) Find the mean donation in each year. b) Evaluate the headline. c) Give a more accurate summary of what changed.

Hints

- Separate a total from a per-person average. - Use the relevant denominator for the word “generosity.” - Several components can change in opposite directions while the total rises.

Solution

1. The first-year mean is \(\frac{20000}{100}=\$200\) per donor. 2. The second-year mean is \(\frac{40000}{400}=\$100\) per donor. 3. Total donations doubled, but the average donation per donor fell by \(\$100\), or \(50\%\), so the headline's claim about individual generosity is unsupported. 4. A more accurate summary is that many more donors produced a higher total despite a lower average gift.

Answer

a) \(\$200\) and \(\$100\) per donor. b) The headline is misleading; average giving fell by \(50\%\). c) Total donations doubled because the donor count quadrupled, while the mean gift was cut in half.
55015411
Study A selects a simple random sample of students and records their sleep time and grades without assigning a treatment. Study B recruits volunteers and randomly assigns them to two study schedules. a) Which study can support a population-wide association for the population represented by its sampling frame? b) Which study can support a causal comparison for its participants? c) Explain why neither study automatically supports both conclusions.

Hints

- Use the method of selecting participants to judge generalization. - Use the method of assigning conditions to judge causation. - Do not treat random sampling and random assignment as interchangeable.

Solution

1. Study A can support a population-wide association because it uses a simple random sample, but it cannot establish causation because sleep time was not assigned. 2. Study B can support a causal comparison for its participants because the schedules were randomly assigned, but volunteers may not represent a larger student population. 3. Random sampling supports generalization, while random assignment supports causal comparison. Each study includes only one of these design features.

Answer

a) Study A. b) Study B. c) Random sampling supports generalization; random assignment supports causation. Neither study has both features.
55015511
The residual plot comes from a fitted linear model. a) Describe the pattern of the residuals around zero. b) Is there clear evidence in this plot that a linear model misses a curved pattern? c) State one additional feature of the data that should still be checked before using the model for prediction.
Figure for problem 550155

Hints

- Look for a systematic shape rather than expecting every residual to equal zero. - Compare this pattern with an arch, trend, or funnel shape. - A satisfactory residual shape is one model check, not a complete guarantee.

Solution

1. The residuals are scattered above and below zero without a clear curved or steadily increasing pattern. 2. No. This plot does not show clear evidence of missing curvature, so the residual pattern is consistent with a linear model. 3. Other checks include unusual points, changing residual spread, whether predictions stay within the observed x-range, and whether the data-collection design supports the intended use.

Answer

a) The residuals show no clear systematic pattern around zero. b) No; the plot is consistent with a linear relationship. c) For example, check for unusual points, changing spread, or extrapolation beyond the observed range.
55015611
The following steps describe a randomization test, but they are out of order. A) Count how many simulated differences are at least as extreme as the observed difference. B) Randomly shuffle the treatment labels and calculate a simulated difference. C) Calculate the observed difference between the treatment groups. D) Repeat the shuffle-and-calculate process many times. a) Put the steps in a logical order. b) State how the simulated tail proportion is calculated.

Hints

- The observed statistic must be known before simulated results can be compared with it. - One shuffled assignment produces one value in the no-effect distribution. - The final proportion uses extreme simulated trials as the numerator.

Solution

1. First calculate the observed difference, so C comes first. 2. Shuffle the treatment labels and calculate one simulated difference, so B comes next. 3. Repeat that process many times, so D follows. 4. Finally count the simulated differences at least as extreme as the observed difference, so A is last. 5. The tail proportion is the number of simulated differences at least as extreme as the observed difference divided by the total number of shuffled trials.

Answer

a) C, B, D, A. b) Divide the number of simulated results at least as extreme as the observed result by the total number of randomization trials.
54934511
A news article reports that months with higher ice-cream sales also have more drowning incidents and concludes that buying ice cream increases drowning risk. a) Is the reported study evidence observational or experimental? b) Identify a plausible confounding variable. c) Explain why the association does not justify the causal conclusion. d) State one study design that could investigate the relationship more appropriately without deliberately exposing people to danger.

Hints

- Ask whether researchers assigned anyone to buy ice cream. - Look for a third factor that changes both variables in the same months. - A causal claim requires ruling out alternative explanations for the association.

Solution

1. The evidence is observational because no treatment was randomly assigned. 2. Temperature or season is a plausible confounding variable because warm weather can increase both ice-cream sales and swimming exposure. 3. The observed association can be explained by the confounder, so it does not isolate an effect of buying ice cream on drowning risk. 4. A safer design could analyze individual-level observational data while adjusting or comparing within temperature, season, location, and swimming-exposure groups.

Answer

a) Observational. b) Temperature or season. c) A common cause can produce both higher ice-cream sales and more swimming-related exposure, so causation is not established. d) Use a carefully controlled observational analysis that accounts for weather and swimming exposure.
54934611
A school recruits \(120\) volunteers for a study of a new study-planning app. The volunteers are randomly assigned to use the app or not use it. The app group scores an average of \(3.2\) points higher, and a randomization simulation gives a tail proportion of \(0.010\). a) What does random assignment allow the researchers to conclude? b) What does the volunteer sample prevent them from concluding? c) Interpret the simulated tail proportion.

Hints

- Treat random assignment and random sampling as serving different inferential purposes. - Use the tail proportion to judge compatibility with a no-effect model. - Match the scope of the conclusion to how participants entered the study.

Solution

1. Random assignment makes the treatment groups comparable on average, so the small tail proportion supports a causal conclusion about the app for the participating students in this experiment. 2. Because participants volunteered rather than being randomly sampled, the result cannot automatically be generalized to all students at the school. 3. Under a no-effect randomization model, only about \(1.0\%\) of simulated assignments produced a difference at least as favorable as \(3.2\) points.

Answer

a) The study supports a causal effect of the app for the participating students in this experiment. b) It does not justify generalizing the effect to every student at the school. c) A difference at least this large occurred in about \(1\%\) of no-effect simulations.
54934911
A delivery company analyzes \(50{,}000\) shipments. A new routing rule lowers mean delivery time by \(0.2\) minute, and the result is reported as statistically significant with a simulated tail probability below \(0.001\). The company had decided in advance that a reduction smaller than \(2\) minutes would not justify implementation costs. a) Explain why the result can be statistically significant. b) Is the result practically significant under the company's rule? c) What two quantities should the decision report include?

Hints

- Separate evidence that an effect is nonzero from evidence that it is large enough to matter. - Use the threshold chosen before seeing the data. - A complete report should describe magnitude as well as uncertainty.

Solution

1. The very large sample makes the estimate highly precise, so even a small difference can be unlikely under a no-effect model. 2. The reduction of \(0.2\) minute is below the prespecified \(2\)-minute practical threshold, so it is not practically significant for the decision. 3. The report should include both the effect size, \(0.2\) minute, and its uncertainty or simulation-based significance measure.

Answer

a) A huge sample can make a very small effect distinguishable from zero. b) No; \(0.2\) minute is below the \(2\)-minute practical threshold. c) Report the effect size and a measure of its uncertainty or statistical evidence.
54935311
The residual plot shows residuals from a fitted linear model, ordered by increasing \(x\). a) Describe the residual pattern. b) Is a linear model appropriate? Explain. c) What kind of model feature should be considered instead?
Figure for problem 549353

Hints

- Look at the signs and sizes in sequence rather than averaging them. - A useful residual plot should not leave a visible structure unexplained. - Match the shape of the residual pattern to a missing feature in the model.

Solution

1. The residuals begin negative, become positive in the middle, and return to negative, forming a systematic curved pattern. 2. A linear model is not appropriate because residuals should fluctuate without a clear pattern around zero; the model underpredicts in the middle and overpredicts at both ends. 3. A curved relationship, such as a quadratic model, should be considered and checked against the data.

Answer

a) The residuals form an arch-shaped pattern. b) No. Their systematic pattern shows that the line misses curvature. c) Consider a curved model, such as a quadratic, and recheck residuals.
54935511
The bar chart compares support for two options. A commentator says Option B has “five times as much support” because its visible bar is five times taller above the baseline shown. a) Read the actual percentages. b) Find the difference in percentage points and the relative increase from A to B. c) Explain why the commentator's comparison is misleading.
Figure for problem 549355

Hints

- Use the axis values, not the apparent bar-height ratio. - Calculate both an absolute percentage-point difference and a relative change. - Compare the visible bar heights with the actual percentages they represent.

Solution

1. Option A has \(48\%\) support and Option B has \(52\%\). 2. The difference is \(52-48=4\) percentage points. The relative increase is \(\frac{52-48}{48}\approx0.0833=8.33\%\). 3. The y-axis begins at \(47\%\) rather than \(0\%\), so the visible bar heights are \(48-47=1\) and \(52-47=5\). The second visible height is five times the first, but that comparison uses distances above an arbitrary truncated baseline, not the support percentages.

Answer

a) A: \(48\%\); B: \(52\%\). b) Difference: \(4\) percentage points; relative increase: about \(8.33\%\). c) The visible heights are \(1\) and \(5\) percentage points above a truncated \(47\%\) baseline. The fivefold visual comparison is not a comparison of the actual support percentages.
54935711
Across counties, those with more public-library visits per resident tend to have higher average reading scores. A columnist concludes that any individual who visits a library more often will have a higher reading score. a) Identify the error in using county-level data to make an individual claim. b) Give two reasons the county-level association may not hold for individuals. c) What type of data would be needed to evaluate the individual claim?

Hints

- Check whether the evidence and the conclusion refer to the same unit: counties or individuals. - Group averages can conceal very different within-group relationships. - The desired claim requires measurements connected at the individual level.

Solution

1. The conclusion incorrectly transfers an association between county averages to individual people. 2. Counties differ in income, school funding, age distribution, and other features; county averages also do not link each person's library use to that person's score. 3. Individual-level data pairing each person's library use with reading outcomes, along with an appropriate design to address confounding, are needed.

Answer

a) County-level averages do not establish the same relationship for individuals. b) County differences can confound the relationship, and aggregate averages do not pair individual use with individual scores. c) Collect linked individual-level data and account for confounding or use a suitable experiment.
54936011
A cross-sectional survey finds that teenagers with more daily screen time report higher anxiety. A headline says, “Screen time causes anxiety.” a) Give one plausible reverse-causation explanation. b) Give one plausible confounder. c) Explain what the survey can legitimately conclude.

Hints

- Ask whether the proposed outcome could also influence the proposed cause. - Look for a third factor connected to both measured variables. - Match the strength of the conclusion to the observational design.

Solution

1. Teenagers experiencing anxiety may turn to screens more often, so anxiety could partly influence screen time. 2. Social isolation, sleep disruption, family stress, or school workload could affect both screen use and anxiety. 3. The survey supports an association at one point in time but does not determine causal direction or isolate a screen-time effect.

Answer

a) Anxiety may lead some teenagers to use screens more. b) For example, sleep disruption or social isolation could affect both. c) The survey shows an association, not a causal direction.
54936211
A degree-9 polynomial passes exactly through all \(10\) training data points, giving zero training error. On \(5\) new test points from the same process, its average error is \(18\) units. A quadratic model has training error \(3\) units and test error \(4\) units. a) Which model should be preferred for prediction? b) Explain why zero training error can be a warning sign rather than proof of superiority. c) Name the modeling problem shown by the degree-9 polynomial.

Hints

- Judge prediction using data not used to fit the model. - A flexible curve can explain both signal and accidental irregularities. - Compare performance on training and test data separately.

Solution

1. The quadratic should be preferred because its test error of \(4\) units is much smaller than \(18\) units. 2. A highly flexible model can reproduce random noise and peculiarities of the training data that do not repeat in new data. 3. The degree-9 polynomial is overfit: it has excellent training fit but poor generalization.

Answer

a) The quadratic model. b) Perfect training fit may capture noise rather than the underlying relationship. c) Overfitting.
54936311
A coach selects the \(20\) players with the lowest free-throw percentages after one unusually poor week and gives them a new warm-up routine. Their average rises the next week, but there is no comparison group. Explain how regression toward the mean could produce the increase and describe a stronger design.

Hints

- Notice that selection used an extreme first measurement. - Ask what may happen on a second measurement without any treatment. - A comparison group should experience the same timing and selection process.

Solution

1. The players were selected for extreme low results that partly reflect temporary random variation. 2. On a later week, the unusually poor variation is unlikely to repeat as strongly, so percentages can move closer to typical levels even if the routine has no effect. 3. A stronger design randomly assigns similarly low-performing players to the new routine or a comparison routine and compares later changes.

Answer

The selected extreme lows can be followed by less extreme results simply because temporary poor variation does not repeat; this is regression toward the mean. Randomly assign comparable low-performing players to treatment and control routines to test the effect.
54936411
Among \(1000\) transactions, \(50\) are fraudulent and \(950\) are legitimate. A model labels every transaction “legitimate” and is advertised as \(95\%\) accurate. a) Verify the accuracy. b) Find the percentage of fraudulent transactions the model detects. c) Explain why accuracy alone is a misleading performance measure here. d) Name one additional performance quantity that should be reported.

Hints

- Count correct predictions separately for each actual class. - Examine performance on the rare outcome rather than only the total. - A useful metric should reflect the kind of error that matters for the task.

Solution

1. The model correctly labels all \(950\) legitimate transactions, so accuracy is \(\frac{950}{1000}=95\%\). 2. It detects \(0\) of \(50\) fraudulent transactions, so its fraud-detection rate is \(0\%\). 3. The classes are highly imbalanced, so a model can obtain high overall accuracy by always predicting the common class while failing completely on the important rare class. 4. The report should include the fraud-detection rate, false-alarm rate, or a full confusion matrix.

Answer

a) \(95\%\). b) \(0\%\). c) The common legitimate class dominates the accuracy calculation and hides total failure on fraud. d) Report fraud-detection rate and false-alarm rate, or the full confusion matrix.
54936611
A screening model is tested on \(100\) positive cases and \(900\) negative cases. <table><thead><tr><th>Threshold</th><th>Positive cases flagged</th><th>Negative cases flagged</th></tr></thead><tbody><tr><td>Low</td><td>\(90\)</td><td>\(180\)</td></tr><tr><td>High</td><td>\(65\)</td><td>\(45\)</td></tr></tbody></table> a) Find the detection rate and false-alarm rate for each threshold. b) Describe the tradeoff when the threshold is raised. c) Explain why the “best” threshold depends on context.

Hints

- Use actual positives as the denominator for detection and actual negatives for false alarms. - Compare both rates before judging a threshold. - A threshold decision reflects the costs of different mistakes, not only mathematical accuracy.

Solution

1. At the low threshold, the detection rate is \(\frac{90}{100}=90\%\), and the false-alarm rate is \(\frac{180}{900}=20\%\). 2. At the high threshold, the detection rate is \(\frac{65}{100}=65\%\), and the false-alarm rate is \(\frac{45}{900}=5\%\). 3. Raising the threshold reduces false alarms but misses more positive cases. 4. The best threshold depends on the relative costs of missed positives, false alarms, follow-up resources, and the intended use.

Answer

a) Low: \(90\%\) detection and \(20\%\) false alarms. High: \(65\%\) detection and \(5\%\) false alarms. b) The higher threshold lowers both detections and false alarms. c) The choice depends on the consequences and costs of the two error types.
54936711
An employee salary survey receives answers from \(800\) of \(1000\) selected employees. The respondents' mean salary is \(\$62{,}000\). Later payroll records show that the \(200\) nonrespondents have a mean salary of \(\$120{,}000\). a) Find the mean salary for all \(1000\) selected employees. b) Quantify the bias in the respondent-only mean relative to the full selected sample. c) Explain what the missing-data pattern suggests.

Hints

- Combine group totals rather than averaging the two group means equally. - Weight each mean by its group size. - Ask whether response status is associated with the measured variable.

Solution

1. The combined salary total is \(800\cdot62{,}000+200\cdot120{,}000=73{,}600{,}000\) dollars. 2. The full-sample mean is \(\frac{73600000}{1000}=\$73{,}600\). 3. The respondent-only mean is \(73{,}600-62{,}000=\$11{,}600\) too low. 4. Nonresponse is related to salary, so treating the observed salaries as representative of all selected employees is inappropriate.

Answer

a) \(\$73{,}600\). b) The respondent-only mean underestimates by \(\$11{,}600\). c) The missingness is related to the outcome, producing nonresponse bias.
54936911
The boxplots compare delivery times for two services. Both have a mean of \(30\) minutes. A report concludes that the services are equally reliable because their means match. a) Compare the medians and interquartile ranges. b) Which service is more consistent? c) Explain why equal means do not imply equal reliability.
Figure for problem 549369

Hints

- Compare both centers and spreads shown by the boxes. - Consistency is reflected by how tightly outcomes cluster. - One summary statistic cannot determine an entire distribution.

Solution

1. Service A has median \(30\) and interquartile range \(32-28=4\) minutes. Service B has median \(25\) and interquartile range \(45-15=30\) minutes. 2. Service A is more consistent because its central half and full whisker range are much narrower. 3. Equal means describe only one measure of center; Service B's broad, skewed distribution can have the same mean while producing far less predictable delivery times.

Answer

a) A: median \(30\), IQR \(4\). B: median \(25\), IQR \(30\). b) Service A. c) Reliability depends on variability and distribution shape, not only the mean.
54937011
In a randomized experiment with only \(20\) participants, the treatment group's baseline score mean is \(70\), while the control group's baseline mean is \(62\). A critic says randomization failed because the means are not equal. a) Explain why unequal baseline means can occur after valid randomization. b) What should researchers check before interpreting the treatment outcome? c) Does the imbalance prove the assignment was manipulated? Explain.

Hints

- Randomization creates probabilities, not identical group summaries. - Small samples allow larger chance imbalances. - Evaluate both the assignment procedure and the effect of the baseline difference.

Solution

1. Random assignment balances groups on average over many experiments but does not guarantee exact equality in every small sample. 2. Researchers should report baseline differences, check whether they are influential, and consider an analysis of change or an adjustment planned for baseline score. 3. The imbalance alone does not prove manipulation; the assignment procedure should be audited, but chance variation is a plausible explanation.

Answer

a) Small randomized groups can differ by chance even under a valid procedure. b) Report and account for the baseline difference when evaluating outcomes. c) No. Unequal means are possible under honest random assignment.
54937211
A classifier is evaluated on two groups. <table><thead><tr><th>Group</th><th>Actual positives</th><th>Positives correctly flagged</th><th>Actual negatives</th><th>Negatives incorrectly flagged</th></tr></thead><tbody><tr><td>A</td><td>\(100\)</td><td>\(90\)</td><td>\(100\)</td><td>\(10\)</td></tr><tr><td>B</td><td>\(20\)</td><td>\(10\)</td><td>\(180\)</td><td>\(10\)</td></tr></tbody></table> a) Find the overall accuracy for each group. b) Find the missed-positive rate for each group. c) Explain why equal accuracy does not imply equal performance across groups.

Hints

- First recover correctly classified negatives from the false-positive counts. - Use actual positives as the denominator for missed-positive rate. - A shared overall percentage can be built from very different kinds of errors.

Solution

1. Group A has \(90\) correct positives and \(90\) correct negatives, so accuracy is \(\frac{180}{200}=90\%\). 2. Group B has \(10\) correct positives and \(170\) correct negatives, so accuracy is also \(\frac{180}{200}=90\%\). 3. Group A misses \(10\) of \(100\) positives, a \(10\%\) missed-positive rate. Group B misses \(10\) of \(20\), a \(50\%\) rate. 4. Different class proportions allow the same total accuracy to hide a much worse error rate on positive cases in Group B.

Answer

a) Both groups have \(90\%\) accuracy. b) Group A: \(10\%\); Group B: \(50\%\). c) Accuracy combines error types and is affected by group class proportions, so it can conceal unequal missed-positive rates.
54937411
A travel-time model is trained on all available trips. Its accuracy is then reported using only trips made on clear weekdays, even though it will be used on weekends and during rain or snow. a) Identify the evaluation-set problem. b) Explain how the reported accuracy could be overly optimistic. c) Describe a more appropriate validation plan.

Hints

- Compare the conditions in the test set with the conditions in actual use. - A performance estimate is only as representative as the data used to calculate it. - Important operating conditions should appear in validation and in separate checks.

Solution

1. The evaluation set does not represent the conditions under which the model will be deployed. 2. Clear weekdays may have more regular traffic and fewer unusual delays, so performance on them can exceed performance during harder weekend and weather conditions. 3. Validation should use held-out trips sampled across the intended deployment distribution, with separate performance summaries by weather, day type, route, and other important conditions.

Answer

a) The validation data are unrepresentative of deployment conditions. b) The selected trips are easier and less variable than the full use case. c) Test on held-out data covering all relevant conditions and report subgroup performance.
54937611
A treatment has different estimated effects in two population groups. Group 1 makes up \(80\%\) of the population and has an estimated effect of \(+10\) units. Group 2 makes up \(20\%\) and has an estimated effect of \(-5\) units. a) Find the population-weighted average effect. b) Explain why reporting only the average effect can be misleading. c) What additional result should accompany the average?

Hints

- Weight each group effect by its share of the population. - An average can combine effects with opposite signs. - Decision-making may require knowing who benefits and who may be harmed.

Solution

1. The weighted average effect is \(0.80\cdot10+0.20\cdot(-5)=8-1=7\) units. 2. The positive average hides that the treatment may be harmful for Group 2, so the same conclusion does not apply uniformly to every subgroup. 3. The report should include group-specific effect estimates, uncertainty, group sizes, and evidence about whether the difference between subgroup effects is reliable.

Answer

a) \(+7\) units. b) The average conceals a negative estimated effect in Group 2. c) Report subgroup effects with uncertainty and evaluate the evidence for effect differences.
54938311
A treatment reduces the event rate from \(2\%\) to \(1\%\). An advertisement says, “The treatment cuts risk by \(50\%\).” a) Verify the relative-risk reduction. b) Find the absolute risk reduction in percentage points. c) For \(10{,}000\) similar people, estimate the difference in event counts. d) Evaluate whether the advertisement gives a complete picture.

Hints

- Calculate change once relative to the original risk and once directly on the percentage scale. - Translate both rates into counts for a common population size. - A relative statement can sound large when the baseline rate is small.

Solution

1. The relative reduction is \(\frac{2\%-1\%}{2\%}=\frac{1}{2}=50\%\). 2. The absolute reduction is \(2\%-1\%=1\) percentage point. 3. Without treatment, the expected count is \(10{,}000\cdot0.02=200\); with treatment, it is \(10{,}000\cdot0.01=100\), a difference of \(100\) events. 4. The relative claim is mathematically correct but incomplete without the starting risk and absolute change.

Answer

a) \(50\%\) relative reduction. b) \(1\) percentage point absolute reduction. c) About \(100\) fewer events per \(10{,}000\) people. d) The claim is correct but should also report the baseline risk and absolute reduction.
54938411
A model estimates that the mean daily demand next month will be between \(490\) and \(510\) units with \(95\%\) confidence. Historical day-to-day demand varies widely, and a prediction interval for one future day is \((420, 580)\). a) Explain the different targets of the two intervals. b) Why is it incorrect to tell a manager that one day's demand will be between \(490\) and \(510\) with \(95\%\) confidence? c) Which interval is relevant for staffing a single day?

Hints

- Identify whether each interval concerns an average or one observation. - Individual outcomes vary even if the mean were known exactly. - Match the interval target to the operational decision.

Solution

1. The confidence interval \((490, 510)\) estimates the population or future-period mean demand, while the prediction interval \((420, 580)\) describes uncertainty for one individual future day's demand. 2. Individual days vary around the mean, so uncertainty for one day includes both uncertainty in the mean and natural day-to-day variation. 3. The prediction interval \((420, 580)\) is relevant for planning a single day's staffing.

Answer

a) The narrow interval estimates mean demand; the wide interval predicts one future observation. b) A single day has additional natural variability not represented by the confidence interval for the mean. c) Use \((420, 580)\).
54938611
An app reports that monthly active users rose \(40\%\), from \(600\) in February to \(840\) in March. January had \(1000\) users, February was affected by a two-day outage, and March of the previous year had \(830\) users. a) Verify the advertised February-to-March increase. b) Find the January-to-March change and the year-over-year March change. c) Evaluate the choice of comparison window.

Hints

- Calculate each change relative to its own starting value. - Compare periods that differ in whether an unusual disruption occurred. - A claim can be numerically correct yet misleading because of endpoint selection.

Solution

1. The February-to-March increase is \(\frac{840-600}{600}=0.40=40\%\). 2. The January-to-March change is \(\frac{840-1000}{1000}=-0.16=-16\%\). 3. The year-over-year March change is \(\frac{840-830}{830}\approx0.0120=1.20\%\). 4. The advertised comparison starts from an unusually low outage month and therefore exaggerates ordinary growth; broader and comparable periods give a different picture.

Answer

a) \(40\%\). b) January to March: \(-16\%\); year-over-year March: about \(+1.20\%\). c) The selected outage month creates a misleadingly low baseline and overstates growth.
54938711
The residual plot comes from a fitted linear model. a) What happens to the residual spread as \(x\) increases, and what model feature is questionable? b) How should prediction uncertainty differ across the displayed \(x\)-ranges? c) Why is one constant “plus or minus \(3\)” prediction rule inappropriate?
Figure for problem 549387

Hints

- Compare the center and spread of residuals separately. - Prediction uncertainty should reflect the local size of typical errors. - A single error margin assumes similar variability everywhere.

Solution

1. The residuals remain centered near zero, but their spread increases as \(x\) increases. Constant residual variability is therefore questionable. 2. Predictions should be most precise at low \(x\), less precise in the middle, and least precise at high \(x\). 3. A constant \(\pm3\) rule is wider than needed for many low-\(x\) cases but far too narrow for high-\(x\) cases, so it misrepresents changing uncertainty.

Answer

a) Residual variability increases with \(x\), so constant spread is questionable. b) Prediction intervals should widen as \(x\) increases. c) One fixed margin ignores the strong change in residual spread across the domain.
54938811
A crop experiment estimates that a fertilizer increases yield by \(5\) bushels per acre in dry conditions and by \(20\) bushels per acre in wet conditions. A report averages the results and says the fertilizer “adds \(12.5\) bushels per acre in all conditions.” a) Identify the problem with the report. b) What do the two effects show about how the fertilizer result depends on moisture? c) State a more accurate conclusion.

Hints

- Compare the effect across levels of the environmental condition. - A single effect is questionable when the change depends strongly on another condition. - Context-specific estimates can be more informative than an unweighted average.

Solution

1. The report treats a context-dependent effect as one constant effect and therefore misrepresents both dry and wet conditions. 2. The fertilizer effect changes with moisture conditions rather than remaining constant. 3. A more accurate conclusion is that the estimated increase is about \(5\) bushels per acre when dry and \(20\) when wet; any overall average must be weighted by the frequency of those conditions and should not replace the separate effects.

Answer

a) The average is incorrectly applied uniformly to two conditions with very different effects. b) The fertilizer effect depends on whether conditions are dry or wet. c) Report the separate \(5\)- and \(20\)-bushel effects and, if needed, a properly weighted overall average.
54939311
A forecasting model has mean absolute error \(12\) units on a test period. A simple baseline that predicts the value from the same month one year earlier has mean absolute error \(8\) units. The model's developers emphasize that its \(R^2\) is \(0.90\) and call it highly accurate. a) Which method predicts better on the stated error measure? b) Why should the simple baseline be included in the evaluation? c) Explain why a high \(R^2\) does not override worse test error.

Hints

- Use the performance measure tied directly to the prediction goal. - A complex method should be compared with a simple credible alternative. - Goodness-of-fit summaries and out-of-sample error answer different questions.

Solution

1. The seasonal baseline predicts better because its mean absolute error is \(8\) rather than \(12\) units. 2. A baseline shows whether the complex model improves on a readily available prediction rule that captures seasonality. 3. \(R^2\) describes variation explained under a particular fit, while test error directly measures out-of-sample prediction accuracy on the chosen scale; the model performs worse for the stated prediction goal.

Answer

a) The seasonal baseline. b) It establishes a meaningful minimum performance standard for the complex model. c) The model's higher-level fit summary does not compensate for larger errors on unseen data.
54935011
The histogram shows a randomization distribution of mean differences under a no-effect model. The observed difference is \(8.0\) units. Of \(1000\) simulated differences, \(12\) were at least \(8.0\). a) Estimate the one-sided tail probability. b) State what the tail probability means in this study. c) Does it equal the probability that the no-effect model is true? Explain.
Figure for problem 549350

Hints

- Use extreme simulated outcomes divided by all simulated outcomes. - Keep the conditioning direction clear: the model is assumed when the probability is calculated. - Distinguish probability of data under a model from probability of a model given data.

Solution

1. The estimated tail probability is \(\frac{12}{1000}=0.012\). 2. If the no-effect model and randomization process were appropriate, about \(1.2\%\) of assignments would produce a mean difference at least as large as \(8.0\). 3. The tail probability is calculated assuming the no-effect model; it is not the probability that the model itself is true.

Answer

a) \(0.012\). b) Under the no-effect model, a result at least this large occurs in about \(1.2\%\) of randomizations. c) No. It is a probability of simulated data under the model, not a probability assigned to the model.
54935111
Researchers test a treatment on \(20\) unrelated outcomes and highlight the only outcome with a simulated tail probability below \(0.05\), which is \(0.030\). Under a model where the treatment affects none of the outcomes, assume each test has a \(0.05\) false-flag probability and the tests are independent. a) Find the probability of at least one false flag among \(20\) tests. b) Explain why the single value \(0.030\) is weaker evidence than it would be in a study with one prespecified outcome. c) Identify one better reporting practice.

Hints

- Use the complement of having no false flags in any test. - Count how many opportunities the researchers had to find an unusual result. - Good reporting prevents readers from seeing only the most favorable analysis.

Solution

1. The probability of no false flags is \((0.95)^{20}\approx0.3585\). 2. The probability of at least one false flag is \(1-0.3585\approx0.6415\). 3. Searching across many outcomes makes at least one apparently unusual result common even when no outcomes are affected. 4. Researchers should prespecify primary outcomes, report all tested outcomes, or use a method that accounts for multiple comparisons.

Answer

a) \(1-(0.95)^{20}\approx0.6415\). b) With twenty opportunities, a small tail probability can arise through selection of the most favorable result. c) Prespecify and report primary outcomes, disclose all tests, or adjust for multiple comparisons.
54935411
A university reports admission results for two departments. <table><thead><tr><th>Department</th><th>Women admitted</th><th>Women applicants</th><th>Men admitted</th><th>Men applicants</th></tr></thead><tbody><tr><td>A</td><td>\(80\)</td><td>\(100\)</td><td>\(18\)</td><td>\(20\)</td></tr><tr><td>B</td><td>\(2\)</td><td>\(20\)</td><td>\(60\)</td><td>\(100\)</td></tr></tbody></table> a) Compare admission rates within each department. b) Compare the overall admission rates. c) Explain how the overall comparison reverses the within-department comparisons.

Hints

- Compute rates using applicants within the same department as denominators. - Then recompute after pooling each gender across departments. - Compare how each group is weighted toward the high-rate and low-rate departments.

Solution

1. In Department A, the rates are \(\frac{80}{100}=80\%\) for women and \(\frac{18}{20}=90\%\) for men. In Department B, they are \(\frac{2}{20}=10\%\) for women and \(\frac{60}{100}=60\%\) for men. 2. Overall, women have \(\frac{82}{120}\approx68.3\%\), while men have \(\frac{78}{120}=65\%\). 3. More women applied to high-admission Department A, while more men applied to low-admission Department B. Different group weights reverse the aggregated comparison.

Answer

a) Men have higher admission rates in both departments: \(90\%\) versus \(80\%\) in A and \(60\%\) versus \(10\%\) in B. b) Overall, women have about \(68.3\%\) admitted and men \(65\%\). c) The applicant groups are distributed very differently across departments with very different base admission rates.
54935611
Participants are randomly assigned to an exercise program or a control group. By the end, \(30\%\) of the exercise group and \(5\%\) of the control group have dropped out. Among completers, the exercise group has a lower mean blood pressure. a) Why does random assignment at the start not completely protect the completer-only comparison? b) Give a mechanism that could bias the observed effect. c) State one analysis or reporting step that would improve evaluation of the result.

Hints

- Random assignment balances groups before later events occur. - Ask who is missing from each final comparison and why. - A trustworthy report should show how different reasonable assumptions about missing outcomes affect the conclusion.

Solution

1. Differential dropout can destroy the comparability created by random assignment among the participants who remain. 2. If exercise participants with little improvement or difficulty following the program are more likely to leave, the completers can make the program look more effective than it was for everyone assigned. 3. The report should compare dropout reasons and baseline traits, include outcomes for all assigned participants when possible, and show how conclusions change under reasonable assumptions about the missing outcomes.

Answer

a) Unequal attrition can make the remaining groups systematically different. b) For example, exercise participants who do not improve may be more likely to drop out. c) Analyze all assigned participants when possible and report how the conclusion changes under reasonable assumptions about missing outcomes.
54935811
From \(2005\) to \(2025\), streaming subscriptions and average college-textbook prices both increased. A regression between the annual series has \(R^2=0.96\), and a report claims streaming subscriptions caused textbook prices to rise. Explain why the claim is unsupported and identify a more meaningful modeling step.

Hints

- Separate goodness of fit from evidence about cause. - Ask what third variable changes systematically with both annual series. - Account for the shared time trend before interpreting the relationship.

Solution

1. \(R^2\) describes the strength of a linear association; it does not establish a causal mechanism or rule out confounding. 2. Both variables have strong upward time trends, so later years are high in both series even if the variables are unrelated. 3. A better analysis would account for time trend and inflation, compare year-to-year changes after adjusting for inflation and use predictors connected to a plausible textbook-price mechanism.

Answer

The high \(R^2\) can be caused by a shared upward time trend and does not establish causation. A more meaningful analysis should account for the time trend and inflation, then evaluate variables tied to a plausible pricing mechanism.
54935911
Study 1 estimates an effect of \(5\) units with a \(95\%\) confidence interval of \((2, 8)\). A larger replication estimates an effect of \(-1\) unit with interval \((-4, 2)\). A commentator says the second study proves the first study was fraudulent. a) Explain why that conclusion is not justified by the intervals. b) Compare the ranges of plausible effects. c) State two reasonable next steps for evaluating the disagreement.

Hints

- An estimate from one sample is not expected to match another exactly. - Compare both the centers and the full plausible ranges. - Investigate methodological differences before making claims about conduct.

Solution

1. Different random samples and study conditions can produce different estimates; disagreement does not by itself establish misconduct. 2. Study 1 supports positive effects from \(2\) to \(8\), while Study 2 includes negative, zero, and small positive effects from \(-4\) to \(2\). They share the boundary value \(2\) but give substantially different centers. 3. Researchers should compare protocols, populations, measurements, and attrition, and combine or repeat well-designed studies when appropriate.

Answer

a) Sampling variation or design differences can explain disagreement; the intervals alone do not show fraud. b) The first supports a positive effect, while the replication supports values from moderately negative through small positive. c) Audit design differences and conduct or combine additional carefully designed replications.
54936511
A risk model assigns a \(30\%\) event probability to each of \(200\) cases. The event actually occurs in \(90\) of those cases. a) Find the observed event rate. b) Evaluate the model's calibration for this risk group. c) Explain why this calculation alone does not show whether the model ranks high-risk and low-risk cases well.

Hints

- Compare the predicted percentage with the observed relative frequency. - State both the direction and size of the discrepancy. - Probability accuracy and ranking performance are different model properties.

Solution

1. The observed event rate is \(\frac{90}{200}=0.45=45\%\). 2. The observed rate is \(45\%-30\%=15\) percentage points higher than predicted, so the model underpredicts risk for this group and is poorly calibrated here. 3. Calibration compares predicted probabilities with observed frequencies; ranking ability requires comparing whether cases with higher assigned risks actually tend to experience more events than cases with lower assigned risks.

Answer

a) \(45\%\). b) The model underpredicts by \(15\) percentage points for this group. c) Calibration within one risk group does not measure how well the model orders cases across different risk levels.
54936811
A randomized trial finds no clear overall treatment effect. After seeing the data, analysts examine \(12\) demographic subgroups and report that the treatment appears effective in one subgroup with a tail probability of \(0.04\). a) Why should the subgroup claim be treated cautiously? b) What additional evidence would strengthen it? c) Distinguish a prespecified subgroup analysis from a data-discovered subgroup analysis.

Hints

- Count how many subgroup opportunities were examined before one was highlighted. - Timing matters: ask whether the subgroup was chosen before or after seeing results. - Exploratory findings become stronger when independently replicated.

Solution

1. Searching many subgroups creates multiple opportunities for a chance result, and the subgroup was selected after inspecting the data. 2. The claim would be stronger with a planned subgroup comparison and replication in an independent study designed to include that subgroup. 3. A prespecified analysis is planned before outcomes are examined; a data-discovered analysis is exploratory and requires more cautious interpretation and confirmation.

Answer

a) The result may be a chance finding from multiple after-the-fact subgroup searches. b) Use a prespecified subgroup comparison and independent replication. c) Prespecified analyses test planned questions; data-discovered analyses generate hypotheses that need confirmation.
54937111
An analyst checks an experiment's result every day and stops as soon as the simulated tail probability falls below \(0.05\). Under a true no-effect model, \(142\) of \(1000\) simulated experiments using this rule eventually stop with a “significant” result. a) Estimate the false-positive rate of the stopping rule. b) Compare it with the intended \(5\%\) rate. c) Explain why repeated unplanned checking changes the inference.

Hints

- Use falsely stopped simulations divided by all no-effect simulations. - Compare both the absolute and multiplicative increase over the target rate. - A decision rule includes when the data are examined and when collection stops.

Solution

1. The simulated false-positive rate is \(\frac{142}{1000}=0.142=14.2\%\). 2. This is \(14.2\%-5.0\%=9.2\) percentage points above the intended rate and is about \(\frac{14.2}{5}=2.84\) times as large. 3. Each additional look creates another opportunity for random fluctuation to cross the threshold; the original tail-probability interpretation assumes a fixed analysis plan.

Answer

a) \(14.2\%\). b) It is \(9.2\) percentage points higher, about \(2.84\) times the intended rate. c) Repeated optional checking adds chances to stop on a random extreme result.
54937311
A city introduces a road-safety law. Crashes fall from \(100\) to \(70\) in the city. During the same period, crashes in a similar comparison city without the law fall from \(110\) to \(88\). a) Find the before-after change in each city. b) Use the comparison city's change to form an adjusted difference in changes. c) Explain why the raw decline of \(30\) crashes may overstate the law's effect.

Hints

- Compute each location's change using the same subtraction order. - Subtract the comparison change from the treated change. - Use the untreated location to estimate what might have happened anyway.

Solution

1. The treated city changes by \(70-100=-30\) crashes. The comparison city changes by \(88-110=-22\) crashes. 2. The adjusted difference in changes is \((-30)-(-22)=-8\) crashes. 3. A broad trend reduced crashes even without the law, as shown by the comparison city. The raw decline attributes that shared trend to the law, while the adjusted comparison suggests an additional decline of about \(8\) crashes.

Answer

a) Treated city: \(-30\); comparison city: \(-22\). b) \(-8\) crashes. c) Most of the decline may reflect a trend affecting both cities rather than the law alone.
54937711
A study records \(10\) daily observations from each of \(30\) students and analyzes the \(300\) observations as if they came from \(300\) unrelated students. a) Identify the violated model assumption. b) Explain how the mistake affects estimated uncertainty. c) Give two valid ways to restructure the analysis or design.

Hints

- Identify which observations share the same source. - Repeated measurements do not add as much new information as measurements from new independent units. - Either change the unit of analysis or use a model that represents the grouping.

Solution

1. Observations from the same student are likely correlated, violating the assumption that all \(300\) observations are independent. 2. Treating repeated observations as independent overstates the amount of separate information and usually makes standard errors and intervals too small. 3. The analysis can summarize each student first or use a method that treats observations from the same student as related; the design can also sample more students rather than only more days per student.

Answer

a) Independence is violated within each student's repeated measurements. b) Uncertainty is understated because the effective sample size is less than \(300\). c) Analyze student-level summaries or use a method that accounts for repeated observations from each student, and increase the number of students when possible.
54937811
A data set has \(100\) missing scores. An analyst replaces every missing value with the observed mean of \(75\) and then fits a model as if all values were observed. a) Explain how this replacement affects the data's mean. b) Explain how it affects the data's variability. c) Why can confidence intervals from the completed data be too narrow?

Hints

- Consider what happens when many new values are placed exactly at the center. - A filled-in value is not the same as an observed value. - Valid uncertainty should include uncertainty about what was missing.

Solution

1. Replacing missing values by the observed mean leaves the completed-data mean at \(75\) if the observed mean was \(75\). 2. Every imputed value lies exactly at the center, so the completed data have artificially reduced variability. 3. The method treats guessed values as known and ignores uncertainty about missing outcomes, causing model uncertainty and confidence-interval width to be understated.

Answer

a) It preserves the observed mean of \(75\). b) It artificially reduces variability by adding many central values. c) The analysis ignores uncertainty in the imputed values and therefore appears too precise.
54937911
A model predicts whether a student will pass a final exam. One predictor is “course completed,” a status entered only after final grades are posted. The model has \(99\%\) test accuracy on a randomly split historical data set. a) Identify the modeling flaw. b) Explain why a random train-test split does not solve it. c) Describe how predictors and validation should be redesigned.

Hints

- Ask when each predictor becomes available relative to the prediction decision. - A held-out set is only useful if it represents the information available in real use. - Validation should preserve time order when the application predicts the future.

Solution

1. The predictor contains information recorded after the outcome and therefore leaks the answer into the model. 2. Both randomly split sets contain the leaked variable, so the model can exploit future information in training and testing even though that information is unavailable at prediction time. 3. Predictors must be restricted to information available before the exam, and validation should reproduce the intended timeline, such as training on earlier terms and testing on a later term.

Answer

a) Outcome or future-information leakage. b) The leaked variable appears in both random splits, so the test is still unrealistic. c) Use only pre-exam predictors and validate on data separated according to the real prediction timeline.
54938011
A loan-default model had \(92\%\) accuracy when developed in \(2018\). Using the same threshold on \(2026\) applications, accuracy is \(76\%\), and applicants now have different income patterns and loan types. a) Give a statistical reason the old performance estimate may no longer apply. b) What evidence should be checked before continued use? c) Describe one responsible response to the decline.

Hints

- A model's evidence is tied to the population and period on which it was evaluated. - Compare current inputs and errors with the development data. - Model evaluation should continue after deployment rather than occur only once.

Solution

1. The population and relationship between predictors and outcomes may have shifted over time, so the 2018 validation distribution no longer represents 2026 applications. 2. The lender should check current calibration, error rates, subgroup performance, predictor distributions, and whether outcome definitions or economic conditions changed. 3. The model should be recalibrated, retrained, replaced, or restricted after validation on recent representative data, with continuing performance monitoring.

Answer

a) The applicant population and the relationships used by the model have changed, so the old validation is no longer reliable. b) Check current calibration, error rates, subgroup results, and predictor shifts. c) Revalidate on recent data and recalibrate, retrain, replace, or limit the model as needed.
54938211
In a randomized trial, \(100\) people are assigned to a program and \(100\) to control. Only \(60\) assigned to the program actually participate. The mean outcome is \(74\) for everyone assigned to the program and \(70\) for everyone assigned to control. Actual participants have mean \(78\). a) Find the effect estimate based on original random assignment. b) Why is comparing actual participants with the control group potentially biased? c) Which comparison best preserves the benefit of randomization?

Hints

- Keep the groups defined by the random mechanism. - A choice made after assignment can be related to the outcome. - Distinguish the effect of being offered a program from the effect among self-selected users.

Solution

1. The assignment-based effect estimate is \(74-70=4\) units. 2. People who choose to participate may differ in motivation, availability, or baseline condition from nonparticipants and controls, so the \(78-70=8\) comparison mixes treatment with self-selection. 3. Comparing all people according to original assignment preserves randomized group comparability and estimates the effect of offering or assigning the program.

Answer

a) \(4\) units. b) Actual participation is self-selected after assignment, so the \(8\)-unit comparison is confounded. c) Compare all \(200\) people by their original randomized assignment.
54938511
A screening test has a \(90\%\) positive rate among people with a condition and a \(5\%\) positive rate among people without it. In the screened population, \(2\%\) have the condition. a) For \(10{,}000\) people, find the expected numbers of true positives and false positives. b) Among positive tests, estimate the percentage who actually have the condition. c) Explain why “\(90\%\) sensitivity” does not mean a person with a positive result has a \(90\%\) chance of having the condition.
Figure for problem 549385

Hints

- Begin by splitting the population according to the base rate. - Apply each conditional positive rate within the appropriate branch. - Check which event appears after the conditioning bar in each probability.

Solution

1. Of \(10{,}000\) people, \(200\) have the condition and \(9800\) do not. 2. The expected true positives are \(200\cdot0.90=180\). The expected false positives are \(9800\cdot0.05=490\). 3. The percentage with the condition among all positive tests is \(\frac{180}{180+490}\approx0.2687=26.87\%\). 4. Sensitivity conditions on actually having the condition, while the requested probability conditions on receiving a positive test; the low base rate and false positives matter.

Answer

a) \(180\) true positives and \(490\) false positives. b) About \(26.87\%\). c) Sensitivity is \(P(+\mid C)\), not \(P(C\mid +)\); the population base rate changes the latter probability.
54938911
A modeling team fits \(50\) candidate models and repeatedly checks performance on the same test set. It selects the model with the highest test accuracy and reports that accuracy as an unbiased estimate of future performance. a) Explain why the test set is no longer functioning as an independent test. b) What direction of bias is likely in the selected accuracy? c) Describe a better model-development and evaluation process.

Hints

- Track whether information from the supposed test set changed any modeling decision. - Selecting the maximum among many noisy estimates favors unusually high values. - Reserve some data until every model choice is complete.

Solution

1. Repeatedly using test results to choose among models makes information from the test set part of model selection. 2. The highest observed test accuracy includes favorable random variation, so it is likely biased upward relative to future performance. 3. Use training data for fitting, separate validation data for model choices, and a final untouched test set for one-time evaluation.

Answer

a) Test performance influenced model selection, so the set is no longer independent of development. b) The reported accuracy is likely too high. c) Separate training, validation, and final untouched testing stages.
54939011
Students are randomly assigned by classroom to receive a new tutoring program or the usual program. Students in neighboring classrooms share tutoring materials and study together after school. a) Identify the problem this behavior creates for keeping the treatment and control conditions separate. b) Explain how the sharing could affect the estimated treatment difference. c) Suggest one design or analysis change.

Hints

- Ask whether one participant's assigned treatment can change another participant's experience. - Consider what happens when the control group receives part of the active program. - Choose assignment units that reduce contact across treatment conditions.

Solution

1. The study's assumption that one group's treatment does not affect outcomes in the other group is threatened by spillover or interference. 2. Control students may receive part of the tutoring benefit, making the groups more similar and often reducing the estimated difference; other social effects could also complicate direction. 3. Researchers could assign more widely separated schools or groups, measure how much sharing occurs, and report how the results change when cross-group exposure is considered.

Answer

a) Spillover makes the treatment and control conditions overlap. b) Shared materials can contaminate the control condition and usually shrink the observed contrast. c) Assign more separated groups and measure or account for cross-group exposure.
54939111
Fifty matched pairs of participants are formed using similar baseline scores. Within every pair, one participant is randomly assigned to Treatment A and the other to Treatment B. The mean of the \(50\) within-pair differences, A minus B, is \(3.2\) points. The standard deviation of the within-pair differences is \(4\) points, while individual outcome scores have a standard deviation of \(15\) points. An analyst discards the pair labels and compares the two treatment groups as unrelated samples. a) Does discarding pair labels remove the random assignment? Explain. b) Why is the unpaired analysis likely less precise? c) What statistic and analysis should use the design most effectively?

Hints

- Separate how treatment was assigned from how outcomes are later analyzed. - Matching is useful when partners share variation that can cancel in a difference. - Use the same observational unit in the analysis that the randomization created.

Solution

1. The original within-pair random assignment still occurred, so discarding labels does not erase the random treatment mechanism or necessarily change the treatment-group mean difference. 2. Matching made partners similar at baseline, and the within-pair differences vary by only \(4\) points compared with \(15\)-point variation among individual outcomes. Ignoring pairs fails to remove that shared baseline variation and therefore loses precision. 3. The analysis should use the \(50\) within-pair differences, centered at the observed mean difference of \(3.2\) points, with a method designed for paired data.

Answer

a) No. Treatment was still randomly assigned within pairs, although the analysis ignores useful design information. b) Ignoring matching leaves large between-participant variation in the error instead of canceling it within pairs. c) Analyze the \(50\) A-minus-B pair differences, whose mean is \(3.2\) points.
54939211
Two models fit the historical data about equally well. Model A forecasts next year's demand as \(120\) units, while Model B forecasts \(180\) units because they make different assumptions about whether recent growth will continue. A report presents only Model A's forecast. a) Identify the omitted source of uncertainty. b) Why is a narrow interval around \(120\) potentially misleading? c) Describe a more responsible report.

Hints

- Separate uncertainty inside one chosen model from uncertainty about choosing the model itself. - Good historical fit does not make two different future assumptions equivalent. - Show how the conclusion changes under other plausible model assumptions.

Solution

1. The report omits model uncertainty: different plausible model structures and assumptions produce very different forecasts. 2. A narrow interval conditional on Model A can describe random error within that model while ignoring the larger uncertainty about which model is appropriate. 3. A responsible report should show forecasts from both justified models, explain their assumptions, evaluate them on relevant validation data, and show how the forecast changes under the different plausible models.

Answer

a) Uncertainty about which model and assumptions are appropriate. b) It can look precise while ignoring another plausible forecast far from \(120\). c) Report alternative models, assumptions, validation performance, and sensitivity of conclusions to model choice.
54939411
A study recruits \(300\) volunteers from one fitness center and randomly assigns them to an energy drink or a placebo. Fifteen people leave the drink group and two leave the placebo group. Among completers, the drink group has a mean heart rate \(4.0\) beats per minute higher. A randomization simulation gives a tail proportion of \(0.020\), and a \(95\%\) confidence interval for the mean difference is \((0.5, 7.5)\) beats per minute. A headline says, “The drink raises every adult's heart rate by exactly \(4\) beats per minute.” a) Evaluate the evidence for a causal effect. b) Evaluate generalization to all adults. c) Explain two problems with the word “exactly” and the claim about every adult. d) Explain how differential dropout affects the result. e) Write a conclusion supported by the study.

Hints

- Evaluate random assignment, sampling, missing outcomes, and uncertainty separately. - Distinguish a sample average from a population effect and from every individual's response. - A strong conclusion should state both the supported effect and the study's scope limitations.

Solution

1. Random assignment initially supports causal inference, and the small tail proportion indicates that the observed completer difference is unusual under the no-effect randomization model. However, differential attrition weakens the randomized comparison, so the completer-only result is not a clean causal estimate. 2. Volunteers from one fitness center are not a random sample of all adults, so broad population generalization is unsupported. 3. The \(4.0\) value is a sample mean difference, not an exact population effect; the confidence interval shows plausible mean effects from \(0.5\) to \(7.5\). An average effect also does not imply the same change for every individual. 4. The much higher dropout in the drink group can make completers systematically different and weaken the comparability created by random assignment. 5. A supported conclusion is that the study found a higher mean heart rate among drink-group completers and statistical evidence inconsistent with a no-effect model, but the effect size is uncertain, differential attrition weakens causal interpretation, and the result cannot be generalized to all adults.

Answer

a) Random assignment initially supports causation, but differential attrition means the completer-only comparison is not a clean causal estimate. b) The volunteer sample from one fitness center does not support generalization to all adults. c) \(4.0\) is an estimate with interval \((0.5, 7.5)\), and a mean effect is not an identical individual effect. d) Unequal dropout can make the remaining randomized groups systematically different. e) The drink-group completers had a higher mean heart rate, but the effect size is uncertain, differential attrition weakens causal interpretation, and generalization to all adults is unsupported.

All problems may be used, copied and printed free of charge for school and tutoring, including paid tutoring. Commercial adaptations as well as publication or redistribution on the internet are not permitted.