Aimathic
Login | English | Deutsch

Free math worksheets

Build your own math worksheets from 28,000 problems for grades 3 to 12, from fractions to calculus. Every problem includes step-by-step solutions.

Correlation

Click problems to add them to your worksheet.

53943212
For a sample measuring weekly rehearsal hours and tempo-accuracy score, the correlation is \(r = 0.91\). Interpret the value and state one conclusion that is not justified by correlation alone.

Hints

- Use the sign of \(r\) to determine whether the variables tend to move in the same or opposite directions. - Use how close \(|r|\) is to \(1\) to judge linear strength. - Distinguish a statement about a tendency in the sample from a claim that rehearsal causes the change.

Solution

1. Because \(r\) is positive, greater weekly rehearsal time tends to be associated with higher tempo-accuracy scores. 2. Because \(|r| = 0.91\) is close to \(1\), the linear association is strong. 3. Correlation alone does not show that increasing rehearsal time causes a higher tempo-accuracy score.

Answer

There is a strong positive linear association: performers who rehearse more tend to have higher tempo-accuracy scores. The correlation alone does not establish that more rehearsal causes higher accuracy.
53943312
For a sample measuring air temperature and natural-gas use, the correlation is \(r = -0.74\). Interpret the value and state one conclusion that is not justified by correlation alone.

Hints

- Read the negative sign as movement in opposite directions. - Judge strength from the magnitude \(0.74\), not from the sign. - Keep an association statement separate from a claim that temperature itself causes all changes in gas use.

Solution

1. Because \(r\) is negative, higher air temperatures tend to be associated with lower natural-gas use. 2. Because \(|r| = 0.74\), the linear association is moderately strong. 3. Correlation alone does not prove that changing air temperature directly causes the observed change in natural-gas use.

Answer

There is a moderately strong negative linear association: higher temperatures tend to accompany lower natural-gas use. The correlation alone does not establish a cause-and-effect relationship.
53943412
For a sample measuring tree trunk diameter and tree age, the correlation is \(r = 0.18\). Interpret the value and state one conclusion that is not justified by correlation alone.

Hints

- Use the sign of \(0.18\) for direction. - Compare the magnitude \(0.18\) with \(0\) and \(1\) to judge linear strength. - Avoid turning a sample association between tree measurements into a causal claim.

Solution

1. The positive sign means larger trunk diameters tend to be associated with greater tree ages in this sample. 2. Because \(|r| = 0.18\) is close to \(0\), the linear association is weak. 3. The correlation does not prove that increasing trunk diameter causes a tree to become older or that age alone causes the observed diameter.

Answer

There is a weak positive linear association between trunk diameter and tree age. The correlation alone does not establish a cause-and-effect relationship.
53943512
For a sample measuring ticket price and number of tickets sold, the correlation is \(r = -0.96\). Interpret the value and state one conclusion that is not justified by correlation alone.

Hints

- Interpret the negative sign as opposite movement between price and sales. - Use the closeness of \(0.96\) to \(1\) to judge strength. - Consider whether factors such as event popularity could affect sales as well as price.

Solution

1. The negative sign means higher ticket prices tend to be associated with fewer tickets sold. 2. Because \(|r| = 0.96\) is close to \(1\), the linear association is strong. 3. Correlation alone does not prove that changing price caused the entire change in ticket sales; other factors may differ across observations.

Answer

There is a strong negative linear association: higher ticket prices tend to accompany fewer tickets sold. The correlation alone does not prove that price caused the change in sales.
53943612
For a sample measuring river speed and downstream canoe distance traveled in a fixed time, the correlation is \(r = 0.63\). Interpret the value and state one conclusion that is not justified by correlation alone.

Hints

- Use the positive sign to determine whether river speed and distance tend to increase together. - Treat \(0.63\) as the magnitude when judging strength. - Separate the observed tendency from a claim that river speed is the only cause of the distance traveled.

Solution

1. The positive sign means faster river speeds tend to be associated with greater downstream canoe distances in the fixed time. 2. Because \(|r| = 0.63\), the linear association is moderate. 3. Correlation alone does not prove that river speed by itself caused the observed distances; paddling effort and other conditions may also vary.

Answer

There is a moderate positive linear association: faster river speeds tend to accompany greater downstream distances in the fixed time. The correlation alone does not establish causation.
53943712
For a sample measuring number of practice attempts and completion time, the correlation is \(r = 0.02\). Interpret the value and state one conclusion that is not justified by correlation alone.

Hints

- Focus on how close \(0.02\) is to \(0\), rather than emphasizing its positive sign. - Remember that correlation measures linear association only. - Do not use a near-zero correlation to claim either causation or complete independence.

Solution

1. Because \(r = 0.02\) is extremely close to \(0\), the sample shows essentially no linear association between practice attempts and completion time. 2. The tiny positive sign is not meaningful evidence of a positive relationship. 3. Correlation alone also cannot rule out a nonlinear association or establish any cause-and-effect conclusion.

Answer

The sample shows essentially no linear association between practice attempts and completion time. The correlation does not rule out a nonlinear pattern and does not justify a causal conclusion.
53945712
A student says, “Since \(r = 0\) for two quantitative variables, the variables have no relationship of any kind.” Identify and correct the error.

Hints

- Recall that correlation measures only straight-line association. - Think of a symmetric U-shaped pattern whose positive and negative linear tendencies cancel. - Replace “no relationship” with the narrower claim supported by \(r = 0\).

Solution

1. A correlation of \(r = 0\) means the data have no linear association. 2. Correlation does not measure nonlinear association. 3. The variables could still have a curved or some other nonlinear relationship.

Answer

The error is treating zero correlation as no relationship. \(r = 0\) indicates no linear association, but a nonlinear relationship may still be present.
53945812
A student discussing distance run and recovery heart rate says, “The correlation is \(1.08\), showing an extremely strong association.” Identify and correct the error.

Hints

- Recall the complete allowable interval for a correlation coefficient. - Compare \(1.08\) with the largest possible value. - An invalid coefficient cannot support a valid strength interpretation.

Solution

1. Every correlation coefficient must satisfy \(-1 \le r \le 1\). 2. Because \(1.08 > 1\), the reported value cannot be a valid correlation. 3. The data or calculation must be checked before the association’s direction or strength can be interpreted.

Answer

\(1.08\) is not a possible correlation because \(-1 \le r \le 1\). The correlation must be recomputed before interpreting the association.
54805912
The paired observations are \((1, 9)\), \((2, 6)\), and \((3, 3)\). A student says, “The correlation cannot be negative because every \(x\)-value and every \(y\)-value is positive.” Compute the correlation and correct the student’s reasoning.

Hints

- Plot or inspect the ordered pairs from left to right. - Correlation direction is about how the variables change together, not about the signs of the raw measurements. - Check whether the points lie exactly on a straight line.

Solution

1. The three points lie exactly on the decreasing line \(y=12-3x\). 2. Therefore, the correlation is \(r=-1\). 3. The sign of correlation describes whether larger-than-average values of one variable tend to occur with larger-than-average or smaller-than-average values of the other; it does not depend on whether the raw measurements themselves are positive or negative.

Answer

\(r=-1\). Correlation is negative because the variables have a perfect decreasing linear relationship, even though all of the observed values are positive.
54811512
Study A reports a correlation of \(r=0.64\) between two quantitative variables. Study B reports \(r=-0.64\) for a different pair of quantitative variables measured in completely different units. Compare the strength and direction of the two linear associations.

Hints

- Separate the magnitude of a correlation from its sign. - Decide which part describes strength and which part describes direction. - Recall whether correlation carries measurement units.

Solution

1. The magnitude \(|r|\) describes the strength of the linear association, while the sign describes its direction. 2. Both correlations have magnitude \(0.64\), so they represent the same linear-association strength. 3. Study A has a positive direction, while Study B has a negative direction. 4. Because correlation is unitless, the different measurement units do not prevent this comparison of linear strength.

Answer

The two associations have equal linear strength because both have \(|r|=0.64\). Study A is positive, while Study B is negative.
53943812
The correlation between distance from a motion sensor in feet and recorded signal strength is \(r = 0.29\). The analyst converts the explanatory variable from feet to meters. What is the new correlation? Explain.

Hints

- Write the unit conversion as multiplying every distance by the same positive number. - Correlation is computed from standardized values, which are unchanged by positive rescaling. - Check whether either the direction or the relative spacing of the points changes.

Solution

1. Converting feet to meters multiplies every explanatory-variable value by the same positive constant. 2. Positive rescaling does not change standardized values or the direction and strength of a linear association. 3. Therefore, the new correlation is \(r = 0.29\).

Answer

\(r = 0.29\). Converting feet to meters is a positive rescaling, so correlation does not change.
53943912
The correlation between elevation and average air temperature is \(r = -0.64\). The analyst converts the response variable from degrees Celsius to degrees Fahrenheit. What is the new correlation? Explain.

Hints

- Express the Fahrenheit conversion as a multiplication followed by an addition. - Determine whether the multiplier is positive or negative. - Correlation is unchanged by shifts and positive rescaling of either variable.

Solution

1. Converting Celsius to Fahrenheit uses \(F = \frac{9}{5}C + 32\). 2. Adding a constant and multiplying by a positive constant do not change standardized values or correlation. 3. Therefore, the new correlation is \(r = -0.64\).

Answer

\(r = -0.64\). The Celsius-to-Fahrenheit conversion is a positive affine transformation, so correlation is unchanged.
53944012
The correlation between fabric weight and shipping cost is \(r = 0.27\). The analyst multiplies every response value by \(-1\). What is the new correlation? Explain.

Hints

- Decide what reflecting all response values across zero does to an upward or downward trend. - The transformation changes direction but not how tightly the points follow the pattern. - Keep the magnitude \(0.27\) and determine the new sign.

Solution

1. Multiplying one variable by a negative constant reverses the direction of the linear association. 2. The transformation does not change the magnitude of correlation. 3. Therefore, the new correlation is \(r = -0.27\).

Answer

\(r = -0.27\). Multiplying one variable by \(-1\) reverses the sign of correlation but leaves its magnitude unchanged.
53944112
The correlation between rainfall and reservoir inflow is \(r = 0.45\). The analyst interchanges the explanatory and response variables. What is the new correlation? Explain.

Hints

- Compare the standardized-product formula before and after swapping the variables. - Multiplication is commutative, so each paired product is unchanged. - Swapping axes does not reflect or rescale either variable.

Solution

1. Correlation treats the two variables symmetrically: \(r_{xy} = r_{yx}\). 2. Interchanging the axes changes neither the direction nor the strength of the linear association. 3. Therefore, the new correlation is \(r = 0.45\).

Answer

\(r = 0.45\). Correlation is symmetric, so interchanging the explanatory and response variables does not change it.
53944212
The correlation between stage-light angle and illuminated area is \(r = 0.77\). The analyst adds \(100\) to every explanatory-variable value. What is the new correlation? Explain.

Hints

- Decide whether adding \(100\) changes the spacing between any two angle values. - A horizontal shift does not reflect or stretch the scatterplot. - Correlation depends on standardized positions, not the variable’s zero point.

Solution

1. Adding \(100\) shifts every explanatory value by the same amount. 2. A shift changes the mean but leaves all standardized values and the scatterplot’s shape unchanged. 3. Therefore, the new correlation is \(r = 0.77\).

Answer

\(r = 0.77\). Adding a constant to one variable does not change correlation.
53944312
The correlation between roasting time and coffee-bean mass is \(r = -0.73\). The analyst multiplies each variable by a negative constant. What is the new correlation? Explain.

Hints

- Track what one negative multiplier would do to the correlation sign. - Apply the sign change once for each of the two variables. - Rescaling by nonzero constants does not change the magnitude of \(r\).

Solution

1. Multiplying one variable by a negative constant reverses the sign of correlation. 2. Because both variables are multiplied by negative constants, the two sign reversals cancel. 3. The magnitude is unchanged, so the new correlation is \(r = -0.73\).

Answer

\(r = -0.73\). Reflecting both variables reverses the axes twice, so the correlation’s sign and magnitude are unchanged.
53944512
Use the scatterplot of arm span and height. Is \(r\) most plausibly near \(-1\), near \(0\), or near \(1\)? Explain.
Figure for problem 539445

Hints

- Use the left-to-right direction of the plotted points to choose the sign of \(r\). - Judge the magnitude by how tightly the points follow a straight line. - A value near \(1\) represents a very strong positive linear association.

Solution

1. The points rise from left to right, so the correlation is positive. 2. The points cluster tightly around an increasing straight line, indicating a strong linear association. 3. Therefore, \(r\) is most plausibly near \(1\).

Answer

\(r\) should be near \(1\) because the points follow a strong positive linear pattern.
53944612
Use the scatterplot of trail steepness and hiking speed. Is \(r\) most plausibly near \(-1\), near \(0\), or near \(1\)? Explain.
Figure for problem 539446

Hints

- Determine the sign from whether hiking speed rises or falls as trail steepness increases. - Judge the magnitude by how tightly the plotted points follow a straight line. - Combine the negative direction with a magnitude close to \(1\).

Solution

1. The points fall from left to right, so the correlation is negative. 2. The points cluster tightly around a decreasing straight line, indicating a strong linear association. 3. Therefore, \(r\) is most plausibly near \(-1\).

Answer

\(r\) should be near \(-1\) because the points follow a strong negative linear pattern.
53944712
Use the scatterplot of shoe length and height. Is \(r\) most plausibly a large negative value, near \(0\), or a large positive value? Explain.
Figure for problem 539447

Hints

- Look for the overall left-to-right direction rather than requiring every point to rise. - Use the amount of vertical scatter to judge whether the magnitude is near \(0\) or near \(1\). - Combine the slight positive direction with a small magnitude.

Solution

1. The points have only a slight upward tendency, so the direction is weakly positive. 2. The points form a diffuse cloud rather than clustering tightly around a line. 3. Therefore, \(r\) should be a small positive value, which is closer to \(0\) than to \(1\).

Answer

\(r\) should be small and positive. The direction is slightly positive, but the wide scatter makes the linear association weak.
53945012
Use technology to compute the correlation for the paired data on crowd size in hundreds and concession wait in minutes. Round to three decimal places, then interpret its direction and strength. <table> <thead><tr><th>Crowd size in hundreds</th><th>Concession wait in minutes</th></tr></thead> <tbody> <tr><td>1</td><td>12.30</td></tr> <tr><td>2</td><td>14.23</td></tr> <tr><td>3</td><td>16.83</td></tr> <tr><td>4</td><td>18.67</td></tr> <tr><td>5</td><td>21.07</td></tr> <tr><td>6</td><td>23.17</td></tr> </tbody> </table>

Hints

- Keep each crowd size paired with the wait time in the same table row. - Use the correlation command rather than fitting a nonlinear model. - After calculating \(r\), use its sign for direction and its distance from \(0\) for linear strength.

Solution

1. Enter the crowd-size values and corresponding wait times in paired technology lists. 2. The computed correlation is \(r = 0.9993316\ldots\), which rounds to \(r \approx 0.999\). 3. The positive sign and magnitude close to \(1\) indicate a very strong positive linear association.

Answer

\(r \approx 0.999\); the data show a very strong positive linear association.
53945112
Use technology to compute the correlation for the paired data on sunlight hours and net electricity drawn from the grid in kilowatt-hours. Negative grid values mean electricity was exported. Round to three decimal places, then interpret the correlation’s direction and strength. <table> <thead><tr><th>Sunlight hours</th><th>Net grid electricity in kilowatt-hours</th></tr></thead> <tbody> <tr><td>2</td><td>5.40</td></tr> <tr><td>3</td><td>2.47</td></tr> <tr><td>4</td><td>0.87</td></tr> <tr><td>5</td><td>-2.27</td></tr> <tr><td>6</td><td>-4.27</td></tr> <tr><td>7</td><td>-6.87</td></tr> </tbody> </table>

Hints

- Treat negative electricity values as valid numerical observations representing net export. - Preserve the row-by-row pairing when entering the two lists. - Interpret the negative sign separately from the magnitude near \(1\).

Solution

1. Enter the sunlight-hour values and corresponding net-grid values as paired data. 2. The computed correlation is \(r = -0.9977906\ldots\), which rounds to \(r \approx -0.998\). 3. The negative sign and magnitude close to \(1\) indicate a very strong negative linear association.

Answer

\(r \approx -0.998\); the data show a very strong negative linear association.
53945612
Two samples measuring pages in a manual and reading time have correlations \(r_1 = -0.82\) and \(r_2 = 0.76\). Which sample has the stronger linear association? Compare direction as well as strength.

Hints

- Ignore the signs temporarily and compare \(0.82\) with \(0.76\) for strength. - Restore each sign when describing direction. - A negative correlation can be stronger than a positive correlation.

Solution

1. Compare strength using absolute values: \(|r_1| = 0.82\) and \(|r_2| = 0.76\). 2. Because \(0.82 > 0.76\), the first sample has the stronger linear association. 3. The first association is negative, while the second is positive.

Answer

The first sample has the stronger linear association. Its association is negative, while the second sample’s association is positive.
54797712
A data set contains paired measurements, but every value of the explanatory variable is exactly \(5.0\). A student says, “The correlation must be \(0\) because there is no upward or downward trend.” Is the student correct? Explain.

Hints

- Check whether both variables have nonzero spread before interpreting correlation. - Think about what happens when all horizontal coordinates are identical. - Distinguish an undefined statistic from a statistic whose value is zero.

Solution

1. Correlation requires variation in both quantitative variables. 2. Here, the explanatory variable has standard deviation \(0\). 3. Because the correlation calculation divides by the explanatory-variable standard deviation, the correlation is undefined rather than \(0\).

Answer

No. The correlation is undefined because the explanatory variable has no variation.
54798312
Software gives a correlation of \(r=0.9964\) for two quantitative variables. A report rounds this to \(r=1.00\) and then states, “The data have a perfect positive linear association.” Evaluate the statement.

Hints

- Use the unrounded value when deciding whether a boundary value has actually been reached. - Recall what a correlation of exactly \(1\) means geometrically. - Distinguish “very close to perfect” from “perfect.”

Solution

1. The unrounded correlation \(0.9964\) is less than \(1\), so the association is not mathematically perfect. 2. Rounding to two decimal places hides the small difference from \(1\). 3. The data have an extremely strong positive linear association, but a perfect association would require every point to lie exactly on an increasing line.

Answer

The statement is too strong. The association is extremely strong and positive, but \(r=0.9964\) is not a perfect correlation.
54800512
A regression analysis reports a correlation of \(r=0.98\) between two quantitative variables. A student says, “The model is \(98\%\) accurate, so individual predictions should be within about \(2\%\) of the observed values.” Evaluate the statement.

Hints

- Recall what correlation measures and whether it has units or percent-error meaning. - Distinguish linear association from the size of individual prediction errors. - Think about what quantity is used to examine how far observations fall from fitted values.

Solution

1. Correlation measures the direction and strength of a linear association; it is not a percent-accuracy measure for individual predictions. 2. The corresponding coefficient of determination is \(r^2=(0.98)^2=0.9604\), so about \(96.04\%\) of the sample variation in the response is explained by the fitted linear relationship. 3. Neither \(r\) nor \(r^2\) states that individual predictions are within a fixed percentage of observed values. Prediction accuracy depends on the residuals and their scale.

Answer

The statement is incorrect. \(r=0.98\) indicates a very strong positive linear association, not \(98\%\) prediction accuracy. Here \(r^2=0.9604\), but individual prediction errors must be assessed from residuals rather than from \(r\) alone.
54801512
Sample A contains \(12\) paired observations and has correlation \(r=0.65\). Sample B contains \(400\) paired observations and also has correlation \(r=0.65\). A student says, “The association in Sample B is stronger because the sample is much larger.” Evaluate the statement, distinguishing the strength of the observed linear association from the amount of evidence available about a population relationship.

Hints

- Separate what the value of \(r\) describes from what sample size affects. - Ask whether the numerical correlation changes just because more observations are present. - Think about precision of population conclusions as a different issue from descriptive association strength.

Solution

1. The sample correlation \(r\) describes the direction and strength of the observed linear association, so both samples have the same reported sample linear strength: \(r=0.65\). 2. The larger sample size does not make the numerical sample correlation stronger. 3. A larger sample can provide more precise and stable information about a population relationship, so Sample B can support stronger statistical evidence even though the observed correlation coefficient is the same.

Answer

The student is incorrect about sample association strength: both samples have the same reported linear association, \(r=0.65\). The larger sample in B can provide more precise evidence about the population, but it does not make \(r=0.65\) a stronger sample correlation.
54803012
Software reports \(r^2=0.64\) for a simple linear regression, and the fitted slope is negative. Find the sample correlation \(r\). Explain why \(r^2\) alone is not enough to determine the sign.

Hints

- Recover the magnitude of the correlation from the coefficient of determination. - Use the fitted line’s direction to determine the missing sign. - Remember what information is lost when a signed quantity is squared.

Solution

1. Since \(r^2=0.64\), the magnitude of the correlation is \(|r|=\sqrt{0.64}=0.80\). 2. In simple linear regression, the correlation and the least-squares slope have the same sign. 3. Because the fitted slope is negative, \(r=-0.80\). 4. Squaring removes the sign, so \(r^2\) alone gives only the magnitude of \(r\), not its direction.

Answer

\(r=-0.80\). The value \(r^2=0.64\) gives \(|r|=0.80\), and the negative slope determines that the correlation is negative.
54803912
Data Sets A and B both have sample correlation \(r=0.70\). A student says, “Because the correlations are equal, the two data sets have the same association.” Evaluate the statement using the scatterplots.
Figure for problem 548039

Hints

- Compare the forms and unusual features in the two panels before using the shared numerical summary. - Recall which aspect of association correlation is designed to summarize. - Ask what information about curves and isolated points can be lost in a single coefficient.

Solution

1. The value \(r=0.70\) summarizes the direction and strength of linear association, not the complete shape of a scatterplot. 2. Data Set A forms a roughly oval cloud around an increasing straight-line pattern. 3. Data Set B has a pronounced curved pattern among the main observations and one isolated point. 4. Equal correlations can therefore occur with very different forms and unusual features. Correlation should be interpreted together with the scatterplot.

Answer

The statement is incorrect. Both data sets have the same linear correlation, but their forms and unusual features differ substantially. Data Set A is roughly linear, while Data Set B is curved and contains an isolated point.
54805212
Consider the paired data <table> <thead><tr><th>\(x\)</th><th>\(y\)</th></tr></thead> <tbody> <tr><td>\(0\)</td><td>\(1\)</td></tr> <tr><td>\(1\)</td><td>\(2\)</td></tr> <tr><td>\(2\)</td><td>\(4\)</td></tr> <tr><td>\(3\)</td><td>\(8\)</td></tr> <tr><td>\(4\)</td><td>\(16\)</td></tr> </tbody> </table> Use technology to compute the correlation and round to three decimals. Then explain why the result does not mean the data follow an exact linear relationship.

Hints

- Compute the linear correlation from the paired values. - Look at how the response changes from one row to the next. - A strong correlation can occur for a systematic pattern that is not exactly linear.

Solution

1. The correlation for the five pairs is \(r\approx0.933\), indicating a strong positive linear association. 2. The response values double whenever \(x\) increases by \(1\), so the points follow the exact exponential pattern \(y=2^x\). 3. The association is therefore perfectly systematic but nonlinear, which is why the correlation is strong yet less than \(1\).

Answer

\(r\approx0.933\). The association is strong and positive, but the points follow the exact exponential pattern \(y=2^x\), not a straight line.
54807212
A data set of paired quantitative measurements has correlation \(r=-0.71\). An analyst randomly shuffles the order of the rows but keeps the two values in each row together. What is the correlation after the row order is shuffled? Explain.

Hints

- Ask whether any \(x\)-value has been paired with a different \(y\)-value. - Distinguish changing row order from changing the pairings themselves. - A summary of paired data does not depend on which row is listed first.

Solution

1. Correlation depends on the set of paired observations, not on the order in which the rows are stored or displayed. 2. Shuffling whole rows preserves every original ordered pair. 3. Therefore, all deviations, products, and standardized pairings used in the correlation are unchanged as a collection.

Answer

The correlation remains \(r=-0.71\). Reordering complete paired observations does not change the association.
54808012
Two quantitative variables have correlation \(r=-0.43\). The analyst converts every \(x\)-value to its z-score and every \(y\)-value to its z-score. What is the correlation between \(z_x\) and \(z_y\)? Explain.

Hints

- Break standardization into a shift and a rescaling. - Check the sign of the scaling factor used for a standard deviation. - Recall which linear changes leave correlation unchanged.

Solution

1. Standardizing \(x\) subtracts \(\bar x\) and divides by the positive standard deviation \(s_x\). 2. Standardizing \(y\) likewise subtracts \(\bar y\) and divides by the positive standard deviation \(s_y\). 3. Shifts and positive rescalings do not change correlation, so the direction and strength of the linear association remain the same.

Answer

The correlation remains \(r=-0.43\).
54809012
Two data sets use the same four \(x\)-values, \(1, 2, 3, 4\), and the same four \(y\)-values, \(1, 2, 3, 4\). Data set A pairs them as \((1, 1), (2, 2), (3, 3), (4, 4)\). Data set B pairs them as \((1, 4), (2, 3), (3, 2), (4, 1)\). Compare the correlations. What does this show about using the separate distributions of \(x\) and \(y\) to determine correlation?

Hints

- Plot or mentally arrange the ordered pairs for each data set. - Check whether each set lies on an increasing or decreasing straight line. - Notice what changed between the data sets even though the individual value lists did not.

Solution

1. In data set A, all points lie exactly on an increasing line, so \(r=1\). 2. In data set B, all points lie exactly on a decreasing line, so \(r=-1\). 3. The two data sets have identical separate lists of \(x\)-values and \(y\)-values, but different pairings. 4. Therefore, the marginal distributions alone do not determine correlation; the pairing of the measurements matters.

Answer

Data set A has \(r=1\), while data set B has \(r=-1\). The separate distributions of \(x\) and \(y\) are not enough to determine correlation because correlation depends on how the values are paired.
54811912
For \(n=5\) paired observations, each variable has been converted to sample z-scores. The sum of the five products \(z_{x,i}z_{y,i}\) is \(2.8\). Find the sample correlation.

Hints

- Use the relationship between correlation and the products of paired sample z-scores. - Remember that sample standardization leads to a denominator of \(n-1\). - Check that the resulting value lies between \(-1\) and \(1\).

Solution

1. With sample z-scores, the sample correlation is \(r=\frac{1}{n-1}\sum z_{x,i}z_{y,i}\). 2. Here, \(n-1=4\), so \(r=\frac{2.8}{4}=0.70\).

Answer

The sample correlation is \(r=0.70\).
54814112
A distance variable is supposed to be recorded in miles, but half of its entries were accidentally entered in kilometers without conversion. An analyst computes the correlation between this mixed-unit column and travel time. Why is the resulting correlation not meaningful as a description of the intended data?

Hints

- Distinguish converting an entire variable from converting only some observations. - Correlation is unit-invariant only under one consistent linear rescaling of the whole variable. - Ask whether equal numerical differences currently represent equal physical distance across all rows.

Solution

1. Correlation is unchanged by a single consistent positive unit conversion applied to every value of a variable. 2. Here, however, only part of the column uses a different scale, so the recorded numbers do not represent one consistent quantitative variable. 3. The mixed units distort relative distances among observations and can change the apparent paired pattern with travel time. 4. The entries must be converted to one common unit before computing correlation.

Answer

The correlation is not meaningful because the distance column mixes two scales within one variable. Convert all distances to a common unit first, then compute the correlation.
54815112
Two quantitative variables have sample correlation \(r=0.80\), with sample standard deviations \(s_x=3\) and \(s_y=5\). Find the sample covariance between \(x\) and \(y\).

Hints

- Relate covariance to correlation and the two sample standard deviations. - Rearrange the standardization formula to isolate covariance. - Unlike correlation, covariance carries units from both variables.

Solution

1. Correlation standardizes covariance: \(r=\frac{s_{xy}}{s_xs_y}\). 2. Therefore, \(s_{xy}=rs_xs_y=0.80\cdot3\cdot5=12\).

Answer

The sample covariance is \(12\) in the product units of \(x\) and \(y\).
53944412
Use the scatterplot of outside temperature and a building’s heating-and-cooling energy use. Is \(r\) most plausibly near \(-1\), near \(0\), or near \(1\)? Explain why association strength and \(|r|\) can differ.
Figure for problem 539444

Hints

- Identify the overall shape before deciding whether a straight line summarizes it well. - Compare the direction of the points on the left side with the direction on the right side. - Correlation measures straight-line association, not closeness to every possible curve.

Solution

1. The points follow a clear U-shaped pattern, so the variables have a strong nonlinear association. 2. Correlation measures linear association. The decreasing trend on the left side is balanced by the increasing trend on the right side. 3. Therefore, \(r\) is plausibly near \(0\), even though the overall association is strong.

Answer

\(r\) should be near \(0\). The U-shaped association is strong, but its two sides have opposite linear directions that balance.
53944812
Use the scatterplot of fertilizer amount and days to germination. Would \(r\) necessarily be close to \(-1\)? Explain why the association can be strong even when \(|r|\) is not close to \(1\).
Figure for problem 539448

Hints

- Determine both the direction and the form of the plotted pattern. - Decide whether one straight line would fit the entire curve equally well. - Correlation measures the strength of a linear pattern, not closeness to every possible curve.

Solution

1. The plotted points follow a tight decreasing curve, so the variables have a strong overall association. 2. The curve flattens as fertilizer amount increases, so a straight line does not match the form perfectly. 3. Therefore, \(r\) should be negative but need not be close to \(-1\), because correlation measures only linear association.

Answer

No. \(r\) should be negative, but it need not be close to \(-1\). The points follow a strong decreasing curve rather than a straight line.
53944912
An observational study finds \(r = -0.81\) between weekly recreational screen time and nightly sleep duration. A report says, “Reducing screen time will therefore increase sleep duration.” Evaluate the claim and name a plausible third variable that could help explain the association.

Hints

- First interpret what the negative correlation says about the observed tendency. - Check whether participants were randomly assigned different amounts of screen time. - Look for a factor that could affect both computer use and available sleep time.

Solution

1. The value \(r = -0.81\) describes a strong negative linear association in the observed sample. 2. Because the study is observational, the correlation alone does not establish that changing screen time causes a change in sleep duration. 3. A plausible confounding variable is after-school workload: a heavier workload could increase computer use and also reduce sleep.

Answer

The causal claim is not justified by the observational correlation alone. After-school workload is a plausible confounding variable because it could be related to both screen time and sleep duration.
53945212
Use the scatterplot of distance from a city center and commute time. What would most likely happen to the positive correlation if the far-right point were removed? Explain.
Figure for problem 539452

Hints

- Compare the far-right observation with the direction of the main group of points. - Consider how much that point expands the range of distance values. - A high-leverage point that follows the trend often increases the magnitude of correlation.

Solution

1. The far-right point has high leverage because its distance value is far from the other values. 2. The point lies along the extension of the increasing trend. 3. Removing a high-leverage point that reinforces the trend would most likely decrease the magnitude of the positive correlation.

Answer

The positive correlation would most likely become weaker because the far-right high-leverage point extends and reinforces the increasing trend.
53945312
Use the scatterplot of paint age and color fading. What would most likely happen to the positive correlation if the far-right point were removed? Explain.
Figure for problem 539453

Hints

- Compare the far-right point with the increasing pattern formed by the main group. - A point far from the other explanatory-variable values can have substantial leverage. - Decide whether removing a point that conflicts with the trend should strengthen or weaken \(r\).

Solution

1. The far-right point has high leverage because its paint-age value is far from the rest. 2. It lies well below and opposes the main increasing trend. 3. Removing this influential point would most likely increase the positive correlation and make the remaining linear pattern appear stronger.

Answer

The positive correlation would most likely increase because removing the far-right high-leverage point eliminates an observation that strongly opposes the trend.
53945412
Use the scatterplot of number of online ads and weekly sales. The red cross marks a newly added observation. How could the added point change the correlation? Explain.
Figure for problem 539454

Hints

- First describe the direction of the blue-dot cloud without the red-cross observation. - Note that the new observation is extreme in both the horizontal and vertical directions. - Consider how one high-leverage point can affect the standardized products used in correlation.

Solution

1. The original blue-dot cloud has little linear association, so its correlation is near \(0\). 2. The added red-cross point is far to the right and high above the cloud, so it has high leverage. 3. That one point can create a noticeable positive correlation even though the main group has little linear association.

Answer

The added high-leverage point can create a misleading positive correlation even though the original cloud has little linear association.
53945512
Use the scatterplot of wave height and time to complete a fixed surfing course. The open circle shows a point’s original position on the curve, and the red cross shows the same point after it was moved upward. How would this change most likely affect the correlation? Explain.
Figure for problem 539455

Hints

- Compare the moved point’s horizontal position with the center and ends of the observed range. - Small vertical changes at central x-values usually have less influence than changes to high-leverage points. - Correlation still summarizes only the linear part of the curved pattern.

Solution

1. The moved point has low leverage because its wave-height value is near the middle of the observed range. 2. The vertical move is small relative to the full pattern, so it changes the point’s agreement with the decreasing curve only slightly. 3. The correlation would therefore change only modestly, while the overall nonlinear decreasing form remains.

Answer

The correlation would most likely change only modestly because the point is near the center of the x-range and is moved only a small amount; the nonlinear decreasing pattern remains.
54739112
For \(Y=X^2\), the three displayed \((X,Y)\) outcomes are equally likely. a) Show that \(X\) and \(Y\) have correlation \(0\). b) Let \(A=\{X>0\}\) and \(B=\{Y>0\}\). Determine whether \(A\) and \(B\) are independent. c) Explain what this example shows about zero correlation and independence.
Figure for problem 547391

Hints

- Compute \(E(X)\), \(E(Y)\), and \(E(XY)\) from the three equally likely values. - List which values of \(X\) satisfy each event. - Compare the joint event probability with the product of the two marginal probabilities.

Solution

1. \(E(X)=0\), \(E(Y)=\frac23\), and \(E(XY)=E(X^3)=0\). Therefore \(\operatorname{Cov}(X,Y)=0-0\cdot\frac23=0\). Both variables have positive variance, so their correlation is \(0\). 2. \(P(A)=\frac13\), \(P(B)=\frac23\), and \(P(A\cap B)=\frac13\). 3. Since \(\frac13\ne\frac13\cdot\frac23=\frac29\), the events are not independent. 4. Zero correlation rules out linear association but does not imply independence of the variables or of events formed from them.

Answer

a) \(\operatorname{Corr}(X,Y)=0\). b) The events are not independent because \(P(A\cap B)=\frac13\ne\frac29=P(A)P(B)\). c) Zero correlation does not imply independence.
54794712
A manufacturing study records temperature deviation from a target setting, \(x\), and a defect score, \(y\), for five runs. <table> <thead><tr><th>Temperature deviation \(x\)</th><th>Defect score \(y\)</th></tr></thead> <tbody> <tr><td>\(-2\)</td><td>\(4\)</td></tr> <tr><td>\(-1\)</td><td>\(1\)</td></tr> <tr><td>\(0\)</td><td>\(0\)</td></tr> <tr><td>\(1\)</td><td>\(1\)</td></tr> <tr><td>\(2\)</td><td>\(4\)</td></tr> </tbody> </table> Use technology to compute the correlation coefficient. Then explain why the result does not mean the variables are unrelated.

Hints

- Compute the statistic from the paired observations rather than judging only from the table. - Think about what type of association the correlation coefficient summarizes. - Look for a pattern that a straight line would fail to capture.

Solution

1. The correlation coefficient for the five ordered pairs is \(r=0.000\). 2. The points follow the exact nonlinear pattern \(y=x^2\). 3. Correlation measures linear association, so \(r=0\) indicates no linear association here, not no relationship of any kind.

Answer

\(r=0.000\). The variables have a perfect U-shaped nonlinear relationship even though their linear correlation is zero.
54796212
A data set has correlation \(r=0.62\). An analyst creates a larger file by duplicating every ordered pair exactly once, so each original observation now appears twice. The analyst claims that the correlation should become larger because the sample size doubled. Determine the new correlation and evaluate the claim.

Hints

- Ask whether the geometry of the point cloud changes when every point is copied. - Think about what happens to centered sums when every contribution is repeated equally. - Separate sample size from the amount of genuinely new information.

Solution

1. Duplicating every pair leaves the \(x\)- and \(y\)-means unchanged. 2. Every centered product and centered square contribution is duplicated, so the quantities in the correlation calculation are all multiplied by the same factor. 3. Their ratio is unchanged, so the correlation remains \(0.62\). 4. The larger row count does not create new information because the observations were copied rather than newly sampled.

Answer

The correlation remains \(r=0.62\). Duplicating every observation does not strengthen the linear association or add new information.
54796712
Two positive variables have correlation \(r=0.78\). A student replaces every \(x\)-value with \(\sqrt{x}\) and says, “The correlation must still be \(0.78\) because the transformation preserves the order of the observations.” Evaluate the claim.

Hints

- Distinguish transformations that only shift or rescale an axis from transformations that change spacing nonuniformly. - Correlation measures linear association, not just rank order. - Ask whether the transformed point cloud must have the same shape.

Solution

1. Correlation is unchanged by positive linear transformations such as changing units or adding a constant. 2. The transformation \(x\mapsto\sqrt{x}\) is nonlinear. 3. A nonlinear transformation changes relative spacing and can make a relationship more or less linear, so the new correlation cannot be determined from the original \(r\) alone.

Answer

The claim is incorrect. The new correlation need not equal \(0.78\); a nonlinear transformation can change correlation even if it preserves order.
54799712
For a data set with nonconstant explanatory values, the correlation between \(x\) and the observed response \(y\) is \(r=-0.62\). A least-squares regression line with a negative slope is fitted, and each observation is assigned its predicted value \(\hat y\). What is the correlation between \(x\) and the predicted values \(\hat y\)? Explain why it differs from the correlation between \(x\) and the observed responses.

Hints

- Think about the geometric location of all pairs \((x, \hat y)\). - Correlation reaches an extreme value when points lie exactly on a nonhorizontal straight line. - Distinguish fitted values from the original observed responses.

Solution

1. Every predicted value has the form \(\hat y=a+bx\) with the same negative slope \(b\). 2. Therefore, the ordered pairs \((x, \hat y)\) lie exactly on a decreasing straight line. 3. For nonconstant \(x\)-values, an exact decreasing linear relationship has correlation \(-1\). 4. The observed responses do not all lie on the fitted line, so their correlation with \(x\) is only \(-0.62\).

Answer

The correlation between \(x\) and \(\hat y\) is \(-1\). The fitted values lie exactly on a decreasing straight line, whereas the observed responses include deviations from that line.
54802312
A technician records the following paired values for a calibration setting \(x\) and an instrument response \(y\): <table> <thead><tr><th>\(x\)</th><th>Recorded \(y\)</th></tr></thead> <tbody> <tr><td>\(1\)</td><td>\(2\)</td></tr> <tr><td>\(2\)</td><td>\(4\)</td></tr> <tr><td>\(3\)</td><td>\(6\)</td></tr> <tr><td>\(4\)</td><td>\(8\)</td></tr> <tr><td>\(5\)</td><td>\(50\)</td></tr> </tbody> </table> Use technology to compute the correlation. A data check then shows that the last response was entered incorrectly and should be \(10\), not \(50\). Recompute the correlation after the correction and explain what the change illustrates. Round the first correlation to three decimals.

Hints

- Compute the correlation using all recorded pairs first. - Then change only the verified data-entry error and recompute. - Compare how closely each version of the data follows a straight line.

Solution

1. With the recorded value \(50\), the correlation is \(r\approx0.781\). 2. After replacing \(50\) with \(10\), the points are \((1, 2)\), \((2, 4)\), \((3, 6)\), \((4, 8)\), and \((5, 10)\), which lie exactly on an increasing straight line. 3. The corrected correlation is therefore \(r=1\). 4. The large change shows that correlation can be sensitive to an unusual or erroneous observation, especially in a small data set.

Answer

Before correction, \(r\approx0.781\). After correction, \(r=1\). The change illustrates that one unusual data value can substantially affect correlation in a small sample.
54804612
A district computes, for each of \(25\) schools, the school’s average weekly study time and average exam score. Across the \(25\) school averages, the correlation is \(r=0.88\). A student says, “Therefore, the correlation between individual students’ study time and exam score must also be about \(0.88\).” Evaluate the statement.

Hints

- Identify the observational unit used to compute the reported correlation. - Compare that unit with the unit in the student’s conclusion. - Averaging within groups removes within-group variation, which can change an association.

Solution

1. The reported correlation is calculated from \(25\) school-level averages, so its observational units are schools. 2. An individual-student correlation would be calculated from student-level pairs and can have a different value because averaging changes the variation within and between schools. 3. A strong association among group averages therefore does not determine the strength of the association among individuals.

Answer

The statement is incorrect. \(r=0.88\) describes the association among school averages, not among individual students. The student-level correlation could be different and cannot be inferred from the group-level correlation alone.
54809612
Two data sets both have correlation \(r=0.60\). In data set A, \(s_x=2\) and \(s_y=6\). In data set B, \(s_x=6\) and \(s_y=2\). Must their least-squares regression slopes be equal? Use the summary statistics to justify your answer.

Hints

- Correlation has no units, while a regression slope depends on the scales of both variables. - Relate slope to correlation and the two standard deviations. - Compare the scale ratios in the two data sets.

Solution

1. Correlation is standardized, but the least-squares slope also depends on the relative scales of the two variables. 2. For data set A, \(b=0.60\cdot\frac{6}{2}=1.8\). 3. For data set B, \(b=0.60\cdot\frac{2}{6}=0.2\). 4. Therefore, equal correlations do not imply equal regression slopes.

Answer

No. Data set A has slope \(1.8\), while data set B has slope \(0.2\). The same correlation can correspond to different slopes when the variable spreads differ.
54812612
For the three paired observations \((1, 1), (2, 3), (3, 2)\), compute the sample correlation \(r\).

Hints

- Find the mean of each variable before comparing paired deviations. - Use the centered products to measure whether the variables tend to vary together. - Standardize the covariance by the two sample standard deviations.

Solution

1. The means are \(\bar x=2\) and \(\bar y=2\). 2. The centered values are \((-1, 0, 1)\) for \(x\) and \((-1, 1, 0)\) for \(y\). 3. The sample standard deviations are \(s_x=1\) and \(s_y=1\), and the sample covariance is \(\frac{(-1)\cdot(-1)+0\cdot1+1\cdot0}{2}=0.5\). 4. Therefore, \(r=\frac{0.5}{1\cdot1}=0.50\).

Answer

The sample correlation is \(r=0.50\).
54795712
A sample has correlation \(r=0.72\) between two quantitative variables. A new observation is added whose coordinates are exactly the original sample means \((\bar x, \bar y)\). What happens to the correlation? Explain without recomputing it from all observations.

Hints

- Think about the deviations of the new point from the existing means. - Ask whether the new point changes either center. - Consider how a point with zero centered deviations contributes to a linear-association measure.

Solution

1. Adding \((\bar x, \bar y)\) does not change either sample mean because the new point is located at the existing mean of each variable. 2. Relative to those means, the new point has deviations \(0\) and \(0\), so it adds nothing to the centered cross-product sum or to either centered sum of squares. 3. In the correlation formula, the centered cross-product sum and both centered sums of squares are therefore unchanged, so the correlation remains \(0.72\).

Answer

The correlation remains \(r=0.72\).
54798812
A company tests a tuning setting \(x\) and a response score \(y\) on two different machines. <table> <thead><tr><th>Machine</th><th>\(x\)</th><th>\(y\)</th></tr></thead> <tbody> <tr><td>A</td><td>\(1\)</td><td>\(8\)</td></tr> <tr><td>A</td><td>\(2\)</td><td>\(9\)</td></tr> <tr><td>A</td><td>\(3\)</td><td>\(10\)</td></tr> <tr><td>B</td><td>\(7\)</td><td>\(1\)</td></tr> <tr><td>B</td><td>\(8\)</td><td>\(2\)</td></tr> <tr><td>B</td><td>\(9\)</td><td>\(3\)</td></tr> </tbody> </table> Use technology to find the correlation for each machine separately and for all six observations pooled together. Round the pooled correlation to three decimals. Explain why the pooled result does not have to resemble either within-machine correlation.

Hints

- Compute the linear association within each group before combining the observations. - Compare where the two groups sit relative to one another in the coordinate plane. - The correlation of pooled data is not determined by averaging the separate correlations.

Solution

1. For Machine A, the three points lie exactly on \(y=x+7\), so \(r_A=1\). 2. For Machine B, the three points lie exactly on \(y=x-6\), so \(r_B=1\). 3. Pooling all six observations gives \(r\approx-0.880\). 4. The machine groups occupy different regions of the scatterplot. Machine B has larger \(x\)-values but much smaller \(y\)-values than Machine A, so the between-group pattern dominates the pooled correlation even though each within-group relationship is perfectly positive.

Answer

Machine A: \(r=1\). Machine B: \(r=1\). Pooled data: \(r\approx-0.880\). Correlation depends on the full pattern of points. Pooling groups with different locations can produce a very different overall linear association from the association within each group.
54813512
A data set contains the three points \((0, 0), (1, 2), (2, 1)\), whose correlation is \(0.50\). The point \((0, 0)\) is then duplicated once, while the other two points remain single observations. Does the correlation stay the same? Compute the new correlation.

Hints

- Contrast duplicating one observation with duplicating every observation equally. - Recompute the means because the duplicated point changes the center of the data. - Use the centered paired deviations to find the new correlation.

Solution

1. Duplicating only one point changes that point's weight in the paired data, so correlation need not remain unchanged. 2. For the four points, \(\bar x=\bar y=0.75\). 3. The centered cross-product sum is \(1.75\), and each centered square sum is \(2.75\). 4. Therefore, \(r=\frac{1.75}{\sqrt{2.75\cdot2.75}}=\frac{1.75}{2.75}\approx0.636\).

Answer

No. The new correlation is approximately \(r=0.636\). Duplicating one observation changes its weight and can change correlation.

All problems may be used, copied and printed free of charge for school and tutoring, including paid tutoring. Commercial adaptations as well as publication or redistribution on the internet are not permitted.