Aimathic
Login | English | Deutsch

Free math worksheets

Build your own math worksheets from 28,000 problems for grades 3 to 12, from fractions to calculus. Every problem includes step-by-step solutions.

Learning from data

Click problems to add them to your worksheet.

53960312
A wildlife center has records for \(642\) rescued birds. A researcher selects \(50\) records and notes the recovery time for each bird. In this context, what does one datum represent, and what is the data set?

Hints

- Identify what is recorded for one selected bird. - Then describe the collection formed by repeating that recording for every selected bird.

Solution

1. One datum is the recorded recovery time for one selected bird. 2. The data set is the collection of the \(50\) recorded recovery times.

Answer

One datum is one bird’s recovery time; the data set is all \(50\) recovery times.
53960612
A district contains \(N=12{,}480\) high school students. A survey is completed by \(n=625\) of them. Interpret \(N\) and \(n\) in context.

Hints

- Connect each symbol to either the full group or the observed subset. - Check which symbol denotes the size of each group rather than a measured variable.

Solution

1. \(N=12{,}480\) is the population size: all high school students in the district. 2. \(n=625\) is the sample size: the students whose survey responses were collected.

Answer

\(N=12{,}480\) is the population size, and \(n=625\) is the sample size.
53961512
A factory produced \(24{,}000\) ceramic tiles in one week. Inspectors test \(160\) tiles. Which statement correctly describes the sample? a) Every tile produced that week b) The \(160\) tested tiles c) The measurements recorded on the tested tiles Explain.

Hints

- Separate the selected items from the information recorded about them. - Match each answer choice to population, sample, or data before choosing.

Solution

1. A sample is the subset of items selected from the population. 2. Therefore, the sample is the \(160\) tested tiles; the measurements are the data obtained from them.

Answer

b) The \(160\) tested tiles are the sample.
53961712
A study of community gardens records the number of pounds of produce harvested by each of \(35\) selected plots. Give an example of one datum and describe the complete data set without inventing values.

Hints

- Describe the form of the information rather than making up a numerical observation. - Tie one datum to one observational unit and include the measurement unit.

Solution

1. One datum is the harvest weight, in pounds, from one selected plot. 2. The data set is the collection of all \(35\) recorded harvest weights.

Answer

One datum is one selected plot’s harvest weight; the data set is all \(35\) such weights.
53962712
A city reports that the mean age of all firefighters employed by the city is \(38.6\) years. Is \(38.6\) a statistic or a parameter? Explain.

Hints

- Identify whether the numerical summary comes from a sample or the whole population. - A parameter describes an entire defined population.

Solution

1. The mean describes all firefighters employed by the city, which is the population of interest. 2. Therefore, \(38.6\) years is a parameter.

Answer

It is a parameter because it summarizes the entire population of city firefighters.
54857712
A sculpture park has exactly \(73\) permanent sculptures. Staff record the material of every permanent sculpture before updating the park guide. Is this data collection a sample or a census? Identify the population and justify your classification.

Hints

- Compare the observed group with the full group named in the study. - Decide whether any population members were left unobserved. - Use the classification that applies when the entire population is recorded.

Solution

1. The population is all \(73\) permanent sculptures in the park. 2. Staff record information from every member of that population. 3. Therefore, the data collection is a census.

Answer

It is a census. The population is all \(73\) permanent sculptures, and every sculpture in that population is included.
53960412
A streaming service has \(8{,}400{,}000\) active accounts. It wants to study how many hours users watched last week. Explain why a statistical study based on a sample may be preferable to collecting data from every active account.

Hints

- Compare the practical cost of studying the entire population with studying a subset. - Consider whether a well-chosen subset can still provide information about the full account population.

Solution

1. The population is extremely large. 2. Processing and analyzing every account may require unnecessary time and resources. 3. A suitably selected sample can provide information about the larger population.

Answer

The population is very large, so a sample can provide useful information with much less time and effort than a census.
53960712
A museum wants to know how long visitors spend in a new exhibit. Its plan states the population and the sample but does not state what will be recorded from each selected visitor. Identify the missing component needed to answer the question.

Hints

- Ask what piece of information must be collected from each sampled individual. - Name both the characteristic and a unit that would make observations comparable.

Solution

1. The study needs a datum from each selected visitor. 2. The missing information is each visitor’s time spent in the exhibit, measured in a specified unit such as minutes.

Answer

The plan must specify that the time each selected visitor spends in the exhibit will be recorded, such as in minutes.
53961212
A study uses a sample of \(120\) restaurant receipts from April to investigate tipping behavior at one restaurant. A student says the population is “the \(120\) receipts.” Correct the statement.

Hints

- Distinguish the records actually selected from the larger collection they represent. - The population should use the same place and time boundaries as the sampling frame.

Solution

1. The \(120\) receipts are the sample. 2. The population is all receipts, or all customer transactions represented by receipts, at that restaurant during April.

Answer

The \(120\) receipts are the sample; the population is all receipts at the restaurant during April.
53961812
A student wants to study the typical length of songs played by a campus radio station. The station has played \(11{,}600\) songs this semester. Explain what must be sampled and what datum should be collected from each sampled unit.

Hints

- Identify the items that make up the population and the characteristic needed from each one. - The sampling units should be songs because song length is the recorded characteristic.

Solution

1. A subset of the songs played this semester should be sampled. 2. For each sampled song, its length should be recorded in a consistent unit, such as seconds.

Answer

Sample songs played this semester and record the duration of each sampled song, such as in seconds.
53961912
Put these components in a logical order for planning a statistical study: analyze the data, define the investigative question, collect data from a sample, identify the population.

Hints

- Think about what must be decided before any observations can be collected. - Arrange the components so that each later step uses decisions made earlier.

Solution

1. Define the investigative question. 2. Identify the population addressed by that question. 3. Collect data from a sample of the population. 4. Analyze the collected data.

Answer

Define the investigative question; identify the population; collect data from a sample; analyze the data.
53962312
In a sample of \(250\) registered voters, \(58\%\) support a proposed bond. The true percentage among all registered voters in the county is unknown. Identify the statistic and the parameter.

Hints

- Identify whether the numerical summary comes from a sample or the whole population. - Trace the stated percentage back to the group from which it was calculated.

Solution

1. The sample percentage, \(58\%\), is a statistic. 2. The unknown percentage among all registered voters in the county is the population parameter.

Answer

Statistic: \(58\%\). Parameter: the true countywide percentage of registered voters who support the bond.
53963012
A sample of \(40\) light bulbs has a mean lifetime of \(1184\,\text{hours}\). The manufacturer wants to estimate the mean lifetime of every bulb of that model. Identify the statistic, parameter, and observational units.

Hints

- Match each term to the sample, the population, or the items measured. - The individual bulb is measured, while the mean summarizes a group of bulbs.

Solution

1. The statistic is the sample mean, \(1184\,\text{hours}\). 2. The parameter is the unknown mean lifetime of all bulbs of that model. 3. The observational units are the individual sampled light bulbs.

Answer

Statistic: \(1184\,\text{hours}\). Parameter: the population mean lifetime. Observational units: individual sampled bulbs.
53963812
A retailer samples \(300\) of its online orders from June and finds that \(4.7\%\) were returned. An employee says, “The parameter is \(4.7\%\).” Explain the error.

Hints

- Trace where the numerical value came from. - The unknown population percentage is the parameter the sample statistic is intended to estimate.

Solution

1. The value \(4.7\%\) was calculated from the sample, so it is a statistic. 2. The parameter is the unknown percentage of all the retailer’s June online orders that were returned.

Answer

\(4.7\%\) is a statistic; the parameter is the unknown return percentage for all the retailer’s June online orders.
54857112
An online archive studies how many document pages are viewed during a visit. Its sample contains data from \(900\) website sessions created by \(520\) distinct visitors. For this investigative question, identify the observational units and the sample size. Explain why the number of distinct visitors is not the sample size.

Hints

- Match one recorded value to the item or event that produced it. - Decide whether the question concerns people or separate visits. - Count observational units rather than unique identities.

Solution

1. The variable is pages viewed during one website session, so each session is an observational unit. 2. The sample contains \(900\) sessions, giving sample size \(n=900\). 3. Some visitors created more than one session, so the \(520\) distinct people do not equal the number of observed visits.

Answer

The observational units are website sessions, and the sample size is \(n=900\). The \(520\) visitors are not the sample size because a visitor may contribute more than one session.
54857212
A community college is considering three questions about its tutoring center. a) How many chairs are in the tutoring center today? b) What proportion of tutoring visits this semester lasted more than \(30\) minutes? c) What time does the tutoring center close on Friday? Which question requires a statistical investigation? Explain what feature makes it statistical.

Hints

- Decide which question expects variability among observations. - Separate a fixed fact from a question answered by collecting many data values. - Identify what would be measured repeatedly.

Solution

1. Questions a) and c) ask for fixed facts that have one determined answer. 2. Question b) requires data from many tutoring visits because visit duration varies from one visit to another. 3. The proportion must be calculated from the collection of visit-duration data.

Answer

b) requires a statistical investigation because tutoring-visit durations vary, so data from multiple visits are needed to determine the proportion lasting more than \(30\) minutes.
54857512
A county has \(96\) public playgrounds. The mean age of all \(96\) playgrounds is \(11.4\) years. A random sample of \(24\) playgrounds also happens to have mean age \(11.4\) years. Identify the parameter and the statistic, and explain why they are different even though their numerical values are equal.

Hints

- Trace each mean back to the group from which it was calculated. - Classify the summaries by their source rather than by their numerical size. - Consider whether equal numbers must represent the same statistical role.

Solution

1. The population mean age of all \(96\) playgrounds is the parameter. 2. The sample mean age of the \(24\) selected playgrounds is the statistic. 3. A parameter and a statistic are distinguished by the group summarized, not by whether their numerical values happen to match.

Answer

Parameter: the mean age of all \(96\) playgrounds, \(11.4\) years. Statistic: the mean age of the \(24\) sampled playgrounds, also \(11.4\) years. They differ because one summarizes the population and the other summarizes a sample.
54858012
A clinic data table has one row for each of \(75\) patients and six recorded variables in each row. A student says the sample size is \(450\) because the table contains \(75\cdot6=450\) entries. Correct the student’s reasoning.

Hints

- Determine what one row represents. - Separate the number of observational units from the number of variables recorded per unit. - Recall what the symbol \(n\) counts.

Solution

1. Sample size counts observational units, not the total number of recorded cells. 2. Each row represents one patient, so there are \(75\) observational units. 3. The six variables provide multiple data values about each patient but do not create additional patients.

Answer

The sample size is \(n=75\), not \(450\). The table contains six variables for each of \(75\) patient observational units.
54858312
During one week in October, a bike-repair shop records a median completion time of \(2.6\) days for \(41\) completed repairs. Write a complete interpretation of this statistic that does not overstate the population or time period represented.

Hints

- Name the group and time period from which the number was calculated. - Interpret the statistic using its position in the ordered data. - Avoid extending the conclusion beyond the observations described.

Solution

1. The value \(2.6\) days is the sample median for the \(41\) repairs completed during the stated October week. 2. A median interpretation states that about half of those recorded completion times were at or below \(2.6\) days and about half were at or above it. 3. The statistic does not by itself describe repairs from other weeks or all repairs the shop will complete.

Answer

Among the \(41\) repairs completed during that week in October, about half had completion times at or below \(2.6\) days and about half had completion times at or above \(2.6\) days. This statement does not automatically extend to other weeks.
54858912
A file of race-completion times contains the values \(142\), \(151\), \(163\), and \(170\), but the file does not state a unit. Explain why the data cannot be interpreted correctly until the unit is known. Give two possible interpretations of the value \(142\).

Hints

- Ask what information gives a number its real-world scale. - Imagine attaching two plausible time units to the same recorded value. - Consider whether conclusions would change under those interpretations.

Solution

1. A numerical datum needs a measurement unit to connect it to the real-world variable. 2. The value \(142\) could represent \(142\) seconds, or \(142\) minutes, among other possibilities. 3. Those interpretations describe very different completion times, so the unit must be recovered before analysis is reported.

Answer

The values are incomplete without a unit. For example, \(142\) could mean \(142\) seconds or \(142\) minutes, which are not equivalent race times.
54859112
A museum records the time visitors spend in its new gallery. A curator asks, “Do visitors spend more time in the new gallery than in the permanent gallery?” The available file contains times only for the new gallery. Explain why the file cannot answer the comparative question and identify the missing data.

Hints

- Identify every group named in the question. - Check whether the file contains observations for each group. - State what additional distribution is needed for the comparison.

Solution

1. The question compares two gallery-time distributions. 2. The file supplies observations for only the new gallery, so there is no reference distribution for the permanent gallery. 3. Comparable visit-time data for the permanent gallery are also required.

Answer

The file cannot answer the question because it includes only one comparison group. Visit-time data for the permanent gallery are missing.
54860312
A rail operator asks, “What is the distribution of arrival delays for trains on Route K?” The available schedule lists only planned arrival times, not actual arrival times. Explain why the schedule cannot answer the question and identify the required data for each train trip.

Hints

- Define what must be compared to obtain one delay. - Check which part of that comparison is absent. - State the complete datum needed for each trip.

Solution

1. A delay compares an actual arrival time with the corresponding planned arrival time. 2. The schedule supplies only the planned times, so no trip-level delay can be calculated. 3. Each train trip needs both its planned and actual arrival times, or a directly recorded arrival delay.

Answer

The schedule alone cannot provide delays. Each Route K trip needs a matched planned and actual arrival time, or a recorded delay value.
54856912
Before collecting data, a school’s theater club chose this investigative question: “What is the typical number of hours members spend rehearsing each week?” After seeing that several members reported very large values, the club proposed a follow-up question: “Why do a few members rehearse much longer than everyone else?” Explain why the follow-up question should not replace the original question for this study, and state one appropriate way to use the collected data.

Hints

- Distinguish a question chosen before seeing the data from one suggested by an observed pattern. - Compare the information needed to describe unusually large values with the information needed to explain their causes. - Think about how an exploratory finding can guide a later investigation without replacing the original analysis.

Solution

1. The original question was specified before the results were examined, so it remains the primary question for this study. 2. The follow-up question was prompted by an observed feature and asks about causes, but rehearsal-hour values alone do not identify why the large values occurred. 3. The club may report the original descriptive analysis and treat the new question as an exploratory question for a later study that collects possible explanatory variables.

Answer

The original question should remain the primary question because it was set before the data were examined. The new question is a reasonable exploratory follow-up, but the recorded rehearsal hours alone cannot establish causes. The club should describe the distribution for the original question and use the unexpected values to motivate a later study.
54857012
A transit agency records the number of riders in each train car whenever the train reaches a station. An analyst wants to answer, “What is the typical length of an individual rider’s trip?” Explain why the recorded counts are not sufficient, and identify the additional information needed.

Hints

- Identify the observational unit required by the question. - Ask whether the current records connect the beginning and end of the same observation. - State what one complete datum would need to contain.

Solution

1. The car counts show how many riders are present at each station, but they do not link a rider’s boarding station to that rider’s exit station. 2. Trip length is defined for an individual ride, so the data must identify where each observed ride begins and ends, or directly record its distance or number of stops.

Answer

The counts are not sufficient because they do not track individual rides from boarding to exit. The agency needs each ride’s starting and ending stations, or an equivalent trip-length measurement.
54857312
A public library wants to estimate the proportion of all attempts to borrow a digital book that fail because every license is in use. The library’s current file contains only completed digital checkouts. Can that file answer the investigative question? Explain and identify what records are missing.

Hints

- Compare the population named in the question with the observations included in the file. - Determine whether both possible outcomes are represented. - Identify the denominator required for the requested proportion.

Solution

1. The population of interest is all attempts to borrow a digital book, including successful and unsuccessful attempts. 2. The current file contains only successful attempts, so failures are absent and the desired proportion cannot be calculated. 3. The library must record every borrowing attempt and whether it resulted in a completed checkout or a license-unavailable failure.

Answer

No. The file excludes failed attempts, so it cannot provide the requested proportion. Records of all borrowing attempts and their outcomes are needed.
54857412
A college wants to study the number of credits carried by students enrolled on September \(15\). Its file includes everyone who registered at any time during the semester, including students who withdrew before September \(15\) and students who enrolled afterward. Explain why the file does not match the population in the investigative question, and state which records should be included.

Hints

- Locate the time condition that defines membership in the population. - Compare that condition with the rule used to build the file. - Decide which records fall outside the stated boundary.

Solution

1. The population is defined by enrollment status on September \(15\), not by whether a person registered at any point during the semester. 2. Students who withdrew before that date and students who enrolled later are outside the population. 3. The file should contain only students enrolled on September \(15\), with each student’s credit load on that date.

Answer

The file uses the wrong time boundary. It should include exactly the students enrolled on September \(15\) and record each one’s number of credits on that date.
54857612
A greenhouse sensor records temperature every \(10\) minutes. A researcher’s question is, “What was the highest temperature on each day in July?” The raw file has one row for every sensor reading. For the final data set that answers the question, what should the observational units and data values be?

Hints

- Identify the phrase that tells how many final values the question requires. - Distinguish raw sensor readings from the summarized observations used in the analysis. - Decide what one completed row of the analysis data set should represent.

Solution

1. The question asks for one result for each day, so the observational units in the final data set are the days in July. 2. All readings within a day must be combined to find that day’s maximum temperature. 3. Each datum in the final data set is the highest recorded temperature for one day.

Answer

The observational units should be the days in July, and each datum should be that day’s maximum recorded temperature.
54857812
A student proposes the question, “Are the new study rooms better than the old study rooms?” Explain why this is not yet a well-defined investigative question. Rewrite it as one measurable statistical question about noise.

Hints

- Identify the word whose meaning could differ from person to person. - Replace the vague judgment with a characteristic that can be recorded consistently. - Include the groups and conditions needed for a fair comparison.

Solution

1. The word “better” does not identify a specific characteristic to measure or a defined comparison. 2. A measurable question must specify the observational units, the variable, and the groups or conditions being compared. 3. One valid revision is: “How does the distribution of sound levels, in decibels, during occupied hours compare between the new and old study rooms?”

Answer

“Better” is undefined, so the proposed question does not specify what data are needed. One valid revision is: “How does the distribution of sound levels, in decibels, during occupied hours compare between the new and old study rooms?”
54857912
A marine center asks, “What is the typical recovery time for all sea turtles treated this year?” Its data file contains recovery times only for green sea turtles, although the center also treated loggerhead and leatherback turtles. Explain why the file cannot answer the stated question and give two valid ways to resolve the mismatch.

Hints

- Compare the group named in the question with the group represented in the file. - Look for members of the population who have no data. - Consider changing either the data collection or the scope of the question.

Solution

1. The stated population includes all sea turtles treated this year, but the file represents only green sea turtles. 2. One resolution is to collect recovery times for every treated species so the data match the original population. 3. Another resolution is to narrow the investigative question to green sea turtles treated this year.

Answer

The file does not represent the full stated population. The center can either add data for the other treated species or revise the question so it concerns only green sea turtles treated this year.
54858112
A company wants to study the variability in customer wait times among individual service calls. Its monthly report gives only one average wait time for each of \(18\) branch offices. Explain why those \(18\) averages cannot reveal the distribution of individual call wait times.

Hints

- Identify the level at which the question asks about variation. - Compare that level with what one reported value represents. - Consider what information is lost when many observations are reduced to one average.

Solution

1. Each reported value summarizes an entire branch rather than representing one service call. 2. Many different call-level distributions can have the same branch average. 3. To study variability among individual calls, the company needs the wait time for each sampled call, not only branch averages.

Answer

The branch averages hide the variation among calls within each branch. Individual call wait times are needed to describe the distribution of call-level waits.
54858212
Two ovens bake the same type of ceramic piece. A technician records the firing temperature for each piece but accidentally removes the oven label from every record. Can the resulting data answer, “How do firing-temperature distributions compare between Oven A and Oven B?” Explain.

Hints

- List the pieces of information needed to form each comparison group. - Determine whether every recorded value can still be linked to its source. - Distinguish analyzing pooled data from comparing labeled groups.

Solution

1. The temperature values remain available, but the contextual variable identifying the oven is missing. 2. Without that label, each observation cannot be assigned to Oven A or Oven B. 3. The combined values can describe overall firing temperatures, but they cannot support a comparison between the two ovens.

Answer

No. The oven label is required to separate the observations into the two groups. Without it, only the combined temperature distribution can be studied.
54858412
An outdoor festival wants to estimate the proportion of attendees who used public transportation. Its ticket file has one row for each online purchaser, but a purchaser may buy tickets for several people. Explain why purchasers are not the correct observational units for the stated question and identify the appropriate units.

Hints

- Match the denominator of the requested proportion to the units in the data. - Check whether one database row can stand for more than one person in the population. - Identify who actually has the characteristic being measured.

Solution

1. The question concerns individual attendees, not ticket transactions or purchasers. 2. One purchaser can represent several attendees whose transportation choices may differ. 3. The observational units should be individual festival attendees, with transportation method recorded for each attendee.

Answer

Purchasers are not the correct units because one purchase can cover several attendees. The observational units should be individual attendees.
54858512
A sample contains \(240\) streetlights. Inspection results are missing for \(18\) of them, and \(40\) of the \(222\) streetlights with recorded results failed. Calculate the recorded-result completion rate and the failure rate among streetlights with known results. Explain why a report should state both denominators.

Hints

- Separate the question about available results from the question about failures. - Identify the appropriate denominator for each percentage. - Do not silently treat missing outcomes as successful inspections.

Solution

1. The recorded-result completion rate is \(\frac{222}{240}=0.925=92.5\%\). 2. The failure rate among known results is \(\frac{40}{222}\approx0.1802\approx18.0\%\). 3. The first denominator describes data completeness, while the second describes failures only among streetlights whose outcomes are known.

Answer

The completion rate is \(92.5\%\). The failure rate among known results is approximately \(18.0\%\). Both denominators should be reported because missing inspection outcomes are not known passes or failures.
54858612
A researcher writes only, “What is the distribution of monthly electricity use?” Identify two important boundaries missing from the investigative question and rewrite it as a well-defined question about apartments in one building.

Hints

- Ask who or what will supply each observation. - Look for a time boundary that makes the measurement comparable. - Include a unit so the recorded variable is clear.

Solution

1. The question does not identify which observational units belong to the population. 2. It also does not specify the month or time period for the measurement. 3. One valid revision is: “What is the distribution of electricity use, in kilowatt-hours, among occupied apartments in Building C during January?”

Answer

The population and time period are missing. One valid revision is: “What is the distribution of electricity use, in kilowatt-hours, among occupied apartments in Building C during January?”
54858712
A warehouse wants the proportion of shipped items that were damaged. Its file has one row per order and records only whether the order contained at least one damaged item. Explain why this order-level file cannot determine the item-level damage proportion, and state what additional counts are required.

Hints

- Identify the observational units named in the requested proportion. - Compare those units with what one file row represents. - Determine the numerator and denominator that the desired proportion needs.

Solution

1. The requested proportion uses individual shipped items as observational units. 2. An order may contain several items, and the current indicator does not show how many items were damaged or how many items were shipped. 3. The warehouse needs the total number of shipped items and the number of damaged items, ideally recorded at the item level.

Answer

The file summarizes orders rather than items, so it cannot provide an item-level damage proportion. The total number of shipped items and the number of damaged items are required.
54858812
A survey export contains a column labeled “time,” but the data dictionary is missing. The values could be elapsed seconds used to complete the survey or clock times when responses were submitted. Explain why the distribution should not be analyzed until the variable definition is recovered, and name two pieces of metadata that the definition must include.

Hints

- Consider whether one column label can refer to more than one mathematical quantity. - Ask what information is needed to attach meaning to a numerical value. - Separate a variable’s name from its operational definition.

Solution

1. Elapsed duration and clock time are different variables with different meanings, units, and useful summaries. 2. The numerical values cannot be interpreted or graphed correctly without knowing which quantity was recorded. 3. The data dictionary should identify the measured quantity and its unit or time convention, including how the value was calculated or encoded.

Answer

Analysis must wait because “time” does not identify the variable’s meaning. The metadata should state what the values measure and the unit or clock convention used, together with the recording rule.
54859012
Two airports provide files labeled “canceled flights.” Airport A counts only flights canceled on the day of departure. Airport B also counts flights canceled several days in advance. An analyst wants to combine the files and describe one cancellation distribution. Explain the problem and what must happen before the data are combined.

Hints

- Compare the exact rule used to include an observation in each file. - Decide whether identical labels guarantee identical measurements. - State what consistency is required for a combined data set.

Solution

1. The two files use different operational definitions for the same label. 2. Their recorded values are therefore not directly comparable because the variable is measured under different rules. 3. The airports must adopt one common definition and recode or recollect the data consistently before combining them.

Answer

The label does not represent the same variable in both files. A common definition of “canceled flight” must be applied consistently before the data can be combined.
54859212
A city wants to study the change in annual water use for each municipal building after conservation upgrades. The current file contains only each building’s water use after the upgrades. Explain why the file cannot answer the question and describe the paired information needed.

Hints

- Determine how a change value is defined for one observational unit. - Check whether both required time points are present. - Explain why the measurements must be linked to the same building.

Solution

1. A change for one building requires two measurements for that same building. 2. The current file contains only the after-upgrade value, so a before-to-after difference cannot be calculated. 3. Each building needs both its before-upgrade and after-upgrade annual water use, linked by building.

Answer

The file lacks a baseline for each building. It must contain matched before- and after-upgrade water-use values for every observed building.
54859412
A forestry study planted \(600\) seedlings and, three years later, measured the heights of the \(418\) seedlings still alive. A report claims that the height data describe all \(600\) originally planted seedlings after three years. Explain the flaw in that statement.

Hints

- Compare the original group with the group that could actually be measured later. - Identify which observational units have no response value. - Limit the conclusion to the population represented by the available data.

Solution

1. Height after three years was recorded only for seedlings that survived. 2. The \(182\) seedlings that died have no three-year height measurements and may differ systematically from the survivors. 3. The observed height distribution describes the surviving seedlings, not all originally planted seedlings.

Answer

The data describe only the \(418\) surviving seedlings. They cannot represent the three-year heights of all \(600\) originally planted seedlings because the nonsurvivors are missing.
54859512
A fitness app stores \(0\) in the “daily exercise minutes” field both when a user completed no exercise and when the device failed to sync. Explain why the distribution cannot be interpreted correctly from this field alone and identify the additional variable the data set needs.

Hints

- Distinguish a measured value of zero from the absence of a measurement. - Consider how combining the two meanings changes the graph at zero. - Identify a separate field that would resolve the ambiguity.

Solution

1. A true value of \(0\) is an observed outcome, while a sync failure means the exercise value is unknown. 2. Using the same code for both conditions mixes valid zeros with missing data and can inflate the apparent frequency at \(0\). 3. The file needs a separate synchronization or measurement-status variable so true zeros can be distinguished from missing observations.

Answer

The value \(0\) has two incompatible meanings, so the zero frequency is ambiguous. Add a status indicator that distinguishes a valid zero from a failed or missing measurement.
54859612
A file contains the daily high temperatures for \(30\) consecutive days, but the rows were randomly reordered and the dates were removed. A researcher wants to find the longest run of consecutive days with a high above \(90\,{}^\circ\text{F}\). Explain why the values alone are insufficient and identify the information that must be restored.

Hints

- Decide whether the requested result depends only on the set of values or also on their order. - Consider what random reordering preserves and what it destroys. - Name the metadata needed to reconstruct adjacency.

Solution

1. A run of consecutive days depends on the chronological order of the observations. 2. Randomly reordered temperatures preserve the one-variable distribution but destroy the adjacency relationships between days. 3. The date or original sequence position for every temperature must be restored before the longest run can be determined.

Answer

The temperature values preserve the distribution but not which days were consecutive. Restore each observation’s date or chronological position before finding the longest run.
54859712
A shipping company combines pickup and delivery timestamps from regional offices. Some offices record local time, while others record Coordinated Universal Time, and the file does not identify which convention each office used. Why should delivery durations not be calculated yet, and what correction is required?

Hints

- Determine what must be consistent before subtracting two time records. - Consider how identical clock readings can represent different moments in different locations. - State the data-cleaning step needed before analysis.

Solution

1. A duration requires pickup and delivery timestamps expressed on the same time scale. 2. Mixing unidentified local times with Coordinated Universal Time can create incorrect elapsed times. 3. The company must identify each timestamp’s time zone and convert all timestamps to one common standard before calculating durations.

Answer

Delivery durations should not be calculated until every timestamp’s time zone is known and all times are converted to a common standard.
54859812
Halfway through a customer study, a survey changed its satisfaction scale from \(1\)–\(5\) to \(1\)–\(10\), but the final file places all ratings in one column without identifying the scale used. Explain why the raw ratings should not be analyzed together and what information is needed.

Hints

- Compare what a particular numerical rating means under each scale. - Decide whether equal recorded numbers represent equal satisfaction levels. - Identify the contextual label needed for each response.

Solution

1. The same recorded number has different meanings on the two scales. 2. Combining the raw values would treat non-equivalent ratings as though they were measured identically. 3. Each record needs a scale-version label, after which the ratings can be analyzed separately or converted using a justified common definition.

Answer

The ratings are not directly comparable because they come from different scales. Each response must be linked to its scale version before the data are separated or validly converted.
54859912
A vehicle-speed file combines records from two countries. Some speeds are in miles per hour and others are in kilometers per hour, but the unit label was lost. Explain why the numerical values cannot be safely combined and what must be recovered before analysis.

Hints

- Ask whether the same numerical value represents the same physical speed in both systems. - Identify the information needed to convert a measurement. - State the consistency required before forming one distribution.

Solution

1. Miles per hour and kilometers per hour use different numerical scales for the same physical speed. 2. Without a unit label, a recorded value cannot be interpreted or converted correctly. 3. The unit for each observation must be recovered, and all speeds must then be converted to one common unit.

Answer

The values cannot be safely combined because their units are unknown. Each observation’s unit must be recovered before all speeds are converted to a common scale.
54860012
In a waiting-time file, any wait longer than \(90\) minutes is recorded only as “\(90+\).” Can the file be used to determine the exact maximum wait and the exact mean wait? Explain what the censored records do and do not reveal.

Hints

- Separate a lower bound from an exact measurement. - Check which requested summaries require the actual value of every observation. - Identify what information remains known for each capped record.

Solution

1. Each “\(90+\)” record establishes only that the wait exceeded \(90\) minutes. 2. The exact largest wait cannot be identified because the values above \(90\) are unknown. 3. The exact mean also cannot be calculated because the censored observations do not have exact numerical values.

Answer

No. The file shows only that certain waits exceeded \(90\) minutes. It cannot provide the exact maximum or exact mean without the actual values of those waits.
54860112
A database join produces \(1260\) rows for \(900\) unique library loans because loans with multiple authors appear once for each author. An analyst wants the distribution of checkout duration per loan. Explain why using all \(1260\) rows would be incorrect and how the data should be organized.

Hints

- Identify what one observation should represent for the question. - Determine why some real-world units appear in several rows. - Remove repeated representations without deleting distinct loans.

Solution

1. The observational unit is a library loan, not a loan-author combination. 2. Repeated rows would give extra weight to loans associated with multiple authors. 3. The final data set should contain one checkout-duration value for each of the \(900\) unique loans.

Answer

Using all \(1260\) rows would count some loans more than once. The analysis file should have one row and one checkout-duration value per unique loan, for sample size \(n=900\).
54860212
A city calculates a monthly recycling rate by dividing recycled material by total waste collected in the same month. During a file merge, each month’s recycled amount was paired with the total waste amount from the following month. Explain why the resulting rates are invalid and state the key field that must be used to repair the merge.

Hints

- Check whether the numerator and denominator describe the same time period. - Identify which field should uniquely align the two files. - Consider what real-world quantity a mismatched ratio would represent.

Solution

1. A rate requires a numerator and denominator referring to the same population and time period. 2. Pairing values from different months creates ratios that do not describe any month’s recycling performance. 3. The files must be joined using a common month identifier so each recycled amount is matched with total waste from that same month.

Answer

The rates are invalid because their numerators and denominators refer to different months. Rejoin the files using the month as the matching key.
54859312
An air-quality monitor records one value every minute from \(6\) a.m. to \(10\) p.m. but only one value per hour overnight. An analyst makes an unweighted distribution of all recorded values and calls it “the distribution of air quality over a full day.” Explain the weighting problem and give one way to make equal amounts of clock time contribute equally.

Hints

- Compare how many rows are contributed by daytime and overnight hours. - Decide whether the intended observational units are readings or equal units of time. - Consider a data structure in which every hour contributes the same amount.

Solution

1. Daytime hours contribute many more recorded values than overnight hours, so the unweighted distribution describes recorded readings rather than equal portions of time. 2. Daytime conditions are overrepresented because of the higher sampling frequency. 3. One solution is to summarize the data into equal-length time intervals, such as one value per hour, before forming the distribution. Equivalent time-based weights could also be used.

Answer

The unweighted data overrepresent daytime because readings are taken more frequently then. Use one summary per equal time interval, or apply time-based weights, so each hour contributes equally.

All problems may be used, copied and printed free of charge for school and tutoring, including paid tutoring. Commercial adaptations as well as publication or redistribution on the internet are not permitted.