5135556
A computer analyzes an English-language news article containing \(400\) letters. The letter E appears \(52\) times, N appears \(28\) times, and Q appears once.
a) Find the relative frequencies of E, N, and Q. Give each result as a percent.
b) A reference model for this activity uses \(12.7\%\) as the long-run proportion of E in English text. Compare your result from part a with this value and give one possible reason for the difference.
Hints
- Divide each letter count by the total number of letters.
- How do you convert a decimal to a percent?
- Why might one article differ from a long-run language model?
Solution
1. a) The relative frequency of E is \(\frac{52}{400} = 0.13 = 13\%\).
2. The relative frequency of N is \(\frac{28}{400} = 0.07 = 7\%\).
3. The relative frequency of Q is \(\frac{1}{400} = 0.0025 = 0.25\%\).
4. b) The observed E frequency of \(13\%\) is close to the model value of \(12.7\%\). A particular article can differ because it is a limited sample and its topic and word choices affect letter counts.
Answer
a) E: \(13\%\); N: \(7\%\); Q: \(0.25\%\)
b) The observed \(13\%\) is slightly above \(12.7\%\). Sampling variation and the article’s vocabulary can explain the difference.
