Text pre-processing 4. Exercise 6
Comparing frequency lists with spaCy
Choose two texts containing (at least 100 words each). The texts should be in the same language and should be comparable in some meaningful way.
You may use texts from a public source, your own writing, or ideally samples of the corpus that you plan to analyze for your final project. Choose a spaCy pipeline that supports the language of both texts. You can browse available pipelines on the spaCy Models page. For a Korean example, see Applying the workflow to another language.
Part 1. Select two texts
Record the following information for both texts:
- file name
- language
- source
- number of words or other relevant text units
- reason for comparing the two texts
For a language that does not normally use spaces to separate words, explain how you counted the relevant text units.
Create two file-path variables, one for each text. Then use open() with UTF-8 encoding and read each file into a separate string variable.
Check the first 100 characters of each string and count whitespace-separated words with .split(). Do the two texts meet the minimum length requirement?
Part 2. Process both texts with spaCy
Install spaCy and download one pipeline that supports the language of both texts. Use the same pipeline for both texts so that their analyses are comparable.
Import spaCy, load the selected pipeline, and create one Doc object for each text. Record the model name, language, and pipeline components. Which annotations will you need for this exercise?
Inspect the first 20 tokens from each Doc. For each token, display its original form, lemma, and part-of-speech tag. Exclude spaces from this inspection.
Part 3. Create frequency lists
Review the extraction function from 5-3. Write a function that receives a Doc and returns a list of lowercase lemmas, excluding spaces and punctuation.
Call the function for both documents and use Counter to create one lemma-frequency list for each text.
For each text, report:
- lemma instances: the total number of lemmas, including repetitions;
- lemma types: the number of unique lemmas.
Part 4. Compare the frequency lists
Use Counter.most_common(20) to print the 20 most frequent lemmas in each text. Compare the lists qualitatively. Which lemmas appear in both texts, and which seem specific to one text?
Because the texts may have different lengths, raw counts are not enough. Calculate lemma frequencies per 1,000 lemma instances using count / total lemma instances * 1,000.
Use this function skeleton as a hint:
def normalized_frequency(frequency, total_units, number_of_units=1000):
normalized = {}
for unit, count in frequency.items():
# Store the normalized frequency for this unit.
normalized[unit] = ...
return normalized
Apply the function separately to the two lemma-frequency lists. Use the number of retained lemma instances in each text as the denominator.
Create a set containing the union of the lemmas found in both texts. For each lemma, retrieve its normalized frequency in Text A and Text B. Use 0 when a lemma is absent from one text, and round the displayed values to two decimal places.
You may focus this comparison on the most frequent units or on units that appear noticeably more often in one text than in the other. Explain why you selected those units.
Part 5. Report your observations
Submit a short descriptive report containing:
- the research question or reason for comparing the texts
- the text sources, languages, and sizes
- the spaCy pipeline name and components
- lemma instances and types for each text
- the 20 most frequent lemmas in each text
- selected normalized frequencies per 1,000 lemma instances
- two observations about similarities and differences between the texts
Submit the two text files (or their source links), your code, the lemma frequency lists, and the report.