Text pre-processing 3. Frequency calculation
Frequency calculation
Corpus analysis often begins by counting how many times a particular linguistic unit occurs. We will compare:
- token instances, which are individual occurrences of word forms;
- token types, which are unique word forms;
- lemma types, which group word forms assigned the same lemma.
Load the pipeline
Import spaCy and Python’s Counter class:
import spacy
from collections import Counter
Load the English transformer pipeline if nlp_en has not already been created in the current Python session. Starting a new terminal usually means starting a new Python session, so you may need to run this code again:
nlp_en = spacy.load("en_core_web_trf")
Read the corpus file
Download freq_test.txt if you do not already have the practice corpus. The file is also available in the GitHub repository.
Note: Save the downloaded file as ./corpus/freq_test.txt to your current working directory. Create a corpus folder if it does not already exist.
with open(
"./corpus/freq_test.txt", "r", encoding="utf-8"
) as f:
raw = f.read()
print(raw[:50])
print(raw[:100])
Process the text with spaCy:
doc = nlp_en(raw)
Extract tokens and lemmas
The following functions extract lowercase tokens and lemmas while excluding whitespace and punctuation:
def extract_tokens(doc):
token_list=[]
for token in doc:
if not token.is_space and not token.is_punct:
token_list.append(token)
return token_list
def extract_lemmas(doc):
lemma_list=[]
for token in doc:
if not token.is_space and not token.is_punct:
lemma_list.append(token.lemma_)
return lemma_list
Apply the functions:
tokens = extract_tokens(doc)
lemmas = extract_lemmas(doc)
print(tokens[:10])
print(lemmas[:10])
Calculate token and lemma frequencies
Python’s Counter class counts how many times each item occurs:
token_frequency = Counter(tokens)
lemma_frequency = Counter(lemmas)
Use most_common() to display units in descending order of frequency:
print(token_frequency.most_common(10))
print(lemma_frequency.most_common(10))
To display the 10 least frequent units, sort the (unit, frequency) pairs by frequency in ascending order:
print(sorted(token_frequency.items(), key=lambda item: item[1])[:10])
print(sorted(lemma_frequency.items(), key=lambda item: item[1])[:10])
Save a frequency list as a CSV file
A frequency list can be saved as a .csv file for later analysis.
Import Python’s built-in csv module:
import csv
Define a function that saves a Counter object in descending order of frequency:
def save_frequency_csv(frequency, filepath, unit_name):
with open(
filepath,
"w",
encoding="utf-8",
newline=""
) as f:
writer = csv.writer(f)
writer.writerow([unit_name, "frequency"])
writer.writerows(frequency.most_common())
The function takes three arguments:
- `frequency`: a `Counter` object containing units and their frequencies
- `filepath`: the path and filename of the CSV file to create
- `unit_name`: the label for the first column (e.g., `token`, `lemma`)
The first row contains column names. The remaining rows contain each unit and its frequency, ordered from most frequent to least frequent by `most_common()`.
Save the token-frequency list:
save_frequency_csv(
token_frequency,
"token_frequency.csv",
"token"
)
###### To-Do ######
# Save the lemma frequency list
##################