Back to Module 5

Frequency calculation

Corpus analysis often begins by counting how many times a particular linguistic unit occurs. We will compare:

  • token instances, which are individual occurrences of word forms;
  • token types, which are unique word forms;
  • lemma types, which group word forms assigned the same lemma.

Load the pipeline

Import spaCy and Python’s Counter class:

import spacy
from collections import Counter

Load the English transformer pipeline if nlp_en has not already been created in the current Python session. Starting a new terminal usually means starting a new Python session, so you may need to run this code again:

nlp_en = spacy.load("en_core_web_trf")

Read the corpus file

Download freq_test.txt if you do not already have the practice corpus. The file is also available in the GitHub repository.

Note: Save the downloaded file as ./corpus/freq_test.txt to your current working directory. Create a corpus folder if it does not already exist.

with open(
    "./corpus/freq_test.txt", "r", encoding="utf-8"
) as f:
    raw = f.read()

print(raw[:50])
print(raw[:100])

Process the text with spaCy:

doc = nlp_en(raw)

Extract tokens and lemmas

The following functions extract lowercase tokens and lemmas while excluding whitespace and punctuation:

def extract_tokens(doc):
    token_list=[]

    for token in doc:
        if not token.is_space and not token.is_punct:
            token_list.append(token)
    
    return token_list


def extract_lemmas(doc):
    lemma_list=[]

    for token in doc:
        if not token.is_space and not token.is_punct:
            lemma_list.append(token.lemma_)
    
    return lemma_list

Apply the functions:

tokens = extract_tokens(doc)
lemmas = extract_lemmas(doc)

print(tokens[:10])

print(lemmas[:10])

Calculate token and lemma frequencies

Python’s Counter class counts how many times each item occurs:

token_frequency = Counter(tokens)
lemma_frequency = Counter(lemmas)

Use most_common() to display units in descending order of frequency:

print(token_frequency.most_common(10))
print(lemma_frequency.most_common(10))

To display the 10 least frequent units, sort the (unit, frequency) pairs by frequency in ascending order:

print(sorted(token_frequency.items(), key=lambda item: item[1])[:10])
print(sorted(lemma_frequency.items(), key=lambda item: item[1])[:10])

Save a frequency list as a CSV file

A frequency list can be saved as a .csv file for later analysis.

Import Python’s built-in csv module:

import csv

Define a function that saves a Counter object in descending order of frequency:

def save_frequency_csv(frequency, filepath, unit_name):
    with open(
        filepath,
        "w",
        encoding="utf-8",
        newline=""
    ) as f:
        writer = csv.writer(f)

        writer.writerow([unit_name, "frequency"])
        writer.writerows(frequency.most_common())
The function takes three arguments:

- `frequency`: a `Counter` object containing units and their frequencies
- `filepath`: the path and filename of the CSV file to create
- `unit_name`: the label for the first column (e.g., `token`, `lemma`)

The first row contains column names. The remaining rows contain each unit and its frequency, ordered from most frequent to least frequent by `most_common()`.

Save the token-frequency list:

save_frequency_csv(
    token_frequency,
    "token_frequency.csv",
    "token"
)

###### To-Do ######
# Save the lemma frequency list
##################

Next: Exercises