Text pre-processing 2. Lemmatization
Lemmatization
Lemmatization converts an inflected word form into a base or dictionary form called a lemma. For example:
writes -> write
wrote -> write
writing -> write
We will continue using spaCy pipeline for the lemmatization.
English lemmatization
Process an English sentence using the English pipeline:
lemma_text = "I am running because I want to eat some ice cream!"
lemma_doc = nlp_en(lemma_text)
Each Token object stores several linguistic annotations:
for token in lemma_doc:
print(
token.text,
token.lemma_,
token.pos_,
token.tag_,
token.morph
)
The attributes used here are:
token.text: the original word form;token.lemma_: the lemma assigned by spaCy;token.pos_: the universal part-of-speech category;token.tag_: a more language-specific part-of-speech tag;token.morph: morphological features associated with the token.
For example, the surface form running should be associated with the lemma run.
To inspect only the word form and lemma:
for token in lemma_doc:
print(token.text, "->", token.lemma_)
Extracting English lemmas
To extract the lemmas as a list:
lemmas = []
for token in lemma_doc:
lemmas.append(token.lemma_)
print(lemmas)
To exclude spaces and punctuation:
lemmas2=[]
for token in lemma_doc:
if not token.is_space and not token.is_punct:
lemmas2.append(token.lemma_)
print(lemmas2)
Sometimes you might prefer to save the lemmas in lowercase.
lemmas3 = []
for token in lemma_doc:
if not token.is_space and not token.is_punct:
lemmas3.append(token.lemma_.lower()) #.lower() was the method that we've learned earlier.
print(lemmas3)
Inspecting a document
The same attributes can be inspected throughout an English document:
longer_text=
"""
Second Language Studies (SLS) is an interdisciplinary field that addresses the learning, teaching, and use of second (or multiple) languages from educational, linguistic, psychological, sociological, and anthropological perspectives. Our undergraduate and graduate programs are relevant to students interested in ESL and TESOL. Our programs also go well beyond those areas and include teaching and researching languages other than English.
"""
longer_doc = nlp_en(longer_text)
def inspect_doc(doc):
for token in doc:
if not token.is_space:
print(
token.text,
token.lemma_,
token.pos_,
sep="\t"
)
inspect_doc(english_doc)
A reusable lemma-extraction function
def extract_lemmas(text, nlp):
doc = nlp(text)
lemma_list = []
for token in doc:
if not token.is_space and not token.is_punct:
lemma_list.append(token.lemma_.lower())
return lemma_list
longer_text_lemmas = extract_lemmas(longer_text, nlp_en)
print(longer_text_lemmas)
print(len(longer_text_lemmas)) #counts all lemmas, including repeated lemmas
print(len(set(longer_text_lemmas))) #counts the number of unique lemmas
Context-sensitive lemmatization in English
Lemmatization is not simply the removal of suffixes. The assigned lemma may depend on the token’s grammatical context and part of speech.
Consider the word saw:
text = "I saw a saw."
doc = nlp_en(text)
for token in doc:
print(
token.text,
token.lemma_,
token.pos_
)
The two occurrences have the same surface form but different grammatical functions. The first is a verb and the second is a noun.