English POS tagging
English POS tagging
(Review)
Install spaCy from a terminal or command prompt:
pip install -U spacy
Then download the English transformer pipeline:
python -m spacy download en_core_web_trf
Load the pipeline
Import spaCy and load the pipeline:
import spacy
nlp_en = spacy.load("en_core_web_trf")
Process an English sentence
Create an English text:
english_text = (
"What is a camel's favorite day of the week?")
Process the text with the pipeline:
english_doc = nlp_en(english_text)
Print the tokenized text with a loop:
for token in english_doc:
print(token.text)
Inspecting part-of-speech annotations
We already know that each Token stored multiple linguistic annotations:
for token in english_doc:
print(
token.text,
token.lemma_,
token.pos_,
token.tag_,
token.morph,
sep="\t"
)
The important attributes in this tutorial are:
token.pos_: a coarse-grained UPOS tag;token.tag_: a fine-grained XPOS tag;token.morph: morphological features predicted for the token (if this information was trained)
Creating and ssaving an analysis table
First, import pandas
import pandas as pd
Note. If you see an error saying that “No module name …”, exit from python and install pakcages (i.e., pip install pandas); in this case, you might need to import spacy and the model again!
Note. Pandas is a libary for organizing/analyzing tabular data. Here pd is a commonly used abbrevication for pandas.
Create a DataFrame containing the token-level annotations:
english_annotations = pd.DataFrame(
[ { "token": token.text,
"lemma": token.lemma_,
"upos": token.pos_,
"tag": token.tag_,
"morph": str(token.morph),
}
for token in english_doc
]
)
Save the DataFrame as a CSV file
english_annotations.to_csv(
"english_annotations.csv",
index=False, # prevents from adding row numbers as a separate column
encoding="utf-8-sig", # helps presenting non-English characters
)
Extracting nouns
english_text2 = "Maria bought a book at the university library."
english_doc2 = nlp_en(english_text2)
Extract the lemmas of all common nouns:
english_nouns = [
token.lemma_.lower()
for token in english_doc2
if token.pos_ == "NOUN"
]
print(english_nouns)
Counting part-of-speech categories
Use Counter to calculate the frequency of each UPOS category:
from collections import Counter
english_pos_counts = Counter(
token.pos_
for token in english_doc
if not token.is_space and not token.is_punct
)
print(english_pos_counts)
To display the categories from most to least frequent:
for pos, frequency in english_pos_counts.most_common():
print(pos, frequency)