Back to Module 7

English POS tagging

(Review)

Install spaCy from a terminal or command prompt:

pip install -U spacy

Then download the English transformer pipeline:

python -m spacy download en_core_web_trf

Load the pipeline

Import spaCy and load the pipeline:

import spacy

nlp_en = spacy.load("en_core_web_trf")

Process an English sentence

Create an English text:

english_text = (
    "What is a camel's favorite day of the week?")

Process the text with the pipeline:

english_doc = nlp_en(english_text)

Print the tokenized text with a loop:

for token in english_doc:
    print(token.text)

Inspecting part-of-speech annotations

We already know that each Token stored multiple linguistic annotations:

for token in english_doc:
    print(
        token.text,
        token.lemma_,
        token.pos_,
        token.tag_,
        token.morph,
        sep="\t"
    )

The important attributes in this tutorial are:

  • token.pos_: a coarse-grained UPOS tag;
  • token.tag_: a fine-grained XPOS tag;
  • token.morph: morphological features predicted for the token (if this information was trained)

Creating and ssaving an analysis table

First, import pandas

import pandas as pd 

Note. If you see an error saying that “No module name …”, exit from python and install pakcages (i.e., pip install pandas); in this case, you might need to import spacy and the model again!

Note. Pandas is a libary for organizing/analyzing tabular data. Here pd is a commonly used abbrevication for pandas.

Create a DataFrame containing the token-level annotations:

english_annotations = pd.DataFrame( 
    [ { "token": token.text,
    "lemma": token.lemma_,
    "upos": token.pos_,
    "tag": token.tag_,
    "morph": str(token.morph),
    }
    for token in english_doc 
    ] 
)

Save the DataFrame as a CSV file

english_annotations.to_csv(
    "english_annotations.csv", 
    index=False, # prevents from adding row numbers as a separate column
    encoding="utf-8-sig", # helps presenting non-English characters
)

Extracting nouns

english_text2 = "Maria bought a book at the university library."

english_doc2 = nlp_en(english_text2)

Extract the lemmas of all common nouns:

english_nouns = [
    token.lemma_.lower()
    for token in english_doc2
    if token.pos_ == "NOUN"
]

print(english_nouns)

Counting part-of-speech categories

Use Counter to calculate the frequency of each UPOS category:

from collections import Counter

english_pos_counts = Counter(
    token.pos_
    for token in english_doc
    if not token.is_space and not token.is_punct
)

print(english_pos_counts)

To display the categories from most to least frequent:

for pos, frequency in english_pos_counts.most_common():
    print(pos, frequency)

Next: English POS tagging exercise