Back to Module 5

Tokenization

In this tutorial, we will use spaCy for English linguistic tokenization. spaCy provides language-specific processing pipelines for tasks such as tokenization, POS tagging, morphological analysis, lemmatization, dependency parsing, and named entity recognition.

Installing spaCy

Note1. I might encourage you to make a new virtual environment for a new project. If you don’t remember how to do this, please check here.

Note2. This should be done in your terminal, NOT in your python script (which ends in .py).

First, install spaCy:

pip install spacy

You must also install a trained pipeline. In this module, we will use the English transformer pipeline:

python3 -m spacy download en_core_web_trf

Transformer models produce contextualized representations that can support accurate contextual analyses. They require more memory and processing time than smaller pipelines.

Loading the English pipeline

Note3. Now, let’s creat a python script (e.g., eng_tok.py).

import spacy

nlp_en = spacy.load("en_core_web_trf")

print(nlp_en.lang) #language

The nlp_en object is a spaCy Language object. It contains the English tokenizer and the other processing components included in the pipeline.

From a string to a Doc

When we apply a spaCy Language object to a raw string, spaCy returns a Doc object:

text = "Second Language Studies is an interdisciplinary field!"

doc = nlp_en(text) # check python3 -m pip install "numpy<2" for numpy issue

print(type(text)) # print the type of text
print(type(doc)) # print the type of doc

The basic processing workflow is:

raw string -> spaCy Language pipeline -> Doc

A doc stores the original text together with its tokens and linguistic annotations.

print(doc.text)

for token in doc:
    print(token)
    # make sure that you have an empty line here to drag and run line (shift+enter)

Each item produced by the loop is a spaCy Token object.

Inspecting tokens

Each token in a Doc has several attributes:

for token in doc:
    print(
        token.text,
        token.idx,
        token.is_punct,
        token.is_space
    )

The attributes used here are:

  • token.text: the original token string;
  • token.idx: the starting character position of the token;
  • token.is_punct: whether the token is punctuation;
  • token.is_space: whether the token consists of whitespace.

To store the tokens in a Python list:

tokens = []

for token in doc:
    tokens.append(token.text)

print(tokens)
print(type(tokens))

English tokenization

Consider contractions and possessive forms in English:

english_text = "I can't attend today's meeting." # can't, today's

english_doc = nlp_en(english_text)


###### To-Do ######
# 1. Extract all the tokens in english_text
# 2. Put them into a new list, "tokens2"
# 3. print the list "tokens2"
##################

spaCy does not simply split the text at every space. It applies English-specific tokenization rules to identify token boundaries. For example, a contraction such as can't may be divided into more than one token because it contains multiple grammatical elements.

Next: Lemmatization