Text pre-processing 1. Tokenization
Tokenization
In this tutorial, we will use spaCy for English linguistic tokenization. spaCy provides language-specific processing pipelines for tasks such as tokenization, POS tagging, morphological analysis, lemmatization, dependency parsing, and named entity recognition.
Installing spaCy
Note1. I might encourage you to make a new virtual environment for a new project. If you don’t remember how to do this, please check here.
Note2. This should be done in your terminal, NOT in your python script (which ends in .py).
First, install spaCy:
pip install spacy
You must also install a trained pipeline. In this module, we will use the English transformer pipeline:
python3 -m spacy download en_core_web_trf
Transformer models produce contextualized representations that can support accurate contextual analyses. They require more memory and processing time than smaller pipelines.
Loading the English pipeline
Note3. Now, let’s creat a python script (e.g., eng_tok.py).
import spacy
nlp_en = spacy.load("en_core_web_trf")
print(nlp_en.lang) #language
The nlp_en object is a spaCy Language object. It contains the English tokenizer and the other processing components included in the pipeline.
From a string to a Doc
When we apply a spaCy Language object to a raw string, spaCy returns a Doc object:
text = "Second Language Studies is an interdisciplinary field!"
doc = nlp_en(text) # check python3 -m pip install "numpy<2" for numpy issue
print(type(text)) # print the type of text
print(type(doc)) # print the type of doc
The basic processing workflow is:
raw string -> spaCy Language pipeline -> Doc
A doc stores the original text together with its tokens and linguistic annotations.
print(doc.text)
for token in doc:
print(token)
# make sure that you have an empty line here to drag and run line (shift+enter)
Each item produced by the loop is a spaCy Token object.
Inspecting tokens
Each token in a Doc has several attributes:
for token in doc:
print(
token.text,
token.idx,
token.is_punct,
token.is_space
)
The attributes used here are:
token.text: the original token string;token.idx: the starting character position of the token;token.is_punct: whether the token is punctuation;token.is_space: whether the token consists of whitespace.
To store the tokens in a Python list:
tokens = []
for token in doc:
tokens.append(token.text)
print(tokens)
print(type(tokens))
English tokenization
Consider contractions and possessive forms in English:
english_text = "I can't attend today's meeting." # can't, today's
english_doc = nlp_en(english_text)
###### To-Do ######
# 1. Extract all the tokens in english_text
# 2. Put them into a new list, "tokens2"
# 3. print the list "tokens2"
##################
spaCy does not simply split the text at every space. It applies English-specific tokenization rules to identify token boundaries. For example, a contraction such as can't may be divided into more than one token because it contains multiple grammatical elements.