English POS tagging exercise
POS analysis of multiple files
Download sample.zip. In this task, you will:
- download the sample corpus to your local computer;
- process all
.txtfiles using spaCy; - count the UPOS and XPOS tags in each tex (both raw and normalized frequency);
- organize the results in pandas DataFrames;
- save the results as CSV files.
This would be a part of Assignment 3.
Download the sample corpus
After extracting the ZIP file, specify the local folder containing the .txt files:
from pathlib import Path
sample_folder = Path("YOUR LOCAL PATH")
text_files = sorted(sample_folder.glob("*.txt"))
print(f"Number of text files: {len(text_files)}")
for text_file in text_files:
print(text_file.name)
Expected output:
Number of text files: 5
sample1.txt
sample2.txt
sample3.txt
sample4.txt
sample5.txt
Process all text files
Read each .txt file and process the texts using nlp_en.pipe():
texts = [
text_file.read_text(encoding="utf-8")
for text_file in text_files
]
docs = list(nlp_en.pipe(texts))
Collect token-level POS annotations
For the frequency analysis, create another DataFrame that excludes spaces and punctuation:
token_annotations = []
for text_file, doc in zip(text_files, docs):
for token in doc:
if not token.is_space and not token.is_punct:
token_annotations.append(
{
"file": text_file.name,
"token": token.text,
"lemma": token.lemma_,
"upos": token.pos_,
"xpos": token.tag_,
}
)
The resulting table contains the following columns:
file: the source filename;token: the original word or punctuation mark;lemma: the base form of the token;upos: the universal POS tag;xpos: the detailed English POS tag.
print(f”Number of token annotations: {len(token_annotations)}”) print(token_annotations[:5])
anno_df = pd.DataFrame(token_annotations) #save into dataframe
print(anno_df.head())
Count UPOS tags in each text
Group the data by filename and UPOS tag:
upos_counts = (
anno_df
.groupby(["file", "upos"])
.size()
.reset_index(name="count")
)
print(upos_counts.head())
Normalize UPOS frequencies
Now try counting the normalized frequency of each UPOS tag (because the total number of tokens may vary across texts). You may use the same equation that we pracited in the earlier moduel (i.e., per 1,000 lemma instances)
Create a wide-format UPOS table
To show one text per row and one UPOS category per column:
upos_table = (
upos_counts
.pivot(
index="file",
columns="upos",
values="count",
)
.fillna(0)
.astype(int)
.reset_index()
)
print(upos_table.head())
Count XPOS tags in each text
Group the data by filename and XPOS tag, as you did for UPOS.
Normalize XPOS frequencies
Now try counting the normalized frequency of each XPOS tag (because the total number of tokens may vary across texts).
Create a wide-format XPOS table
Following the UPOS example, show one text per row and one XPOS tag per column.
Submission
Please submit the following files with the following naming conventions:
- python code: en_pos.py
- Token-level annotations: token_annotations.csv
- UPOS normalized counts: upos_counts.csv
- XPOS normalized counts: xpos_counts.csv
- Wide-format UPOS table: all_texts_upos_counts.csv
- Wide-format XPOS table: all_texts_xpos_counts.csv
- Final compressed folder: lastname_firstname_pos_analysis.zip