Back: English POS tagging

POS analysis of multiple files

Download sample.zip. In this task, you will:

  1. download the sample corpus to your local computer;
  2. process all .txt files using spaCy;
  3. count the UPOS and XPOS tags in each tex (both raw and normalized frequency);
  4. organize the results in pandas DataFrames;
  5. save the results as CSV files.

This would be a part of Assignment 3.


Download the sample corpus

After extracting the ZIP file, specify the local folder containing the .txt files:

from pathlib import Path 

sample_folder = Path("YOUR LOCAL PATH") 

text_files = sorted(sample_folder.glob("*.txt")) 

print(f"Number of text files: {len(text_files)}") 

for text_file in text_files: 
    print(text_file.name)

Expected output:

Number of text files: 5
sample1.txt
sample2.txt
sample3.txt
sample4.txt
sample5.txt

Process all text files

Read each .txt file and process the texts using nlp_en.pipe():

texts = [
    text_file.read_text(encoding="utf-8")
    for text_file in text_files
]

docs = list(nlp_en.pipe(texts))

Collect token-level POS annotations

For the frequency analysis, create another DataFrame that excludes spaces and punctuation:

token_annotations = []

for text_file, doc in zip(text_files, docs):
    for token in doc:
        if not token.is_space and not token.is_punct:
            token_annotations.append(
                {
                    "file": text_file.name,
                    "token": token.text,
                    "lemma": token.lemma_,
                    "upos": token.pos_,
                    "xpos": token.tag_,
                }
            )

The resulting table contains the following columns:

  • file: the source filename;
  • token: the original word or punctuation mark;
  • lemma: the base form of the token;
  • upos: the universal POS tag;
  • xpos: the detailed English POS tag.

print(f”Number of token annotations: {len(token_annotations)}”) print(token_annotations[:5])

anno_df = pd.DataFrame(token_annotations) #save into dataframe
print(anno_df.head())

Count UPOS tags in each text

Group the data by filename and UPOS tag:

upos_counts = (
    anno_df
    .groupby(["file", "upos"])
    .size()
    .reset_index(name="count")
)

print(upos_counts.head())

Normalize UPOS frequencies

Now try counting the normalized frequency of each UPOS tag (because the total number of tokens may vary across texts). You may use the same equation that we pracited in the earlier moduel (i.e., per 1,000 lemma instances)


Create a wide-format UPOS table

To show one text per row and one UPOS category per column:

upos_table = (
    upos_counts
    .pivot(
        index="file",
        columns="upos",
        values="count",
    )
    .fillna(0)
    .astype(int)
    .reset_index()
)
print(upos_table.head())

Count XPOS tags in each text

Group the data by filename and XPOS tag, as you did for UPOS.

Normalize XPOS frequencies

Now try counting the normalized frequency of each XPOS tag (because the total number of tokens may vary across texts).

Create a wide-format XPOS table

Following the UPOS example, show one text per row and one XPOS tag per column.


Submission

Please submit the following files with the following naming conventions:

  • python code: en_pos.py
  • Token-level annotations: token_annotations.csv
  • UPOS normalized counts: upos_counts.csv
  • XPOS normalized counts: xpos_counts.csv
  • Wide-format UPOS table: all_texts_upos_counts.csv
  • Wide-format XPOS table: all_texts_xpos_counts.csv
  • Final compressed folder: lastname_firstname_pos_analysis.zip

Next: English dependency parsing