Back: CoNLL-U format

Dependency parsing with multiple texts

Now, you’re ready to apply the dependency parsing from the previous modules to multiple texts. Use the same five sample texts from the English POS tagging exercises. Follow the same steps from that activity to load the .txt files and process them with nlp_en.pipe() so that you have:

text_files
docs # saved all the outputs from five sample texts

(If you’re not sure about what this means, revisit English POS tagging exercises and repeat the process through Process all text files.)

Inspect dependency annotations across texts

Choose one of the sample texts and inspect its dependency annotations:

doc = docs[0] #sample1.txt

for token in doc:
    if not token.is_space:
        print(
            token.text,
            token.pos_,
            token.head.text,
            token.dep_,
            sep="\t"
        )

Create CoNLL-U files

Use the doc_to_conllu() function from the CoNLL-U format module to convert each sample text into CoNLL-U format.

First, create a folder for the output files:

from pathlib import Path

output_folder = Path("dependency_output")
output_folder.mkdir(parents=True, exist_ok=True) #makes a new folder

Then convert and save all five texts:

for text_file, doc in zip(text_files, docs):

    conllu_text = doc_to_conllu(
        doc,
        text_label=text_file.stem
    )

    output_path = (
        output_folder /
        f"{text_file.stem}.tsv"
    )

    output_path.write_text(
        conllu_text,
        encoding="utf-8"
    )

    print(f"Saved: {output_path.name}")

Your output folder should contain:

sample1.tsv
sample2.tsv
sample3.tsv
sample4.tsv
sample5.tsv

Check your output

The CoNLL-U content is plain text, so saving it with a .tsv extension lets you open it easily with common programs. For example:

  • Excel: open the .tsv file (double-click it, or use File → Open in Excel). The tab-separated CoNLL-U fields should appear as separate columns.
  • VSCode, TextEdit, or Notepad: open the file as plain text.

Open at least one `.tsv` file and check that:

1. each token is represented using the ten CoNLL-U columns;
2. `HEAD` contains the syntactic head ID;
3. `DEPREL` contains the dependency relation;
4. sentences are separated by a blank line.

Save your inspection results in a plain-text file named `inspection_results.docx/pdf/txt/md`. For one of the `.tsv` files, include:

* the filename you inspected;
* the first sentence's `# sent_id` and `# text` lines;
* a yes/no answer for each of the five checks above;
* one token from that sentence, with its `HEAD` and `DEPREL` values, and your judgment of whether the parser's analysis is linguistically correct (and, if not, what you think the correct `HEAD` or `DEPREL` should be, with a brief reason). Please be critical here!


# Submission

Please submit the following files with the following naming conventions:

- python code: end_dep.py
- Dependency annotation file 1: sample1.tsv
- Dependency annotation file 2: sample2.tsv
- Dependency annotation file 3: sample3.tsv
- Dependency annotation file 4: sample4.tsv
- Dependency annotation file 5: sample5.tsv
- Inspection results: inspection_results.docx/pdf/txt/md
- Final compressed folder: `lastname_firstname_dependency_analysis.zip`

Back to Module 7