English dependency parsing exercise
Dependency parsing with multiple texts
Now, you’re ready to apply the dependency parsing from the previous modules to multiple texts. Use the same five sample texts from the English POS tagging exercises. Follow the same steps from that activity to load the .txt files and process them with nlp_en.pipe() so that you have:
text_files
docs # saved all the outputs from five sample texts
(If you’re not sure about what this means, revisit English POS tagging exercises and repeat the process through Process all text files.)
Inspect dependency annotations across texts
Choose one of the sample texts and inspect its dependency annotations:
doc = docs[0] #sample1.txt
for token in doc:
if not token.is_space:
print(
token.text,
token.pos_,
token.head.text,
token.dep_,
sep="\t"
)
Create CoNLL-U files
Use the doc_to_conllu() function from the CoNLL-U format module to convert each sample text into CoNLL-U format.
First, create a folder for the output files:
from pathlib import Path
output_folder = Path("dependency_output")
output_folder.mkdir(parents=True, exist_ok=True) #makes a new folder
Then convert and save all five texts:
for text_file, doc in zip(text_files, docs):
conllu_text = doc_to_conllu(
doc,
text_label=text_file.stem
)
output_path = (
output_folder /
f"{text_file.stem}.tsv"
)
output_path.write_text(
conllu_text,
encoding="utf-8"
)
print(f"Saved: {output_path.name}")
Your output folder should contain:
sample1.tsv
sample2.tsv
sample3.tsv
sample4.tsv
sample5.tsv
Check your output
The CoNLL-U content is plain text, so saving it with a .tsv extension lets you open it easily with common programs. For example:
- Excel: open the
.tsvfile (double-click it, or use File → Open in Excel). The tab-separated CoNLL-U fields should appear as separate columns. - VSCode, TextEdit, or Notepad: open the file as plain text.
Open at least one `.tsv` file and check that:
1. each token is represented using the ten CoNLL-U columns;
2. `HEAD` contains the syntactic head ID;
3. `DEPREL` contains the dependency relation;
4. sentences are separated by a blank line.
Save your inspection results in a plain-text file named `inspection_results.docx/pdf/txt/md`. For one of the `.tsv` files, include:
* the filename you inspected;
* the first sentence's `# sent_id` and `# text` lines;
* a yes/no answer for each of the five checks above;
* one token from that sentence, with its `HEAD` and `DEPREL` values, and your judgment of whether the parser's analysis is linguistically correct (and, if not, what you think the correct `HEAD` or `DEPREL` should be, with a brief reason). Please be critical here!
# Submission
Please submit the following files with the following naming conventions:
- python code: end_dep.py
- Dependency annotation file 1: sample1.tsv
- Dependency annotation file 2: sample2.tsv
- Dependency annotation file 3: sample3.tsv
- Dependency annotation file 4: sample4.tsv
- Dependency annotation file 5: sample5.tsv
- Inspection results: inspection_results.docx/pdf/txt/md
- Final compressed folder: `lastname_firstname_dependency_analysis.zip`