Back: English POS tagging exercise

English dependency parsing

In the previous activity, we used spaCy to obtain tokenization, lemmas, and POS annotations. The same pipeline also provides dependency annotations, which represent grammatical relationships between words. We will focus particularly on:

  • HEAD: the word that a token depends on
  • DEPREL: the grammatical relationship between the token and its head

(Note that I use dependency parsing as a general term for both the automatic assignment of dependency relations to words (dependency tagging) and the analysis of the resulting dependency structure. This is because, in practice, dependency analysis involves both identifying dependency labels and examining the head–dependent relationships they represent.)

Loading the pipeline

import spacy

nlp_en = spacy.load("en_core_web_trf")

Inspecting dependency annotations

Process a sentence:

text = "The students analyzed the data carefully."

doc = nlp_en(text)

Inspect the dependency information:

for token in doc:
    print(
        token.text,
        token.pos_,
        token.head.text,
        token.dep_,
        sep="\t"
    )

The important attributes are:

  • token.text: the current token
  • token.pos_: its UPOS category
  • token.head.text: its syntactic head
  • token.dep_: its dependency relation to the head

For example:

students    NOUN    analyzed    nsubj
data        NOUN    analyzed    dobj
carefully   ADV     analyzed    advmod

Now, instead of this boring sentence, try your own sample sentence and inspect the result.


Next: CoNLL-U format