Back: English dependency parsing

CoNLL-U format

It is useful to become familiar with the CoNLL-U format, a widely used format for representing token-level linguistic annotations in NLP. Each token is represented on a separate line, with columns for information such as the token form, lemma, POS tags, morphological features, syntactic head, and dependency relation.

In CoNLL-U, the syntactic head of each token is represented by its token ID. For example:

1   Students   student   NOUN   NNS   Number=Plur   2   nsubj   _   _
2   analyzed   analyze   VERB   VBD   Tense=Past    0   ROOT    _   _
3   data       datum     NOUN   NNS   Number=Plur   2   dobj    _   _

Here:

  • token 1 (Students) depends on token 2 (analyzed);
  • token 3 (data) also depends on token 2;
  • HEAD = 0 indicates that token 2 is the root of the sentence.

Creating CoNLL-U output

Use the following function to convert a spaCy Doc into CoNLL-U-style output:

def doc_to_conllu(doc):
    lines = []
    for sent in doc.sents:
        for token in sent:
            if token.dep_ == "ROOT":
                head_id = 0
            else:
                head_id = token.head.i - sent.start + 1
            feats = str(token.morph)
            if feats == "":
                feats = "_"
            columns = [
                str(token.i - sent.start + 1),  # ID (resets per sentence)
                token.text,                     # FORM
                token.lemma_,                   # LEMMA
                token.pos_,                     # UPOS
                token.tag_,                     # XPOS
                feats,                          # FEATS
                str(head_id),                   # HEAD (relative to sentence)
                token.dep_,                     # DEPREL
                "_",                            # DEPS
                "_"                             # MISC
            ]
            lines.append("\t".join(columns))
        lines.append("") 
    return "\n".join(lines)

Try it:

doc = nlp_en(
    "The students carefully analyzed the new dataset."
)

print(doc_to_conllu(doc))

Next: English dependency parsing exercise