CoNLL-U format
Back: English dependency parsing
CoNLL-U format
It is useful to become familiar with the CoNLL-U format, a widely used format for representing token-level linguistic annotations in NLP. Each token is represented on a separate line, with columns for information such as the token form, lemma, POS tags, morphological features, syntactic head, and dependency relation.
In CoNLL-U, the syntactic head of each token is represented by its token ID. For example:
1 Students student NOUN NNS Number=Plur 2 nsubj _ _
2 analyzed analyze VERB VBD Tense=Past 0 ROOT _ _
3 data datum NOUN NNS Number=Plur 2 dobj _ _
Here:
- token
1(Students) depends on token2(analyzed); - token
3(data) also depends on token2; HEAD = 0indicates that token2is the root of the sentence.
Creating CoNLL-U output
Use the following function to convert a spaCy Doc into CoNLL-U-style output:
def doc_to_conllu(doc):
lines = []
for sent in doc.sents:
for token in sent:
if token.dep_ == "ROOT":
head_id = 0
else:
head_id = token.head.i - sent.start + 1
feats = str(token.morph)
if feats == "":
feats = "_"
columns = [
str(token.i - sent.start + 1), # ID (resets per sentence)
token.text, # FORM
token.lemma_, # LEMMA
token.pos_, # UPOS
token.tag_, # XPOS
feats, # FEATS
str(head_id), # HEAD (relative to sentence)
token.dep_, # DEPREL
"_", # DEPS
"_" # MISC
]
lines.append("\t".join(columns))
lines.append("")
return "\n".join(lines)
Try it:
doc = nlp_en(
"The students carefully analyzed the new dataset."
)
print(doc_to_conllu(doc))