Back: English dependency parsing exercise

Packages beyond spaCy

Choosing an NLP package

spaCy is not the only option for linguistic annotation. Other Python packages can be used for tokenization, POS tagging, lemmatization, and dependency parsing, for example:

  • Stanza (Stanford NLP Group): pretrained pipelines for many languages, trained on Universal Dependencies treebanks;
  • Trankit: transformer-based multilingual pipelines;
  • UDPipe (Charles University): lightweight UD-based pipelines;

Different packages can give different tokenization, tag sets, and accuracy for the same text, so you should check the following before using a model:

  1. Model: Which model or package are you loading, and what version? Which processors does it include (tokenizer, tagger, lemmatizer, parser)?
  2. Training data: Which corpus or treebank was the model trained on (e.g., a UD treebank such as GSD or Kaist)? What genre is it, and how similar is it to your own texts?
  3. Performance: What accuracy (e.g., UPOS, XPOS, lemma, LAS) is reported on the model’s page or documentation? On which test data? Performance on your own data is likely to be lower.
  4. Annotation scheme: Which UPOS and language-specific (XPOS) tag sets are used, and how are tokens segmented?
  5. Sample annotation: Before processing an entire corpus, annotate a small sample and inspect the results manually. Check whether the tokenization, POS tags, lemmas, and dependency relations are linguistically reasonable and appropriate for your research question.

Back to Module 7