Back: L2SCA

Introduction

LxGrTgr (Lexicogrammatical Tagger) identifies lexicogrammatical complexity features in English texts. It uses the spaCy en_core_web_trf model to analyze a text and assign complexity tags to relevant words and structures.

English text (.txt)
	|
	v
spaCy + LxGrTgr
	|
	v
Tokens with complexity tags
	|
	v
Counts of selected features

LxGrTgr can tag a single string, write tagged output to a file, or process a folder of texts. Each output row includes a token, its lemma, and its complexity tag. Detailed tag descriptions and annotation guidelines are available online.

1. Install LxGrTgr

LxGrTgr requires Python 3, spaCy, and spaCy’s English transformer model. In a terminal, install the package and model:

pip install lxgrtgr
python -m spacy download en_core_web_trf

Before using LxGrTgr, check that spaCy and the model work correctly:

import spacy

nlp = spacy.load("en_core_web_trf")
doc = nlp("This is a sample sentence.")

for token in doc:
    print(token.text, token.lemma_, token.pos_, token.dep_)

This tutorial follows the project’s readme.md. Read a package’s README before using it: it normally contains the required setup, expected input, and known limitations.

2. Tag a sample text

Create a file named tag_sample.py and add the following code. It uses the example sentence from the LxGrTgr README:

import lxgrtgr as lxgr

sample = "This is a very important opportunity that only comes once in a lifetime."
tagged_sample = lxgr.tag(sample)
lxgr.printer(tagged_sample)

Run it in the terminal:

python tag_sample.py

The output shows token ID, word, lemma, and complexity tag. For example, LxGrTgr tags very as rb+jjrbmod and important as attr+npremod in this sentence.

To save the tagged result as a tab-separated file, add this line after lxgr.printer(tagged_sample):

lxgr.writer("tagged_sample.tsv", tagged_sample)

3. Analyze a folder of texts

To tag all .txt files in a folder, use tagFolder(). The following commands use the text files in this tutorial repository’s corpus folder and write tagged files to a new tagged_corpus folder:

import lxgrtgr as lxgr

lxgr.tagFolder("corpus/", "tagged_corpus/")
counts = lxgr.countTagsFolder("tagged_corpus/")
lxgr.writeCounts(counts, "lxgrtgr_counts.tsv")

lxgrtgr_counts.tsv contains feature counts for each document. By default, writeCounts() normalizes counts per 10,000 words. Use normed=False to write raw counts instead.

4. Inspection

Automated tags should be checked rather than accepted without question. Choose one tag from the output of tag_sample.py or a tagged corpus file and inspect it.

  1. Define the tag. Use the LxGrTgr documentation to write a short definition and identify the linguistic structure it represents.

  2. Identify it yourself. Read the original text and find every example of that structure. Record the words or phrases that you think should receive the tag.

  3. Compare with LxGrTgr. Compare your list with the tagged output. Identify matches, missing tags, and unexpected tags. Consider whether part of speech, dependency structure, punctuation, or an ambiguous expression may explain any difference.


Next: Exercises