Measuring complexity 3. LxGrTgr
Introduction
LxGrTgr (Lexicogrammatical Tagger) identifies lexicogrammatical complexity features in English texts. It uses the spaCy en_core_web_trf model to analyze a text and assign complexity tags to relevant words and structures.
English text (.txt)
|
v
spaCy + LxGrTgr
|
v
Tokens with complexity tags
|
v
Counts of selected features
LxGrTgr can tag a single string, write tagged output to a file, or process a folder of texts. Each output row includes a token, its lemma, and its complexity tag. Detailed tag descriptions and annotation guidelines are available online.
1. Install LxGrTgr
LxGrTgr requires Python 3, spaCy, and spaCy’s English transformer model. In a terminal, install the package and model:
pip install lxgrtgr
python -m spacy download en_core_web_trf
Before using LxGrTgr, check that spaCy and the model work correctly:
import spacy
nlp = spacy.load("en_core_web_trf")
doc = nlp("This is a sample sentence.")
for token in doc:
print(token.text, token.lemma_, token.pos_, token.dep_)
This tutorial follows the project’s readme.md. Read a package’s README before using it: it normally contains the required setup, expected input, and known limitations.
2. Tag a sample text
Create a file named tag_sample.py and add the following code. It uses the example sentence from the LxGrTgr README:
import lxgrtgr as lxgr
sample = "This is a very important opportunity that only comes once in a lifetime."
tagged_sample = lxgr.tag(sample)
lxgr.printer(tagged_sample)
Run it in the terminal:
python tag_sample.py
The output shows token ID, word, lemma, and complexity tag. For example, LxGrTgr tags very as rb+jjrbmod and important as attr+npremod in this sentence.
To save the tagged result as a tab-separated file, add this line after lxgr.printer(tagged_sample):
lxgr.writer("tagged_sample.tsv", tagged_sample)
3. Analyze a folder of texts
To tag all .txt files in a folder, use tagFolder(). The following commands use the text files in this tutorial repository’s corpus folder and write tagged files to a new tagged_corpus folder:
import lxgrtgr as lxgr
lxgr.tagFolder("corpus/", "tagged_corpus/")
counts = lxgr.countTagsFolder("tagged_corpus/")
lxgr.writeCounts(counts, "lxgrtgr_counts.tsv")
lxgrtgr_counts.tsv contains feature counts for each document. By default, writeCounts() normalizes counts per 10,000 words. Use normed=False to write raw counts instead.
4. Inspection
Automated tags should be checked rather than accepted without question. Choose one tag from the output of tag_sample.py or a tagged corpus file and inspect it.
-
Define the tag. Use the LxGrTgr documentation to write a short definition and identify the linguistic structure it represents.
-
Identify it yourself. Read the original text and find every example of that structure. Record the words or phrases that you think should receive the tag.
-
Compare with LxGrTgr. Compare your list with the tagged output. Identify matches, missing tags, and unexpected tags. Consider whether part of speech, dependency structure, punctuation, or an ambiguous expression may explain any difference.