Back to Module 8

Introduction

L2 Syntactic Complexity Analyzer (L2SCA) automatically measures the syntactic complexity of English writing. Given a UTF-8 plain-text file, it uses the Stanford Parser and Tregex to count nine syntactic structures and calculate 14 complexity indices. It is based on Lu (2010).

English text (.txt)
	|
	v
Stanford Parser + Tregex
	|
	v
Counts: words, sentences, clauses, T-units, ...
	|
	v
14 complexity measures (for example, MLS, MLT, C/S)

L2SCA reports counts for words (W), sentences (S), verb phrases (VP), clauses (C), T-units (T), dependent clauses (DC), complex T-units (CT), coordinate phrases (CP), and complex nominals (CN). Its indices include mean length of sentence (MLS), mean length of T-unit (MLT), mean length of clause (MLC), clauses per sentence (C/S), clauses per T-unit (C/T), dependent clauses per clause (DC/C), and complex nominals per clause (CN/C).

1. Download and install L2SCA

Download L2SCA 4.2.2. The tool runs on macOS and Linux and requires Java and Python 3. A machine with at least 2 GB of memory is recommended.

You can extract the downloaded archive manually by double-clicking it in Finder. Alternatively, to practice using the terminal, move to the folder containing the archive and run:

tar -xzf L2SCA-2023-08-15.tgz
cd L2SCA-2023-08-15

This tutorial follows the instructions in README-L2SCA.txt. (Note. Most Python packages include a README.md or similar file with essential setup and usage information. It may be tempting to skip it (just as we often skip the manual for a newly purchased device and start using it right away) but reading it before running a new tool is usually worth the time.)

2. Analyze the sample texts

L2SCA includes two sample texts, sample1.txt and sample2.txt, in its samples-L2SCA folder. Run the folder analyzer to analyze both files at once. The input directory path must end with a slash:

python3 analyzeFolder.py samples-L2SCA/ samples-L2SCA/samples_testing.csv

Open samples-L2SCA/samples_testing.csv to inspect the results. It contains one row for each sample text, with the structure counts and 14 complexity indices. Compare it with the included samples-L2SCA/samples_output to confirm that the analyzer ran correctly.

3. Inspection

Automatic analysis should be checked rather than accepted without question. Choose two of L2SCA’s 14 indices and inspect it using sample1.txt, sample2.txt, and samples_testing.csv.

  1. Define the index. Find its full name and write a short definition. Identify the structures used to calculate it. For example, mean length of sentence (MLS) is $W / S$: the number of words divided by the number of sentences.

  2. Count it yourself. Read both sample texts and manually count the relevant structures. Record the values you used and calculate the index for each text. For an index expressed as a ratio, use the same numerator and denominator shown in its abbreviation.

  3. Compare with L2SCA. Locate the same index in the sample1.txt and sample2.txt rows of samples-L2SCA/samples_testing.csv, then compare the values with your calculations. Explain whether the results match. If they differ, check whether tokenization, sentence boundaries, punctuation, or a syntactic structure that is difficult to identify manually could explain the difference.

References

  • Lu, X. (2010). Automatic analysis of syntactic complexity in second language writing. International Journal of Corpus Linguistics, 15(4), 474-496.

  • Lu, X. (2014). Computational methods for corpus annotation and analysis. Springer.