Back to Module 5

Applying the workflow to another language (e.g., Korean)

The main lessons in Module 5 use English so that the workflow is easy to follow. The same workflow can be applied to another language, but the tools and the interpretation of their output must match that language.

This page applies the tokenization, lemmatization, and frequency workflow to Korean using spaCy’s ko_core_news_lg pipeline.

Install the Korean spaCy pipeline

Install spaCy if it is not already installed:

pip install spacy

Download the Korean spaCy pipeline:

python -m spacy download ko_core_news_lg

ko_core_news_lg is the largest official Korean pipeline and includes Korean-specific tokenization and trained components for linguistic annotation.

You can find pipelines for other languages on the spaCy Models page. Check the language, pipeline name, and available components before installing a model. Then replace ko_core_news_lg in the download and spacy.load() commands with the pipeline you selected.

Load the Korean spaCy pipeline

import spacy

nlp_ko = spacy.load("ko_core_news_lg")

print("Language:", nlp_ko.lang)
print("Pipeline components:", nlp_ko.pipe_names)

Process a Korean sentence:

korean_text = "학생들은 과제를 작성하고 있었습니다."
korean_doc = nlp_ko(korean_text)

print(type(korean_doc))

Korean tokenization

A Korean space-delimited unit, often called an eojeol, may contain several grammatical elements. Inspect the token boundaries:

for token in korean_doc:
    print(
        token.text,
        token.idx,
        token.is_punct,
        token.is_space
    )

The resulting tokens may not correspond directly to units separated by spaces. Token boundaries should be interpreted in relation to the Korean pipeline rather than assumed to represent individual morphemes.

Korean morphological analysis and lemmatization

Inspect the annotations assigned to each token:

for token in korean_doc:
    if not token.is_space and not token.is_punct:
        print(
            "Token:", token.text,
            "\nLemma:", token.lemma_,
            "\nUniversal POS:", token.pos_,
            "\nKorean tag:", token.tag_,
            "\nMorphology:", token.morph,
            "\n"
        )

The attributes can be interpreted as follows:

  • token.lemma_: the lemma representation assigned by the lemmatizer;
  • token.pos_: a universal part-of-speech category;
  • token.tag_: a Korean-specific, fine-grained part-of-speech analysis;
  • token.morph: morphological information predicted for the token.

The value of token.tag_ may contain multiple tags joined with +. This indicates that the token contains more than one morphological element. The morphologizer does not necessarily create a separate spaCy Token for every morpheme, so do not assume that one spaCy token equals one Korean morpheme.

Extract Korean lemmas for frequency analysis:

korean_lemmas = [
    token.lemma_.lower()
    for token in korean_doc
    if not token.is_space and not token.is_punct
]

print(korean_lemmas)

Korean token and lemma frequencies

The frequency-counting workflow can be reused after processing Korean text with the Korean pipeline:

from collections import Counter

korean_tokens = [
    token.text
    for token in korean_doc
    if not token.is_space and not token.is_punct
]

token_frequency = Counter(korean_tokens)
lemma_frequency = Counter(korean_lemmas)

print("20 most common tokens:")
print(token_frequency.most_common(20))

print("\n20 most common lemmas:")
print(lemma_frequency.most_common(20))

Back to Module 5