Text pre-processing 1. Tokenization 2. Lemmatization 3. Subword tokenization 4. Frequency calculation 5. Exercises