Get in Touch
 Duration 21 hours

Course Outline

Comprehensive training outline

  1. Introduction to NLP
    • Foundations of NLP
    • Overview of NLP Frameworks
    • Commercial use cases of NLP
    • Techniques for scraping data from the web
    • Utilizing various APIs to collect text data
    • Managing and storing text corpora, including content and associated metadata
    • Benefits of Python and an NLTK introduction
  2. Practical Understanding of a Corpus and Dataset
    • The necessity of a corpus
    • Methods for Corpus Analysis
    • Categories of data attributes
    • Variations in file formats for corpora
    • Preparatory steps for datasets in NLP applications
  3. Understanding the Structure of a Sentences
    • Core components of NLP
    • Principles of natural language understanding
    • Morphological analysis: stems, words, tokens, and speech tags
    • Syntactic analysis
    • Semantic analysis
    • Strategies for handling ambiguity
  4. Text data preprocessing
    • Corpus: raw text
      • Sentence-level tokenization
      • Stemming techniques for raw text
      • Lemmatization of raw text
      • Elimination of stop words
    • Corpus: raw sentences
      • Word-level tokenization
      • Word-level lemmatization
    • Manipulating Term-Document and Document-Term matrices
    • Tokenizing text into n-grams and sentences
    • Customized and practical preprocessing workflows
  5. Analyzing Text data
    • Fundamental NLP features
      • Parsers and parsing mechanisms
      • POS tagging and tagger implementation
      • Name entity recognition
      • Working with N-grams
      • The Bag of Words model
    • Statistical aspects of NLP
      • Linear algebra concepts applied to NLP
      • Probabilistic theory in NLP
      • TF-IDF weighting
      • Vectorization techniques
      • Use of Encoders and Decoders
      • Data normalization
      • Application of Probabilistic Models
    • Advanced feature engineering and NLP
      • Introduction to word2vec
      • Internal components of the word2vec model
      • Underlying logic of the word2vec model
      • Extensions of the word2vec concept
      • Practical applications of the word2vec model
    • Case study: Applying the Bag of Words model for automatic text summarization using simplified and full Luhn's algorithms
  6. Document Clustering, Classification and Topic Modeling
    • Document clustering and pattern detection (including hierarchical clustering, k-means, and other methods)
    • Document comparison and classification using TFIDF, Jaccard index, and cosine similarity
    • Classifying documents with Naïve Bayes and Maximum Entropy models
  7. Identifying Important Text Elements
    • Dimensionality reduction: Principal Component Analysis, Singular Value Decomposition, and non-negative matrix factorization
    • Topic modeling and information retrieval via Latent Semantic Analysis
  8. Entity Extraction, Sentiment Analysis and Advanced Topic Modeling
    • Distinguishing positive vs. negative sentiment and intensity
    • Item Response Theory
    • Applying part-of-speech tagging to identify people, places, and organizations in text
    • Advanced topic modeling techniques: Latent Dirichlet Allocation
  9. Case studies
    • Extracting insights from unstructured user reviews
    • Classifying and visualizing sentiment in product review data
    • Analyzing search logs for usage patterns
    • Implementing text classification
    • Performing topic modelling

Requirements

A foundational understanding of NLP principles and an appreciation for how AI is applied in business contexts

Number of participants


Price per participant

Testimonials (1)

Upcoming Courses

Related Categories