Skip to content

Input vectorization

One-Hot Encoding

Technique where each word is represented by a vector with high bit corresponding to the word’s index in the vocabulary

Pros - Simplicity - Compatibility Cons - High Dimensional - Loss of Semantic Information - Sparsity

Bag of Words

Vector representing the frequency of words, disregarding grammar and word order.

Pros - Simple - Provides a clear understanding of text Cons - Ignores the order and context of words - High dimensional - Fails to capture the semantic meaning between words

Term Frequency-Inverse Document Frequency (TF-IDF)

Weighs the frequency of words by their importance across documents.

TF = Number of times term t appears in document d / Total number of terms in document d

IDF = log(Total number of documents / Number of documents containing term t)

TF-IDF = TF * IDF

Pros - Simple - Importance Weighing - Improves document relevance

Cons

  • High dimensional sparse vectors
  • No context capture
  • Treats synonyms as separate

Count Vectoriser

Focuses on counting the occurrences of each word in the document. It converts a collection of text documents to a matrix of token counts where each elements represents the count of a word in a specific document.

Word Embedding

Dense vector representations in a continuous vector space where semantically similar words are located close to each other. Captures the semantic relationships between words.

Pros

  • Captures the semantic meaning and relationships
  • Dense representations are computationally efficient

Cons - Requires large collection of samples for efficient embedding

Technique Accuracy Computation Time Memory Usage Applicability
Bag of Words (BoW) Low to Moderate Low High Simple text classification tasks
TF-IDF Moderate Moderate High Text classification, information retrieval, keyword extraction
Count Vectorizer Low to Moderate Low High Tasks focusing on word frequency
Word Embeddings High High Moderate to High Sentiment analysis, named entity recognition, machine translation