1. Home
  2. AI & Machine Learning
  3. Naive Bayes Classifier

Naive Bayes Classifier

Count how often each word appears in spam and in normal mail, then use Bayes’ theorem to score a new message. Simple, fast and surprisingly accurate.

Interactive 3DBeginner11 min readAI/MLUpdated

Drag to rotate · Right-drag to pan · Click, then scroll to zoom · Space play · ←→ step

What's happening

Pseudocode

    Try this in the 3D model

    • Press Classify for “free money meeting”. Which word pulls the result towards not-spam?
    • Try “project meeting now”. “now” is a spam word, so why does the message still come out as not spam?
    • Find a word that appears only in spam training messages. Why is its not-spam probability not exactly 0?
    • Type your own message using the known words and predict the verdict before pressing Classify.

    Bayes’ theorem in one line

    You get an email that says “win free money”. How likely is it to be spam? Bayes’ theorem turns the question around:

    P(spam | words) = P(words | spam) · P(spam) / P(words)
    • P(spam), the prior: how common spam is in general.
    • P(words | spam), the likelihood: how typical these words are for spam. We can measure this by counting words in spam we have already seen.
    • P(words) is the same for every class, so we can skip it and simply normalise at the end.

    The naive assumption

    P(words | spam) for a whole message is hard to estimate, because we would need to have seen this exact message before. Naive Bayes assumes every word is independent given the class, so the likelihood becomes a simple product:

    P(win, free, money | spam) ≈ P(win | spam) × P(free | spam) × P(money | spam)

    That is obviously not true of real language, but the classifier only has to rank the classes correctly, and for that the shortcut works remarkably well.

    Training = counting

    From the 8 training messages in the 3D model (after dropping tiny words like “a”, “at” and “the”), spam contains 12 words and not-spam 11, with a vocabulary of 14 different words.

    P(word | class) = (count of word in class + 1) / (words in class + vocabulary size)

    The +1 is Laplace smoothing. Without it, a word that never appeared in spam would get probability 0, and one such word would wipe out the whole spam score.

    Word Count in spam P(word | spam) Count in not spam P(word | not spam)
    free 2 3/26 ≈ 0.115 0 1/25 = 0.040
    money 2 3/26 ≈ 0.115 0 1/25 = 0.040
    meeting 0 1/26 ≈ 0.038 2 3/25 = 0.120

    Classifying “free money meeting”

    • score(spam) = 0.5 × 0.115 × 0.115 × 0.038 ≈ 0.000256
    • score(not spam) = 0.5 × 0.040 × 0.040 × 0.120 ≈ 0.000096
    • P(spam | message) = 0.000256 / (0.000256 + 0.000096) ≈ 72.7 % → spam

    “meeting” pulls towards not-spam, but “free” and “money” together win.

    Code

    from collections import Counter
    
    STOP = {"a", "at", "the", "an", "to", "of", "and", "is", "in", "on"}
    words = lambda s: [w for w in s.lower().split() if w not in STOP]
    
    train = [("win money now", 1), ("free money offer", 1), ("win a free prize", 1),
             ("cheap offer now", 1), ("meeting at noon", 0), ("project meeting tomorrow", 0),
             ("lunch at noon tomorrow", 0), ("send the project notes", 0)]
    
    counts = {1: Counter(), 0: Counter()}
    for msg, label in train:
        counts[label].update(words(msg))
    vocab = set(counts[0]) | set(counts[1])
    
    def p_spam(msg):
        score = {}
        for c in (1, 0):
            total = sum(counts[c].values())
            s = sum(1 for _, label in train if label == c) / len(train)    # prior
            for w in words(msg):
                if w in vocab:
                    s *= (counts[c][w] + 1) / (total + len(vocab))         # Laplace smoothing
            score[c] = s
        return score[1] / (score[1] + score[0])
    
    print(round(p_spam("free money meeting"), 3))     # 0.727
    print(round(p_spam("lunch tomorrow at noon"), 3)) # 0.047

    With scikit-learn the same model is CountVectorizer followed by MultinomialNB.

    Where is it used?

    • Spam filtering. Paul Graham’s 2002 essay A Plan for Spam popularised Bayesian filters, and they are still part of many mail systems.
    • Sentiment analysis (positive or negative review) and topic classification of news articles.
    • Medical diagnosis from independent symptoms, and as a quick baseline before trying heavier models.

    Naive Bayes vs other classifiers

    Naive Bayes Logistic regression Decision tree
    Training Counting, very fast Iterative optimisation Greedy splits
    Needs much data No Some Some
    Handles thousands of features ✅ easily ✅ ⚠️ slower
    Probabilities Often over-confident Well calibrated Coarse
    Main assumption Features independent Linear in log-odds Axis-aligned splits

    Common mistakes

    • Forgetting smoothing, so one unseen word makes a probability exactly 0.
    • Multiplying hundreds of tiny probabilities directly. Use logarithms to avoid underflow.
    • Forgetting the prior, which matters when one class is much more common than the other.
    • Trusting the output probability too much. Naive Bayes ranks well, but its probabilities tend to be too extreme.

    Complexity at a glance

    Case / operationTimeWhy
    Training (n messages, total length L)O(L)Just count words, one pass.
    Classifying a message of m wordsO(m · k)k classes, one lookup per word and class.
    Model sizeO(V · k)One probability per word per class.
    Extra spaceO(V · k)

    Quick check

    Test yourself — pick an answer to see if you got it.

    1. Why is it called “naive” Bayes?

    2. A word never appeared in any spam training message. What does Laplace smoothing do?

    3. Spam scores 0.0003 and not-spam scores 0.0001. What is P(spam | message)?

    4. Why do real implementations add log-probabilities instead of multiplying probabilities?

    Saved only in this browser — no account needed.
    Spotted a mistake or a bug in the 3D model?

    Report a mistake

    in Naive Bayes Classifier. Thank you — every report makes the lesson better for the next reader.

    We'll also include a link to the step of the 3D model you're on and your browser type, so we can reproduce it.