Skip to main content

Bayes’ Theorem and the Naive Bayes Algorithm

A transcription of the handwritten Bayes notes

1 Topics

  • Conditional probability
  • Independent events
  • Conditionally independent events
  • Bayes’ theorem
  • The naive Bayes algorithm

2 Independent events

Events AA and BB are independent if

P(A∩B)=P(A)P(B), P(A \cap B) = P(A)P(B),

or, equivalently,

P(B∣A)=P(B). P(B \mid A) = P(B).

The square sample space SS is divided so that AA is its left half and BB is its lower half. Therefore, A∩BA \cap B is exactly the lower-left quarter of the sample space. Vertical hatching identifies AA, horizontal hatching identifies BB, and the intersection is crosshatched.

A square divided into four equal quarters. Vertical blue hatching marks event A across the left half. Horizontal gold hatching marks event B across the lower half. Both textures cross in the lower-left quarter, marking the intersection.
Figure 1: A occupies the vertically hatched left 50%, B occupies the horizontally hatched lower 50%, and their crosshatched intersection occupies the lower-left 25%.

3 Conditional probability

The probability of BB given AA is

P(B∣A)=P(A∩B)P(A). P(B \mid A) = \frac{P(A \cap B)}{P(A)}.

In this Venn diagram, SS is the sample space and conditioning on AA reduces the sample space to AA. The darker overlap is A∩BA \cap B, while BcB^c denotes the region outside BB.

A rectangular sample space S containing overlapping sets B and A. The overlap is labeled A intersection B, the area outside B is labeled B complement, and A is identified as the reduced sample space.
Figure 2: Conditioning on A makes A the reduced sample space; the favorable region for B is A intersection B.

It follows that

P(A∩B)=P(B∣A)P(A) P(A \cap B) = P(B \mid A)P(A)

or

P(A∩B)=P(A∣B)P(B). P(A \cap B) = P(A \mid B)P(B).

Note that

P(A∩B)=P(B∩A). P(A \cap B) = P(B \cap A).

4 Independence in naive Bayes

In the application of the naive Bayes algorithm presented in the book, independence would mean that all words in the corpus are independent. This is not what is assumed in naive Bayes. What is assumed is conditional independence: the words are independent given the class, spam or ham.

The words are therefore not generally independent without conditioning:

P(W1,¬W2,¬W3,W4)≠P(W1)P(¬W2)P(¬W3)P(W4). P(W_1, \neg W_2, \neg W_3, W_4) \ne P(W_1)P(\neg W_2)P(\neg W_3)P(W_4).

The original notes also flag an error in the denominator in the first edition of the referenced book.

5 Conditionally independent events

Events B1B_1 and B2B_2 are conditionally independent given AA if

P(B1,B2∣A)=P(B1∣A)P(B2∣A). P(B_1, B_2 \mid A) = P(B_1 \mid A)P(B_2 \mid A).

This does not imply the unconditional-independence condition

P(B1,B2)=P(B1)P(B2). P(B_1, B_2) = P(B_1)P(B_2).

6 Conditional independence in the SMS data

In the naive Bayes application to the SMS data, conditional independence means that all words are independent given the class.

  • The words in the spam dictionary are assumed independent given the class spam.
  • The words in the ham dictionary are assumed independent given the class ham.
  • Both assumptions are conditional on knowing the class, ham or spam.

All the words in the combined dictionary are not assumed to be unconditionally independent.

7 Bayes’ theorem

Bayes’ theorem can be written as

P(B∣A)⏟posterior=P(B∩A)P(A)=P(A∣B)⏞likelihoodP(B)⏞priorP(A). \underbrace{P(B \mid A)}_{\text{posterior}} = \frac{P(B \cap A)}{P(A)} = \frac{ \overbrace{P(A \mid B)}^{\text{likelihood}} \overbrace{P(B)}^{\text{prior}} }{P(A)}.

For a binary event BB, the denominator can be expanded using the law of total probability:

P(B∣A)=P(A∣B)P(B)P(A∣B)P(B)+P(A∣Bc)P(Bc). P(B \mid A) = \frac{P(A \mid B)P(B)} {P(A \mid B)P(B) + P(A \mid B^c)P(B^c)}.

Because the denominator is constant when comparing possible values of BB for the same observed AA,

P(B∣A)∝P(A∣B)P(B), P(B \mid A) \propto P(A \mid B)P(B),

or, in words,

posterior∝likelihood×prior. \text{posterior} \propto \text{likelihood} \times \text{prior}.

8 The naive Bayes algorithm

The naive Bayes learner is trained by constructing a likelihood table for the appearance of words. In the book example, the vocabulary consists of W1,W2,W3,W_1, W_2, W_3, and W4W_4.

When a new message is received, the posterior probability is calculated to determine whether the message is more likely to be spam or ham, given the likelihood of the words found in the message text.

Suppose the observed message contains W1W_1 and W4W_4, but not W2W_2 or W3W_3.

8.1 Score for spam

Using conditional independence,

P(spam∣W1,¬W2,¬W3,W4)∝P(W1∣spam)P(¬W2∣spam)×P(¬W3∣spam)P(W4∣spam)P(spam)∝(420)(1020)(2020)(1220)(20100)=0.012. \begin{aligned} P(\text{spam} \mid W_1, \neg W_2, \neg W_3, W_4) &\propto P(W_1 \mid \text{spam}) P(\neg W_2 \mid \text{spam}) \\ &\quad \times P(\neg W_3 \mid \text{spam}) P(W_4 \mid \text{spam})P(\text{spam}) \\ &\propto \left(\frac{4}{20}\right) \left(\frac{10}{20}\right) \left(\frac{20}{20}\right) \left(\frac{12}{20}\right) \left(\frac{20}{100}\right) \\ &= 0.012. \end{aligned}

8.2 Score for ham

Similarly,

P(ham∣W1,¬W2,¬W3,W4)∝P(W1∣ham)P(¬W2∣ham)×P(¬W3∣ham)P(W4∣ham)P(ham)∝(180)(6480)(7180)(2380)(80100)≈0.002. \begin{aligned} P(\text{ham} \mid W_1, \neg W_2, \neg W_3, W_4) &\propto P(W_1 \mid \text{ham}) P(\neg W_2 \mid \text{ham}) \\ &\quad \times P(\neg W_3 \mid \text{ham}) P(W_4 \mid \text{ham})P(\text{ham}) \\ &\propto \left(\frac{1}{80}\right) \left(\frac{64}{80}\right) \left(\frac{71}{80}\right) \left(\frac{23}{80}\right) \left(\frac{80}{100}\right) \\ &\approx 0.002. \end{aligned}

Because

0.0120.002=6, \frac{0.012}{0.002} = 6,

the message is six times more likely to be spam than ham and is therefore classified as spam.

9 Source

This document is a typeset transcription of BayesNotes.pdf. Minor punctuation and capitalization have been standardized, and the mathematical content has been preserved. The independent-events diagram uses R’s grid graphics, and the conditional-probability diagram is generated with the VennDiagram R package.