Bayes’ Theorem and the Naive Bayes Algorithm
A transcription of the handwritten Bayes notes
1 Topics
- Conditional probability
- Independent events
- Conditionally independent events
- Bayes’ theorem
- The naive Bayes algorithm
2 Independent events
Events and are independent if
or, equivalently,
The square sample space is divided so that is its left half and is its lower half. Therefore, is exactly the lower-left quarter of the sample space. Vertical hatching identifies , horizontal hatching identifies , and the intersection is crosshatched.
3 Conditional probability
The probability of given is
In this Venn diagram, is the sample space and conditioning on reduces the sample space to . The darker overlap is , while denotes the region outside .
It follows that
or
Note that
4 Independence in naive Bayes
In the application of the naive Bayes algorithm presented in the book, independence would mean that all words in the corpus are independent. This is not what is assumed in naive Bayes. What is assumed is conditional independence: the words are independent given the class, spam or ham.
The words are therefore not generally independent without conditioning:
The original notes also flag an error in the denominator in the first edition of the referenced book.
5 Conditionally independent events
Events and are conditionally independent given if
This does not imply the unconditional-independence condition
6 Conditional independence in the SMS data
In the naive Bayes application to the SMS data, conditional independence means that all words are independent given the class.
- The words in the spam dictionary are assumed independent given the class spam.
- The words in the ham dictionary are assumed independent given the class ham.
- Both assumptions are conditional on knowing the class, ham or spam.
All the words in the combined dictionary are not assumed to be unconditionally independent.
7 Bayes’ theorem
Bayes’ theorem can be written as
For a binary event , the denominator can be expanded using the law of total probability:
Because the denominator is constant when comparing possible values of for the same observed ,
or, in words,
8 The naive Bayes algorithm
The naive Bayes learner is trained by constructing a likelihood table for the appearance of words. In the book example, the vocabulary consists of and .
When a new message is received, the posterior probability is calculated to determine whether the message is more likely to be spam or ham, given the likelihood of the words found in the message text.
Suppose the observed message contains and , but not or .
8.1 Score for spam
Using conditional independence,
8.2 Score for ham
Similarly,
Because
the message is six times more likely to be spam than ham and is therefore classified as spam.
9 Source
This document is a typeset transcription of BayesNotes.pdf. Minor punctuation and capitalization have been standardized, and the mathematical content has been preserved. The independent-events diagram uses R’s grid graphics, and the conditional-probability diagram is generated with the VennDiagram R package.