--- title: "Bayes' Theorem and the Naive Bayes Algorithm" subtitle: "A transcription of the handwritten Bayes notes" format: html: embed-resources: true toc: true toc-depth: 2 number-sections: true html-math-method: mathml editor: visual --- ## Topics - Conditional probability - Independent events - Conditionally independent events - Bayes' theorem - The naive Bayes algorithm ## Independent events Events $A$ and $B$ are independent if $$ P(A \cap B) = P(A)P(B), $$ or, equivalently, $$ P(B \mid A) = P(B). $$ The square sample space $S$ is divided so that $A$ is its left half and $B$ is its lower half. Therefore, $A \cap B$ is exactly the lower-left quarter of the sample space. Vertical hatching identifies $A$, horizontal hatching identifies $B$, and the intersection is crosshatched. ```{r} #| label: fig-independent-events #| fig-cap: "A occupies the vertically hatched left 50%, B occupies the horizontally hatched lower 50%, and their crosshatched intersection occupies the lower-left 25%." #| fig-alt: "A square divided into four equal quarters. Vertical blue hatching marks event A across the left half. Horizontal gold hatching marks event B across the lower half. Both textures cross in the lower-left quarter, marking the intersection." #| fig-width: 6 #| fig-height: 6 #| echo: false independent_figure <- grid::grid.grabExpr({ grid::pushViewport( grid::viewport( width = grid::unit(0.84, "snpc"), height = grid::unit(0.84, "snpc") ) ) grid::grid.rect( x = 0.5, y = 0.5, width = 1, height = 1, gp = grid::gpar(fill = "#F8FAFC", col = NA) ) grid::grid.rect( x = 0.25, y = 0.5, width = 0.5, height = 1, gp = grid::gpar(fill = "#60A5FA", col = "#1D4ED8", alpha = 0.50, lwd = 2) ) grid::grid.rect( x = 0.5, y = 0.25, width = 1, height = 0.5, gp = grid::gpar(fill = "#FBBF24", col = "#B45309", alpha = 0.50, lwd = 2) ) a_hatch_x <- seq(0.035, 0.465, by = 0.035) grid::grid.segments( x0 = a_hatch_x, x1 = a_hatch_x, y0 = 0.02, y1 = 0.98, gp = grid::gpar(col = "#1D4ED8", alpha = 0.55, lwd = 1) ) b_hatch_y <- seq(0.035, 0.465, by = 0.035) grid::grid.segments( x0 = 0.02, x1 = 0.98, y0 = b_hatch_y, y1 = b_hatch_y, gp = grid::gpar(col = "#B45309", alpha = 0.55, lwd = 1) ) grid::grid.text( "A\n50%", x = 0.25, y = 0.75, gp = grid::gpar(fontsize = 14, fontface = "bold") ) grid::grid.text( "B\n50%", x = 0.75, y = 0.25, gp = grid::gpar(fontsize = 14, fontface = "bold") ) grid::grid.text( "A \u2229 B\n25%", x = 0.25, y = 0.25, gp = grid::gpar(fontsize = 14, fontface = "bold") ) grid::grid.text( "S", x = 0.94, y = 0.94, gp = grid::gpar(fontsize = 15, fontface = "bold") ) grid::grid.rect( x = 0.5, y = 0.5, width = 1, height = 1, gp = grid::gpar(fill = NA, col = "#334155", lwd = 2.5) ) grid::popViewport() }, width = 6, height = 6) grid::grid.newpage() grid::grid.draw(independent_figure) ``` ## Conditional probability The probability of $B$ given $A$ is $$ P(B \mid A) = \frac{P(A \cap B)}{P(A)}. $$ In this Venn diagram, $S$ is the sample space and conditioning on $A$ reduces the sample space to $A$. The darker overlap is $A \cap B$, while $B^c$ denotes the region outside $B$. ```{r} #| label: fig-conditional-probability #| fig-cap: "Conditioning on A makes A the reduced sample space; the favorable region for B is A intersection B." #| fig-alt: "A rectangular sample space S containing overlapping sets B and A. The overlap is labeled A intersection B, the area outside B is labeled B complement, and A is identified as the reduced sample space." #| fig-width: 7 #| fig-height: 5 #| echo: false conditional_venn <- VennDiagram::draw.pairwise.venn( area1 = 65, area2 = 35, cross.area = 15, category = c("B", "A"), euler.d = TRUE, scaled = TRUE, fill = c("#FBBF24", "#60A5FA"), alpha = c(0.35, 0.60), col = c("#B45309", "#1D4ED8"), lwd = 2, label.col = rep("transparent", 3), cex = rep(1, 3), cat.cex = 1.4, cat.fontface = "bold", cat.pos = c(-35, 35), cat.dist = c(0.06, 0.06), ind = FALSE, margin = 0.14 ) conditional_figure <- grid::grid.grabExpr({ grid::grid.rect( x = 0.5, y = 0.5, width = 0.92, height = 0.82, gp = grid::gpar(fill = "#F8FAFC", col = "#334155", lwd = 2) ) grid::grid.draw(conditional_venn) grid::grid.rect( x = 0.5, y = 0.5, width = 0.92, height = 0.82, gp = grid::gpar(fill = NA, col = "#334155", lwd = 2) ) grid::grid.text( "S", x = 0.91, y = 0.86, gp = grid::gpar(fontsize = 15, fontface = "bold") ) grid::grid.text( "A \u2229 B", x = 0.50, y = 0.50, gp = grid::gpar(fontsize = 14, fontface = "bold") ) grid::grid.text( expression(B^c), x = 0.88, y = 0.25, gp = grid::gpar(fontsize = 15) ) grid::grid.text( "Reduced sample space: A", x = 0.50, y = 0.13, gp = grid::gpar(fontsize = 13, fontface = "italic", col = "#1D4ED8") ) }, width = 7, height = 5) grid::grid.newpage() grid::grid.draw(conditional_figure) ``` It follows that $$ P(A \cap B) = P(B \mid A)P(A) $$ or $$ P(A \cap B) = P(A \mid B)P(B). $$ Note that $$ P(A \cap B) = P(B \cap A). $$ ## Independence in naive Bayes In the application of the naive Bayes algorithm presented in the book, independence would mean that all words in the corpus are independent. This is **not** what is assumed in naive Bayes. What is assumed is **conditional independence**: the words are independent given the class, spam or ham. The words are therefore not generally independent without conditioning: $$ P(W_1, \neg W_2, \neg W_3, W_4) \ne P(W_1)P(\neg W_2)P(\neg W_3)P(W_4). $$ The original notes also flag an error in the denominator in the first edition of the referenced book. ## Conditionally independent events Events $B_1$ and $B_2$ are conditionally independent given $A$ if $$ P(B_1, B_2 \mid A) = P(B_1 \mid A)P(B_2 \mid A). $$ This does not imply the unconditional-independence condition $$ P(B_1, B_2) = P(B_1)P(B_2). $$ ## Conditional independence in the SMS data In the naive Bayes application to the SMS data, conditional independence means that all words are independent given the class. - The words in the spam dictionary are assumed independent given the class **spam**. - The words in the ham dictionary are assumed independent given the class **ham**. - Both assumptions are conditional on knowing the class, ham or spam. All the words in the combined dictionary are not assumed to be unconditionally independent. ## Bayes' theorem Bayes' theorem can be written as $$ \underbrace{P(B \mid A)}_{\text{posterior}} = \frac{P(B \cap A)}{P(A)} = \frac{ \overbrace{P(A \mid B)}^{\text{likelihood}} \overbrace{P(B)}^{\text{prior}} }{P(A)}. $$ For a binary event $B$, the denominator can be expanded using the law of total probability: $$ P(B \mid A) = \frac{P(A \mid B)P(B)} {P(A \mid B)P(B) + P(A \mid B^c)P(B^c)}. $$ Because the denominator is constant when comparing possible values of $B$ for the same observed $A$, $$ P(B \mid A) \propto P(A \mid B)P(B), $$ or, in words, $$ \text{posterior} \propto \text{likelihood} \times \text{prior}. $$ ## The naive Bayes algorithm The naive Bayes learner is trained by constructing a likelihood table for the appearance of words. In the book example, the vocabulary consists of $W_1, W_2, W_3,$ and $W_4$. When a new message is received, the posterior probability is calculated to determine whether the message is more likely to be spam or ham, given the likelihood of the words found in the message text. Suppose the observed message contains $W_1$ and $W_4$, but not $W_2$ or $W_3$. ### Score for spam Using conditional independence, $$ \begin{aligned} P(\text{spam} \mid W_1, \neg W_2, \neg W_3, W_4) &\propto P(W_1 \mid \text{spam}) P(\neg W_2 \mid \text{spam}) \\ &\quad \times P(\neg W_3 \mid \text{spam}) P(W_4 \mid \text{spam})P(\text{spam}) \\ &\propto \left(\frac{4}{20}\right) \left(\frac{10}{20}\right) \left(\frac{20}{20}\right) \left(\frac{12}{20}\right) \left(\frac{20}{100}\right) \\ &= 0.012. \end{aligned} $$ ### Score for ham Similarly, $$ \begin{aligned} P(\text{ham} \mid W_1, \neg W_2, \neg W_3, W_4) &\propto P(W_1 \mid \text{ham}) P(\neg W_2 \mid \text{ham}) \\ &\quad \times P(\neg W_3 \mid \text{ham}) P(W_4 \mid \text{ham})P(\text{ham}) \\ &\propto \left(\frac{1}{80}\right) \left(\frac{64}{80}\right) \left(\frac{71}{80}\right) \left(\frac{23}{80}\right) \left(\frac{80}{100}\right) \\ &\approx 0.002. \end{aligned} $$ Because $$ \frac{0.012}{0.002} = 6, $$ the message is six times more likely to be spam than ham and is therefore classified as **spam**. ## Source This document is a typeset transcription of [BayesNotes.pdf](BayesNotes.pdf). Minor punctuation and capitalization have been standardized, and the mathematical content has been preserved. The independent-events diagram uses R's `grid` graphics, and the conditional-probability diagram is generated with the [`VennDiagram`](https://CRAN.R-project.org/package=VennDiagram) R package.