Naive Bayes Classifiers Approach:compute the posterior probability P(C | A1, A)priorvalues of C using the Bayes theoremP(AA :..A IC)P(CP(CIAA ...A) :P(AA 1..A)Choosevalje ofCthatmaximizesposterioriCA1, A2, ..., An)likelihoodEquivalent to choosing value of C that maximizesP(A1, A2, .., AnlC) P(C).evidence How to estimate P(At, A2, ..., An / C)?
Naïve Bayes Classifiers ◼ Approach: ◆ compute the posterior probability P(C | A1 , A2 , ., An ) for all values of C using the Bayes theorem ◆ Choose value of C that maximizes P(C | A1 , A2 , ., An ) ◆ Equivalent to choosing value of C that maximizes P(A1 , A2 , ., An |C) P(C) ◼ How to estimate P(A1 , A2 , ., An | C )? ( ) ( | ) ( ) ( | ) 1 2 1 2 1 2 n n n P A A A P A A A C P C P C A A A = posteriori likelihood prior evidence
Naive Bayes Classifiers Assume independence among attributes A, when class isgiven:P(A1, A2, ..., An IC) = P(A/ C) P(A2l C)... P(Anl C)Can estimate P(A;I C) for all A; and CjThis is a simplifying assumption which may be violatedin reality The classifier that uses the Naive Bayes assumption andcomputes the MAP hypothesis is called Naive BayesclassifierCNaive Bayes = arg max P(c)P(x |c) = arg max P(c)I I P(a, Ic)MAP:maximumaposterior
◼ Assume independence among attributes Ai when class is given: ◆ P(A1 , A2 , ., An |C) = P(A1 | Cj ) P(A2 | Cj ). P(An | Cj ) ◆ Can estimate P(Ai | Cj ) for all Ai and Cj . ◆ This is a simplifying assumption which may be violated in reality ◼ The classifier that uses the Naïve Bayes assumption and computes the MAP hypothesis is called Naïve Bayes classifier = = i i c c Naive Bayes c arg max P(c)P(x | c) arg max P(c) P(a | c) Naïve Bayes Classifiers MAP: maximum a posterior