The Beta distribution serves as a conjugate prior to both Binomial and Bernoulli likelihoods because when you multiply the Binomial likelihood (proportional to θ^Z × (1-θ)^(n-Z)) by the Beta prior (proportional to θ^(a-1) × (1-θ)^(b-1)), the resulting posterior distribution maintains the same functional form as the prior, specifically proportional to θ^(a+Z-1) × (1-θ)^(n+b-Z-1), which is itself a Beta distribution with updated parameters a' = a + Z and b' = n + b - Z.
Beta Conjugate Prior to Binomial and Bernoulli Likelihoods
Added:in this video I want to provide a proof that the beta distribution is actually a conjugate prior to the Beni and the binomial distributions so first of all let's write out what actually is a beta distribution so a beta distribution is defined by two parameters which I'm going to call A and B here and it's equal to Theta to the power A-1 * 1 - Theta to^ B-1 all divided through by a function a normalizing function which is called a beta function which again takes these two parameters as the input but the important thing to notice about this sort of denominator here is that this denominator here is merely a constant which ensures that the beta distribution when you integrate from three Theta equal 0 to Theta equal 1 actually integrates to one so that it is in fact a valid probability distribution so what we could actually do is we could rewrite this as just being equal to some constant time Theta to^ a * 1 - Theta to the power B minus one and that should also be a minus one as well so to prove that the beta is actually a conjugate prior to the B and the binomial distribution we're first of all going to need our basian formula so what we're going to do is just write that down which is just that the probability of theta given the data and I'm not going to condition on model Choice here but you could do I'm sort of implicitly doing that that's equal to the probability of the data given a choice of theta in other words the likelihood times the probability of theta which is just our prior divided through by the probability of the data and what we do is we notice that the denominator here is independent of choice of theta so it doesn't depend on Theta so what we can actually do is we can rewrite this as sort of just being proportional to the numerator so the probability of the data given choice of theta times the probability of theta the prior and we're just sort of forgetting about the denominator again because again it is a sort of normalizing constant and importantly it is the same for different values of theta now what we're going to do is we're going to put in the beta distribution for our prior and for the likelihood what we're going to do is we're going to insert our binomial and our sort of Bia distribution so first of all we need to write down what the formula for that is and we know that for a Beni and a binomial distribution that the probability of the data given choice of theta is equal to something times so I'm just going to write that again as proportional to Theta to the^ Z time 1us Theta the^ n minus Z where Zed here I've chosen as to be the sum of our sort of individual trials or individual flips of a coin if you like from I = 1 to n so then if we multiply this by our prior which is just the probability of theta which we know is proportional to Theta to^ a -1 * 1 - Theta to^ B-1 we multiply these two things together then we're going to get that the posterior distribution the probability Theta given choice of data is going to be proportional to if we just do the sort of multiplication here Theta to the power a + z -1 * 1 - Theta the^ n + b - Z minus one and we note that this now has exactly the same form as our original beta distribution if you compare the numerator here with that which we've obtained at the bottom here these essentially are of the same form except now we have a new value of a which I'm going to call a primed which is equal to a plus Z and we've got a new value of B which I'm going to call B primed which is equal to n + B minus Z and we know that the posterior distribution is a valid probability distribution so it has to integrate to one so we're going to need to divide this expression through by a normalizing factor in order for it to integrate to one and we already know what that normalizing factor is going to be be is just going to be this beta function of our new a and our new B hence we find that our posterior distribution the probability of theta given choice of data is equal to Theta to the power a primed -1 * 1 - Theta the^ B Prim minus1 all divided through by our beta function of a primed and B primed in other words we appr proved that the posterior distribution is a beta distribution and hence the beta distribution is conjugate to both the Beni and the binomial likelihoods
Up Next

Bayesian Inference: Deriving the Posterior Mean & Precision
@JingchenMonikaHu
852 views•2019-02-12

Gain Recalibration in Hippocampal Path Integration: Math Theory
@1024kyz
144 views•2020-07-02

Conjugate Priors Explained: Definitions, Examples & Why They Matter
@oxeduc4209
46.5K views•2014-08-12

The Mathematical Impossibility of Accurate World Maps
@Vox
23.3M views•2016-12-02
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Mathematics



















![[Bayesian inference for a proportion] Updating the Beta prior part 1](https://i.ytimg.com/vi_webp/bfsUzReHDAg/maxresdefault.webp)






















