Bayesian Model Selection

Bayesian evidence, Bayes factors and Occam's razor in model selection.

Any given data set can, in principle, be explained by many different models. One of the key questions underlying science is that of model selection: how do we choose between competing theories that purport to explain the observed data? The great paradigm shifts in science fall squarely within this domain. With so many models available, overfitting remains a real concern, which makes choosing the most suitable model essential in statistics.

In the context of astronomy - as with most areas of science - the next two decades will see a massive increase in data volume through large surveys such as the Square Kilometre Array (SKA) and the Large Synoptic Survey Telescope (LSST). Robust statistical analysis to perform model selection at scale will be a critical factor in the success of such future surveys.

Even with very good data, we may not know when to stop fitting. Two competing models may fit the data equally well, so how do we choose the more appropriate one? The answer is to prefer the simpler model, a principle known as Occam’s razor. A complex model that explains the data only slightly better than a simpler one should be penalised for the additional parameters it introduces, since extra parameters reduce predictive power. In this context, Bayesian model selection is becoming an increasingly important tool for determining whether the introduction of a new parameter is justified by the data.

In this post, we focus on the Bayesian evidence, also known as the model likelihood or the marginal likelihood, and denoted by $\mathcal{Z}$. It is the probability of the data, $\mathcal{D}$, given a model, $\mathcal{M}$, and is obtained by averaging the likelihood over the prior:

\[\mathcal{Z} = \mathcal{P}(\mathcal{D}\,|\,\mathcal{M}) = \int \mathcal{P}(\mathcal{D}\,|\,\boldsymbol{\theta},\,\mathcal{M})\,\mathcal{P}(\boldsymbol{\theta}\,|\,\mathcal{M})\,d\boldsymbol{\theta}\]

This averaging is what builds Occam's razor into the evidence. A model with many parameters, or with very broad priors, spreads its predictions over a wide range of possible datasets, so it assigns relatively little probability to the one we actually observed. It is rewarded only if its extra flexibility improves the fit enough to compensate.

To compare two models, $\mathcal{M}_{1}$ and $\mathcal{M}_{2}$, we apply Bayes' theorem to the models themselves. The ratio of their posterior probabilities splits neatly into two parts:

\[\textrm{posterior odds} = \textrm{Bayes factor} \times \textrm{prior odds}\]

The prior odds express how much we favoured one model over the other before seeing the data. The Bayes factor, $B_{12}$, is simply the ratio of the two evidences, and captures how the data change those odds. With non-committal priors on the models, the posterior odds equal the Bayes factor. The larger $B_{12}$, the stronger our belief that $\mathcal{M}_{1}$ is the better model; if $B_{12}$ is below one, $\mathcal{M}_{2}$ is preferred. The Bayes factor is typically interpreted using the Jeffreys scale (see Kass et al., 1995), an empirically determined scale shown in the table below.

$\textrm{ln }\left(B_{12}\right)$$B_{12}$Evidence against $\mathcal{M}_{2}$
0 to 11 to 3Inconclusive
1 to 33 to 20Weak Evidence
3 to 520 to 150Moderate Evidence
>5>150Strong Evidence