《The Elements of Statistical Learning: Data Mining, Inference, and Prediction》 · 第 2 版 · 2.6 · 第 31-32 页
A more interesting example is the multinomial likelihood for the regression function Pr(G|X) for a qualitative output G. Suppose we have a model Pr(G = Gk|X = x) = pk,θ(x), k = 1, . . . , K for the conditional probability of each class given X, indexed by the parameter vector θ. Then the og-likelihood (also referred to as the cross-entropy) is
$$
L(theta) = \sum_{i=1}^N \log p_{g_i,\theta}(\theta_i) (2.36)
$$
and when maximized it delivers values of θ that best conform with the data in this likelihood sense.
由 Euler2023 提交