学习问题的表示VladimirN.VapnikThe Nature ofStatistical Learning TheoryChapter 1Setting of the Learning ProblemIn this book we consider the learning problem as a problem of finding adesired dependence using a lamited number of observations
xGY学习问题的表示LMy(i)Agenerator (G)of random vectors x ERn,drawn iudependentlyfrom a fixed but unknown probability distribution function F(),(ii) A supervisor (S) who returns an output value y to every input vectorX, according to a conditional distrihution function F(yr), also fixedbut unknown.(ii) A learning machine (LM) capable of implementing a set of functionsf(a,α), α E A, where A is a set of parametersThe selection of the desired function is based qn a training set. of inde-pendent and identically distributed (i.i.d.) observations drawn according toF(a,y) = F(r)F(yla):(1.1)(a1,y1),...,(ze,ye)
In order to choose the best available approximation to the supervisor'sresponse, onemeasures theloss,or discrepancy,L(y,f (x,a))between theresponse y of thesupervisorto a given input a and the responsef(s,a)provided by thelearning machine. Consider the expected value of the lossgivenbytheriskfunctaonalR(α) =L(y, f(a,a))dF(a,y).(1.2)The goal is to find the function f(a,co) that minimizes the risk functionalR(α)(overtheclass of functionsf(a,a)aEA)inthesituationwherethe joint probability distribution function F(α,y) is unknownl and the onlyavailable information is contained in the training set (1.1).Pattern Recognition[oify=f(a,a)(1.3)L(y, f(a,a) =1ifyf(a,a)(1.4)Regression EstimationL(y,f(a,a))= (y-f(a,a)2(1.5)L(p(r,a) = _logp(z,a)
1.4THEGENERALSETTINGOFTHELEARNINGPROBLEMThe general setting of the learning problem can be described as follows.Let the probability measure F(z) bedefined on the space Z. Consider theset offunctionsQ(z,a),αEA.Thegoal istominimizetheriskfunctionalR(a) m / Q(z,)dF(z), αE A,(1.6)wherethe probability measure F(z)is unknown, but an i.i.d. sample(1.7)Z1,-+*,2eis given.The learning problems considered above are particular cases of this gen-eral problemofmimimizingtheriskfunctional (1.6)onthebasisofempiricaldata(1.7), where z describes a pair (,y) and Q(z,a) is the specific lossfunction (e.g-, one of (1.3), (1.4), or (1.5)). In the following we will de-scribe the resultsobtained for thegeneral statement of theproblem.Toapply them to specific problems, one has to substitute the correspondinglossfunctions intheformulasobtained
1.5The EmpiricalRisk Minimization (ERM)Inductive Principle(i) The risk functional R(aα)is replaced by the so-called empirical riskfunctionaltQ(4,0)(1.8)Remp(a) =1=1constructed on the basis of the training set (1.7)(ii) One approximates the function Q(z,ao) that minimizes risk (1.6) bythe function Q(z, ae) minimizing the empirical risk (1.8)This principle is called the Empirical Risk Minimization inductiveprinciple(ERMprinciple)We say that an inductiveprinciple definesa learningprocess if forany given set of observations the learning machine chooses theapproximation using this inductive principle.In learning theory theERMprincipleplaysacrucialrole
1.5 The Empirical Risk Minimization (ERM) Inductive Principle This principle is called the Empirical Risk Minimization inductive principle (ERM principle). We say that an inductive principle defines a learning process if for any given set of observations the learning machine chooses the approximation using this inductive principle. In learning theory the ERM principle plays a crucial role