Data Transformation and FeatureSelection/ExtractionDataMining:Conceptsand2026年9月16日Techniques
2026年9月16日 Data Mining: Concepts and Techniques 1 Data Transformation and Feature Selection/Extraction
Cbntinuous Attribute TemperatureOutlook Tempreature Humidity WindyClassN40 highfalseSunnyN37 hightruesunnyP34 highfalseovercastP26 highfalserainPrain15 normalfalseN13 normalraintrueP17normaltrueovercastN28 highfalsesunnyP25 normalfalsesunnyP23rainnormalfalseP27truenormalsunnyP22 hightrueovercastP40 normalfalseovercastN31 highraintrue
Outlook Tempreature Humidity Windy Class Sunny 40 high false N sunny 37 high true N overcast 34 high false P rain 26 high false P rain 15 normal false P rain 13 normal true N overcast 17 normal true P sunny 28 high false N sunny 25 normal false P rain 23 normal false P sunny 27 normal true P overcast 22 high true P overcast 40 normal false P rain 31 high true N Continuous Attribute Temperature
DiscretizationThree types of attributes:NominalvaluesfromanunorderedsetExample:attribute"outlook"fromweatherdataValues:"sunny","overcast",and"rainy"Ordinalvalues from an ordered setExample:attribute"temperature"inweatherdataValues:"hot">"mild">"cool"Continuous real numbersDiscretization:divide the range of a continuous attribute into intervalsSome classification algorithms only accept categorical attributes.Reduce data size by discretizationSupervised (entropy) vs. Unsupervised (binning)3DataMining:ConceptsandTechniques
Data Mining: Concepts and Techniques 3 Discretization ◼ Three types of attributes: ◼ Nominal — values from an unordered set ◼ Example: attribute “outlook” from weather data ◼ Values: “sunny”,”overcast”, and “rainy” ◼ Ordinal — values from an ordered set ◼ Example: attribute “temperature” in weather data ◼ Values: “hot” > “mild” > “cool” ◼ Continuous — real numbers ◼ Discretization: ◼ divide the range of a continuous attribute into intervals ◼ Some classification algorithms only accept categorical attributes. ◼ Reduce data size by discretization ◼ Supervised (entropy) vs. Unsupervised (binning)
Simple Discretization Methods: BinningEqual-width (distance) partitioning: It divides the range into Nintervals of equal size:uniformgridif A and Bare the lowest and highest values of theattribute, the width of intervals will be: W = (B-A)/NThe most straightforwardBut outliers may dominate presentation: Skewed data is nothandled well.Equal-depth (frequency) partitioning:It divides the range into N intervals, each containingapproximay samenumberof samplesDataMining:ConceptsandTechniques
Data Mining: Concepts and Techniques 4 Simple Discretization Methods: Binning ◼ Equal-width (distance) partitioning: ◼ It divides the range into N intervals of equal size: uniform grid ◼ if A and B are the lowest and highest values of the attribute, the width of intervals will be: W = (B –A)/N. ◼ The most straightforward ◼ But outliers may dominate presentation: Skewed data is not handled well. ◼ Equal-depth (frequency) partitioning: ◼ It divides the range into N intervals, each containing approximay same number of samples
Histograms40 A popular data reductiontechnique35Divide data into buckets30and store average (sum)25for each bucketCan be constructed20optimally in one15dimension using dynamic10programmingRelated to quantization5problems.+?C00000000000000000900000000800006000010000S00000L5DataMining:ConceptsandTechnigues
Data Mining: Concepts and Techniques 5 Histograms ◼ A popular data reduction technique ◼ Divide data into buckets and store average (sum) for each bucket ◼ Can be constructed optimally in one dimension using dynamic programming ◼ Related to quantization problems. 0 5 1 0 1 5 2 0 2 5 3 0 3 5 4 0 10000 20000 30000 40000 50000 60000 70000 80000 90000 100000