Wide character vs. Multi-byte charactersText information needs to be represented by theright data types.- Multi byte characters: data are processed on a per-bytebasis: Big5, GB, EUC, even UTF-8- Wide characters: Fixed-byte encoding and no testing ofhigh bit is needed.Processing representation for wide characters:- Big Endian vS. Little Endian? Data type dependent: only for wide characters? System architecture dependent·Distinction: OxFEFF for Big Endian and OxFFFE forLittle EndianLecture4
Lecture4 Wide character vs. Multi-byte characters • Text information needs to be represented by the right data types. − Multi byte characters: data are processed on a per-byte basis: Big5, GB, EUC, even UTF-8 − Wide characters: Fixed-byte encoding and no testing of high bit is needed. • Processing representation for wide characters: − Big Endian vs. Little Endian • Data type dependent: only for wide characters • System architecture dependent • Distinction: 0xFEFF for Big Endian and 0xFFFE for Little Endian
Character Input Input method: A scheme of mapping charactersfrom their external representations to the internalcodepoints used in computer systemsClassification of input methods:-Images:· Off-line character recognition (Optical characterrecognition) On-line character recognition- Speech: voice recognition- Character features: Keyboard input based on glyphshapes and pronunciations.Lecture4
Lecture4 Character Input • Input method: A scheme of mapping characters from their external representations to the internal codepoints used in computer systems. • Classification of input methods: − Images: • Off-line character recognition (Optical character recognition) • On-line character recognition − Speech: voice recognition − Character features: Keyboard input based on glyph shapes and pronunciations
Character Input Based on ImagesOptical Character Recognition (via image, off-line ): Written material --> scanner --> bitmap image file (e.gTIFF, JPEG) --> characters (represented by an internalcode)- very difficult for unrestricted handwritten characters,commercially viable for printed materials and acuracydepends on printing quality- Degree of difficulty increases when the total number ofcharacters to be recognized increasesOn-line character Recognition (by pen writing devices):- Handwriting information capture (pen-in, pen-out, pen-movement, on-line) --> Stroke information (preprocessing with noise reduction) --> Searching for thecharacter based on the sequence of strokes.commercially viableLecture4
Lecture4 Character Input Based on Images • Optical Character Recognition (via image, off-line ): − Written material -> scanner -> bitmap image file (e.g. TIFF, JPEG) -> characters (represented by an internal code) − very difficult for unrestricted handwritten characters, commercially viable for printed materials and acuracy depends on printing quality − Degree of difficulty increases when the total number of characters to be recognized increases • On-line character Recognition (by pen writing devices): − Handwriting information capture (pen-in, pen-out, penmovement, on-line) -> Stroke information (pre processing with noise reduction) -> Searching for the character based on the sequence of strokes. − commercially viable
Speech Recognition (by voice input):- Capture speech by microphones --> speech signalsegmentation --> speech signal converted tophonetic transcription --> phonetic spellingconverted to internal code.- becoming commercially viable, problem withnon-native speaker, conversion from colloguialto written text- more affordable and getting common in the next5-10yrsLecture4
Lecture4 • Speech Recognition (by voice input): − Capture speech by microphones -> speech signal segmentation -> speech signal converted to phonetic transcription -> phonetic spelling converted to internal code. − becoming commercially viable, problem with non-native speaker, conversion from colloquial to written text − more affordable and getting common in the next 5-10yrs
Keyboard based Input method: an encoding method whichmaps a sequence ofkeystrokes (with a predefined keyboardlayout) to an internal code of a character.- Conceptually, an input method can be considered as amapping table with two columns: 1st column X is asequence of keys, 2nd column Y is the correspondinginternal code- Uniqueness requirement: for any two internalcodepoints Y, and Y, if Y, + Y, then X, + X.Input methods are normally language (script) dependent:- Input for Chinese and Greek Letters in GB are twodifferent input methods and are thus separately invokedLecture4
Lecture4 • Keyboard based Input method: an encoding method which maps a sequence of keystrokes (with a predefined keyboard layout) to an internal code of a character. − Conceptually, an input method can be considered as a mapping table with two columns: 1st column X is a sequence of keys, 2nd column Y is the corresponding internal code. − Uniqueness requirement: for any two internal codepoints Yi and Yj , if Yi ≠ Yj then Xi ≠ Xj . • Input methods are normally language (script) dependent: − Input for Chinese and Greek Letters in GB are two different input methods and are thus separately invoked