Error-tolerant finite-state recognition with applications to morphological analysis and spelling correction

Oflazer, K.

Error-tolerant finite-state recognition with applications to morphological analysis and spelling correction

dc.citation.epage	89	en_US
dc.citation.issueNumber	1	en_US
dc.citation.spage	73	en_US
dc.citation.volumeNumber	22	en_US
dc.contributor.author	Oflazer, K.	en_US
dc.date.accessioned	2016-02-08T10:51:04Z
dc.date.available	2016-02-08T10:51:04Z	en_US
dc.date.issued	1996	en_US
dc.department	Department of Computer Engineering	en_US
dc.description.abstract	This paper presents the notion of error-tolerant recognition with finite-state recognizers along with results from some applications. Error-tolerant recognition enables the recognition of strings that deviate mildly from any string in the regular set recognized by the underlying finite-state recognizer Such recognition has applications to error-tolerant morphological processing, spelling correction and approximate string matching in information retrieval. After a description of the concepts and algorithms involved, we give examples from two applications: in the context of morphological analysis, error-tolerant recognition allows misspelled input word forms to be corrected and morphologically analyzed concurrently. We present an application of this to error-tolerant analysis of the agglutinative morphology of Turkish words. The algorithm can be applied to morphological analysis of any language whose morphology has been fully captured by a single (and possibly very large) finite-state transducer, regardless of the word formation processes and morpholographemic phenomena involved. In the context of spelling correction, error-tolerant recognition can be used to enumerate candidate correct forms from a given misspelled string within a certain edit distance. Error-tolerant recognition can be applied to spelling correction for any language, if (a) it has a word list comprising all inflected forms, or (b) its morphology has been fully described by a finite-state transducer. We present experimental results for spelling correction for a number of languages. These results indicate that such recognition works very efficiently for candidate generation in spelling correction for many European languages (English, Dutch, French, German, and Italian, among others) with very large word lists of root and inflected forms (some containing well over 200,000 forms), generating all candidate solutions within 10 to 45 milliseconds (with an edit distance of 1) on a SPARCStation 10/41. For spelling correction in Turkish, error-tolerant recognition operating with a (circular) recognizer of Turkish words (with about 29,000 states and 119,000 transitions) can generate all candidate words in less than 20 milliseconds, with an edit distance of 1.	en_US
dc.identifier.eissn	1530-9312	en_US
dc.identifier.issn	0891-2017	en_US
dc.identifier.uri	http://hdl.handle.net/11693/25833	en_US
dc.language.iso	English	en_US
dc.publisher	MIT Press	en_US
dc.source.title	Computational Linguistics	en_US
dc.title	Error-tolerant finite-state recognition with applications to morphological analysis and spelling correction	en_US
dc.type	Article	en_US

Files

Original bundle

Now showing 1 - 1 of 1

Name:: Error-tolerant Finite-state Recognition with Applications to Morphological Analysis and Spelling Correction.pdf
Size:: 999.71 KB
Format:: Adobe Portable Document Format
Description:: Full printable version

Download

Collections

Scholarly Publications - Computer Engineering