Paper: Multext-East: Parallel and Comparable Corpora and Lexicons for Six Central and Eastern European Languages

ACL ID C98-1049
Title Multext-East: Parallel and Comparable Corpora and Lexicons for Six Central and Eastern European Languages
Venue International Conference on Computational Linguistics
Session Main Conference
Year 1998
Authors

The EU Copernicus project Multext-East has created a multi-lingual corpus of text and speech data, covering the six languages of the project: Bulgarian, Czech, Estonian, Hungarian, Romanian, and Slovene. In addition, wordform lexicons for each of the languages were developed. The corpus includes a parallel component consisting of Orweli's Nineteen Eighty-Four, with versions in all six languages tagged for part-of-speech and aligned to English (also tagged for POS). We describe the encoding format and data architecture designed especially for this corpus, which is generally usable for encoding linguistic corpora. We also describe the methodology for the development of a harmonized set of morphosyntactic descriptions (MSDs), which builds upon the scheme for western European...