28.6.11

Tilde Translator for Android OS and iOS

Tilde Translator now is also available for Android OS and iOS.

Features:
• Single words are translated based on entries in the Tilde Dictionary
• Text translation is done by the machine translation tool
• Directions: English-Latvian, Latvian-English, Latvian-Russian
• History of translations is available offline
• Transliteration option for easy text input if Latvian keyboard layout is not available
• Interface in Latvian and in English

To run the application, you need an active Internet connection.

9.6.11

LetsMT! project presented at EU projects exhibition at EAMT2011


The European Association for Machine Translation (EAMT) organises their yearly conferences regularly in May. This year the venue was Faculty of Arts, Katholieke Universitet Leuven in Belgium. This two day conference started on 2011-05-30 and on 2011-05-31 there was an exhibition of European projects related to machine translation. Since this audience is considered to be a natural lieu for LetsMT! project, our poster and flyers were presented there and Raivis Skadinš was the presenter. Since our project raised quite an interest, he had answer a lot of questions from the participants of the conference.

More about LetsMT! project here: http://www.letsmt.eu/

2.6.11

CHAT workshop


I was delighted to have the honour to lead the first CHAT workshop on creation, harmonization and application of terminology resources held on May 11, 2011 at the University of Latvia, in Riga, Latvia. It was co-located with the 18th Nordic Conference of Computational Linguistics, NODALIDA 2011.

Terminology plays an extremely important role in the translation and localization industry, as well as in natural language processing. Different national and international activities have been undertaken to create terminology resources and apply them in computer-assisted and machine translation tools (e.g. TTC: Terminology Extraction, Translation Tools and Comparable Corpora). Another issue is the consolidation and harmonization of dispersed multilingual terminology resources. An important step towards this direction is the federated approach of providing access to content from multiple data sources, such as EuroTermBank, as well as international activities of providing common language resources and their applications CLARA and open linguistic infrastructures (META-NORD: Baltic and Nordic Branch of the European Open Linguistic Infrastructure), to serve the needs of industry and research communities in language resources, including terminologies.

The main idea of CHAT was to focus on the fostering the cooperation between the European projects and research and development activities in the area of terminology, and bring together academic and industrial researchers, as well as attract and involve postgraduate students and young researchers.

Altogether, 11 papers were accepted to CHAT for presentations which cover various topics on:
• automated approaches to terminology extraction and creation of terminology resources
• compiling multilingual terminology
• ensuring interoperability and harmonization of terminology resources
• integrating these resources in natural language processing applications
• distributing and sharing terminological data and some others

CHAT was a joint effort of our colleagues from:
• Tilde (Latvia)
• Norwegian School of Economics and Business Administration (Norway)
• The Seventh Framework Programme TTC project
• The Seventh Framework Programme CLARA project
• The Competitiveness and Innovation Programme META-NORD project
• and Programme Committee members from ten countries throughout the world (see CHAT homepage)

Two invited speakers kindly accepted our invitation to give their keynote presentations. They were Prof. Gerhard Budin (University of Vienna, Austrian Academy of Sciences) and Prof. Emmanuel Morin (University of Nantes, Computer Science laboratory of the Nantes-Atlantique region of France). Prof. Gerhard Budin gave a keynote presentation on “Terminology Resource Development in Global Domain Communities” with an overview of “Practical Experiences, Case Studies and Conclusions for Future Projects”. Prof. Emmanuel Morin gave a keynote presentation on “Bilingual Terminology Extraction from Comparable Corpora”.

We had two paper presentation sessions (9 papers) and a demonstration session (2 demos).

Overall, the workshop was truly multilingual, multicultural and multidomain! I hope the participants found the workshop interesting and useful for their further research in the development of terminology resources and services of the future, had fruitful discussions and revealed promising perspectives, and simply spent nice time during that day.

CHAT proceedings were published in the electronic repository of the University of Tartu Library as NEALT Proceedings Series vol. 12 and can be found at http://dspace.utlib.ee/dspace/handle/10062/16956. There are 12 files in the volume: 11 papers as separate files + one file (Proceedings) that contains contents, preface, program of the workshop and all the 11 papers.
NODALIDA 2011 in pictures: http://www.lumii.lv/nodalida2011/photos.html
CHAT 2011 in pictures: https://picasaweb.google.com/115946514697851463753/CHAT2011?feat=directlink, the quality of the photos leaves much to be desired, but still it’s better than nothing.

On behalf of the workshop organizing committee I would like to express our gratitude to the NODALIDA conference for this opportunity to co-locate the workshop as a satellite event. The University of Latvia for hosting this workshop and the Institute of Mathematics and Computer Science, in particular, for their cooperation and efforts during the workshop organization. The invited speakers for their responsiveness and substantial input to the programme. Co-organizers of the event for our close collaboration and Programme Committee members for their time and attention during the preparation of the workshop and review process, in particular. All the participants for their interesting papers and presentations. My colleagues from Tilde for their assistance, advice and support! Prof. Mare Koit, Editor-in-Chief of the NEALT Publication Series at University of Tartu, for her cooperation and producing the electronic proceedings.


CHAT 2011 homepage.
NODALIDA 2011 homepage.


Tatiana Gornostay
CHAT PC chair

28.2.11

Tilde Translator Presented at Research Workshop: Machine Translation and Morphologically-rich Languages


In the end of January, Tilde machine translation system developers participated in the research workshop which was devoted to machine translation challenges working with morphology-rich languages. During the workshop, presentation was also given about the English-Latvian and English-Lithuanian machine translation systems developed by Tilde.

Many studies are made worldwide about machine translation but, unfortunately, most often they tell about problems related to translations into English. For instance, very good systems have been developed to translate from Chinese into English, from Arabic into English, etc. However, the methods that are applied to create such systems are not equally effective for translation from English into other languages because English is a morphologically very simple language. The traditional statistical methods that work well when translating into English do not work good enough, if translation is made into a language in which words are inflected or declined, where the word sequence in a sentence is relatively free, where words must be coordinated for the same gender, number or case endings.

Read more about the research workshop on: http://cl.haifa.ac.il/MT/

1.10.10

Importance of Parallel Text in Machine Translation

Modern machine translation systems learn from existing translations how to translate. translate.tilde.com also is such a system — a statistical machine translation system computing probabilities based on existing translations. Then, the probabilities of the translations are used in translating.
For a computer to be able to compute the probabilities of translations, we need the so-called parallel text, i.e., a text in one language with a corresponding translation in the other language; sentences of both texts should be aligned, i.e., we must know which sentence corresponds to which translation. The larger is the size of parallel texts available, the higher is the level of a machine translator you can train. Therefore, the parallel text is very significant in the development of machine translation.

As a result, globally today various projects and activities take place with the aim to collect as many as possible parallel texts for both improving the machine translation and increasing the productivity of human translators. Tilde is also involved in a number of such activities. They include research projects, development projects of new services, and just good initiatives. This time, I'll tell about one of such initiatives, next time — about research and other projects we are involved in.

Tilde in collaboration with other companies such as Adobe, Oracle, Sun, Intel, Microsoft, etc., is one of the founders of the international TAUS Data Association (TAUS DA). The organization was created with the aim to share parallel text resources among those who have large parallel text resources available to them. The TAUS DA database contains translations from different organizations, including companies, EU institutions, individual translators. And these translations are very useful both in increasing the productivity of translators and in improving the machine translation systems. Tilde, too, has made available a big part of translations done in Tilde, and now they are included in the TAUS DA database.

Currently, these are just the first steps of real application (for details, see: http://www.tausdata.org/index.php/visitor-center/use-cases). Tilde also together with Adobe took part in an experiment organized by the TAUS DA with the aim to find out whether it is possible to create a customized machine translation system over a very short period of time (24 h) based on the TAUS DA data that would assist in translating real software interface and documentation. The answer is: yes, within 24 hours you can create a machine translation system based on the TAUS DA data that provides good English–Latvian translations of Adobe texts.

24.10.09

Welcome to Tilde MT team blog!

Tilde has been working in a field of language technologies for many years. We started with Latvian spellchecker, electronic dictionaries, grammar checker and other useful tools. Our goal is to give small languages the same support of language technologies as there are for big languages.
We started our machine translation project several years ago. We released so called foreign language reading tool in 2005. Although it was not MT system, it used MT elements and it evolved into our first English-Latvian MT system released in 2008. That was a rule based MT system, we put a lot of effort in it and I think the result was not bad. Later we released the Latvian-Russian MT on the same platform. Now we are working also on statistical MT and we are releasing our first free online English-Latvian MT system these days. It is at your disposal now. Try it, use it and let us know your opinion. We have set up this blog to share our thoughts about MT with you and to hear your thoughts. Your thoughts are important to us.

Raivis Skadiņš,
Chief Software Architect, MT team leader