Abstract
The aim of the Fira de Recerca en Directe 2015 (Live Research Fair) held at the Barcelona science museum CosmoCaixa is to encourage young people in our country to take up scientific careers. At each stand, students were presented with a scientific enigma and were encouraged to find the answer to it through conversations with researchers and through experiments using laboratory equipment. At the Tradumàtica Research Group’s stand, we presented research we have been carrying out in Machine Translation (MT) and MT postediting. Traditionally, scientific outreach programs have a unidirectional format from researchers to the citizenry. However, our experiences at the Fair on the perception young people have of MT clearly show a bidirectional interaction between researchers and the young people, and this experience helped us come up with ideas to improve our research.
Keywords
Thanks for doing this. It is very useful. (Teenager commenting on machine translation research)
In recent decades, research groups have made important moves in the direction of establishing links with the society that finances them through their taxes and donations. Scientific knowledge transfer and outreach programs may also lead to entrepreneurship. The mass media do great work in disseminating science out to young people, but live events provide deeper experiential immersions in science. In this context, the Tradumàtica Research Group at the Department of Translation, Interpreting and Asian Studies (Autonomous University of Barcelona) has been involved in active long-term collaborations in the world of scientific outreach. One recent example is the 13th Fira de Recerca en Directe (Live Research Fair), held at the CosmoCaixa (Science Museum) in Barcelona from April 8 to 11, 2015.
The aim of the Research Fair is to encourage young people in our country to take up scientific careers, thus increasing the number of researchers. Moreover, young people get a chance to experience science as an amusing activity that they can participate in and gain direct satisfaction from. In the words of scientific journalist Vladimir de Semir (2014), The aim of scientific outreach is to allow citizens to access new opportunities for personal development, basically in the work area, as well as allowing them to participate with a sufficiently critical spirit in the social, ethical and political debate that major advances in science and technology open up.
At the various stands in the Fair, students were presented with different scientific projects by several research groups. At the Tradumàtica Research Group’s stand, we presented research we have been carrying out in statistical machine translation (SMT) and MT postediting. Our current funded research project ProjecTA aims to compare the computer-assisted human translation process with the SMT process to improve SMT output. Introducing MT in the professional translation process requires new skills from professional translators: developing digital data mining, managing million-word plurilingual corpora, building linguistic resources for MT, training for SMT systems, preediting text, postediting text, and providing feedback for the SMT system. Apart from the technological aspects, our research on MT also takes into account human factors related to freelance professional translators and in-house translators working in companies. From several surveys and focus groups examining the scope of the ProjecTA project, we observed that many translation companies were reluctant to adopt MT. In contrast, nonprofessional MT users at the Fair had a very different and a more open-minded attitude. By discussing SMT with students at the Fair, bidirectional interaction took place: We introduced them to MT research and they helped us come up with ideas to improve our research.
Context: The Live Research Fair
The Live Research Fair is an exhibition highlighting some of the research projects being carried out in research institutes and centers in Catalonia. It is aimed at both research groups (who have the opportunity of showing off their work to young people) and students and their secondary school teachers, who see it as an opportunity to motivate their students through conversations with young researchers from a wide variety of disciplines. At each stand, students were presented with a scientific enigma and were encouraged to find the answer to it through conversations with researchers and through experiments using laboratory equipment. The invitation to resolve the scientific enigma was accompanied by posters illustrating the basic theory and describing the project and the research group. In addition, each group brought instruments and objects that allowed students empirically approach the research subject.
In this year’s Fair, participants totaled 880 students from ESO (Obligatory Secondary Education) and baccalaureate levels, ranging in age from 15 to 18 years. Sixty-four teachers from 20 different educational centers also attended. Those who attended visited 12 different stands on health sciences, genetics, neurobiotechnology, radiology, pharmacology, chemistry, photonic sciences, linguistics, archaeology, climate science, and environmental monitoring. 1 Another stand presented the winners of the Secondary School Research Program 2014-2015 sponsored by the Scientific Park of Barcelona and the Catalonia-La Pedrera Foundation. In total, there were 86 researchers.
The event lasted 3 days; for 6 hours a day from 10 a.m. to 3 p.m. At the Fair, everything is controlled and organized to the millimeter. Each research group had a 6 m2 stand and access to computers and audiovisual material, supplied by the organizers, besides the group’s own lab equipment. Students arrived in groups of six or seven and were at each stand for between 10 and 15 minutes. In that time, the posters were explained to them and then the young people participated in a practical activity related to the research—all in line with a constructivist philosophy, which asserts that knowledge is gained through practical experience.
Example of a Scientific Outreach Activity on Machine Translation
At the UAB Tradumàtica Research Group’s stand, we presented research we have been carrying out to MT and postediting. Like the other research groups, we had three posters through which to present our research on the question, “Machine Translation: Does It Work?”
The approximate number of students who participated in our activity was 350. This figure is significant enough to provide some interesting observations.
The Perception Young People Have of Machine Translation
At different moments during the presentation, the young people were asked certain questions: Do you use machine translation? What use do you make of it? Which translators do you think use it? Do you know how machine translation works? Have you been satisfied with the results? How do you see the future of machine translation?
Practically all of the young people responded in the affirmative to the first question, although the frequency of such use oscillated between “sometimes” and “always.” The majority affirmed that they had used MT to translate texts that they did not understand, or to be surer of their content (class materials, textbooks, and exercises), or else for song lyrics, emails, Facebook messages, and so on. Numerous students added that they use such translators as dictionaries, because they are “more practical and faster” than using online dictionaries.
The young people at the Fair declared that they had used both direct translation (into their mother tongue) and inverse translation (into a foreign-language) with equal frequency. In fact, some students explained that given the “shortcomings” or “faults” obtained, they often carried out both operations on the same text (e.g., English → Catalan and then Catalan → English) as a system for testing accuracy.
The pairs of languages most used were English-Catalan and English-Spanish and vice versa. Some way behind were French-Catalan or Spanish and then Arab-Catalan or Spanish. Minority uses included German-Catalan or Spanish and Italian-Catalan or Spanish. The greater use of English is due to the fact that it is currently the language that is offered and demanded in the majority of educational centers outside universities.
The vast majority of students knew about and used Google Translate. Although a few did mention machine translators offered by generalist newspapers or other software online.
Despite their daily use of MTs, the students had little idea of how they worked. Some “imagined” that they are fitted with grammar rules and enormous dictionaries with “all the words in alphabetical order.” None of them had ever heard of statistical translation, so we explained it to them during the activity.
Responses to questions about the degree of satisfaction on the one hand and the future of MT on the other were probably the most interesting.
The great majority of young people told us that machine translators did not work well and they particularly noted their “lack of naturalness” and/or “coherence” in many phrases along with errors in verb, article, and proper nouns use. In other words, in general, they are conscious of the shortcomings and clearly identify errors. Despite this, they consider such translators as important tools and are convinced that “probably” or “surely” in the future all translation will be MT with the exception of “some literary works.” In this latter case, they noted that it was necessary to take into account “the context of the work.” It can be said that in general students prioritized the advantages of MT (speed, ease of use, multiplicity of languages) over the inconveniences. To improve the results they took into account preediting and postediting.
In short, these young people demonstrated considerable confidence that future research would “get over the problems” and “limitations” that MT currently presents. For the boys and girls who came to our stand, MT is not a myth but rather an everyday reality.
Simultaneously with this dialogue with the students, a practical session divulging the reality took place.
Ways to Explain Machine Translation and Postediting
The content of the posters, drawn up on the basis of a corporate template and following suggestions about content, had to be clear, didactic, and enjoyable.
The first poster dealt with MT as a communication experience related with HT (Human Translation) and other translation processes. The working of SMT was also explained. It was stressed that in human translation the translator has to understand the text in order to be able to translate it, while in SMT, the system “translates” without understanding what it is translating. In the latter, searches are made for the most frequent equivalences in documents introduced into databases, many of these being parallel texts (existing translations from one language into another). The automatic system statistically calculates the probabilities that a particular combination of words will be translated in a particular way into the other language. The SMT system used in our lab is Moses, an open-source project started in 2005 by Philipp Koehn and Hieu Hoang, and later EU funded, which incorporates contributions from many sources. For any language pair, translation models are trained automatically. Once we have a trained model, a search algorithm quickly finds the highest probability translation among the exponential number of choices. As is clearly explained on the Moses website, the two main components in Moses are the training pipeline and the decoder, though there are other elements. In the training pipeline, raw data (parallel and monolingual) are selected and turned into an MT model. Raw data are then tokenized and word-aligned. Word alignments are used to extract phrase–phrase translations and statistics are used to estimate probabilities. The decoder is a single C++ application that will translate the source sentence into the target language. Finally, in the tuning step the different statistical models used are weighted against each other to produce the best possible translations. These technical aspects were summarized to the students without going deeply into questions affecting SMT research such as referred by Turchi, De Bie, and Cristianini (2008): How far can this representation take us towards the target of achieving human-quality translations? Are the current limitations due to the approximation error of this representation, or to a lack of sufficient training data? How much space for improvement is there, given new data or new statistical estimation methods or given different models with different complexities?
In the second poster, the activities of the research group were described along with the challenges of postediting machine translations. That is to say editing and revising text carried out by professional human posteditors on material churned out by an MT.
To exemplify this, we showed visitors an erroneous headline from a bilingual edition of a newspaper that was translated using MT and postediting. We explained to them that the example contained two types of problems: mistakes by the machine translator, unable to resolve semantic ambiguities, and negligence by the human posteditor, unable to detect the MT mistakes.
We also showed the students phrases from more distant languages such as Chinese or Russian, translated by Google Translate, into Spanish. In deciding which fragments were better translated, we could infer which were the more frequently recurring phrases within the texts stored in the translator’s database.
The third poster summarized the main linguistic obstacles to MT (polysemy, idiomatic expressions, terminology, etc.). It also proposed future lines of research such as increasing the presence of minority and less translated languages on the web, MT into many languages simultaneously, translating from audio, and so on.
Hands-On
The practical part that we had prepared involved four activities that we deemed appropriate for different profiles of students (age, interest, attention span, etc.).
The first was a 2-minute competition detecting linguistic errors and postediting problems in MT. This was to make students aware of the need for postediting in MT and to help them understand the responsibility of the translator in the final result. The examples, chosen from real-world texts, often made the audience laugh. The competition model we chose pleased the students because it appealed to their competitiveness.
The second activity consisted of noting eloquent examples of the dangers of a bad translation. Among others, we demonstrated three possible translations of a single medical prospectus.
In the third activity, lasting 5 minutes, we showed the translation of a fragment of a famous pop song. Side-by-side with the original chorus, in English, was the MT of it and the official version in Spanish, which was an adaptation that rhymed. The need for human intervention in adapting the latter was obvious. Students also appreciated the use of MT systems such as that known as human-aided machine translation, where the output of MT systems can save time and money by providing draft translations that are then postedited for publication (Hutchins, 2010).
The fourth activity consisted of demonstrating how we can all improve MT databases if the system allows feedback. In this way, users become active participants in improving the MT system. For instance, by sending correct postedited translations to Google Translate through the instruction “Not Correct?—Send.”
By the end of the experience, the students learnt how SMT really works, understood the value of human intervention, tested their SMT postediting skills, and learnt how to improve an MT system.
Conclusions: What Have We Researchers Learned?
First, it has given us an idea for future experiments. In previous academic research, we have carried out tests on postediting using professional translators and translation students. In these tests, linguistic knowledge was an important factor to be taken into account. At the Fair we have been able to observe the relationship of our field (MT) with a collective from the general public, young people with no professional training. Some of the attitudes of the people we came into contact with have made us think about three characteristics that affect the postediting of MT and which can influence its effectiveness—beyond mere linguistic knowledge. These include motivation (the fruit of a desire to win a competition), concentration span versus instinctive impulse when spotting errors, and finally, knowledge of the world that allows users to contextualize errors correctly. Often, the students who asked questions during the explanation of the posters were also the fastest in the postediting competition.
This suggested to us the carrying out of exploratory scientific studies into these personal abilities in professional translators and posteditors, since the abilities that characterize a good translator may not always coincide with those that a good posteditor of MT possesses. There are currently excellent professional translators who are reluctant to work with machine translations because they lack motivation due to the difficulties involved in improving them. These are often translators who like to consider all the options and hesitate before taking decisions and selecting the exact word or phrase.
Second, we were witnesses to the real spontaneous experiences of young people interacting with MT and their ability to correctly reformulate translations due to their open minds, free of the prejudices of previous generations toward the limitations of MT in its current state.
Last but not least, in the Fair we engaged in a dialogue with users/citizens that has encouraged us to explore the possibilities of citizen science being applied to MT. Big data and crowd sourcing approaches could be explored to get SMT postedited texts in two priority areas: less-translated languages and specific purposes domains. In the area of less-translated languages, the number of quality corpora from languages that nowadays have a minor presence on the net could be improved by educating citizens in improving MT quality. In the area of specific purposes domains, for instance, partnerships could be established with other research groups working on health, environment, genetics, and so on, which could provide corpora around specific domains for specialized MT systems. Then trained posteditors could postedit texts for those specialized fields in order to have more reliable corpora. Multilingual scientific communication could be improved through quality MT postedited by specialists.
Footnotes
Acknowledgements
We would like to thank the Barcelona Scientific Park and its co-organizers (La Caixa Foundation and Barcelona City Hall) for the opportunity to participate in the 13th Live Research Fair.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the Ministerio de Economía y Competitividad del Gobierno de España. Programa Estatal de Investigación, Desarrollo e Innovación Orientada a los Retos de la Sociedad (Grant Number FFI2013-46041-R).
