Computer-Aided Translation Technology: A Practical Introduction Lingua Inglese Università degli Studi di Roma Tor Vergata 17 pag. Document shared on www.docsity.com Downloaded by: naomi.pelliccioni1 (naomi.pelliccioni0500@gmail.com) INTRODUCTION CAT technology can be understood to include any type of computeryzed tool that translators use to help them do their job. This could encompass tools such as word processors, grammar checkers, e-mail, and the World Wide Web (WWW). While these are certainly valuable, possibly even indispensable, tools for a modern day translator. 1. WHY DO TRANSLATORS NEED TO LEARN ABOUT TECHNOLOGY? Kingscott and Haynes both issue warnings that the pace of change is beginning to accellerate. They foresee a dramatic increase in the use of CAT tools and note that this increase will be needs.driven, rather than research-driven. What does it mean? In a global marketplace, companies are finding out that failure to translate results in a loss of international sales. A clear example of this trend can be found in the software localization industry. The term “localization” refers to the process of costumizing or adapting a product for a target language and culture. Thibodeau (2000) noticed that American software companies often report international revenues exceeding 50% of total sales. The main reason for localizing a product is economic. But why? A product that is not making profit in the domestic market may perform better in another market. A non-localized product is less likely to survive over the long run localization can extend a product’s life cycle. Most products can be more profitable overseas because these markets often supports higher prices. In order to stay competitive and increase profits, companies in a variety of fields localize their products and web sites. Many companies now aim at launching a product or a web site and its accompanying document at the same time, a practice known as simultaneous shipping, or “simship”. Localization results in higher translation costs and a slower time-to-market; and time is money. Therefore the translator is sometimes required to work very quickly (and cheaply). Shaler feels that by the time they graduate, translation students must be aware of the wide variety of translation tools available and have had some exposure to a representative selection of these tools. In addition, they should have learned about the financial and operational implications of the introduction and use of translation tools in a traditional translation environment. Shaler and Kingscott are not the only ones who feel that technology deserves a more prominent place in the translation curriculum. Austermühl (2001): familiarity with CAT technology is becoming a prerequisite for professional translators. These skills are in great demand in the translation marketplace. Kenny (1999): graduates who are conversant in CAT technology are a real advantage when it comes to working in highly technologized translation environment, such as the software industry and the organizations of the European Union. But what impact did CAT technology have on the need for translation professionals? Brooks (2000): the creation of the printing press eliminated the job of the scribe, but it created a larger market for books greater demand for authors, editors and illustrators. By analogy, Document shared on www.docsity.com Downloaded by: naomi.pelliccioni1 (naomi.pelliccioni0500@gmail.com) despite what someone might think, the development of CAT tools has not reduced the needfor language professional. On the contrary, it has created jobs for translators who are skilled at using technology! 1.1 Exploring the impact of technology on translation pedagogy Bowker (2002). Additional benefits to be gained by introducing translation technology into the translator-training curriculum: Exploring the impact of technology on translation pedagogy: ■ Investigating human-machine interaction ■ Learning to evaluate technology ■ Examining how tools can change conventional practices ■ Producing data for empirical investigations ■ Reinforcing basic translation skills Integrating technology into the translation curriculum can have an impact on the way in which translation itself is taught. Scherf (1992): the use of technology has led to an individualization of the teaching process. Students can work at their own pace, while trainers have the opportunity to watch them in the translation process then they can have class-wide discussions Ahrenberg and Merkel (1996): the use of tools such as TM forces students to contemplate issues such as text type and consider the intra- and inter-textual features of texts. 1.2. Investigating human-machine interaction By reflecting on the nature of machines and how they work, students can learn to adjust their expectations about what machines can and cannot do. They can learn to adapt their working practices to maximise the benefits to be gained from using CAT tools. As students use technology, they become more aware of the fact that the computers are not capable of applying intelligence and common sense to a task in the same way that humans would. 1.3. Learning to evaluate technology As students learn to use a particular tool, they should also be encouraged to evaluate that tool in terms of its potential for helping them to complete their task more efficiently. Students can also learn to assess tools in the light of a particular task or project in order to determine which type of tool can best help them carry out that task 1.1.4 Examining how tools can change conventional practices It has been observed that technology can sometimes change the very nature of the task that it was designed to facilitate. EXAMPLES Document shared on www.docsity.com Downloaded by: naomi.pelliccioni1 (naomi.pelliccioni0500@gmail.com) When recording information on term records, translators no longer record just the base form of the term. Instead, they record multiple forms of the term so that they can minimize editing by cutting and pasting the appropriate form directly into the target text . Translators who use TMS tend to formulate their texts in such a way as to maximise their potential for reuse. 1.1.5 Producing data for empirical investigations A by-product of the use of CAT technology is the gradual accumulation of data that can be used for other types of studies Electronic corpora and TMs can provide large quantities of easily accessible data that can be used to study translation: ■ Bilingual parallel corpora can be used to investigate translation strategies and decisions. ■ Trainers can build an archive of students translations, which can be used to guide teaching practices. 1.1.6 Reinforcing basic translation skills Time spent teaching technology does not necessarily take away from time spent on other translation skills. If technology is properly integrated into the translation classroom, it not only allows the students to develop new skills, it also leads to an intensification of the basic translation curriculum. L’Homme (1999): students who use technology to produce term records, find solutions in parallel documentation, or produce actual translations are reinforcing these basic translation skills as well as developing good and realistic working practices that can later be applied in the workplace. 2. CAPTURING DATA IN ELECTRONIC FORM Barnbook (1996): Before you can analyse a text it needs to be in a format in which the computer can recognise it, usually in the format of a standard text file on a storage medium. In order for a translator to take advantage of specialized translation technology, the texts to be processed must be in electronic form. In any case, when data are not currently machinereadable document arrives as a fax or a prontout, they must be converted. Ther are two main ways of doing this: 1. Using a combination o scanning hardware and optical-characterrecognition software; 2. Using voice-recognition technology. Optical-character recognition: OCR software takes the scanned image and, through a process of pattern matching, converts the stored image of the text into a form that is truly machine-readable and can be processed by other software. At its most basic, OCR software examine each character in the scanned image and compares it to a series of character patterns stored in a database. Once all the characters have been processed in this way, the new file can be saved in an appropriate format (e.g. a text file) and opened in an application such as a word processor, where it can be edited. A scanner is a computer peripheral, there are several type of scanner: Document shared on www.docsity.com Downloaded by: naomi.pelliccioni1 (naomi.pelliccioni0500@gmail.com) • Handheld scanner are small and lightweight, and have a limited scanning window. Inexpensive. • Large freestanding units are capable of processing vast amounts of data quickly or accurately but they are quite large and relatively expensive. • Flatbed scanner that looks like a small photocopier, this is the type of scanner that is best suited to the needs of most translators. It is easy to use and can process reasonable amounts of texts fairly quickly and accurately. Once the document is in position, the scanner draws an imaginary grid over the document, dividing it up into tiny segments known as pixels. The greater the number of pixels, the higher the resolution of the scanner. Factors affecting the accuracy of OCR 2.1. • Quality of the hard copy • Layout of the text • Quality of the scanning device • … Typical mistakes: number “5” mistaken for letter “s”, letter “r” and “n” mistakenly combined to form letter “m”, letters “c” and “l” mistakenly combined to form letter “d”. In such circumstances, the OCR software does not pick out all the characters. A technique is to integrate a dictionary checking stage, but that way don’t solve all misinterpretation problems. If an OCR program mistakenly identifies the “e” in read as an “o”, the resulting word will be “road”, which will not be identified as an error during a dictionary check because it is a legitimate word. In future OCR software may need to take even larger contexts into account in order to determine whether a given word makes sense in this larger context. 2.1. Voice recognition Voice recognition, also known as speech recognition, is a technology that allows a user to interact with a computer by speaking to it instead of using a keyboard or a mouse. The user speaks into a microphone linked to a computer. The software acoustically analyses the speech input by breaking down the sounds that the hardware “hears” into smaller, indivisible sounds called phonemes. Then a “best guess” algorithm is used to map the phonemes and syllables into words. The computer’s guesses are then compared against a database of stores word patterns. Voice recognition also uses grammatical context and frequency to predict possible words. These statistic tools reduce the amount of time it takes the software to search through the database but also help differentiate between homophones – words that sound the same but are spelled differently and have different meanings, such as “to”, “too”, or “two”. Document shared on www.docsity.com Downloaded by: naomi.pelliccioni1 (naomi.pelliccioni0500@gmail.com) EXAMPLE: in a context such as the following word being “many”, it may seems a logical “best guess” that the preceding word is “too”, “There are too many people in this room”. Type of word recognition: • A voice-recognition dictation system that is speaker dependent requires a user to “train” the software before it can be used. Training consists of reading through a manufacturer-provided list of word or sentences that contain all the different sounds in a language. The use is required to pronounce each word on the list several times before moving on to the next word. A system that has been trained by user A cannot then be used successfully by user B. • Speaker-indipendent voice-recognition technology, also known as universal voice recognition, does not require user training. The manufaturers have pre-trained the software using speech samples from a wide range of people. But if the user speaks with a foreign accent it may be diffucult to get the system to work properly. Another distinction between voice-recognition system is whether they are command/control system or dictation system. • Command/control system allow users to interact with the computer by giving a limited set of commands, which usually correspond to the commands found on application menus (open, copy, paste ecc.) • Dictation system allow users to enter new data into the computer by dictating insted of typing PROS: • Good for poor typists • Good for those who suffer from a physical or visual impairment • Dictating is normally quicker than typing CONS: • It takes time to edit the text and correct errors made by the VR • VR systems work for specific languages (= money for multiple packages) 3. Corpora and Corpus-analysis tools “Putting a word in context means breathing life into it […] If you want to know how words behave you must study them in their natural environment, and the natural environment of words is text, context” (Roumen & Van der Ster, 1993, in Bowker, 2002). The word “corpus” comes from Latin. […] its sense of “body of a person” started in the mid-fifteenth century and the sense of “collection of facts or things” occurred later in 1727. The year 1956 saw an extension of the meaning to include “the body of written or spoken material upon which a linguistic analysis is based”. Document shared on www.docsity.com Downloaded by: naomi.pelliccioni1 (naomi.pelliccioni0500@gmail.com) Leech (1992): Corpora of text collection had been used by linguists and grammarians for the study of language long before the invention of the computer. “computer corpus linguistics” (and not only “corpus linguistics”) would be a more appropriate term for studies based on language database today. Baker (1995): “Corpus-based translation study” (CTS) is the use of corpus linguistic technologies to inform and elucidate the translation process. What is a CORPUS? “In its broadest sense, a corpus is a collection of texts or utterances that is used as a basis for conducting some type of linguistic investigation” Translators usually: ■ Compile and analyse corpora for terminological researches ■ Consult corpora of parallel texts to produce a TT with the appropriate style, format, terminology and phraseology Electronic corpora A corpus in electronic form is an electronic corpus, and its advantage is that it can be manipulated by a computer (and quickly scanned and analysed!) It must be noted that a corpus is not a random collection of texts. The texts are selected according to explicit criteria in order to be used as a representative sample of a particular language or subset of that language. Representativeness is a define feature for a corpus! Types of corpora Given that corpora are specially designed to meet the needs of the project at hand, there are as many different corpora as there are projects. Nevertheless, it is possible to identify some general characteristics that corpora may have: ■ General/reference vs. specialized corpora ■ Written vs. spoken corpora ■ Synchronic vs. diachronic corpora ■ Monolingual, bilingual or multilingual corpora ■ Comparable vs. parallel corpora ■ Native vs. learner corpora Monolingual and bilingual parallel corpus A parallel corpus is a corpus that contains a collection of source texts in a language A aligned with their translations into language B. Document shared on www.docsity.com Downloaded by: naomi.pelliccioni1 (naomi.pelliccioni0500@gmail.com) A monolingual corpus is one that contains text in a single language. A bilingual parallel corpus (also called “bitext”) can be a very powerful tool for a translator source texts aligned with their translations. Corpus analysis tools A corpus-analysis tool is a software used to access and display information contained in a corpus and typically contain features that allow the user to generate and manipulate: ■ Word-frequency lists ■ Concordances ■ Collocations Word frequency lists The most basic feature for a corpus-analysis tool A word-frequency list allows the user to discover how many different words are in a corpus and how often they appear; and it can be manipulated in many ways: ■ Lemmatized lists (words like “translate” “translated” “translating” are treated as separate forms even though they are related. Related words are grouped under a lemma, the total word count for the lemma is sown in the right-hand column, and the frequency counts for individual words belonging to the lemma are shown in parentheses.) ■ Stop lists (contains any items that a user wants the computer to ignore). Let’s take this example (from Bowker, 2002): “I really like translation because I think that translation is really, really fun”. 13 words the corpus contains 13 tokens but some words appear more than once (I, translation, really), and therefore this corpus contains only 9 different words 9 types In a word-frequency list, the number of tokens is shown beside the type. Concordancers Translators not only have to be able to understand the ST, they also have to produce a TT. Dictionaries are helpful, but in order to be able to determine how terms can be used, it is useful to see them in context, and, preferably, in more than one context. A second feature that is common to most corpus-analysis tools is a concordancer. A tool that retrieves all the occurrences of a particular search pattern in its immediate contexts and displays them in an easy-to-read format. The most common display format is known as KWIC (key word in context) display. All of the search pattern are lined up in the centre of the screen. Typically permit more sophisticated search patterns, allowing functions such as case-sensitive searches (to distinguish between Polish and polish) or more characters in a search strings (print* to retrieve print, printed, printer ecc), and searches using Boolean operators (AND, Document shared on www.docsity.com Downloaded by: naomi.pelliccioni1 (naomi.pelliccioni0500@gmail.com) OR, NOT). Is also a context search, in which another term mut appear within a userspecified distance of the search patterns (context in which printer appears within five words of cartridge). Bilingual concordancers Is a tool that can be used to investigate the contents of a parallel corpus. A parallel corpous is a corpus that contains a collection of source texts in a language A aligned with their translations into language B. Alignment is the process whereby sections of the source text are linked up with their corresponding translations. Alignment can take place at many different levels: text, paragraph, sentence or even words. Bilingual concordancers employ statistical measures in order to try to identify possible equivalents for specific search terms. For the best results the ST and the TT must have a similar structure. Problems include cases in which one sentence in the source text has been translated by two sentences in the target text, or vice versa. Like monolingual concordancers, the bilingual one retrieve all occurences of a particular search pattern in its immediate contexts search pattern can be entered in either language. For example, if the search term “disk” is entered using English as the search language, the tool analyzes all the possible equivalents, that includes, for example, the French terms dique, disquette, lecture cc. We also have bilingual quert in which the user can specify a search term in both languages. Collocations Many corpus-analysis tools have the ability to compute collocations, that is characteristic co-occurrence patterns of words: words that typically “go together”. Because language is not random, certain words tend to cluster together, and some of these clusters form collocations. The formula commonly used for determining the likelihood that two words are collocates is the mutual information formula (MI). The higher the MI, the stronger two words are connected. PROS: • Frequency data can be easily generated • Translators can see terms in a variety of contexts simultaneously • A great number of documents can be quickly consulted CONS: • Availability and copyright can be an issue • Aligning text in the case of bilingual corpora is time-consuming • The user have to develop sensible research strategies • Not all tools come equipped with characters sets for all languages 4.Terminology-Management System Document shared on www.docsity.com Downloaded by: naomi.pelliccioni1 (naomi.pelliccioni0500@gmail.com) A major part in any translation project is identifying equivalents for specialized terms. Subject fields such as computing, manufacturing, law and medicine all have significant amounts of field-specific terminology. Researching the specific terms needed to complete a translation is a time-consuming task. A terminology-management system (TMS) can help with various aspects of the translator’s terminology-related tasks, including the storage, retrieval, and updating of term records. Effective terminology management can: • Help to cut costs • Ensure greater consistency (internally and externally) • Improve linguistic quality • Reduce turnaround times for translation STORAGE TMS acts as a repository for consolidating and storing terminological information for use in future translation projects. IN THE PAST: • TMSs stored information using one-to-one correspondences. • Users could only choose from a predefined set of fields. TODAY: • TMSs use a relational model, which permits mapping in multiple language directions. • TMSs adopt a free entry structure, which allows users to define their own set of fields. RETRIEVAL Once the terminology has been stored, translators need to be able to retrieve this information. A range of retrieval mechanisms is available: • A simple look-up to retrieve an exact match. • Wildcards for truncated search (EX: comput*) • Fuzzy matching techniques. A fuzzy match retrieves terms that are similar to the requested search patterns, but that do not match it exactly. Document shared on www.docsity.com Downloaded by: naomi.pelliccioni1 (naomi.pelliccioni0500@gmail.com) ACTIVE TERMINOLOGY RECOGNITION AND PRE-TRANSLATION A feature of some TMSs, particularly those that operate as part of an integrated package with word processors and translation-memory systems. It is essentially a type of automatic dictionary look-up. Some systems also permit a more automated extension of this feature in which a translator can ask the system to do a sort of pre-translation or batch processing of the text. TERM EXTRACTION Also called “term recognition” or “term identification”. This process can help a translator build a term base more quickly. Even if the extraction is performed by a computer, the resulting list of candidates must be verified by a human. Semi-automatic (computeraided) process. Unlike the word-frequency lists described earlier, term extraction tools attempt to identify multi-word units linguistic approach and statistical approach Linguistic approach to term extraction: the tools tries to identify word combinations that match particular part-of-speech patterns. EX: NOUN+NOUN or ADJECTIVE+NOUN (typical in English) In order to implement such an approach, each word in the text must by tagged with its appropriate part of speech. Then, the tool identifies all the occurrences that match the search. Unfortunately, not all texts can be processed this neatly: 1. Not all the combinations that match the specific patterns can be qualified as terms. 2. Some legitimate terms may be formed according to patterns that have not been pre-programmed into the tool. Adopting a linguistic approach to term extraction, the tool looked for the combinations NOUN+NOUN and ADJ+NOUN Not all the combinations that match the specific patterns can be qualified as terms. EX: NOUN+NOUN or ADJ+NOUN “antivirus software” “current status” NOISE 2. Some legitimate terms may be formed according to patterns that have not been pre-programmed into the tool. EX: NOUN+NOUN or ADJ+NOUN “after-the-fact detection” SILENCE (is legitimate but its pattern is PRE+ART+NOUN+NOUN Statistical approach to term extraction: the most straightforward statistical approach to term extraction is for a tool to look for repeated series of lexical items, specifying a threshold (the number of times a series must be repeated). If the threshold is two, a series of lexical items must appear at least twice to be recognized. Document shared on www.docsity.com Downloaded by: naomi.pelliccioni1 (naomi.pelliccioni0500@gmail.com) Unfortunately, not all repeated series qualify as terms ( NOISE) and not all legitimate candidates are repeated ( SILENCE) Both linguistic and statistical approaches have drawbacks, but there is a clear advantage in adopting the latter, statistical approach is NOT language-dependent while the linguistic approach is. 5. Translation-Memory Systems One of the most important sources of information to which a translator can have access is a large body of previous translations. (Kay and Roesheisen, 1993, in Bowker, 2002) Given the staggering volume of translations produced year after year, it is quite obvious that existing translations contain more solutions to more translation problems than any other available resources. (Isabelle, 1993 in Bowker, 2002) The concept of Translation Memory has existed for some time. The idea originated in the 1970s. What is a Translation Memory? A TM is a type of linguistic database that is used to store source texts and their translation, explicitly aligned. A TM is a type of linguistic database that is used to store source texts and their translation, explicitly aligned. The texts are broken down into short segments that often correspond to sentences (often, but not always!) Translation Unit made up of a source text segment and its translated equivalent. Most simply, a TM can be viewed as a list of source-text segments explicitly aligned with their target text counterparts. Translation Unit a source text segment aligned with its translated equivalent. A TM can be viewed as a list of source-text segments explicitly aligned with their target text counterparts. Does it ring a bell? The resulting structure of a TM is sometimes referred to as a parallel corpus, or bitext. Using a TM system the translator will be able to “recycle” previously translated segments. These systems work automatically comparing new source segments against a database of translations. If a matching segment is found,the system will propose the “old” translation to the user, who will then decide whether to use or discard it. SEGMENT the basic unit in a TM system. But deciding what constitutes a segment isn’t easy! Document shared on www.docsity.com Downloaded by: naomi.pelliccioni1 (naomi.pelliccioni0500@gmail.com) SEGMENT the basic unit in a TM system. In most instances, the basic unit of segmentation in a TM is the sentence, and this is why TM are sometimes called sentence memories. However, not all texts are written in sentence form (e.g. headings, table cells, etc.) Many TM systems allow the user to define other units of segmentation in addition to sentences, which can include sentence fragments or even entire paragraphs. Deciding what constitutes a segment is not a trivial task! It seems easy to decide that full sentences will qualify as segments, but how can a TM system identify sentences? Punctuation such as periods, exclamation points and question marks are typically used to indicate the end of a sentence but what happens in case of an abbreviation such as Mr. or Dr.? Or in case of an ellipsis, which can appear in the middle of a sentences? Some of these problems can be solved incorporating stop lists into the TM systems. Another issue related to segmentation is the fact that the segmentation units used in the ST may not correspond exactly to those used in the TT. This lack of one-to-one correspondence can create difficulties for automatic alignment programs. Most TM systems present the user with a number of different types of segment matches. What is a match? Matches are correspondences between a new SL segment and one or more “old” translations contained in the database. The most common types are: ■ EXACT matches ■ FUZZY matches ■ TERM matches EXACT (also called “perfect” matches) 100% identical, including spelling, punctuation, numbers, even formatting, etc. FUZZY when a fuzzy match is found, it means that in the database there is a segment that is similar to the new one (the similarity can range from 1-99%, but the user can set the sensitivity threshold – the standard is between 50 and 99%). TERM if working in association with a term base (a terminological database), the TM system will compare the single terms contained in the new segment with the ones in the term base. A Translation Memory is essentially a type of database. It is basically a software that allows a user to store and retrieve information. However, as with any database, the information must be provided by the user. Therefore, when the user first purchase a TM system, the database is empty. The system becomes useful when the translator begins to store some data (source and target texts) in the TM. Document shared on www.docsity.com Downloaded by: naomi.pelliccioni1 (naomi.pelliccioni0500@gmail.com) How can we create a Translation Memory? Two main ways: 1. Interactive translation: while we translate the text within the TM system, the new TL segments are fed to and stored in the TM. 2. Post-translation alignment: if we have some source texts and their correspondent translations (translated “in the old fashion”), we can upload them in the TM system, ALIGN them and feed the translations to the TM. TM can be exported and sent! Given that a TM system allows the user to re-use previously translated work, which kind of texts are more suitable for inclusion in a TM? The most suitable texts for a TM are repetitive and highly specialized texts, and texts that will be updated or revised: ■ Text with internal repetitions (the higher the percentage of repetitions, the more desirable it is to use a TMS) ■ Revisions (amended version of a previous text) ■ Recycled texts (sometimes referred to as external repetitions) ■ Updates (e.g. when the client makes changes to a text that you are already translating) According to Bowker (2002), the first thing to take into consideration is that an empty TM is of NO use. The performance of the TM system is dependent on the scope and quality of the existing DB and the quality of the translations store in the DB is dependent on the translator’s skills! PROS: it saves you time CONS: if you can’t use the software properly, it’s time consuming! You will need a few week’s training to be able to use a TM system in a way that it will save you time, instead of making you lose time! PROS: it improves consistency (internal and external) CONS: The rigidity in maintaining the same ST’s order in the TT may affect the naturalness of the translation Document shared on www.docsity.com Downloaded by: naomi.pelliccioni1 (naomi.pelliccioni0500@gmail.com) Different software applications store information in different formats, and Translation Memories are no exception. The format used by any given TM is not necessarily compatible with those of other TMs or TM systems. A standard data-exchange format for TMs was developed through the years to solve this problem TMX. The purpose of TMX is to make it easier to import and export data between different TM systems without losing or distorting information. Further considerations ■ TM systems are often quite expensive (even though the prices have been dropping and a few free systems are emerging) ■ And they tend to need “high” minimum requirements to work properly on a PC (a lot of RAM and a good CPU) ■ Different systems work with different formats (even though some standard are emerging – ex: .TMX for TM) ■ Some languages are easier to process than others (especially when it comes to handle the segmentation) ■ Using TM systems affects payments , as the clients may want to pay less for exact and fuzzy matches (but isn’t it fair, in a way?) ■ A “full” TM is an asset, and issues of ownership may arise 6. Other new technologies and emerging trends 6.1. New attitudes toward translation and translators As the volume of translation increases, translators are depending more and more on technology to help them with their task. In order to maximize the usefulness of technology, it has become necessary to change the way in which the translation process is viewed. In the past was often seen as an “add-on” process, and it was divorced from the principal document-production process. An increasing number of clients are beginnning to implement more stringent writing and style guidelines. By asking writers to use a sort of controlled language, clients can maximize the recyclability of texts during the translation process –> eliminating anapora, using preferred terms, avoiding ellipses, usin punctuation consistently, all this skills can increase the number of matches found by a TM. Tools make easier for translators to begin working on documents even before they have been finalized. If an author updates a document, the translator can run the updated source text through the TM system and i twill quickly identify new or changed segments. 6.2. New types of translation work generated by technology Document shared on www.docsity.com Downloaded by: naomi.pelliccioni1 (naomi.pelliccioni0500@gmail.com) As the number of electronic products and resources increases, there is a growing demand for translators who are able to translate media such as software applications, multimedia products, web pages, and even on-line chat sessions. Software localization which is the adaptation of software package to a target language and culture, is one the fastest-growing translation markets. It involves some additional consideration: 1. One issue that translators must deal with is the physical constrains of the screen space. In some development environments, the width of a menu or a button can be adjusted (to accomodate the longer french term “sauvegarder” as “save”), but in other cases, the width is fixed and the translator must choose a term or a transparent abbreviation that fits within the allocated space. 2. Icons or other visual elements can also be problematic. For example, whereas an icon showing an open door may transparently capture the notion expressed by the corresponding command “exit, the transparency would be lost if the translator chose to translate that command into the target language using an equivalent that meant “quit” or “terminate”. 3. Variables are another element that may cause difficulties for translators. A variable i a character or string of characters that acts as a placeholder and is replaced by anoyher, more meaningful string of characters when the relevant software is running. Es pag 132 non c’ho capito troppo. 4. Standardization. If a particular product belongs to a larger family of products, it may be necessary to standardize the terminology (use exit instead of quit). 6.3. New technology generated by new types of translation work Other types of tools are being developed: • Those that check to see if a translation has been truncated because of space restrictions. • Those created to separate out the material that needs to be translated from the software code itself a translator must be able to identify and translate all relevant material without deleting or distorting the computer code. To this end, different types of software have been developed to assist a translators. Some of the simpler software displays the translatable and the non-translatable material in different colours so that translators will know where to focus their attention. More sophisticated software will protect the tags to ensure that they cannot be deleted or edited. 6.4. Conditions required to ensure the continued succes of CAT tools Document shared on www.docsity.com Downloaded by: naomi.pelliccioni1 (naomi.pelliccioni0500@gmail.com) Haynes suggest that more developers of CAT tools should make their product freely available to translator training institutes, noting that these developers may well benefit financialy when the students graduate and are in a position to influence corporate purchasing decisions regarding translation technology. 6.5. Future developments With regard to TMs in particular, there has been concern among translators that the notion of a “text” has been lost because the tools operae primarily at sentence level. Macklovich and Russel also point out that text has global properties that are not easely associated with its lower level components, including administrative information concerning who originally translated a text the date of the translation, the client and so on. In terms of more general developments, it is likely that there will be continued movement away from stand-alone systems and toward a client-server architectur which facilitates networking, thus making it possible for multiple users to share the same corpora, TMs, or term bases. Other improvments that are underway include extending CAT tools to support a wider variety of languages (thai, urdu) by using encoding methods such as unicode and designing new standards and filters to support a wider variety of file formats, including formats using tags (html, xml) without losing the formatting of original source text. Transrouter will evaluate the uselfulness of available resources in the context of specific translation project based on information supplied by the user (a translator or translation manager) and on the automatic analysis of text characteristics using its component features. These component features will include a cost estimator, a word counter, a sentence-length estimator and a sentence-semplicity checker. Importanti capitoli 4-5 da leggere bene sul libro. Document shared on www.docsity.com Downloaded by: naomi.pelliccioni1 (naomi.pelliccioni0500@gmail.com) Document shared on www.docsity.com Downloaded by: naomi.pelliccioni1 (naomi.pelliccioni0500@gmail.com)
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )