Sonntag, 3. April 2011

Open-Source-Software for translators

by Alexandra Kleijn - originally published in Heise C't Open source

The world of translators is a Windows world: MS Word, and the translation tool SDL Trados are the measure of things. Whoever prefers a different operating system or does not want to bow to the dictates of SDL, has so far had a hard time. But now alternative tools are available - even for Mac and Linux users. And they are in part Open Source.  

Hardly any professional group is as loyal to Microsoft as translators. Not surprising when you consider that their clients provide basic texts in many cases still in the well-known doc format -. The migration to Office 2007 and the still relatively new 2010 - and on XML-based, open file formats - is taking place only slowly.


With the open source office suite OpenOffice, however,  a free alternative to MS Office is available, which is hardly inferior to its competition in terms of its functionality. A major advantage of the open-source suite is the fact that it runs equally well on Windows, Mac OS X and Linux. OpenOffice as an independent office software is a good alternative for a number of users as well.

OpenOffice is open source: the application is available together with its source code and anyone who feels called upon can change the application according to his own ideas. Open-source software may also be in general distributed without any restrictions.  Even if the developer of an open source program is free to ask for a license fee for his application, open-source software is usually free of licensing costs. Open source software is generally based on open standards and data exchange between different programs is easy.


Document formats: MS Word reigns
OpenOffice Writer, the word processing application in OpenOffice, can handle the well-known Doc format, used in  Microsoft  2003 and its earlier versions is. The problem, however, are MS-Office documents with macros, many embedded images and forms. Even with complex formatting, such as a division of the text in columns, must often be handled manually.

It does not look so good for the exchange of documents using the new Microsoft's XML-based Office Open XML
format (OOXML), the default file format for MS Office 2007 and 2010 (file extension. Docx). The original format is often lost when opening docx files in OpenOffice. The way back  is (still) blocked: Documents can't be saved as docx in OpenOffice Writer does not save as docx, at least not under Windows and the Mac - the Linux version of OpenOffice offers this possibility.  This rather limits the  usefulness of OpenOffice as a MS-Office replacement for translators: the docx file received from the client for translation, must be returned  in the old DOC format or in the ODF format, available in OpenOffice.

 
ODF, a standard format  for text documents, approved by ISO, can be opened directly in Word 2007 with Service Pack 2 and in Word 2010. In older MS Office versions one can upgrade to ODF support using the Sun ODF Plug-in. The new owner Oracle will ask for money, but one can still download the plug-in from  Softpedia free of charge. The plug-in knows the functions of the new ODF specification 1.2, Microsoft Office 2007 and 2010 only support ODF 1.1. Mac users are left with nothing: Microsoft Office for Mac 2008 and 2011 do not recognize the ODF format.
  
LibreOffice saves in docx format as well
Since autumn 2010, the OpenOffice offshoot LibreOffice has been courting the attention of users. This so-called fork was created after the Oracle's takeover of the Sun, lead developer of the open source office suite. The current LibreOffice 3.3, based on the same source code as OpenOffice 3.3,  can open as well save docx files. It is plagued, however, by the same formatting problems applies  OpenOffice.
For Mac users Microsoft Office for Mac 2011 is available
since last October. The new version has move closer to the Windows version. It also finally supports Visual Basic for Applications (VBA), so that Office macros should  work across platforms.

Translation Tool Translation Memory
For many  translators a Translation Memory System is indispensable. In the market for TM environments one can meet a lot of providers. In the last few years the Top Dog SDL, the manufacturer of the TMS Trados, has gotten quite some competition breathing down its neck. In Germany, for example,  Across and MemoQ. All these tools are, however, proprietary software. After all, cross-platform TM tools such as Wordfast Pro and Swordfish have broken Microsoft's hegemony somewhat. They are written in Java and run also under Mac OS X and Linux.
 

Few years ago everybody was cooking his own little dish, with all the resulting compatibility issues. The latest trend goes increasingly toward open standards: XLIFF (XML Localization Interchange File Format) for document exchange and TMX (Translation Memory Exchange) for translation memories. Both are based on XML. The eXtensible Markup Language separates the content and the other information such as formatting and meta tags, and has established itself as a standard for cross-platform and cross-program data exchange of all kinds. The situation is by no means optimal: many manufacturers push their own interpretations of these standards - so for many, the "dirty" Bilingual Word document remains the measure of things.


OmegaT 

OmegaT - OpenSource TM-based translation tool
for Windows, Mac OS X and Linux
At the moment the only open source TM tool, ready for the production use, is OmegaT.  It is written in Java  and can work directly on Microsoft's OOXML texts. Documents, created in earlier versions of Office documents , must first be converted either with MS Office 2007/2010/2011 (Mac) or to open office before they can be translated. When converting to OpenOffice, the previously mentioned risk still exists, namely that expensive-looking documents do not survive the transition without any formatting changes.

The project openTM2 has yet to grow beyond the beta stage. The focus here is the open-source implementation of a TM oldie: the IBM Translation Manager. The lofty goal of openTM2 has been to become the reference platform for the translation memory exchange standard TMX. The trial version currently available runs only on Windows. 

...to be continued ... 1/2                  

Translation: smo 

Montag, 7. März 2011

A week of awesome tweets

 Note: this has been posted Jan 16th, 2011 9:04pm somewhere else. I decided to bring it in here, to keep my chickens in one place.

Introduction
One of the features in the new SUMO is the invitation to help Firefox users on Twitter (https://support.mozilla.com/en-US/army-of-awesome). Tweets, mentioning Firefox,  get filtered out from the main stream and the Army of awesome (AoA) gets the opportunity to help.
Here's one typical yelp for help, with user's Firefox evidently falling to pieces:

The careless spelling and the four-letter wording is typical for this environment: I have a problem and I want to have it fixed now. The answer provided is a copy of the 16-odd boilerplate answers available to AoA.
I have collected about a week of answers to find out
  • how many are involved in AoA
  • what are the problems they try to address
  • how successful they are
  • ... just curious
How many are involved?
In the week of 7-15 Jan there were at least 772 tweets with #fxhelp tag; at least, because my own participation in this army has been consistently and completely stricken from the public tag listing, statistics etc since around the middle of December. I have thus added a typical week of my own saved contributions. Note that it's quite possible that the dark number is even bigger;  nobody but Twitter can tell.
How many are involved in the army: there's been 54 contributors. Here's a general picture:

legend: how many participants tweeted how many times; for instance 5 tweeted 10-20x.
On the low side there's a group of casual participants (<10) with 36 members, with 20 out them turning up only once. On the high side there's  the group at the top with 7 participants who twitted considerably more.
The following diagram shows the total number of tweets for every of the five groups above:

The top group, counting 7, accounts for 538 / 853 = ~60% of the total traffic.
What are the problems?
The answers provided indicate what the original question has been. The boilerplate suggestions have the possible problems and suggestions pretty well covered. The suggestions are provided in five groups, tagged with #~ in the pie-chart below:

The support questions were one third of twits and,  surprisingly,  the "no URL provided" group was second biggest at 25%. Here's the inner workings of the Support group:
  •   107 Fix crashes
  •    66 slow Firefox startup / Firefox is slow
  •    53 Firefox does not behave
  •    32 High RAM usage 
  •    16 Quick Firefox fixes
The "Get involved" group on the other side of the spectrum looks as follows:
  • 10 Become a beta tester
  •   9 Report a bug
  •   6 Get involved with Mozilla
  •   3 Mozilla Developer Network 
  •   0 Join Drumbeat 
Conclusion
The question of course remains: does it matter? Is there any traction behind it? Surprisingly I have received a lot of feedback, about 50 since the beginning of January. Negative reactions were just a few, some of them sobering and useful like this one:
  • ... that was lame. give me a link when I have no browser.LOL I rebooted & now i'm good. thanks. 
The majority, however, was worth a smiley each. Here's a few:
  • thanks :) yeah I had a couple to update but it is all good now :) I have bookmarked that link too and I will check it regularly :)
  • My addons/plugins are all up to date. I'll try disabling them to see if any are causing problems. Thanks for the suggestion!
  • Thanks for the suggestion will check it out. Happened when we attached files in google mail.
  • Hey thanks a lot!!! :)
To avoid ending on the high-fiving mote, here's a few thoughts:
  • Tweeting can turn into spam if used in a fire-and-forget mode. Do not forget to check the feedback.
  • The perfect score of 0 for Drumbeat provokes me. To be honest, it is hard to sell Drumbeat to desperate teenagers, who can't tweet to their heart's content because Firefox is giving them hives.
  • Thinking of those 20 single-twit contributors: fact is they tagged their message with fxhelp. Our future warriors? A pool of dig-Mozilla aka 5$ Mozillians? Dig it!

Mittwoch, 23. Februar 2011

I Will Not Be Told: Stephen Fry's Speech At Harvard

(from science 2.0)

"...To be told is to wallow in revealed truths.  Bibles and similar religious texts are all about revealed truths which cannot be questioned, and the origins of which require the readers to make many assumptions.  And it was even worse in the dark times of religious book control and illiteracy in which you might not even be allowed to read the book--you have to get the mediated verbal account from someone supposedly holier than you... Discovered truths, on the other hand, are not told.  Of course somebody could tell you a discovered truth, but if you don't trust them you can question it.  Discovered truths can be discussed.  They can be questioned and tested.... "

Hit me the first time, when reading Wieland - or was it Lessing -. About believing without any question into a (badly written and communicated) narrative of things that happened thousand and some years ago.While having problems with the radio report on the rush hour ("...you just cant believe these dudes anymore...")

That's why my penchant for Sokrates: his middle name was curiousity.

ASK sauthon!

Donnerstag, 8. Juli 2010

In a session on Verbatim, Pootle etc yesterday, Matjaz presented a Jetpack add-on, that allows one to use a translation memory (in a TMX form) during the  Verbatim process. The example showed nicely, how one can take an existing TMX file, list the candidates (100% match and some fuzzies as well) and then drop the selection into the target segment.

Verbatim does not use any translation memory mechanism (as far as I know, not even MO, the nonstandard Pootle's version of a translation memory) to propagate the translations which already exist through the material to be translated. Well, fine, but then the JetPack solution by Matjaz will fix it, will it not?

Please note first, the TMX, to be used in the JetPack, comes from outside. It is not an integral part of the Verbatim / Pootle solution. It is a deus ex machina, a translational band-aid - for now at least, if not for some time coming.


Matjaz showed during the demo of his Jetpack solution, how  "Australia" is translated into "Avstralija" with the help of an existing translation memory. Just before this example, however, there was a case of "Argentina" without any suggestion from TMX. Matjaz dropped in the copy of the source and then proceeded  on to the next entry. But hold on a second: where did the target entry for "Argentina" land?  Just in the PO? How adding the new stuff to TMX? After all  it would not hurt to adhere to the Mozilla spirit and expand the TMX with the working being done. I am pretty certain, however, that in the case of Matjaz' add-on TMX get short-changed. 

Here's the situation in a nutshell: external translational sources are tapped to (help Verbatim) help the localizer do the work. The Verbatim per se is not involved,  as it does not as yet have any ways and means to tap into the possibilities, offered by standardized translation memories. On the other hand, the external sources are on receiving end of the stick in the Jetpack proposal, as their initial investment and the use of their assets, as implemented in the Matjaz' proposal, is not  reciprocated by expanding the memory with new translations.

We all have open and free standards for everybody written all over our hearts, flags and banners. Standardized translation memory - free and open-sourced for all - must be among them.

Vito

PS:  one of the listeners, who was evidently new to the subject of translation memories and specifically to the subject of TMX, commented with visible relief "Oh, fine, it's just an XML file". Oh well... XML file is not much, if anything at all, without a corresponding DTD.

To see DTD for TMX files - and implicitely all the work and drive behind it - see
TMX specification by LISA

Jetpack add-on for Verbatim

at Mozilla Summit 2010, 8.july

In a session on Verbatim, Pootle etc yesterday, Matjaz presented a Jetpack add-on, that allows one to use a translation memory (in a TMX form) during the  Verbatim process. The example showed nicely, how one can take an existing TMX file, list the candidates (100% match and some fuzzies as well) and then drop the selection into the target segment.

Verbatim does not use any translation memory mechanism (as far as I know, not even MO, the nonstandard Pootle's version of a translation memory) to propagate the translations which already exist through the material to be translated. Well, fine, but then the JetPack solution by Matjaz will fix it, will it not?

Please note first, the TMX, to be used in the JetPack, comes from outside. It is not an integral part of the Verbatim / Pootle solution. It is a deus ex machina, a translational band-aid - for now at least, if not for some time coming.


Matjaz showed during the demo of his Jetpack solution, how  "Australia" is translated into "Avstralija" with the help of an existing translation memory. Just before this example, however, there was a case of "Argentina" without any suggestion from TMX. Matjaz dropped in the copy of the source and then proceeded  on to the next entry. But hold on a second: where did the target entry for "Argentina" land?  Just in the PO? How adding the new stuff to TMX? After all  it would not hurt to adhere to the Mozilla spirit and expand the TMX with the working being done. I am pretty certain, however, that in the case of Matjaz' add-on TMX get short-changed. 

Here's the situation in a nutshell: external translational sources are tapped to (help Verbatim) help the localizer do the work. The Verbatim per se is not involved,  as it does not as yet have any ways and means to tap into the possibilities, offered by standardized translation memories. On the other hand, the external sources are on receiving end of the stick in the Jetpack proposal, as their initial investment and the use of their assets, as implemented in the Matjaz' proposal, is not  reciprocated by expanding the memory with new translations.

We all have open and free standards for everybody written all over our hearts, flags and banners. Standardized translation memory - free and open-sourced for all - must be among them.

Vito

PS:  one of the listeners, who was evidently new to the subject of translation memories and specifically to the subject of TMX, commented with visible relief "Oh, fine, it's just an XML file". Oh well... XML file is not much, if anything at all, without a corresponding DTD.

To see DTD for TMX files - and implicitely all the work and drive behind it - see
TMX specification by LISA

Freitag, 23. April 2010

lokalizacija = sl.jar

Če si v (zaenkrat shredder imenovani) instalaciji pogledaš podmapo chrome, boš naletel na datoteko sl.jar, ki vsebuje naše dosedanje delo. Lokalizacija ni nič več in nič manj kot vsebina te knjižnice.

Za novo verzijo je treba (samo in nič več kot to) novo verzijo sl.jar.

Francesco( mozilla l10n blog )
You can simply download the latest nightly and work on the chrome/sl.jar of that package: if you're using Windows (but it's probaly possible for other platforms as well), you can use 7zip's File Manager and change files without unpacking/repacking the .jar file.
Da vidimo ... 7Zip. Zap ... Ha!...  dejansko.

Mittwoch, 3. März 2010

Prevodi vključeni v release proceduro

Oglasil se je Simon Paquet z lepo novico, da je prevode vdelal v sistem - gre za stanje, preden je Aleš R poslal zadnji delež. Tu dnevnik dosedanjih dogodkov

Stanje je videti na naslednjih naslovih:sl shipping 3.0 in sl shipping 3.1
Poleg manjkajočih delov je Simon omenil še naslednje: prevajali smo tudi zadeve, za katere je v komentarjih eksplicitno zahtevano, da se jih pusti pri miru. Moja krivda in v bodoče naloga, da skrbim za te primere.

Testbuild naj bi bil po njegovih besedah na razpolago enkrat danes: TB 3.0 oz. TB 3.1




Mercurial, skladišča, pošte ...

Sistem, ki se v Mozilli uporablja za hranjenje in ažuriranje izvorne kode, se imenuje Mercurial.

Kdor je doslej že delal z SVN ipd., se na Mercurial (ali kratko hg) lahko hitro navadi. Glede na to, da gre za prostokodne projekte, je pri Mozilli vse na dlani:

mozilla.org releases repository ... seveda za branje, za pisanje obstaja procedura, kako se do teh pravic pride.

Še en naslov: poštna stran za "lokalizatorje".

Lep pozdrav in upam da lepa novica za vse.



PS: Ali je mogoče komentirati?

Zaenkrat nobenih komentarjev ali prijavljencev na blog, pri čemer ne vem, ali je krivda nezanimiva ali nesporna vsebina, ali pa nastavitev bloga. Lahko kdo poizkusi s komentarjem?