Key Chinese updated, adding new Pinyin features

The program Key, which offers probably the best support for Hanyu Pinyin of any software and thus deserves praise for this alone, has just come out with an update with even more Pinyin features: Key 5.2 (build: August 21, 2011 — earlier builds of 5.2 do not offer all the latest features).

Those of you who already have the program should get the update, as it’s free. But note that if you update from the site, the installer will ask you to uninstall your current version prior to putting in the update, so make sure you have your validation code handy or you’ll end up with no version at all.

(If you don’t already have Key, I recommend that you try it out. A 30-day free trial version can be downloaded from the site.)

Anyway, here’s some of what the latest version offers:

  • Hanzi-with-Pinyin horizontal layout gets preserved when copied into MS Word documents (RT setting), as well as in .html and .pdf files created from such documents.
  • Pinyin Proofing (PP) assistance: with pinyin text displayed, pressing the PP button on the toolbar will colour the background of ambiguous pinyin passages blue; right-clicking on such a blue-background pinyin passage will display the available options.
  • Copy Special: a highlighted Chinese character passage can be copied & pasted automatically in various permutations.
  • Improved number-measureword system: it now works with Chinese-character, pinyin and Arabic numerals.
  • Showing different tones through coloured characters (Language menu under Preferences).
  • Chengyu (fixed four character expression) spacing logic: automatic spacing according to the pinyin standard (Language menu under Preferences).
  • Option to show tone sandhi on grey background (Language menu under Preferences).
  • Full support of standard pinyin orthography in capitalization and spacing.
  • Automatic glossary building.

Some programs, such as Popup Chinese’s “Chinese converter,” will take Chinese characters and then produce pinyin-annotated versions, with the Pinyin appearing on mouseover. Key, however, offers something extra: the ability to produce Hanzi-annotated orthographically correct Pinyin texts (i.e,, the reverse of the above). If you have a text in Key in Chinese characters, all you have to do is go to File --> Export to get Key to save your text in HTML format.

Here’s a sample of what this looks like.

B?n bi?ozh?n gu?dìngle yòng? Zh?ngwén p?ny?n f?ng’àn? p?nxi? xiàndài Hàny? de gu?zé? Nèiróng b?okuò f?ncí liánxi? f?? chéngy? p?nxi?f?? wàiláicí p?nxi?f?? rénmíng dìmíng p?nxi?f?? bi?odiào f?? yíháng gu?zé d?ng?

Basically, this is a “digraphia export” feature — terrific!

If you want something like the above, you do not have to convert the Hanzi to orthographically correct Pinyin first; Key will do it for you automatically. (I hope, though, that they’ll fix those double-width punctuation marks one of these days.)

Let’s say, though, that you want a document with properly word-parsed interlinear Hanzi and Pinyin. Key will do this too. To do this, a input a Hanzi text in Key, then highlight the text (CTRL + A) and choose Format --> Hanzi with Pinyin / Kanji-Kana with Romaji.

In the window that pops up, choose Hanzi with Pinyin / Kanji-kana with Romaji / Hangul with Romanization from the Two-Line Mode section and Show all non-Hanzi symbols in Pinyin line from Options. The results will look something like this:

GIF of a screenshot from Key, showing an interlinear text with word-parsed Pinyin above Chinese characters. This is an image of the text after being pasted into Microsoft Word.

This can be extremely useful for those authoring teaching materials.

Furthermore, such interlinear texts can be copied and pasted into Word. For the interlinear-formatted copy-and-paste into Word to work properly, Key must be set to rich text format, so before selecting the text you wish to use click on the button labeled RT. (Note yellow-highlighted area in the image below.)

screenshot identifying the location of the button that needs to be pressed to make the text RTF

iOS app for writing Pinyin with tone marks

Those of you who, unlike me, own an iPhone, an iPad, or an iPod Touch may find the new Pinyin Typist Mac application of use.

Taffy of Tailingua had a look at this for me.

I’ve had a play with the Pinyin application and I’m generally quite positive about it. It’s clean, unfussy, and gets the job done. The automatic positioning looks to be flawless (i.e. typing zhuang1 gives you zhu?ng, not zh?ang)…. Overall though I like it, as it does what it set out to do without any showboating or unnecessary steps (excepting apostrophes).

Although I wish the apostrophe and hyphen were right there on the main screen instead of on a secondary one, the program allows people to do what they need to do: type Pinyin with tone marks.

It sells for US$3.99 US$2.99.

[Headline changed from “Mac app for writing Pinyin with tone marks”]

Simplified Chinese characters being purged from Taiwan government sites

Taiwan’s government Web sites have begun removing versions of their content in simplified Chinese characters at the instruction of President Ma Ying-jeou (M? Y?ngji?).

This isn’t just a matter of, say, writing “??” (Taiwan) instead of “??” (which, yes, the government here is encouraging). This is much bigger. Entire pages, entire Web sites even, written in simplified Chinese characters are being eliminated.

The Tourism Bureau, for example, removed the version of its site in simplified Chinese characters from the Web on Wednesday. This comes at a time that the government’s further lifting of restrictions against individual Chinese tourists is aimed at bringing in more travelers from China.

The Presidential Office’s spokesman quoted Ma as saying “To maintain our role as the pioneer in Chinese culture, all government bodies should use traditional Chinese in official documents and on their Web sites, so that people around the world can learn about the beauty of traditional characters.” (Is that what pioneers do? I’ll try to find the original Mandarin-language quote later if I get a chance.)

It’s one thing to urge businesses not to remove traditional Chinese characters and replace them with simplified Chinese characters (as the government did on Tuesday). It’s quite another to remove alternate versions in another script — one that a very sizable target audience would have an easier time with.

During the administration of President Chen Shui-bian the government began adding versions in simplified Chinese characters of the Mandarin texts of official Web sites. The Office of the President was one such site. Now the simplified version is gone. That’s happening across government sites.

Here, for example, are some screen shots I took.

This was the language/script selection at the National Palace Museum‘s Web site as of Thursday morning. (Click to see an image of the entire front page.)
click to see image of entire front page
“????” (ji?nt? Zh?ngwén) is brighter because I had my mouse over it to highlight that text.

And here the language/script selection at the National Palace Museum’s Web site as of Thursday evening:
click to see image of entire front page
As you can see, the choice of viewing the site in simplified Chinese characters has been removed.

Here at Pinyin.Info I often have material in Hanyu Pinyin. So I’m certainly not unsympathetic to the idea that sometimes the medium really is a major part of the message. But I doubt that President Ma’s tough-love approach in this area will accomplish anything useful for Taiwan or the survival of traditional Chinese characters; indeed, I believe it will be counter-productive.

To be more blunt about this, this seems like a really, really bad idea.

Google Translate and romaji revisited

OK, Google has improved its Pinyin converter some, though it still fails in important areas. So that’s the present situation for Google and Mandarin.

How about for Google and Japanese?

Professor J. Marshall Unger of the Ohio State University’s Department of East Asian Languages and Literatures generously agreed to reexamine Google’s performance in conversions to rōmaji (Japanese written in romanization).

Below is his latest evaluation.

For his initial analysis (in December 2009), see Google Translate and rōmaji.

I ran the test passage through Google Translate again. There’s some improvement, but it’s still pretty mediocre.

Original Google Translate
6日午後4時35分ごろ、東京都千代田区皇居外苑の都道(内堀通り)の二重橋前交差点で、中国からの観光客の40代の男性が乗用車にはねられ、全身を強く打って間もなく死亡した。車は歩道に乗り上げて歩いていた男性(69)もはね、男性は頭を強く打って意識不明の重体。丸の内署は、運転していた東京都港区白金3丁目、会社役員高橋延拓容疑者(24)を自動車運転過失傷害の疑いで現行犯逮捕し、容疑を同致死に切り替えて調べている。 6-Nichi gogo 4-ji 35-fun-goro, Tōkyō-to Chiyoda-ku Kōkyogaien no todō (uchibori-dōri) no Nijūbashi zen kōsaten de, Chūgoku kara no kankō kyaku no 40-dai no dansei ga jōyōsha ni hane rare, zenshin o tsuyoku Utte mamonaku shibō shita. Kuruma wa hodō ni noriagete aruite ita dansei (69) mo hane, dansei wa atama o tsuyoku utte ishiki fumei no jūtai. Marunouchi-sho wa, unten shite ita Tōkyō-to Minato-ku hakkin 3-chōme, kaisha yakuin Takahashi nobe Tsubuse yōgi-sha (24) o jidōsha unten kashitsu shōgai no utagai de genkō-han taiho shi, yōgi o dō chishi ni kirikaete shirabete iru.
 同署によると、死亡した男性は横断歩道を歩いて渡っていたところを直進してきた車にはねられた。車は左に急ハンドルを切り、車道と歩道の境に置かれた仮設のさくをはね上げ、歩道に乗り上げたという。さくは歩道でランニングをしていた男性(34)に当たり、男性は両足に軽いけが。 Dōsho ni yoru to, shibō shita dansei wa ōdan hodō o aruite watatte ita tokoro o chokushin shite kita kuruma ni hane rareta. Kuruma wa hidari ni kyū handoru o kiri, shadō to hodō no sakai ni oka reta kasetsu no saku o haneage, hodō ni noriageta toyuu. Saku wa hodō de ran’ningu o shite ita dansei (34) niatari, dansei wa ryōashi ni karui kega.
 同署は、死亡した男性の身元確認を進めるとともに、当時の交差点の信号の状況を調べている。 Dōsho wa, shibō shita dansei no mimoto kakunin o susumeru totomoni, tōji no kōsaten no shingō no jōkyō o shirabete iru.
 現場周辺は東京観光のスポットの一つだが、最近はジョギングを楽しむ人も増えている。 Genba shūhen wa Tōkyō kankō no supotto no hitotsudaga, saikin wa jogingu o tanoshimu hito mo fuete iru.


  • The use of numerals dodges a plethora of errors, but “6-Nichi” is still wrong for Muika.
  • Lots of correct capitalizations have been added, but “uchibori” was missed and “Utte” capitalized by mistake.
  • Some false spaces or lack of spaces persist: “hane rare”, “oka reta”; “hitotsudaga” and “niatari” were correctly hitotsu da ga and ni atari in the original test.
  • Names still get butchered (“hakkin” for Shirogane, “nobe Tsubuse” for Nobuhiro.
  • The needless apostrophe in “ran’ningu” is still there.
  • Interestingly, “toyuu” is a new error: it should be to iu.
  • There’s evidence of some attempt to use hyphens, but why not in “kankō kyaku” or “Nijūbashi zen”?

So, to update: Google gets kudos for conscientiousness, but I stick by my original comments.

For more by Prof. Unger, see’s recommended readings, which includes selections from The Fifth Generation Fallacy: Why Japan Is Betting Its Future on Artificial Intelligence, Literacy and Script Reform in Occupation Japan: Reading Between the Lines, and Ideogram: Chinese Characters and the Myth of Disembodied Meaning.

Banqiao — the Xinbei ways

Xinbei, formerly known as Taipei County and now officially bearing the atrocious English name of “New Taipei City,” has made available an online map of its territory.

Interestingly, the map is available not just in Mandarin with traditional Chinese characters and English with Hanyu Pinyin (most of the time — but more on that soon) but also in Mandarin with simplified Chinese characters. A Japanese interface is also available.

The interface for all versions opens to a map centered on Xinbei City Hall. What struck me upon seeing this for the first time was that, in just one small section, Banqiao is spelled four different ways:

  • Banqiao (Hanyu Pinyin)
  • Panchiao (bastardized Wade-Giles)
  • Ban-Chiau (MPS2, with an added hyphen)
  • Banciao (Tongyong Pinyin)

Click the map to see an enlargement.
click for larger version

I want to stress that these are not typos. These are the result of an inattention to detail that is all too common here.

The spelling for the city, er, district is also wrong in the interface, with Tongyong used. Since Banqiao is the seat of the Xinbei City Government and has more than half a million inhabitants,*, it’s not exactly so obscure that spelling its name correctly should be much of a challenge. Tongyong and other systems also crop up in some other names outside the interface.

It should be admitted, however, that the Xinbei map’s romanization is still better overall than the error-filled mess issued by GooGle.

*: including me

Wenlin releases major upgrade (4.0)

Wenlin logoOne of my favorite programs, Wenlin (which bills itself as “software for learning Chinese”), has just released a major upgrade for both Mac and Windows versions. This doesn’t happen often; it has been three-and-a-half years since the most recent big change was issued (Wenlin 3.4) and heaven only knows how long since 3.0 came out. So, yes, this release has many substantial improvements.

One of the features nearest and dearest to my heart is that Wenlin 4.0 features greatly improved handling of Pinyin. I was among the field testers for the new version, so I’ve already spent a lot of time examining this feature. Here are a few important aspects of this:

  • Conversions from Chinese characters follow Hanyu Pinyin orthography much more closely than before. This is a major change for the better. (There’s still some room for improvement. But I don’t think we’ll have to wait years for this.)
  • In the past, using Wenlin to convert long texts in Chinese characters into Pinyin could be a real chore, with users having to examine example after example of Chinese characters with multiple pronunciations in order to select the proper pronunciation for that particular context. But now users may, if they so desire, tell Wenlin not to ask users for disambiguation input. Of course, that doesn’t mean that Wenlin will always guess right; but many users will be happy that this trade-off allows them to skip the frustration of, for example, having to tell the program over and over and over that, yes, in this case 說 is pronounced shuō rather than shuì.
  • Relative newcomers to Mandarin may appreciate that for common words tone sandhi is indicated in Wenlin with additional marks (a dot or line below the vowel). This feature can also be turned off, for those who want standard Pinyin.

There are, of course, many improvements beyond the area of Pinyin. Here are a few:

  • One limitation of Wenlin 3.x was that its English dictionary wasn’t very large. But Wenlin 4.0 includes not only the ABC Chinese-English Comprehensive Dictionary but also the excellent new ABC English-Chinese, Chinese-English Dictionary (now finally in stock in the printed version).
  • The flashcards are now set up to handle not just individual characters but polysyllabic words.
  • There’s full Unicode Unihan 6.0 support for more than 75,000 Chinese characters.
  • And for those who think 75,000 just isn’t enough, users can now access Wenlin’s CDL technology. Through this, users can create new, variant, and rare characters; moreover, these can be published and shared with other Wenlin users or CDL-friendly devices.
  • Seal script versions of more than 11,000 characters are provided.
  • Wenlin contains an e-edition of the Shuowen Jiezi (Shuōwén Jiězì / 說文解字 / 说文解字).
  • Coders will be interested to know that Wenlin appears to be headed toward becoming open-source.
  • Both Mandarin and English entries are marked with grade levels, which aids learners by indicating relative frequency of use. The levels for Mandarin words are based on the Hanyu Shuiping Kaoshi (Hànyǔ Shǔipíng Kǎoshì / 汉语水平考试 / 漢語水平考試 / HSK).

The full version (i.e., the CD with the program comes in a box and is likely packaged with a hard copy of the manual) is US$199, or US$179 if you download it from the Wenlin Web store. Upgrades from 3.x cost US$49.

For more information, see the summary of features and outline of what’s new in Wenlin 4.0.

screenshot from Wenlin 4.0 -- click for larger version

sg domain names in Chinese characters lag

Between November, 23, 2009, when Singapore first began registering .sg names in Chinese characters, and June 10, 2010, when registrations of Chinese-character .sg domain names opened to all without any additional fee, only 1,024 such names were registered, or just 0.88 percent of all .sg domain names. This apparently includes not just second-level domains (e.g., ??.sg) but also third-level domains (e.g., ??

The percentage will likely rise in the coming months, as the process has only recently opened to everyone on a first-come, first-served basis. But, still, demand for such names in Singapore has so far been underwhelming.

A bit more information:

Registrations were accepted in phases, with registrations for government organizations starting on Nov. 23, 2009. Beginning in January, SGNIC began accepting domain name registrations from trademark holders.

During the third phase, the general public was allowed to register domain names starting on March 25, but applicants were charged a “priority fee” of S$100 (US$72) for each domain name, with domain names sought by several applicants awarded to the highest bidder.

In all three phases, applicants could apply for a domain name made up of Chinese numbers or a name with just one Chinese character for a fee of S$500 [US$360]….

The fourth and final phase began on June 10, with SGNIC accepting domain name applications on a first-come, first-served basis. The S$100 priority fee is no longer required, but applicants are no longer allowed to register domain names using Chinese numbers or names with just one Chinese character….

When IDA announced the introduction of Chinese-language domain names last year, SGNIC said the effort was partly intended to help Singaporean businesses target the Chinese market.

source: Singapore registers 1,000 Chinese-language domain names, IDG News Service, June 23, 2010