back

by Tomte·11y ago·view on hn ↗
Antiqua, Fractura, Schwabacher, Textura and all those other charming variants of writing European languages don't have seperate code points for all those presentational variations.

Do we white Western men happen to discriminate against ourselves?

3 comments
The difference between traditional and simplified Chinese characters is more than simply different fonts. Part of the difficulty is that some simplified characters map to multiple traditional characters, which means that converting from one to the other may be lossy. There's also the Japanese equivalent of simplified characters (shinjitai), many of which differ from their Chinese counterparts, as well as characters that were invented in Japan and may or may not have Chinese equivalents (kokuji).
Simplified and traditional characters have different unicode codepoints. Japanese and mainland simplifications have different codepoints when they differ, and Japanese-only characters of course have their own codepoints.

The argument is about characters like 冷. It is given a single codepoint, but Chinese typefaces draw the bottommost stroke diagonally, and Japanese typefaces tend to draw it vertically. (When writing it by hand, both versions are acceptable in Japanese also, http://detail.chiebukuro.yahoo.co.jp/qa/question_detail/q105... )

Some.

I know that not every letter has been Han-unified, but I don't know the specifics.

Is it possible, that this problem of yours has actually been taken care of?

If not, is it possible that it's just a mistake instead of a big evil conspiracy?

AFAIK it was a historical attempt to save on encoding space back when Unicode had a 16bit fixed width and could only support up to 65K characters.

Now that Unicode has expanded out, I am confused why anyone still defends this practice.

That's just one of the reasons people had (and you can't really expect everyone involved to have totally congruent reasons).

But how about "g"?

Do you really believe the two common variants (one storey/two stories) should have separate code points?

What about German vs. French vs. Danish vs. etc.? All different "g"?

Why? And if not, what is the core difference?

The Cyrillic R (looks like P) and and Greek Roh (also looks like P) and the Latin P are not unified either. Although I think unifying some Greek and Cyrillic letters would have made sense.
Those are all stylistic differences. You can still read any of those fonts. Sure some designer can go overboard and make a font hard to read, but that is just a designer being overly fancy!

Think instead of the difference between Cyrillic and Latin.

Sure if you squint hard enough they both have a common origin (Greek), but you'd be rather annoyed if you setup your phone, selected "English" and the OS used Cyrillic letters to spell out English words.

Likewise you'd be annoyed if Greek letters were used.

And you'd be even more upset if some products decided to use Cyrillic, some Greek, and some used Latin.

Actually, I can't. I can decipher quite a bit with lots of effort, but I wouldn't call that "reading".

And there are very few people around who can read all those (especially Textura) without problems.

On your other example: when I was in Russia, I found those "unknown" letters difficult, but fun. But the letters they share with Latin? I didn't see a difference.

And to your phone example: of course, I'd be annoyed. But that's exactly the point: you have to set up your local system correctly, to your expectations and standards.

I'd be much more annoyed if some web page showed an English text in some Chinese transliteration, just because the author (writing English!) was Chinese. But that's basically what you propose!

Sorry, I truly think you've run up an argumentative dead end.

> On your other example: when I was in Russia, I found those "unknown" letters difficult, but fun. But the letters they share with Latin? I didn't see a difference.

Again, Han unification does this.

Most characters are the same, great! But some are different. Sucks for those that are different.

> And to your phone example: of course, I'd be annoyed. But that's exactly the point: you have to set up your local system correctly, to your expectations and standards.

The problem here is that for almost everything else, Unicode is a mapping from a code point (or a set of code points) to some distinct and unique representation on screen.

Except there are some code points for which that isn't true.

> Sorry, I truly think you've run up an argumentative dead end.

You are not arguing any point, other than "let's keep this historical attempt at saving on encoding space around, even though we have expanded out the encoding space so that we don't need the savings anymore."

That isn't a strong argument.

My argument is "user's don't like this, it upsets them, we shouldn't do it."

If something we do as engineers angers or upsets our users, we are doing it wrong. Flat out.

As far as I know, those "some are different" characters have not been unified, but got separate code points. And that's where your whole argumentation falls apart, unless you claim that the Unicode consortium mis-classified lots of characters.

But then I'd just retreat, because I don't speak any Asian language and cannot verify the claim myself. I can only defer to the experts, and they say that issue has been taken care of.

As far as space savings are concerned, I replied to that in another comment to you. It's not okay to dismiss Han unification as just some space saving sttempt that's not needed anymore. Space saving was one motivation, but according to the consortium not the primary one.

If you don't know the specifics, why the hell are you even arguing this point?
I don't know what?

If you're just here to randomly insult people, please leave.

Update: ah, you just mishandled the threading.

But still... I asked very specific questions, in order to understand your point, and you're very combative. My point stands.