Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> String-matching is really scary in Unicode

Isn't this why NFKC normalization exists?



Unicode has four different normalization forms. Different forms are useful for different intended outcomes.

You, as the programmer, need to understand each of them and why you want to use them.

Brief overview: https://en.wikipedia.org/wiki/Unicode_equivalence#Normalizat... More technical details: http://www.unicode.org/reports/tr15/

I --THINK-- offhand, that NFKC is what you want to use when preparing a password input for processing/comparison (it's lossless, but to a specific point). I also --THINK-- that NFC is the form you want to use when retaining source glyph language distinctions.

From the stackoverflow hits:

https://stackoverflow.com/questions/16173328/what-unicode-no...

I agree with the destructive (pre computation/comparison) operation and that either of the NFKD or NFKC forms should be used (since they destroy non-printing differences for visually compatible characters; a more user friendly approach).

The 'C' forms are always more condensed (accents are packed in to a single character where possible), and thus of higher entropy per input byte. It is my belief that this form is likely to be less susceptible to attacks.

The 'D' forms seem like good choices for /editors/ where the precise nature of a character might be altered by adding or removing accents. (Most human input boxes; during the input/edit process)


> The 'C' forms are always more condensed (accents are packed in to a single character where possible), and thus of higher entropy per input byte. It is my belief that this form is likely to be less susceptible to attacks.

What sort of attacks are you talking about?

> The 'D' forms seem like good choices for /editors/ where the precise nature of a character might be altered by adding or removing accents. (Most human input boxes; during the input/edit process)

Editors shouldn't care about C vs D. The reason being, once you've typed the grapheme cluster, it's supposed to act in an editor as if it's a single "character" regardless of whether it's made from one codepoint or several. This means that if I type é then arrow keys and the delete key will operate it on it exactly the same whether it's composed or decomposed.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: