Writing & text
Why identical-looking names fail an exact text match
Inspect composed and decomposed accents, build a separate comparison key and keep the original spelling when matching imported names.
Reproduce one failed match with harmless text
A directory migration reports that two copies of Café do not match. One ends with the single character U+00E9; the other ends with U+0065 followed by U+0301. They can look alike while an exact string comparison says they differ. A screenshot alone will not show which sequence the export contains.
Copy a synthetic version of each string into KitForma Text to Character Codes and select hexadecimal output. Record the code points rather than retyping the word, because retyping can remove the difference you are investigating. The tool reports Unicode code points beyond the ASCII range; its name does not mean non-English letters are invalid.
Composed: U+0043 U+0061 U+0066 U+00E9
Decomposed: U+0043 U+0061 U+0066 U+0065 U+0301
Exact comparison before normalization: falseBuild a comparison key without replacing the source
Unicode normalization provides specified forms for handling equivalent sequences. NFC is a common candidate when a system wants composed canonical forms. Choose the form explicitly in the data contract and apply it consistently to both sides. Do not select a normalization rule merely because it makes one troublesome row pass.
In a JavaScript test, compare the two synthetic strings after normalize("NFC"). The example should match. Store the raw imported value separately from the comparison key and record the transformation version. If the receiving system already defines its own comparison policy, follow that policy instead of adding an undocumented extra pass.
const a = "Caf\u00e9";
const b = "Cafe\u0301";
a === b; // false
a.normalize("NFC") === b.normalize("NFC"); // trueUnicode Standard Annex 15: normalization forms ↗MDN: String.prototype.normalize ↗
Keep normalization separate from other edits
NFC does not turn every accent into its unaccented letter, make every script look alike or choose a language-specific case policy. Trimming spaces, removing punctuation and changing case are additional transformations. Each can merge values that someone intended to keep distinct.
For Turkish data, write down the case rule deliberately: I and İ do not have the same lowercase behavior under every locale policy. Keep display names intact. If you also need a search-friendly key, give it a different field name and document the extra transformations rather than calling it the original name.
Review collisions before joining records
A comparison key is a search aid. Two people can share the same name even when every code point matches. Match records using an appropriate stable identifier and review duplicate keys before joining or deleting anything. Do not infer identity from a normalized name alone.
Create a small report showing source ID, original name, comparison key and collision count. Include a matching accent pair, two genuinely different names and a repeated name belonging to two different synthetic IDs. Hash Generator can illustrate that exact text bytes differ, but a hash does not replace a matching policy or identify a person.
- Keep the original field and a stable source ID.
- Record the normalization and case policies separately.
- Generate keys using the same policy on both inputs.
- Review every key linked to more than one record.
- Join only when the required identifier and review checks pass.
Q: Does Text to Character Codes normalize my text?
No. Use it to inspect code points. This guide’s normalization example runs through the JavaScript normalize method in a separate test. Do not assume that copying a string between tools normalizes it.
Q: Should I remove accents to fix a failed join?
Not as an automatic repair. Removing marks can erase meaningful spelling differences. First determine whether the failure is canonical composition, case, whitespace or an actual different name; retain the original and use a reviewed matching rule.
KITFORMA
Put it into practice
Reading guides is free and needs no account. Linked tools explain any account or Pro requirements.
Further reading
Sources last reviewed:
KitForma prepared this guide with AI assistance. Examples illustrate a workflow; they are not measurements of real sales, search volume or success. Check the linked sources and the result with your own file.
How to use this guide
Published and maintained by KitForma. The linked references explain the relevant formats and definitions. Examples use sample inputs; they do not establish a speed, quality or compatibility guarantee for your files. Check the tool’s stated limits and inspect your downloaded result.