KKitForma.

Language

EnglishEnglishTürkçeTurkishDeutschGermanBlog unavailable in this language · open toolsEspañolSpanishBlog unavailable in this language · open toolsFrançaisFrenchBlog unavailable in this language · open toolsPortuguêsPortugueseBlog unavailable in this language · open toolsItalianoItalianBlog unavailable in this language · open toolsNederlandsDutchBlog unavailable in this language · open toolsPolskiPolishBlog unavailable in this language · open toolsРусскийRussianBlog unavailable in this language · open toolsУкраїнськаUkrainianBlog unavailable in this language · open toolsSvenskaSwedishBlog unavailable in this language · open toolsNorskNorwegianBlog unavailable in this language · open toolsDanskDanishBlog unavailable in this language · open toolsSuomiFinnishBlog unavailable in this language · open toolsČeštinaCzechBlog unavailable in this language · open toolsRomânăRomanianBlog unavailable in this language · open toolsΕλληνικάGreekBlog unavailable in this language · open toolsالعربيةArabicBlog unavailable in this language · open toolsעבריתHebrewBlog unavailable in this language · open toolsفارسیPersianBlog unavailable in this language · open toolsاردوUrduBlog unavailable in this language · open toolsहिन्दीHindiBlog unavailable in this language · open toolsবাংলাBengaliBlog unavailable in this language · open toolsதமிழ்TamilBlog unavailable in this language · open toolsతెలుగుTeluguBlog unavailable in this language · open toolsमराठीMarathiBlog unavailable in this language · open toolsગુજરાતીGujaratiBlog unavailable in this language · open tools简体中文Chinese SimplifiedBlog unavailable in this language · open tools繁體中文Chinese TraditionalBlog unavailable in this language · open tools日本語JapaneseBlog unavailable in this language · open tools한국어KoreanBlog unavailable in this language · open toolsTiếng ViệtVietnameseBlog unavailable in this language · open toolsไทยThaiBlog unavailable in this language · open toolsBahasa IndonesiaIndonesianBlog unavailable in this language · open toolsBahasa MelayuMalayBlog unavailable in this language · open toolsFilipinoFilipinoBlog unavailable in this language · open toolsKiswahiliSwahiliBlog unavailable in this language · open toolsAfrikaansAfrikaansBlog unavailable in this language · open toolsMagyarHungarianBlog unavailable in this language · open toolsБългарскиBulgarianBlog unavailable in this language · open toolsHrvatskiCroatianBlog unavailable in this language · open toolsSrpskiSerbianBlog unavailable in this language · open toolsSlovenčinaSlovakBlog unavailable in this language · open toolsSlovenščinaSlovenianBlog unavailable in this language · open toolsLietuviųLithuanianBlog unavailable in this language · open toolsLatviešuLatvianBlog unavailable in this language · open toolsEestiEstonianBlog unavailable in this language · open toolsCatalàCatalanBlog unavailable in this language · open toolsEuskaraBasqueBlog unavailable in this language · open tools

Writing & text

Why identical-looking names fail an exact text match

Inspect composed and decomposed accents, build a separate comparison key and keep the original spelling when matching imported names.

Reproduce one failed match with harmless text

A directory migration reports that two copies of Café do not match. One ends with the single character U+00E9; the other ends with U+0065 followed by U+0301. They can look alike while an exact string comparison says they differ. A screenshot alone will not show which sequence the export contains.

Copy a synthetic version of each string into KitForma Text to Character Codes and select hexadecimal output. Record the code points rather than retyping the word, because retyping can remove the difference you are investigating. The tool reports Unicode code points beyond the ASCII range; its name does not mean non-English letters are invalid.

Composed:   U+0043 U+0061 U+0066 U+00E9
Decomposed: U+0043 U+0061 U+0066 U+0065 U+0301
Exact comparison before normalization: false

Build a comparison key without replacing the source

Unicode normalization provides specified forms for handling equivalent sequences. NFC is a common candidate when a system wants composed canonical forms. Choose the form explicitly in the data contract and apply it consistently to both sides. Do not select a normalization rule merely because it makes one troublesome row pass.

In a JavaScript test, compare the two synthetic strings after normalize("NFC"). The example should match. Store the raw imported value separately from the comparison key and record the transformation version. If the receiving system already defines its own comparison policy, follow that policy instead of adding an undocumented extra pass.

const a = "Caf\u00e9";
const b = "Cafe\u0301";
a === b; // false
a.normalize("NFC") === b.normalize("NFC"); // true

Unicode Standard Annex 15: normalization forms ↗MDN: String.prototype.normalize ↗

Keep normalization separate from other edits

NFC does not turn every accent into its unaccented letter, make every script look alike or choose a language-specific case policy. Trimming spaces, removing punctuation and changing case are additional transformations. Each can merge values that someone intended to keep distinct.

For Turkish data, write down the case rule deliberately: I and İ do not have the same lowercase behavior under every locale policy. Keep display names intact. If you also need a search-friendly key, give it a different field name and document the extra transformations rather than calling it the original name.

Review collisions before joining records

A comparison key is a search aid. Two people can share the same name even when every code point matches. Match records using an appropriate stable identifier and review duplicate keys before joining or deleting anything. Do not infer identity from a normalized name alone.

Create a small report showing source ID, original name, comparison key and collision count. Include a matching accent pair, two genuinely different names and a repeated name belonging to two different synthetic IDs. Hash Generator can illustrate that exact text bytes differ, but a hash does not replace a matching policy or identify a person.

  1. Keep the original field and a stable source ID.
  2. Record the normalization and case policies separately.
  3. Generate keys using the same policy on both inputs.
  4. Review every key linked to more than one record.
  5. Join only when the required identifier and review checks pass.

Q: Does Text to Character Codes normalize my text?

No. Use it to inspect code points. This guide’s normalization example runs through the JavaScript normalize method in a separate test. Do not assume that copying a string between tools normalizes it.

Q: Should I remove accents to fix a failed join?

Not as an automatic repair. Removing marks can erase meaningful spelling differences. First determine whether the failure is canonical composition, case, whitespace or an actual different name; retain the original and use a reviewed matching rule.

KITFORMA

Put it into practice

Reading guides is free and needs no account. Linked tools explain any account or Pro requirements.

Further reading

Sources last reviewed:

KitForma prepared this guide with AI assistance. Examples illustrate a workflow; they are not measurements of real sales, search volume or success. Check the linked sources and the result with your own file.

How to use this guide

Published and maintained by KitForma. The linked references explain the relevant formats and definitions. Examples use sample inputs; they do not establish a speed, quality or compatibility guarantee for your files. Check the tool’s stated limits and inspect your downloaded result.

Publisher and project details

Get new guides in your inbox

Join our optional email newsletter for KitForma tools, practical guides and product updates.

Give consent on the separate Brevo form, then confirm the link in your email. This is separate from account and support preferences. Unsubscribe using the link in every newsletter.

Open comments in an ad-free view

KitForma

What would you like to do?

Support center