12k
All articles

Cleaning Up Text People Paste Into Your App

Clean pasted text in web forms with Unicode normalization, format-character stripping, whitespace collapsing, and safe handling of ZWJ and NBSP.

OpenReplay Team
OpenReplay Team
Cleaning Up Text People Paste Into Your App

Text pasted into a web form routinely carries invisible characters that nobody typed and no font draws: zero-width spaces, soft hyphens, byte order marks and non-breaking spaces, all of which survive into your database and quietly break exact-match comparison, length validation and search.

If you have ever chased this bug, you know how it goes. Someone pastes a paragraph of their CV out of a word processor into a job application field, the field looks completely normal on screen, and your validator rejects it. Or it saves fine and then nobody can ever find the record again, because the name in the index contains a character the search box has no way to type.

This article shows what actually lands in the field, explains why patching it one code point at a time never converges, and gives you a four-step cleaning function that handles the whole class, including the invisible characters you must not delete.

Key Takeaways

  • The invisible characters that arrive with pasted text belong to the Unicode general category Format, written \p{Cf} in a JavaScript regular expression, so one category match replaces an ever-growing list of individual code points.
  • normalize("NFC") fixes spelling, not invisibility: it makes two encodings of the same accented letter compare equal and leaves a zero-width space exactly where it was.
  • The non-breaking space at U+00A0 is a space character, not a format character, so it survives a \p{Cf} strip untouched and has to be handled by the whitespace-collapsing step.
  • U+200D ZERO WIDTH JOINER does real work: it binds the parts of a multi-person emoji into one glyph, and joiners carry orthographic meaning in Arabic and Indic scripts.
  • \p{...} only means a Unicode property when the regex carries the u or v flag; without one it is an identity escape for the literal letter p.

What Invisible Characters Actually Arrive in the Field?

A pasted string from a word processor or rich-text editor typically mixes typographic substitutions you can see with format characters you cannot. Curly quotes, en and em dashes, and the ellipsis character are visible and mostly harmless. The non-breaking space, the soft hyphen, the zero-width space and the byte order mark are not visible at all, and they are what breaks equality checks.

Dump the string to code points and the argument makes itself:

function inspect(str) {
  return [...str]
    .map((ch) => ch.codePointAt(0))
    .filter((cp) => cp > 0x7f)
    .map((cp) => "U+" + cp.toString(16).toUpperCase().padStart(4, "0"));
}

const pasted = "\uFEFFSenior\u00A0Engineer\u200B, 2019\u20132024";

console.log(inspect(pasted)); // [ 'U+FEFF', 'U+00A0', 'U+200B', 'U+2013' ]
console.log(pasted.length);   // 28, for 26 characters a human would count

Spread the string rather than calling split(""): the string iterator yields whole code points, while split("") cuts astral characters in half at the UTF-16 boundary.

To see what is really in a string you have in hand, paste it into the invisible character cleaner and read the code points back.

This failure class is defined by being invisible. The field renders correctly, so a screenshot of the bug shows nothing wrong and the reporter cannot describe what they did differently. Session replay is one of the few techniques that surfaces it at all, because replays of abandoned forms show the paste and then the retry loop: someone clearing a field and retyping a value that looked identical to the one that was just rejected.

Why Does a Blocklist of Zero Width Characters Keep Growing?

A character class of specific code points is the wrong shape of fix rather than a badly written one, because it can only ever contain the characters someone has already filed a bug about. You delete U+200B after the first report, add U+FEFF when a CSV import breaks, add U+00AD when a hyphenated surname stops matching, and the class keeps growing because it enumerates members of a set instead of naming the set.

The set has a name. The zero-width space (U+200B), zero-width non-joiner (U+200C), zero-width joiner (U+200D), soft hyphen (U+00AD) and byte order mark (U+FEFF) all carry General_Category=Cf, per the Unicode Character Database, and those assignments have been stable across many versions of the standard. The non-breaking space at U+00A0 is not in that set: it is Zs, a space separator, which is why a category strip alone leaves the most common word-processor artifact in place.

Normalize, Strip, Collapse, Trim

The fix is one pass with four ordered steps: normalize the encoding, remove format characters, collapse every whitespace variant to an ordinary space, then trim.

const FORMAT_CHARS = /[\p{Cf}--[\u200C\u200D]]/gv;
const SPACE_RUN = /[\p{Zs}\t\n\r]+/gu;

export function cleanPastedText(input) {
  return input
    .normalize("NFC")               // one canonical spelling per character
    .replace(FORMAT_CHARS, "")      // BOM, zero-width space, soft hyphen, bidi controls
    .replace(SPACE_RUN, " ")        // NBSP, thin spaces, tabs, newlines -> one space
    .trim();
}

Drop \n\r from SPACE_RUN if the field is a textarea where line breaks are content.

Two things earn their place here. \p{...} carries its Unicode meaning only when the regex is in Unicode-aware mode; drop both the u and the v and the engine reads \p as an escaped literal p, so the pattern compiles, runs, and matches nothing you intended. And the collapse step uses an explicit class built on \p{Zs} rather than \s, so it is clear at the call site that U+00A0 is covered.

Ordering matters. normalize("NFC") resolves encoding, not invisibility, so it leaves a zero-width space in place for the strip step to handle. Stripping before collapsing means the byte order mark is deleted rather than converted into a stray space. And trimming last catches the leading space left behind when a removed format character sat next to one.

What Not to Strip

A blind strip of every format character damages real content. U+200D ZERO WIDTH JOINER is the character that binds emoji into a single glyph: in the Unicode emoji standard, what makes a string an emoji ZWJ sequence in the first place is the presence of a joiner, so taking the joiners out leaves several glyphs where there was one.

const family = "👨‍👩‍👧";

console.log([...family].length);                          // 5
console.log([...family.replace(/\p{Cf}/gu, "")].length);  // 3 -> 👨👩👧

Joiners are not decoration outside emoji either. Arabic writers put a non-joiner between two letters to stop them running together in the cursive way they otherwise would, and the Unicode core specification warns that text stripped of these controls either says something else or stops making sense. In Devanagari, a ZWJ after a virama selects the half-form of a consonant instead of the full conjunct. The rule is to strip the invisible characters that carry no meaning in your field and keep the ones doing structural work.

That is what [\p{Cf}--[\u200C\u200D]] expresses: the v flag adds set operators to character classes, and -- is the one that subtracts. On a runtime without unicodeSets, the u-mode equivalent is /(?![\u200C\u200D])\p{Cf}/gu. Do not set both flags on one regex; they are mutually exclusive.

Why Should You Use NFC and Not NFKC?

NFC resolves the two ways Unicode can spell the same character and changes nothing else. NFKC goes further and rewrites compatibility characters: the ff ligature becomes two f’s and a circled Ⓓ becomes a plain D, as the normalize() examples show. That is a content-changing decision, defined by the Unicode normalization annex as compatibility rather than canonical equivalence, and it is worth making deliberately rather than inheriting as a side effect of cleaning a form field.

One related trap: deleting a soft hyphen is a job for your cleaner, not for normalize("NFKC"). The Unicode operation that maps U+00AD to an empty string is NFKC_Casefold, a different transformation from the one String.prototype.normalize("NFKC") performs.

Run the Cleaner on the Server Too

Cleaning in the browser is a courtesy to the person typing; the normalization your database and your search index depend on has to run on the server, because a request can be sent without ever loading your page. Anything that arrives through a public API, a CSV import, a webhook or a mobile client bypasses the input handler entirely, and a single unclean row is enough to make a unique constraint or an exact-match lookup behave inconsistently. The same function works in Node and in the browser, so run it at the boundary where data enters persistence, and let the client-side call be the fast feedback rather than the guarantee.

Stop adding code points to a character class and start naming the category. Four lines, applied at every entry point, remove the invisible characters that carry no meaning in your field while leaving the ones that hold real text together. Write a fixture that mixes a combining accent, a non-breaking space, a soft hyphen, a byte order mark, a zero-width space and a ZWJ emoji, then assert that your cleaner composes the first, collapses the second to a plain space, deletes the next three and returns the emoji unchanged.

FAQs

Does trim remove a non-breaking space or a zero-width space?

trim removes whitespace and line terminators from both ends only, and the ECMAScript whitespace list includes the non-breaking space U+00A0 and the byte order mark U+FEFF, so both vanish at the edges. It never removes the zero-width space U+200B, which is a format character rather than whitespace, and it never touches any of these characters in the middle of a string.

Does a format character strip also remove emoji variation selectors like U+FE0F?

No. The variation selectors U+FE00 to U+FE0F carry General_Category Mn, nonspacing mark, not Cf, so a format category strip leaves them in place and emoji keep their intended presentation. Only the joiners need a keep-list. Widening a cleaner to strip marks as well would delete variation selectors and every combining accent along with them.

Does stripping format characters also remove right-to-left override characters?

Yes. The bidirectional controls, including U+200E, U+200F, the embeddings and overrides from U+202A to U+202E, and the isolates from U+2066 to U+2069, all carry General_Category Cf, so one category match removes them alongside the zero-width space. That matters for display names and filenames, where a right-to-left override reverses rendered text and can disguise an extension.

Why does a pasted value fail maxlength when it looks short enough?

maxlength counts UTF-16 code units, not visible characters, so every invisible passenger spends budget: a byte order mark or a zero-width space costs one unit, and an emoji built from a surrogate pair costs two. Clean the value before you validate its length, and count with the spread form if the limit is meant to match what people see.

Open-source session replay

Complete picture for complete understanding

Capture every clue your frontend is leaving so you can instantly get to the root cause of any issue with OpenReplay — the open-source session replay tool for developers. Self-host it in minutes, and have complete control over your customer data.

Star on GitHub12k

We use cookies to improve your experience. By using our site, you accept cookies.