How to find hidden characters in text
Hidden characters can be legitimate language controls, formatting marks, or accidental copy-and-paste residue. The reliable way to investigate them is to inspect the actual code points instead of trusting visual appearance.
Concrete code points
For example, U+200B ZERO WIDTH SPACE has no visible width, while U+200D ZERO WIDTH JOINER can connect emoji or shaping sequences. U+202E RIGHT-TO-LEFT OVERRIDE affects display direction and may be useful in a controlled text format, but it deserves review in identifiers or filenames. Paste this example into the inspector:
invoice\u200b-2026\u202e.txtRevealText reports the actual character, name, code point, one-based code-point position, and zero-based UTF-16 index. The visible text map replaces each detected mark with a bracketed [U+XXXX] token while preserving its surrounding context.
Variation selectors are meaningful
Variation selectors such as U+FE0F VARIATION SELECTOR-16 and supplementary selectors from U+E0100 through U+E01EF request a particular glyph presentation. They can be invisible on their own but meaningful in emoji and ideographic sequences. RevealText flags them for inspection; the conservative cleaner does not silently remove them.
A finding is evidence about text structure; it is not proof of author intent, maliciousness, or AI authorship.
Legitimate uses and risks
Joiners, variation selectors, directional marks, and script-specific controls can be required for languages, typography, emoji sequences, or bidirectional text. Removing them blindly can change meaning or presentation. Risks include broken matching, confusing display order, and accidental identifier differences.
Clean conservatively
Review every change. Preserve joiners and variation selectors when they carry meaning, and compare before/after output. Start with Inspect this text locally, then use Clean with reviewable options.
Reference: Unicode UTR #36: Unicode Security Considerations.