File Backdoor via Bidirectional Unicode and Zero-Width Characters
An attacker hides malicious instructions inside invisible Unicode characters (bidirectional text controls or zero-width joiners) embedded in text the AI agent is told to write to a file. A human reviewing the file sees normal-looking code, but the actual bytes saved to disk and later executed contain a different, hidden payload.
How the attack works
The attacker crafts input - directly or via prompt injection - that includes invisible Unicode control characters such as bidirectional overrides or zero-width joiners, encoded raw, escaped in JSON/source, or as HTML entities. The AI agent passes this content through in a file-write tool call without stripping the hidden characters. The file that lands on disk looks clean to a code reviewer but actually contains reordered or injected logic once read by a compiler, interpreter, or shell. This is the same technique behind the public 'Trojan Source' vulnerability class, now delivered through an AI agent's write action instead of a direct commit.
Netzilo detection
Netzilo reports this behaviour when it is observed.
- Signature ID
- 5f2c9a13-8d47-4b6e-a0c2-7e91d4f3b8aa
- Severity
- High
Why it matters
An organization can end up with backdoored source code or configuration files that pass visual review, letting an attacker smuggle in malicious logic that executes later - undermining code review as a security control.
What you can do
- →Run automated scans for invisible Unicode control characters (bidi overrides/embeddings/isolates and zero-width joiners) on any file an AI agent writes, before it is committed or executed.
- →Configure editors, diff tools, and CI pipelines to render or flag these characters visibly instead of hiding them.
- →Restrict AI agents from writing files with non-printable or bidirectional Unicode unless the content is expected (e.g. legitimate localization files with RTL scripts).
- →Treat any zero-width character in an otherwise pure-ASCII file, or any explicit LTR/RTL override, as suspicious and requiring manual review.
Known benign look-alikes
- Localization / i18n resource files (.arb, .json, .properties, .po) that legitimately embed bidi embedding or isolate marks around RTL strings. Mitigated - embedding/isolate marks are only scored when the payload contains no Hebrew/Arabic/Syriac script at all; explicit LTR/RTL OVERRIDES (U+202D/U+202E) are still reported because Unicode UAX#9 best practice never requires an override for ordinary RTL prose.
- Emoji ZWJ sequences (family/profession emoji) and Devanagari, Bengali or Persian text using ZWJ/ZWNJ as real joiners in UI strings or docs. Mitigated - zero-width characters are only scored when the payload is otherwise pure ASCII (no other non-ASCII codepoint survives stripping), so the joiner has nothing legitimate to join.
- Files intentionally starting with a UTF-8 BOM. Mitigated - a leading U+FEFF run is stripped before counting.
- Security test fixtures, linter test corpora and CI checks for Trojan Source that deliberately contain these codepoints. These WILL be reported; the action is report, never block, and the kill chain names the exact file so triage is one glance.
- Minified JavaScript or scraped web content pasted into a file with stray U+200B separators. Mitigated - the zero-width-only path additionally requires a code/script/config target path.