cerbia.core.scanners.InvisibleTextScanner detects Unicode characters that are easy for a person to miss but can change how downstream text is interpreted. It counts zero-width characters, bidi overrides, format controls, tag characters, and private-use or unassigned code points.
| Parameter | Default | Meaning |
|---|---|---|
|
|
Number of detected invisible characters allowed before high-risk handling. |
|
all registered |
Language packs used for suspicious-keyword context after characters are removed. |
|
|
Finding severity. |
|
|
Finding action. |
|
|
Accepted content types; |
No hidden characters yields risk 0.0. Counts at or below threshold produce risk 0.2; larger counts begin at 0.7 + count × 0.03, capped at 1.0. If the cleaned text matches a suspicious keyword pattern, the score receives an additional 0.15 bonus.
scanners:
- scanner: cerbia.core.scanners.InvisibleTextScanner
init_args:
threshold: 3
languages: ["en", "es"]