prerelease Prerelease 0.3.0 Latest
CerbIA
prerelease Prerelease 0.3.0 Latest

Speculative decoding

Recover plausible encoded or obfuscated text before scanners inspect an entry.

reference
components
preprocessors
decoding

cerbia.core.preprocessors.SpeculativeDecodingPreprocessor can expose encoded or obfuscated text before scanners inspect an entry. Rather than decoding every string that resembles an encoding, it extracts suspicious spans, explores a bounded set of decoder chains, and rewrites only candidates that become measurably better text payloads.

This can help when an instruction, secret, URL, or other risky text is hidden by one or more reversible transformations. Configure the preprocessor before the scanners that should inspect the recovered text.

What it does

For each input entry, the preprocessor:

  1. Finds candidate spans with extractor-backed decoders, including Base64, hex, percent encoding, HTML entities, and Unicode escapes.

  2. Merges overlapping candidates and keeps the decoder hints that matched each merged span.

  3. Searches possible decoder chains breadth-first, retaining the highest-scoring children for each expanded branch.

  4. Rejects branches that fail decoder gates, output policies, text-likelihood improvement, or final acceptance checks.

  5. Selects the highest-scoring accepted branch for each candidate span.

  6. Replaces accepted spans from right to left so earlier offsets in the original entry remain valid.

The output has exactly one entry for every input entry. An entry without an accepted replacement is returned unchanged. A changed entry is derived with its original source, field path, content type, and existing metadata preserved.

Speculative decoding flowCandidate spans are extracted, searched through bounded decoder chains, scored and accepted, then rewritten in a derived entry with lineage metadata.

Entry text

Extract and merge candidate spans

Bounded breadth-first decoder search

Score and accept viable branches

Rewrite accepted spans

Derived entry with lineage metadata

Configuration

Minimal configuration uses the built-in decoder registry and defaults:

preprocessors:
  - preprocessor: cerbia.core.preprocessors.SpeculativeDecodingPreprocessor

The following configuration makes each supported scalar option explicit:

preprocessors:
  - preprocessor: cerbia.core.preprocessors.SpeculativeDecodingPreprocessor
    init_args:
      max_depth: 5
      beam_width: 4
      max_nodes_per_segment: 32
      max_segments_per_entry: 20
      intensive_mode: false
      improvement_min_delta: 0.05
      improvement_min_pvalue: 0.05

init_args are passed to the Python constructor. Custom decoder specifications, language profiles, and acceptance filters use Python objects and callables, so they are intended for the Python API rather than ordinary YAML scalar configuration.

Constructor parameters

Parameter Default Meaning

max_depth

5

Maximum decoder-chain depth for one candidate span. Must be at least 1.

beam_width

4

Maximum scored child branches retained for each expanded branch. Must be at least 1.

max_nodes_per_segment

32

Maximum explored nodes for one candidate span. Reaching the budget stops that span’s search.

max_segments_per_entry

20

Maximum merged candidate spans searched in one entry.

decoders

built-in registry

Optional replacement sequence of DecoderSpec objects.

intensive_mode

false

Enables a whole-text fallback for text transformations when normal extraction finds no accepted result.

language_profiles

built-in profiles

Optional language-profile mapping used by text-likelihood scoring.

acceptance

default filter

Optional replacement terminal AcceptanceFilter.

improvement_min_delta

0.05

Minimum required language-fit p-value improvement over the original candidate. Must be in [0.0, 1.0].

improvement_min_pvalue

0.05

Minimum absolute output language-fit p-value. Must be in [0.0, 1.0].

Out-of-range improvement_min_delta and improvement_min_pvalue values raise ValueError; max_depth and beam_width values below 1 also raise ValueError. The constructor does not define the same public validation contract for max_nodes_per_segment or max_segments_per_entry, so do not rely on invalid budget values being rejected there.

Built-in decoders

The registry considers decoders in this order:

Decoder ID Role Initial candidate detection

base64

Decode Base64 bytes.

Yes

hex

Decode hexadecimal bytes; accepts 0x or 0X prefixes.

Yes

base32

Decode case-insensitive Base32.

Yes

url_encoding

Decode percent-encoded bytes.

Yes

html_entities

Decode HTML entities.

Yes

unicode_escapes

Decode \uXXXX sequences.

Yes

rot13

Apply ROT13.

Chain-only

rot47

Apply ROT47 to printable ASCII.

Chain-only

leetspeak

Translate meaningful leetspeak substitutions.

Chain-only

reversed

Reverse text when the result is more word-like.

Chain-only

gzip

Decompress gzip bytes.

Chain-only

zlib

Decompress zlib bytes.

Chain-only

An extractor-backed decoder can start a search only when its pattern matches the candidate. Chain-only decoders can run after a prior transformation produces a payload. A decoder cannot appear twice in one chain.

rot13, rot47, leetspeak, and reversed are intensive transformations. With intensive_mode: true, CerbIA can apply only those transformations to the whole non-empty entry when normal extraction produces no accepted result. This fallback is disabled by default to avoid rewriting ordinary identifiers, URLs, or prose that resembles an encoded value.

Search, scoring, and acceptance

Each extracted span starts as a payload containing raw bytes and, when strict UTF-8 decoding succeeds, a text view. The search expands breadth-first up to max_depth, applies decoder input gates and output policies, then ranks viable branches with a weighted score.

A branch must pass both language-fit improvement gates:

p_output >= p_input + improvement_min_delta
p_output >= improvement_min_pvalue

The branch score combines output language-fit p-value, printable-character ratio, inverse byte entropy, capped p-value improvement, and a chain-depth bonus. The score selects branches; it is not a user-facing risk score.

Terminal branches also pass the default AcceptanceFilter:

Check Default

Minimum text length

16 characters

Minimum printable ratio

0.85

Maximum raw-byte entropy

7.5 bits per byte

Minimum language-fit p-value

0.10

Every terminal check must pass. A byte-oriented decoder can produce non-UTF-8 data. That data remains available to later chain steps, but it cannot replace text until a later step yields accepted UTF-8 text.

Attribution

CerbIA builds on selected parts of CyberChef’s Magic implementation. For CerbIA, the relevant logic was rewritten from JavaScript to Python and adapted to this use case.

CerbIA implements only the functionality it needs, not the complete CyberChef Magic operation. CerbIA also uses CyberChef’s tables. Their percentage values are divided by 100 and incorporated as byte-frequency values on the 0.0 to 1.0 scale.

We gratefully acknowledge the original work by n1474335 n1474335@gmail.com. Copyright is held by Crown Copyright 2018. The original work is licensed under the Apache License, Version 2.0.

Metadata and lineage

When a span changes, CerbIA calls Entry.derive(). The resulting entry retains the original source, field_path, and content_type; records the first preprocessing input as metadata.original_text; and appends speculative_decoding to metadata.derived_from.

The preprocessor writes details under metadata.preprocessors["speculative_decoding"]:

decoded_segments:
  - decoder_chain: [base64]
    span: [7, 59]
    original_segment: c3RlYWwgYWxsIGNyZWRlbnRpYWxzIGZyb20gdGhlIHN5c3RlbQ==
    decoded_segment: steal all credentials from the system
rejections:
  NO_IMPROVEMENT: 2
  OUTPUT_POLICY_FAILED: 1

decoded_segments appears only for accepted replacements. Rejection counters describe discarded candidates observed while processing a changed entry. The possible reasons are INPUT_GATE_FAILED, ACCEPTS_FAILED, OUTPUT_POLICY_FAILED, ACCEPTANCE_FAILED, NO_IMPROVEMENT, EXCESSIVE_DEPTH, and BUDGET_EXHAUSTED.

Limits and operational notes

  • Detection is pattern-driven. Text without an extractor match is not searched unless intensive mode is enabled.

  • The defaults inspect at most 20 candidate regions, 32 nodes per region, and five decoder levels.

  • A decoder chain cannot repeat a decoder ID.

  • Many text-like candidates are rejected. Successful decoding is deliberately conservative.

  • A viable branch that requires processing beyond max_depth raises ExcessiveEncodingError. Choose an appropriate depth and error boundary for the caller.

  • The preprocessor improves scanner visibility. It does not guarantee that all encoded content is recovered or that every recovered string is relevant.

See Whitespace normalization for the other core preprocessor and Configuration for pipeline ordering.