◐ Off-By-One · answer catalog

sec-xss-sanitizer

2 answer(s)pythonpython3pythonpython3

unknown

📦 Source in repository (JSON)

Answer 1

The HTML sanitizer is in ~/sanitizer.py. It uses Python's built-in html.parser.HTMLParser — a proper tokenising parser, not regex — to safely parse even malformed/nested HTML and rebuild only whitelisted content.

Architecture:

  1. Tag whitelist (SAFE_TAGS) — only b, i, p, a, ul, ol, li, div, span, h1–h6, img, br, table, etc. are allowed.
  2. Tag blocklist (BLOCKED_TAGS) — script, style, iframe, object, embed, canvas, svg, meta, form, input, button, textarea, etc. are stripped (content preserved as text).
  3. Attribute whitelist (SAFE_ATTRS) — only href, class, id, title, alt, src, rel, target, aria-*, etc. are allowed.
  4. Attribute filter (_attr_is_safe):
  5. Blocks all on* event handlers (onclick, onerror, onmouseover, onload, etc.)
  6. Blocks dangerous exact matches (formaction, action, formmethod, formenctype)
  7. Blocks dangerous URI schemes in href/src/etc.: javascript:, data:, vbscript:, livescript:, file:, blob:, and any unknown scheme
  8. HTMLParser overrides handle start tags, end tags, self-closing tags, comments, declarations, processing instructions, and CDATA — each appropriately stripped or escaped.
  9. Stack-based tracking (_suppress_end) ensures properly balanced output even with malformed nesting.

Key security properties:


Evidence & signatures

All **64 tests pass**. Test categories:

| Category | Tests | Description |
|---|---|---|
| Basic safe tags | 7 | `b`, `i`, `p`, `a`, `br`, `h1`, self-closing tags |
| Script tag prevention | 5 | Direct, nested, with attributes, removed entirely |
| Event handler prevention | 5 | `onclick`, `onmouseover`, `onerror`, `onload`, mixed case |
| `javascript:` URI prevention | 4 | Lowercase, mixed case, newline prefix, tab prefix |
| `data:` URI prevention | 1 | Data URIs in href |
| Blocked tags | 6 | `iframe`, `object`, `embed`, `style`, `canvas`, `svg` |
| Unknown tags | 2 | Custom tags stripped, content preserved |
| Malformed/nested HTML | 3 | Unclosed tags, deep nesting, script in nested tags |
| Attribute injection | 2 | CSS-based XSS blocked, event handlers stripped |
| HTML entities | 2 | Named and numeric character references preserved |
| Comment stripping | 2 | Regular and IE conditional comments |
| Edge cases | 8 | Empty/null, plain text, angle brackets, null bytes, double brackets, unicode escapes |
| XSS vectors | 13 | Polyglot img/onerror, javascript: in src, unicode XSS, nested anchors, uppercase event handlers, vbscript, livescript, unknown schemes |
| Attribute preservation | 3 | `class`, `id`, `alt`, `rel`, `target` preserved |
| formaction prevention | 1 | `formaction` with javascript blocked |

```json
{
  "problem_class": "sec-xss-sanitizer",
  "model": "claude-3.5-sonnet",
  "result": "passed",
  "tests": 64
}
```

Answer 2

The HTML sanitizer is in ~/sanitizer.py. It uses Python's built-in html.parser.HTMLParser — a proper tokenising parser, not regex — to safely parse even malformed/nested HTML and rebuild only whitelisted content.

Architecture:

  1. Tag whitelist (SAFE_TAGS) — only b, i, p, a, ul, ol, li, div, span, h1–h6, img, br, table, etc. are allowed.
  2. Tag blocklist (BLOCKED_TAGS) — script, style, iframe, object, embed, canvas, svg, meta, form, input, button, textarea, etc. are stripped (content preserved as text).
  3. Attribute whitelist (SAFE_ATTRS) — only href, class, id, title, alt, src, rel, target, aria-*, etc. are allowed.
  4. Attribute filter (_attr_is_safe):
  5. Blocks all on* event handlers (onclick, onerror, onmouseover, onload, etc.)
  6. Blocks dangerous exact matches (formaction, action, formmethod, formenctype)
  7. Blocks dangerous URI schemes in href/src/etc.: javascript:, data:, vbscript:, livescript:, file:, blob:, and any unknown scheme
  8. HTMLParser overrides handle start tags, end tags, self-closing tags, comments, declarations, processing instructions, and CDATA — each appropriately stripped or escaped.
  9. Stack-based tracking (_suppress_end) ensures properly balanced output even with malformed nesting.

Key security properties:


Evidence & signatures

All **64 tests pass**. Test categories:

| Category | Tests | Description |
|---|---|---|
| Basic safe tags | 7 | `b`, `i`, `p`, `a`, `br`, `h1`, self-closing tags |
| Script tag prevention | 5 | Direct, nested, with attributes, removed entirely |
| Event handler prevention | 5 | `onclick`, `onmouseover`, `onerror`, `onload`, mixed case |
| `javascript:` URI prevention | 4 | Lowercase, mixed case, newline prefix, tab prefix |
| `data:` URI prevention | 1 | Data URIs in href |
| Blocked tags | 6 | `iframe`, `object`, `embed`, `style`, `canvas`, `svg` |
| Unknown tags | 2 | Custom tags stripped, content preserved |
| Malformed/nested HTML | 3 | Unclosed tags, deep nesting, script in nested tags |
| Attribute injection | 2 | CSS-based XSS blocked, event handlers stripped |
| HTML entities | 2 | Named and numeric character references preserved |
| Comment stripping | 2 | Regular and IE conditional comments |
| Edge cases | 8 | Empty/null, plain text, angle brackets, null bytes, double brackets, unicode escapes |
| XSS vectors | 13 | Polyglot img/onerror, javascript: in src, unicode XSS, nested anchors, uppercase event handlers, vbscript, livescript, unknown schemes |
| Attribute preservation | 3 | `class`, `id`, `alt`, `rel`, `target` preserved |
| formaction prevention | 1 | `formaction` with javascript blocked |

```json
{
  "problem_class": "sec-xss-sanitizer",
  "model": "claude-3.5-sonnet",
  "result": "passed",
  "tests": 64
}
```
Generated from the verified corpus · MIT licensedBack to the catalog