◐ Off-By-One · answer catalog

python-streaming-json-parser

2 answer(s)pythonpython3pythonpython3

process(obj) # each obj is a fully-parsed Python value

📦 Source in repository (JSON)

Answer 1

The file ~/streaming_json_parser.py implements a streaming JSON array parser with these design decisions:

Core approach — A character-level recursive-descent parser reads the input stream in 8 KB chunks while tracking position in a dynamically-growing buffer. The buffer is compacted (trimming consumed prefix) when it exceeds ~128 KB, keeping memory usage bounded to O(largest element) rather than O(entire file).

State machine depth was unnecessary — Instead of an explicit tokenizer + state machine, the parser uses recursive-descent methods (_parse_value, _parse_object, _parse_array, _parse_string, _parse_number) that each consume exactly the characters they need. Nesting depth is implicit in the call stack. The top-level method _iter_array is a generator that yields each complete element before reading the next.

Buffer management strategy: - _fill(n) reads more data until at least n characters are available past the current position - _peek() returns the current character without consuming it - _advance(n) skips n characters forward; when _pos > 131072, the consumed prefix is discarded to prevent unbounded growth - _skip_ws() advances past ' ', '\t', '\n', '\r'

Malformed input handling: - Every parser method raises ValueError with a descriptive message on unexpected tokens - The CLI entry point catches ValueError / StopIteration, prints the error to stderr, and exits with code 1 — but previously-yielded elements are already available to the caller

Key code structure:

class StreamingJSONParser:
    def __init__(self, stream)                   # store stream, buffer, position
    def __iter__(self):                          # returns _iter_array() generator
    def _iter_array(self):                       # yields one element at a time
    def _parse_value(self):                      # dispatches to type-specific parser
    def _parse_object(self) -> dict              # recursive, returns full dict
    def _parse_array(self) -> list               # recursive, returns full list (nested)
    def _parse_string(self) -> str               # handles \uXXXX + surrogate pairs
    def _parse_number(self) -> int | float       # int if no decimal/exponent
    def _parse_literal(self, expected, value)    # true / false / null
# Usage example:
parser = StreamingJSONParser(sys.stdin)
for obj in parser:
    process(obj)          # each obj is a fully-parsed Python value

Evidence & signatures

**Test suite** — 41 tests in `~/test_streaming.py` covering every aspect:

| Category | Tests | What's verified |
|---|---|---|
| **Basic** | 13 | empty array, int, float, string, bool, null, mixed, nested objects/arrays, whitespace |
| **Numbers** | 5 | negative, large integers, scientific notation, negative exponent, special floats |
| **Strings** | 4 | escape sequences (`\n`, `\"`, `\\`, etc.), unicode `\uXXXX`, all escapes, surrogate pairs |
| **Edge cases** | 9 | empty string/object, deep nesting (100 levels), 10K elements, 200K-char strings, large file simulation (100K ints), lazy evaluation |
| **Malformed** | 10 | missing `]`, extra comma, unclosed string, trailing comma in object, invalid literal, leading zero (0123), trailing dot, bare token, non-array top-level, unexpected end, invalid unicode escape |
| **Round-trip** | 1 | fuzz against `json.loads` for 15 value types |

**Streaming verified:**
- `test_lazy_evaluation` — asserts elements are yielded one-by-one via `next(iter(parser))` without consuming the whole input
- `test_large_file_simulation` — 100,000 elements parsed successfully
- `test_many_elements` — 10,000 objects parsed without memory explosion
- `test_very_large_strings` — 200 KB strings handled correctly across buffer-compaction boundaries

**CLI testing:**
```
$ echo '[42, "hello", true, null, {"x": 1}]' | python3 streaming_json_parser.py
42
"hello"
true
null
{"x": 1}

$ echo '[1, 2, broken]' | python3 streaming_json_parser.py
1
2
PARSE ERROR: unexpected character 'b'
Exit code: 1
```

The parser yields elements 1 and 2 before hitting the error — demonstrating graceful degradation.

**Real streaming pipe — 50,000 objects:**
```
$ python3 -c "
  from streaming_json_parser import StreamingJSONParser
  import io
  buf = io.BytesIO(b'[' + b','.join(
    b'{\"id\":%d,\"name\":\"item-%d\"}' % (i,i) for i in range(50000)
  ) + b']')
  stream = io.TextIOWrapper(buf)
  count = sum(1 for _ in StreamingJSONParser(stream))
  print(count)
"
50000
```

---
{"model": "deepseek-v4-flash", "result": "completed"}

Answer 2

The file ~/streaming_json_parser.py implements a streaming JSON array parser with these design decisions:

Core approach — A character-level recursive-descent parser reads the input stream in 8 KB chunks while tracking position in a dynamically-growing buffer. The buffer is compacted (trimming consumed prefix) when it exceeds ~128 KB, keeping memory usage bounded to O(largest element) rather than O(entire file).

State machine depth was unnecessary — Instead of an explicit tokenizer + state machine, the parser uses recursive-descent methods (_parse_value, _parse_object, _parse_array, _parse_string, _parse_number) that each consume exactly the characters they need. Nesting depth is implicit in the call stack. The top-level method _iter_array is a generator that yields each complete element before reading the next.

Buffer management strategy: - _fill(n) reads more data until at least n characters are available past the current position - _peek() returns the current character without consuming it - _advance(n) skips n characters forward; when _pos > 131072, the consumed prefix is discarded to prevent unbounded growth - _skip_ws() advances past ' ', '\t', '\n', '\r'

Malformed input handling: - Every parser method raises ValueError with a descriptive message on unexpected tokens - The CLI entry point catches ValueError / StopIteration, prints the error to stderr, and exits with code 1 — but previously-yielded elements are already available to the caller

Key code structure:

class StreamingJSONParser:
    def __init__(self, stream)                   # store stream, buffer, position
    def __iter__(self):                          # returns _iter_array() generator
    def _iter_array(self):                       # yields one element at a time
    def _parse_value(self):                      # dispatches to type-specific parser
    def _parse_object(self) -> dict              # recursive, returns full dict
    def _parse_array(self) -> list               # recursive, returns full list (nested)
    def _parse_string(self) -> str               # handles \uXXXX + surrogate pairs
    def _parse_number(self) -> int | float       # int if no decimal/exponent
    def _parse_literal(self, expected, value)    # true / false / null
# Usage example:
parser = StreamingJSONParser(sys.stdin)
for obj in parser:
    process(obj)          # each obj is a fully-parsed Python value

Evidence & signatures

**Test suite** — 41 tests in `~/test_streaming.py` covering every aspect:

| Category | Tests | What's verified |
|---|---|---|
| **Basic** | 13 | empty array, int, float, string, bool, null, mixed, nested objects/arrays, whitespace |
| **Numbers** | 5 | negative, large integers, scientific notation, negative exponent, special floats |
| **Strings** | 4 | escape sequences (`\n`, `\"`, `\\`, etc.), unicode `\uXXXX`, all escapes, surrogate pairs |
| **Edge cases** | 9 | empty string/object, deep nesting (100 levels), 10K elements, 200K-char strings, large file simulation (100K ints), lazy evaluation |
| **Malformed** | 10 | missing `]`, extra comma, unclosed string, trailing comma in object, invalid literal, leading zero (0123), trailing dot, bare token, non-array top-level, unexpected end, invalid unicode escape |
| **Round-trip** | 1 | fuzz against `json.loads` for 15 value types |

**Streaming verified:**
- `test_lazy_evaluation` — asserts elements are yielded one-by-one via `next(iter(parser))` without consuming the whole input
- `test_large_file_simulation` — 100,000 elements parsed successfully
- `test_many_elements` — 10,000 objects parsed without memory explosion
- `test_very_large_strings` — 200 KB strings handled correctly across buffer-compaction boundaries

**CLI testing:**
```
$ echo '[42, "hello", true, null, {"x": 1}]' | python3 streaming_json_parser.py
42
"hello"
true
null
{"x": 1}

$ echo '[1, 2, broken]' | python3 streaming_json_parser.py
1
2
PARSE ERROR: unexpected character 'b'
Exit code: 1
```

The parser yields elements 1 and 2 before hitting the error — demonstrating graceful degradation.

**Real streaming pipe — 50,000 objects:**
```
$ python3 -c "
  from streaming_json_parser import StreamingJSONParser
  import io
  buf = io.BytesIO(b'[' + b','.join(
    b'{\"id\":%d,\"name\":\"item-%d\"}' % (i,i) for i in range(50000)
  ) + b']')
  stream = io.TextIOWrapper(buf)
  count = sum(1 for _ in StreamingJSONParser(stream))
  print(count)
"
50000
```

---
{"model": "deepseek-v4-flash", "result": "completed"}
Generated from the verified corpus · MIT licensedBack to the catalog