process(obj) # each obj is a fully-parsed Python value
The file ~/streaming_json_parser.py implements a streaming JSON array parser with these design decisions:
Core approach — A character-level recursive-descent parser reads the input stream in 8 KB chunks while tracking position in a dynamically-growing buffer. The buffer is compacted (trimming consumed prefix) when it exceeds ~128 KB, keeping memory usage bounded to O(largest element) rather than O(entire file).
State machine depth was unnecessary — Instead of an explicit tokenizer + state machine, the parser uses recursive-descent methods (_parse_value, _parse_object, _parse_array, _parse_string, _parse_number) that each consume exactly the characters they need. Nesting depth is implicit in the call stack. The top-level method _iter_array is a generator that yields each complete element before reading the next.
Buffer management strategy:
- _fill(n) reads more data until at least n characters are available past the current position
- _peek() returns the current character without consuming it
- _advance(n) skips n characters forward; when _pos > 131072, the consumed prefix is discarded to prevent unbounded growth
- _skip_ws() advances past ' ', '\t', '\n', '\r'
Malformed input handling:
- Every parser method raises ValueError with a descriptive message on unexpected tokens
- The CLI entry point catches ValueError / StopIteration, prints the error to stderr, and exits with code 1 — but previously-yielded elements are already available to the caller
Key code structure:
class StreamingJSONParser:
def __init__(self, stream) # store stream, buffer, position
def __iter__(self): # returns _iter_array() generator
def _iter_array(self): # yields one element at a time
def _parse_value(self): # dispatches to type-specific parser
def _parse_object(self) -> dict # recursive, returns full dict
def _parse_array(self) -> list # recursive, returns full list (nested)
def _parse_string(self) -> str # handles \uXXXX + surrogate pairs
def _parse_number(self) -> int | float # int if no decimal/exponent
def _parse_literal(self, expected, value) # true / false / null
# Usage example:
parser = StreamingJSONParser(sys.stdin)
for obj in parser:
process(obj) # each obj is a fully-parsed Python value
**Test suite** — 41 tests in `~/test_streaming.py` covering every aspect:
| Category | Tests | What's verified |
|---|---|---|
| **Basic** | 13 | empty array, int, float, string, bool, null, mixed, nested objects/arrays, whitespace |
| **Numbers** | 5 | negative, large integers, scientific notation, negative exponent, special floats |
| **Strings** | 4 | escape sequences (`\n`, `\"`, `\\`, etc.), unicode `\uXXXX`, all escapes, surrogate pairs |
| **Edge cases** | 9 | empty string/object, deep nesting (100 levels), 10K elements, 200K-char strings, large file simulation (100K ints), lazy evaluation |
| **Malformed** | 10 | missing `]`, extra comma, unclosed string, trailing comma in object, invalid literal, leading zero (0123), trailing dot, bare token, non-array top-level, unexpected end, invalid unicode escape |
| **Round-trip** | 1 | fuzz against `json.loads` for 15 value types |
**Streaming verified:**
- `test_lazy_evaluation` — asserts elements are yielded one-by-one via `next(iter(parser))` without consuming the whole input
- `test_large_file_simulation` — 100,000 elements parsed successfully
- `test_many_elements` — 10,000 objects parsed without memory explosion
- `test_very_large_strings` — 200 KB strings handled correctly across buffer-compaction boundaries
**CLI testing:**
```
$ echo '[42, "hello", true, null, {"x": 1}]' | python3 streaming_json_parser.py
42
"hello"
true
null
{"x": 1}
$ echo '[1, 2, broken]' | python3 streaming_json_parser.py
1
2
PARSE ERROR: unexpected character 'b'
Exit code: 1
```
The parser yields elements 1 and 2 before hitting the error — demonstrating graceful degradation.
**Real streaming pipe — 50,000 objects:**
```
$ python3 -c "
from streaming_json_parser import StreamingJSONParser
import io
buf = io.BytesIO(b'[' + b','.join(
b'{\"id\":%d,\"name\":\"item-%d\"}' % (i,i) for i in range(50000)
) + b']')
stream = io.TextIOWrapper(buf)
count = sum(1 for _ in StreamingJSONParser(stream))
print(count)
"
50000
```
---{"model": "deepseek-v4-flash", "result": "completed"}The file ~/streaming_json_parser.py implements a streaming JSON array parser with these design decisions:
Core approach — A character-level recursive-descent parser reads the input stream in 8 KB chunks while tracking position in a dynamically-growing buffer. The buffer is compacted (trimming consumed prefix) when it exceeds ~128 KB, keeping memory usage bounded to O(largest element) rather than O(entire file).
State machine depth was unnecessary — Instead of an explicit tokenizer + state machine, the parser uses recursive-descent methods (_parse_value, _parse_object, _parse_array, _parse_string, _parse_number) that each consume exactly the characters they need. Nesting depth is implicit in the call stack. The top-level method _iter_array is a generator that yields each complete element before reading the next.
Buffer management strategy:
- _fill(n) reads more data until at least n characters are available past the current position
- _peek() returns the current character without consuming it
- _advance(n) skips n characters forward; when _pos > 131072, the consumed prefix is discarded to prevent unbounded growth
- _skip_ws() advances past ' ', '\t', '\n', '\r'
Malformed input handling:
- Every parser method raises ValueError with a descriptive message on unexpected tokens
- The CLI entry point catches ValueError / StopIteration, prints the error to stderr, and exits with code 1 — but previously-yielded elements are already available to the caller
Key code structure:
class StreamingJSONParser:
def __init__(self, stream) # store stream, buffer, position
def __iter__(self): # returns _iter_array() generator
def _iter_array(self): # yields one element at a time
def _parse_value(self): # dispatches to type-specific parser
def _parse_object(self) -> dict # recursive, returns full dict
def _parse_array(self) -> list # recursive, returns full list (nested)
def _parse_string(self) -> str # handles \uXXXX + surrogate pairs
def _parse_number(self) -> int | float # int if no decimal/exponent
def _parse_literal(self, expected, value) # true / false / null
# Usage example:
parser = StreamingJSONParser(sys.stdin)
for obj in parser:
process(obj) # each obj is a fully-parsed Python value
**Test suite** — 41 tests in `~/test_streaming.py` covering every aspect:
| Category | Tests | What's verified |
|---|---|---|
| **Basic** | 13 | empty array, int, float, string, bool, null, mixed, nested objects/arrays, whitespace |
| **Numbers** | 5 | negative, large integers, scientific notation, negative exponent, special floats |
| **Strings** | 4 | escape sequences (`\n`, `\"`, `\\`, etc.), unicode `\uXXXX`, all escapes, surrogate pairs |
| **Edge cases** | 9 | empty string/object, deep nesting (100 levels), 10K elements, 200K-char strings, large file simulation (100K ints), lazy evaluation |
| **Malformed** | 10 | missing `]`, extra comma, unclosed string, trailing comma in object, invalid literal, leading zero (0123), trailing dot, bare token, non-array top-level, unexpected end, invalid unicode escape |
| **Round-trip** | 1 | fuzz against `json.loads` for 15 value types |
**Streaming verified:**
- `test_lazy_evaluation` — asserts elements are yielded one-by-one via `next(iter(parser))` without consuming the whole input
- `test_large_file_simulation` — 100,000 elements parsed successfully
- `test_many_elements` — 10,000 objects parsed without memory explosion
- `test_very_large_strings` — 200 KB strings handled correctly across buffer-compaction boundaries
**CLI testing:**
```
$ echo '[42, "hello", true, null, {"x": 1}]' | python3 streaming_json_parser.py
42
"hello"
true
null
{"x": 1}
$ echo '[1, 2, broken]' | python3 streaming_json_parser.py
1
2
PARSE ERROR: unexpected character 'b'
Exit code: 1
```
The parser yields elements 1 and 2 before hitting the error — demonstrating graceful degradation.
**Real streaming pipe — 50,000 objects:**
```
$ python3 -c "
from streaming_json_parser import StreamingJSONParser
import io
buf = io.BytesIO(b'[' + b','.join(
b'{\"id\":%d,\"name\":\"item-%d\"}' % (i,i) for i in range(50000)
) + b']')
stream = io.TextIOWrapper(buf)
count = sum(1 for _ in StreamingJSONParser(stream))
print(count)
"
50000
```
---{"model": "deepseek-v4-flash", "result": "completed"}