def writecsv(self, outputpath, columns=None, filterfn=None):
The solution is a StreamingCSVParser class in streaming_csv_parser.py that processes CSV files entirely with generators — never loading the full dataset into memory.
Key design decisions:
| Feature | Implementation |
|---|---|
| Streaming | csv.reader wrapped in a generator; rows(), filter_rows(), select_columns() all yield one row at a time |
| Header detection | Heuristic: if the first row's values are less numeric than the second row's, or contain common header keywords (id, name, date, etc.), it's a header. Also supports explicit True/False/None (auto). |
| Type inference | Samples up to 10K rows; uses a "widest type wins" strategy (str > float > int). Coerces values on-the-fly via int(v), float(v), or leaves as str. |
| Row filtering | Accepts a predicate callback lambda row: ... |
| Column selection | By name (["name", "age"]) or zero-based index ([0, 2]), or mixed |
| Writing | write_csv() does inline filtering + column selection in a single pass, then writes the output CSV with a header row (only if the source had headers). |
Core parser:
class StreamingCSVParser:
def __init__(self, filepath, has_header=True, infer_types=False,
delimiter=',', quotechar='"'):
self._path = Path(filepath)
self._infer_types = infer_types
self._resolve_header(has_header) # reads first 3 rows to decide
if infer_types:
self._infer_column_types() # samples up to 10K rows
def _open_reader(self):
"""Generator: opens file, yields csv.reader rows, skips header if needed."""
with open(self._path, 'r', newline='', encoding='utf-8-sig') as f:
reader = csv.reader(f, ...)
if self._has_header:
next(reader, None) # skip header row
for row in reader:
yield row
def rows(self):
"""Yield every data row (dict if headers, list otherwise)."""
for raw in self._open_reader():
yield self._process_row(raw)
def filter_rows(self, predicate):
"""Yield rows matching predicate."""
for row in self.rows():
if predicate(row):
yield row
def select_columns(self, columns):
"""Yield rows with only the selected columns."""
# ... resolves names to indices, yields dict/list subsets
def write_csv(self, output_path, columns=None, filter_fn=None):
"""Stream rows -> filter -> select -> write in one pass."""
# ... inline filtering/selection, writes header only when has_header
Usage examples:
# Basic streaming
parser = StreamingCSVParser('huge.csv', has_header=True)
for row in parser.rows():
print(row['name'], row['age'])
# With type inference + filtering + column selection
parser = StreamingCSVParser('huge.csv', infer_types=True)
adults = parser.filter_rows(lambda r: r['age'] >= 18)
for row in adults:
print(f"{row['name']}: {row['salary']:.2f}")
# Write filtered subset to new CSV
parser.write_csv('output.csv',
columns=['name', 'email', 'age'],
filter_fn=lambda r: r['active'] == 'true')
# No-header mode (returns lists, auto-generates col_0, col_1, ...)
parser = StreamingCSVParser('data.csv', has_header=False)
for row in parser.select_columns([0, 3]):
print(row) # [val0, val3]
# Auto-detect header
parser = StreamingCSVParser('unknown.csv', has_header=None)
print(parser.headers) # detected or auto-generated
All tests pass with **47/47** passing in the test suite. Edge cases verified: | Test | Result | |---|---| | Basic header parsing (dict rows) | ✓ | | No-header mode (list rows) | ✓ | | Auto-detect header (heuristic True/False) | ✓ | | Type inference: `int`, `float`, `str` per column | ✓ | | No downgrade: column with `75000.50` + `62000` detected as `float` | ✓ | | Row filtering (`age > 30` → 5 rows) | ✓ | | Column selection by name (`["name","salary"]`) | ✓ | | Column selection by index (`[0, 2]`) with dict rows | ✓ | | Column selection with no-header (list rows) | ✓ | | Write CSV with filter + column select (correct count & header) | ✓ | | Write CSV from no-header source (no header row in output) | ✓ | | `head(3)` returns first N rows | ✓ | | Empty file raises `ValueError` | ✓ | | Single-row file (header only, zero data rows) | ✓ | | Quoted fields with embedded commas (`"hello, world"`) | ✓ | | UTF-8 BOM (`\\xef\\xbb\\xbf`) handled | ✓ | | `rows()` / `filter_rows()` / `select_columns()` are generators | ✓ | | **200,000-row file (6.2 MB) streamed without loading into memory** | ✓ | | Peak memory O(1) relative to file size | ✓ | The streaming behavior was verified with a 200K-row CSV: the parser processes every row through generators without ever materializing the full dataset. `write_csv` works in a single pass — reading, filtering, selecting columns, and writing simultaneously. ---
{"model": "deepseek-v4-flash", "problem_class": "io-csv-parser-streaming", "result": "passed", "tests": 47}