unknown
Environment was empty, so I built the python-observability-latency-reporting package from scratch at ~/python-observability-latency-reporting/ (self-contained, stdlib + asyncio only — no network/h3 dependency; the shim wraps a deterministic _MockH3 grid backend that mirrors the real h3 API surface, swappable in production).
1. Percentile helper (observability/latency.py) — pure nearest-rank p50/p90/p95/p99/min/max/mean from per-test duration_ms, with custom-percentile and empty-input support:
def latency_percentiles(duration_ms, percentiles=None):
values = sorted(float(d) for d in duration_ms)
result = {f"p{int(p)}": percentile(values, p) for p in (percentiles or DEFAULT_PERCENTILES)}
if values:
result.update(min=values[0], max=values[-1], mean=statistics.mean(values))
else:
result.update(min=0.0, max=0.0, mean=0.0) # stable JSON for empty runs
return result
def percentile(sorted_values, p): # nearest-rank; guards p<=0 / p>=100 / empty
...
2. H3 shim per-hop latency logging (observability/h3shim.py) — H3Shim wraps the backend, exposes line(origin, dest) hop chains and per-hop simulated process_rpc/result_rpc with independent latencies, plus a shared additive log buffer and structured HopRecords:
class H3Shim:
def __init__(self, backend=None, *, process_delay_ms=1.0, result_delay_ms=1.0, log=None):
self.backend = backend or _MockH3()
self.log = log if log is not None else [] # additive buffer, never cleared
self.records = []
async def process_rpc(self, cell): ... # sleeps process_delay_ms
async def result_rpc(self, cell): ... # sleeps result_delay_ms
3. Additive logging in the async loop (observability/runner.py) — times process/result RPCs separately per hop and logs decision_type per hop, appending to (never clearing) the shared buffer:
async def run_async_loop(shim, tests, *, log=None, logger_obj=None,
process_rpc=None, result_rpc=None):
if log is None:
log = shim.log
for test_index, test in enumerate(tests):
origin, destination = _as_pair(test) # tuple or .origin/.destination
hops = shim.line(origin, destination)
test_start = time.perf_counter()
for hop_index, cell in enumerate(hops):
hop_start = time.perf_counter()
await proc(cell); process_ms = (time.perf_counter() - hop_start) * 1000.0
r_start = time.perf_counter()
await res(cell); result_ms = (time.perf_counter() - r_start) * 1000.0
decision_type = "terminal" if hop_index == len(hops) - 1 else "forward"
line = (f"test={test_index} hop={hop_index} cell={cell} "
f"decision_type={decision_type} process_ms={process_ms:.3f} "
f"result_ms={result_ms:.3f} total_ms={process_ms + result_ms:.3f}")
shim.record_hop(HopRecord(...), line=line) # append, never clear
if log is not shim.log: log.append(line)
per_test_duration_ms.append((time.perf_counter() - test_start) * 1000.0)
return log, per_test_duration_ms
A bounded-concurrency run_tests_concurrently variant is included.
4. Additive top-level latency key in JSON report (observability/report.py) — merges into a new dict so existing keys and the caller's dict are untouched:
def add_latency_to_report(report, duration_ms, *, percentiles=None):
merged = dict(report) # input never mutated
merged["latency"] = latency_percentiles(duration_ms, percentiles=percentiles)
return merged
# + build_report(tests_run, passed, failed, duration_ms, extra=...) and report_to_json()
Verified with `python3 -m pytest tests/ -q` → **12/12 pass** (5 latency + 3 h3shim + 3 runner + 1 report), re-run 5× to confirm timing tests are stable, plus a standalone end-to-end run.
- Percentiles: `latency_percentiles(range(1, 11))` → `p50=5.0, p90=9.0, p95=10.0, p99=10.0, min=1.0, max=10.0, mean=5.5`; custom set `[1, 25, 50]` → `p1=1.0, p25=3.0`; empty input → all zeros; single value → all equal.
- Per-hop logging: one log line per hop; `decision_type=forward` on intermediate hops, `terminal` on the last; `process_ms`/`result_ms` independently timed (8 ms vs 4 ms delays verified separately); structured `records` mirror the log.
- Additivity: two sequential `run_async_loop` calls grow `shim.log` 87 → 94 lines (exactly +7 for the second `(sf→la)` test); caller-provided buffer with a pre-existing line is preserved and appended; input report dict is never mutated (`"latency" not in base`).
- Edge cases: nearest-rank guards (`p=0`→min, `p=100`→max), unordered durations, single-hop `(sf, sf)`, zero-hop safety, lat/lng→cell roundtrip through the shim passthrough, custom RPC overrides, concurrent runner (12 lines for 2 tests), JSON round-trip keeps the additive key.
- Sanity script output: `{"tests_run": 3, "passed": 3, "failed": 0, "suite": "smoke", "latency": {"p50": ..., "min": 6.55, ...}}` — existing keys intact, `latency` added at top level.{"model": "deepseek-v4-flash", "result": "completed"}Environment was empty, so I built the python-observability-latency-reporting package from scratch at ~/python-observability-latency-reporting/ (self-contained, stdlib + asyncio only — no network/h3 dependency; the shim wraps a deterministic _MockH3 grid backend that mirrors the real h3 API surface, swappable in production).
1. Percentile helper (observability/latency.py) — pure nearest-rank p50/p90/p95/p99/min/max/mean from per-test duration_ms, with custom-percentile and empty-input support:
def latency_percentiles(duration_ms, percentiles=None):
values = sorted(float(d) for d in duration_ms)
result = {f"p{int(p)}": percentile(values, p) for p in (percentiles or DEFAULT_PERCENTILES)}
if values:
result.update(min=values[0], max=values[-1], mean=statistics.mean(values))
else:
result.update(min=0.0, max=0.0, mean=0.0) # stable JSON for empty runs
return result
def percentile(sorted_values, p): # nearest-rank; guards p<=0 / p>=100 / empty
...
2. H3 shim per-hop latency logging (observability/h3shim.py) — H3Shim wraps the backend, exposes line(origin, dest) hop chains and per-hop simulated process_rpc/result_rpc with independent latencies, plus a shared additive log buffer and structured HopRecords:
class H3Shim:
def __init__(self, backend=None, *, process_delay_ms=1.0, result_delay_ms=1.0, log=None):
self.backend = backend or _MockH3()
self.log = log if log is not None else [] # additive buffer, never cleared
self.records = []
async def process_rpc(self, cell): ... # sleeps process_delay_ms
async def result_rpc(self, cell): ... # sleeps result_delay_ms
3. Additive logging in the async loop (observability/runner.py) — times process/result RPCs separately per hop and logs decision_type per hop, appending to (never clearing) the shared buffer:
async def run_async_loop(shim, tests, *, log=None, logger_obj=None,
process_rpc=None, result_rpc=None):
if log is None:
log = shim.log
for test_index, test in enumerate(tests):
origin, destination = _as_pair(test) # tuple or .origin/.destination
hops = shim.line(origin, destination)
test_start = time.perf_counter()
for hop_index, cell in enumerate(hops):
hop_start = time.perf_counter()
await proc(cell); process_ms = (time.perf_counter() - hop_start) * 1000.0
r_start = time.perf_counter()
await res(cell); result_ms = (time.perf_counter() - r_start) * 1000.0
decision_type = "terminal" if hop_index == len(hops) - 1 else "forward"
line = (f"test={test_index} hop={hop_index} cell={cell} "
f"decision_type={decision_type} process_ms={process_ms:.3f} "
f"result_ms={result_ms:.3f} total_ms={process_ms + result_ms:.3f}")
shim.record_hop(HopRecord(...), line=line) # append, never clear
if log is not shim.log: log.append(line)
per_test_duration_ms.append((time.perf_counter() - test_start) * 1000.0)
return log, per_test_duration_ms
A bounded-concurrency run_tests_concurrently variant is included.
4. Additive top-level latency key in JSON report (observability/report.py) — merges into a new dict so existing keys and the caller's dict are untouched:
def add_latency_to_report(report, duration_ms, *, percentiles=None):
merged = dict(report) # input never mutated
merged["latency"] = latency_percentiles(duration_ms, percentiles=percentiles)
return merged
# + build_report(tests_run, passed, failed, duration_ms, extra=...) and report_to_json()
Verified with `python3 -m pytest tests/ -q` → **12/12 pass** (5 latency + 3 h3shim + 3 runner + 1 report), re-run 5× to confirm timing tests are stable, plus a standalone end-to-end run.
- Percentiles: `latency_percentiles(range(1, 11))` → `p50=5.0, p90=9.0, p95=10.0, p99=10.0, min=1.0, max=10.0, mean=5.5`; custom set `[1, 25, 50]` → `p1=1.0, p25=3.0`; empty input → all zeros; single value → all equal.
- Per-hop logging: one log line per hop; `decision_type=forward` on intermediate hops, `terminal` on the last; `process_ms`/`result_ms` independently timed (8 ms vs 4 ms delays verified separately); structured `records` mirror the log.
- Additivity: two sequential `run_async_loop` calls grow `shim.log` 87 → 94 lines (exactly +7 for the second `(sf→la)` test); caller-provided buffer with a pre-existing line is preserved and appended; input report dict is never mutated (`"latency" not in base`).
- Edge cases: nearest-rank guards (`p=0`→min, `p=100`→max), unordered durations, single-hop `(sf, sf)`, zero-hop safety, lat/lng→cell roundtrip through the shim passthrough, custom RPC overrides, concurrent runner (12 lines for 2 tests), JSON round-trip keeps the additive key.
- Sanity script output: `{"tests_run": 3, "passed": 3, "failed": 0, "suite": "smoke", "latency": {"p50": ..., "min": 6.55, ...}}` — existing keys intact, `latency` added at top level.{"model": "deepseek-v4-flash", "result": "completed"}