◐ Off-By-One · answer catalog

python-monitoring-grafana-prometheus

2 answer(s)godockergodocker

python-monitoring-grafana-prometheus

📦 Source in repository (JSON)

Answer 1

Built the complete MONITORING-02 layer at ~/dexdat-core (pre-existing FastAPI stack + purely additive monitoring). Pattern followed: zero new metrics code — Prometheus/Grafana consume the existing prometheus_client endpoint at /v1/metrics.

dexdat-core/
├── app/main.py, app/metrics.py      # pre-existing FastAPI app exposing /v1/metrics
├── docker-compose.yml               # base stack (dexdat-core, postgres, consensus-sidecar)
├── docker-compose.monitoring.yml    # ADDITIVE: prometheus 9091:9090, grafana 3030:3000, healthchecks
├── prometheus/
│   ├── prometheus.yml               # job dexdat-core → dexdat-core:8000/v1/metrics, rule_files glob
│   └── rules/dexdat_alerts.yml      # 3 alerts on live-verified metrics
├── grafana/
│   ├── provisioning/datasources/prometheus.yml   # provisioned, uid=dexdat-prometheus
│   ├── provisioning/dashboards/dashboards.yml    # file provider
│   └── dashboards/dexdat-overview.json           # 6 panels, ALL uid-referenced
└── tests/test_config_validity.py    # 12 config-validity tests

Key artifacts (excerpts):

Prometheus scrape — metrics_path: /v1/metrics, targets dexdat-core:8000, rule_files glob mounted read-only.

3 alert rules (all expressions restricted to metric names verified live, none from the documented-absent set): - DexdatHighErrorRate — sum(rate(http_requests_total{status=~"5.."}[5m])) / clamp_min(sum(rate(http_requests_total[5m])), 1e-6) > 0.05, for: 5m, severity: critical - DexdatQueueBacklog — dexdat_queue_depth > 30, for: 2m, severity: warning - DexdatMemoryPressure — process_resident_memory_bytes/1024/1024 > 512, for: 5m, severity: warning

Datasource provisioning — type: prometheus, uid: dexdat-prometheus, url: http://prometheus:9090, isDefault: true. Every dashboard panel and the templating variable use "datasource": {"type": "prometheus", "uid": "dexdat-prometheus"} — no numeric datasource IDs.

Dashboard — dexdat-overview.json, uid dexdat-core-overview, 6 panels (request rate by status, latency p95, jobs in progress, queue depth, RSS MiB, CPU cores).

Additive compose — docker-compose.monitoring.yml defines only prometheus (ports 9091:9090, wget healthcheck on /-/healthy) and grafana (ports 3030:3000, wget healthcheck on /api/health); it shares zero services with the base file, so docker compose -f docker-compose.yml -f docker-compose.monitoring.yml up layers cleanly.

Metric absence documented (README): dexdat_consensus_height lives in the consensus sidecar container; pg_up/pg_stat_* need an undep­loyed postgres_exporter — neither appears in any panel or rule.

Evidence & signatures

1. **Live metric verification before authoring panels** — ran the app, curled `/v1/metrics`, confirmed 17 metric families (incl. `http_requests_total{method,path,status}`, `http_request_duration_seconds_bucket{le}`, `dexdat_jobs_in_progress`, `dexdat_queue_depth`), then generated traffic to confirm label sets. All panels/rules reference only these.
2. **12/12 config-validity tests pass** (`pytest tests/ -q` → `12 passed`) — YAML/JSON validity, exactly 3 rules, all exprs reference verified metrics only, uid-referenced panels, healthchecks + `9091:9090`/`3030:3000` ports, additive-compose check (no service overlap with base, merged service set exact).
3. **promtool validation** — `promtool check config` and `promtool check rules` both SUCCESS against the shipped files.
4. **Real Prometheus end-to-end** (2.53.3 binary, no docker daemon available in this non-root container): scraped the live app — `up{job="dexdat-core"} = 1`, 74 series, `dexdat_queue_depth`/`sum(rate(http_requests_total[1m]))` returned real values; `/api/v1/rules` showed all 3 rules `state=inactive health=ok`, zero eval errors.
5. **Alert firing proven with `promtool test rules`** — all 3 rules fired with correct labels + annotations. This caught a real bug: the initial error-rate rule lacked `sum()`, so PromQL's label-set join divided the 5xx series against *itself* → alert reported 100% instead of the 5xx share. Fixed with `sum(...)/sum(...)`; unit tests re-verified.
6. **All 6 dashboard queries evaluated against live Prometheus** — each panel expr returned data (e.g. p95 = 0.00475s, RSS = 52.98 MiB, queue = 16).
7. **Edge cases tested** — rate-window flatness (flat counters → 0 rate, so no false fires), label-join semantics, promtool annotation matching, host port 9090 conflict (relocated live test to 19090), PEP 668 venv handling.
{"model": "deepseek-v4-flash", "problem_class": "python-monitoring-grafana-prometheus", "result": "passed", "tests": 12}

Answer 2

Built the complete MONITORING-02 layer at ~/dexdat-core (pre-existing FastAPI stack + purely additive monitoring). Pattern followed: zero new metrics code — Prometheus/Grafana consume the existing prometheus_client endpoint at /v1/metrics.

dexdat-core/
├── app/main.py, app/metrics.py      # pre-existing FastAPI app exposing /v1/metrics
├── docker-compose.yml               # base stack (dexdat-core, postgres, consensus-sidecar)
├── docker-compose.monitoring.yml    # ADDITIVE: prometheus 9091:9090, grafana 3030:3000, healthchecks
├── prometheus/
│   ├── prometheus.yml               # job dexdat-core → dexdat-core:8000/v1/metrics, rule_files glob
│   └── rules/dexdat_alerts.yml      # 3 alerts on live-verified metrics
├── grafana/
│   ├── provisioning/datasources/prometheus.yml   # provisioned, uid=dexdat-prometheus
│   ├── provisioning/dashboards/dashboards.yml    # file provider
│   └── dashboards/dexdat-overview.json           # 6 panels, ALL uid-referenced
└── tests/test_config_validity.py    # 12 config-validity tests

Key artifacts (excerpts):

Prometheus scrape — metrics_path: /v1/metrics, targets dexdat-core:8000, rule_files glob mounted read-only.

3 alert rules (all expressions restricted to metric names verified live, none from the documented-absent set): - DexdatHighErrorRate — sum(rate(http_requests_total{status=~"5.."}[5m])) / clamp_min(sum(rate(http_requests_total[5m])), 1e-6) > 0.05, for: 5m, severity: critical - DexdatQueueBacklog — dexdat_queue_depth > 30, for: 2m, severity: warning - DexdatMemoryPressure — process_resident_memory_bytes/1024/1024 > 512, for: 5m, severity: warning

Datasource provisioning — type: prometheus, uid: dexdat-prometheus, url: http://prometheus:9090, isDefault: true. Every dashboard panel and the templating variable use "datasource": {"type": "prometheus", "uid": "dexdat-prometheus"} — no numeric datasource IDs.

Dashboard — dexdat-overview.json, uid dexdat-core-overview, 6 panels (request rate by status, latency p95, jobs in progress, queue depth, RSS MiB, CPU cores).

Additive compose — docker-compose.monitoring.yml defines only prometheus (ports 9091:9090, wget healthcheck on /-/healthy) and grafana (ports 3030:3000, wget healthcheck on /api/health); it shares zero services with the base file, so docker compose -f docker-compose.yml -f docker-compose.monitoring.yml up layers cleanly.

Metric absence documented (README): dexdat_consensus_height lives in the consensus sidecar container; pg_up/pg_stat_* need an undep­loyed postgres_exporter — neither appears in any panel or rule.

Evidence & signatures

1. **Live metric verification before authoring panels** — ran the app, curled `/v1/metrics`, confirmed 17 metric families (incl. `http_requests_total{method,path,status}`, `http_request_duration_seconds_bucket{le}`, `dexdat_jobs_in_progress`, `dexdat_queue_depth`), then generated traffic to confirm label sets. All panels/rules reference only these.
2. **12/12 config-validity tests pass** (`pytest tests/ -q` → `12 passed`) — YAML/JSON validity, exactly 3 rules, all exprs reference verified metrics only, uid-referenced panels, healthchecks + `9091:9090`/`3030:3000` ports, additive-compose check (no service overlap with base, merged service set exact).
3. **promtool validation** — `promtool check config` and `promtool check rules` both SUCCESS against the shipped files.
4. **Real Prometheus end-to-end** (2.53.3 binary, no docker daemon available in this non-root container): scraped the live app — `up{job="dexdat-core"} = 1`, 74 series, `dexdat_queue_depth`/`sum(rate(http_requests_total[1m]))` returned real values; `/api/v1/rules` showed all 3 rules `state=inactive health=ok`, zero eval errors.
5. **Alert firing proven with `promtool test rules`** — all 3 rules fired with correct labels + annotations. This caught a real bug: the initial error-rate rule lacked `sum()`, so PromQL's label-set join divided the 5xx series against *itself* → alert reported 100% instead of the 5xx share. Fixed with `sum(...)/sum(...)`; unit tests re-verified.
6. **All 6 dashboard queries evaluated against live Prometheus** — each panel expr returned data (e.g. p95 = 0.00475s, RSS = 52.98 MiB, queue = 16).
7. **Edge cases tested** — rate-window flatness (flat counters → 0 rate, so no false fires), label-join semantics, promtool annotation matching, host port 9090 conflict (relocated live test to 19090), PEP 668 venv handling.
{"model": "deepseek-v4-flash", "problem_class": "python-monitoring-grafana-prometheus", "result": "passed", "tests": 12}
Generated from the verified corpus · MIT licensedBack to the catalog