metric-sampler-samples-trivial-roots
Verified. Deliverable: ~/metric-sampler-samples-trivial-roots/SOLUTION.md (plus runnable main.go / main_test.go).
The harness's sampler (ORDER BY seq ASC LIMIT N) selected the oldest rows, which on this graph are 60/70 tree roots (parent_id IS NULL). A root has empty expected ancestry, so ScoreAncestry returns 1.0 vacuously. The 100.00% was a property of sample composition, not scorer correctness — starving the scorer (budget 1 vs 20) produced identical output.
starved(budget=1) mean=100.00% comp={Sampled:10 Roots:10 NonRoots:0 DepthScored:0}
normal (budget=20) mean=100.00% comp={Sampled:10 Roots:10 NonRoots:0 DepthScored:0}
BROKEN: starved == normal, so the sampler is not measuring
parent_id IS NOT NULL AND depth > 1.{Sampled, Roots, NonRoots, DepthScored} every run.DepthScored == 0 (sample not discriminating).func (s *Store) SampleFixed(limit, minDepth int) []int {
byParent := map[int][]int{}
for _, id := range s.order {
n := s.nodes[id]
if n.ParentID == noParent || s.Depth(id) <= minDepth {
continue
}
byParent[n.ParentID] = append(byParent[n.ParentID], id)
}
// round-robin one child per parent (deterministic parent order)
parents := sortedKeys(byParent)
var out []int
for len(out) < limit {
progressed := false
for _, p := range parents {
if len(out) >= limit { break }
if len(byParent[p]) > 0 {
out = append(out, byParent[p][0])
byParent[p] = byParent[p][1:]
progressed = true
}
}
if !progressed { break }
}
return out
}
SQL equivalent uses a recursive CTE for depth, filters depth > 1 AND parent_id IS NOT NULL, and ORDER BY row_number() OVER (PARTITION BY parent_id ...) for diversification (full query in the markdown).
go vet ./... clean; go test ./... passes; go run .:
== FIXED sampler: parent_id IS NOT NULL AND depth > 1, diversified ==
starved(budget=1) mean=0.00% comp={Sampled:8 Roots:0 NonRoots:8 DepthScored:8}
normal (budget=20) mean=100.00% comp={Sampled:8 Roots:0 NonRoots:8 DepthScored:8}
OK: starved != normal, the metric is actually measuring
Tests pin: old sample is all-roots/non-discriminating (starved == normal), fixed sample is 0 roots and DepthScored == 8, starved collapses to 0% while normal recovers 100%, and parents are diversified.
Detection one-liner: diff <(STARVED=1 ./harness) <(./harness) && echo "BROKEN: sampler, not code" — identical output means the sample cannot exercise the metric.
# Evidence - Problem class: metric-sampler-samples-trivial-roots - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-20T09:47:56.549Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A measurement harness orders its sample by the row's creation/sequence column ASC. On this graph the OLDEST rows are almost all tree ROOTS (measured: 60 of 70 nodes had a null parent), whose 'expected ancestry' is a single node, so every sampled entry scored full marks and the harness reported a guaranteed 100% accuracy. Falsification that exposed it: the same run under a starved configuration (a 1-token budget that forces ancestry drops) STILL printed 100.00%. RULE: before trusting a metric's headline number, force the condition the metric claims to detect and require the number to move. A metric that cannot be made to go DOWN is not measuring. FIX: sample the entries that actually exercise the metric (require depth > 1 / the discriminating condition) in the query itself, diversify the sample across parents/groups rather than taking the first N from one, report the sample's composition counters (root vs non-root, depth-scored), and emit a loud 'sample not discriminating' warning when no sampled entry can discriminate. Detection one-liner: run the harness twice, once with the feature starved and once normally; identical output = the sampler, not the code, is broken.", "environment": "", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "metric-sampler-samples-trivial-roots", "provider": "openrouter", "solved_at": "2026-09-20T09:47:56.549Z", "version": ""}Verified. Deliverable: ~/metric-sampler-samples-trivial-roots/SOLUTION.md (plus runnable main.go / main_test.go).
The harness's sampler (ORDER BY seq ASC LIMIT N) selected the oldest rows, which on this graph are 60/70 tree roots (parent_id IS NULL). A root has empty expected ancestry, so ScoreAncestry returns 1.0 vacuously. The 100.00% was a property of sample composition, not scorer correctness — starving the scorer (budget 1 vs 20) produced identical output.
starved(budget=1) mean=100.00% comp={Sampled:10 Roots:10 NonRoots:0 DepthScored:0}
normal (budget=20) mean=100.00% comp={Sampled:10 Roots:10 NonRoots:0 DepthScored:0}
BROKEN: starved == normal, so the sampler is not measuring
parent_id IS NOT NULL AND depth > 1.{Sampled, Roots, NonRoots, DepthScored} every run.DepthScored == 0 (sample not discriminating).func (s *Store) SampleFixed(limit, minDepth int) []int {
byParent := map[int][]int{}
for _, id := range s.order {
n := s.nodes[id]
if n.ParentID == noParent || s.Depth(id) <= minDepth {
continue
}
byParent[n.ParentID] = append(byParent[n.ParentID], id)
}
// round-robin one child per parent (deterministic parent order)
parents := sortedKeys(byParent)
var out []int
for len(out) < limit {
progressed := false
for _, p := range parents {
if len(out) >= limit { break }
if len(byParent[p]) > 0 {
out = append(out, byParent[p][0])
byParent[p] = byParent[p][1:]
progressed = true
}
}
if !progressed { break }
}
return out
}
SQL equivalent uses a recursive CTE for depth, filters depth > 1 AND parent_id IS NOT NULL, and ORDER BY row_number() OVER (PARTITION BY parent_id ...) for diversification (full query in the markdown).
go vet ./... clean; go test ./... passes; go run .:
== FIXED sampler: parent_id IS NOT NULL AND depth > 1, diversified ==
starved(budget=1) mean=0.00% comp={Sampled:8 Roots:0 NonRoots:8 DepthScored:8}
normal (budget=20) mean=100.00% comp={Sampled:8 Roots:0 NonRoots:8 DepthScored:8}
OK: starved != normal, the metric is actually measuring
Tests pin: old sample is all-roots/non-discriminating (starved == normal), fixed sample is 0 roots and DepthScored == 8, starved collapses to 0% while normal recovers 100%, and parents are diversified.
Detection one-liner: diff <(STARVED=1 ./harness) <(./harness) && echo "BROKEN: sampler, not code" — identical output means the sample cannot exercise the metric.
# Evidence - Problem class: metric-sampler-samples-trivial-roots - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-20T09:47:56.549Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A measurement harness orders its sample by the row's creation/sequence column ASC. On this graph the OLDEST rows are almost all tree ROOTS (measured: 60 of 70 nodes had a null parent), whose 'expected ancestry' is a single node, so every sampled entry scored full marks and the harness reported a guaranteed 100% accuracy. Falsification that exposed it: the same run under a starved configuration (a 1-token budget that forces ancestry drops) STILL printed 100.00%. RULE: before trusting a metric's headline number, force the condition the metric claims to detect and require the number to move. A metric that cannot be made to go DOWN is not measuring. FIX: sample the entries that actually exercise the metric (require depth > 1 / the discriminating condition) in the query itself, diversify the sample across parents/groups rather than taking the first N from one, report the sample's composition counters (root vs non-root, depth-scored), and emit a loud 'sample not discriminating' warning when no sampled entry can discriminate. Detection one-liner: run the harness twice, once with the feature starved and once normally; identical output = the sampler, not the code, is broken.", "environment": "", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "metric-sampler-samples-trivial-roots", "provider": "openrouter", "solved_at": "2026-09-20T09:47:56.549Z", "version": ""}