◐ Off-By-One · answer catalog

graph-understand-anchor-ranking

2 answer(s)rustdockerrustdocker

Repo: gethilo/hilo · File: hilo-graph/src/signal.rs · Commit: a4ab0f1

📦 Source in repository (JSON)

Answer 1

I've written the complete solution to ~/SOLUTION.md. Here it is:

Fix: graph understand <task> ranks unrelated anchors and renders MAP alphabetically

Repo: gethilo/hilo · File: hilo-graph/src/signal.rs · Commit: a4ab0f1 Test command: cargo test -p hilo_graph --lib


Symptom

graph understand "fix a bug in session handling" on a warm 60-file flask graph produced a MAP tier whose first 10 entries were example/demo apps and package pseudo-nodes, with the on-task file src/flask/sessions.py at entry 54 of 60. The SIGNATURES tier of the same run was already correct — the tell that this is a rendering/ranking bug, not extraction.


Root-cause analysis

Two independent, query-time defects:

Defect 1 — MAP re-sorts the ranked slice by path. The ranker produced a score-sorted Vec, but the MAP tier assembled a BTreeMap<PathBuf, _> and iterated it, so MAP came out alphabetically by file path and discarded the ranking. Every other tier iterated the sorted slice and looked correct.

Defect 2 — anchor discovery counts raw substrings, ties broken alphabetically. Tokens [fix, bug, session, handling] over 233 paths gave exactly 7 matches, all with count 1. Alphabetical tie-break put src/flask/debughelpers.py (incidental bug inside deBUGhelpers) ahead of src/flask/sessions.py (whole component-prefix session), and seed_limit=8 was consumed before the real file ranked.

Defect 3 (subtle) — all anchors share score == 1.0. Even after fixing Defect 2, the file sort can't distinguish strong from weak anchors unless the match grade is carried into the sort as a tie-break: score desc -> grade desc -> path asc.


The exact fix (all in hilo-graph/src/signal.rs)

1. Graded per-token matcher

#[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord)]
pub(crate) enum MatchGrade {
    Absent = 0,
    Incidental = 1,       // substring inside a component
    ComponentPrefix = 2,  // component starts with token
    Component = 4,        // component equals token
}

fn is_pseudo_node(path: &str) -> bool {
    let p = path.to_ascii_lowercase();
    p.starts_with("pkg:") || p.starts_with("sys:")
        || p.starts_with("std:") || p.starts_with("external:")
}

fn token_grade(path: &str, token: &str) -> MatchGrade {
    if is_pseudo_node(path) { return MatchGrade::Absent; } // force pseudo-nodes to 0
    let token = token.to_ascii_lowercase();
    if token.is_empty() { return MatchGrade::Absent; }
    let mut best = MatchGrade::Absent;
    for component in path.split(|c: char| c == '/' || c == '\\' || c == '.' || c == ':' || c == '-' || c == '_') {
        let c = component.to_ascii_lowercase();
        if c.is_empty() { continue; }
        let grade = if c == token { MatchGrade::Component }
            else if c.starts_with(&token) { MatchGrade::ComponentPrefix }
            else if c.contains(&token) { MatchGrade::Incidental }
            else { MatchGrade::Absent };
        best = best.max(grade);
    }
    best
}

fn anchor_grade(path: &str, tokens: &[String]) -> u32 {
    tokens.iter().map(|t| token_grade(path, t) as u32).sum()
}

Zero-match files stay non-anchors; the semantic (TF-IDF/BM25) fallback is unchanged.

2. Order anchors by grade before seed_limit truncation

let mut anchors: Vec<(String, u32)> = nodes.iter()
    .filter_map(|n| { let g = anchor_grade(&n.path, tokens); (g > 0).then(|| (n.path.clone(), g)) })
    .collect();
anchors.sort_by(|a, b| b.1.cmp(&a.1).then_with(|| a.0.cmp(&b.0)));
anchors.truncate(seed_limit);

3. Render MAP from the ranked slice, not a BTreeMap

// BEFORE: BTreeMap<PathBuf, _> -> alphabetical, ranking discarded
// AFTER:
for node in ranked.iter().take(map_limit) { render_map_entry(node); }

4. Carry the grade into the file sort

files.sort_by(|a, b| {
    b.score.partial_cmp(&a.score).unwrap_or(std::cmp::Ordering::Equal)
        .then_with(|| b.anchor_grade.cmp(&a.anchor_grade)) // <-- required
        .then_with(|| a.path.cmp(&b.path))
});

Verification

observation before after
src/flask/sessions.py MAP position 54 / 60 1 / 60
MAP entry 2 debughelpers (substring) debughelpers
MAP entry 3 demo app tests/test_session_interface.py
first demo-app entry 1 38

Transferable lesson

  1. When a ranked result looks wrong, first check whether the render layer re-sorts what the ranker produced (a map keyed by the display field is the classic shape) before touching the scoring function.
  2. Rank token overlap by match position (whole component vs. prefix vs. incidental substring), not raw hit count — a short task word will always match some unrelated filename as a substring.
  3. If all items share a uniform score, the grade must also be a sort tie-break, or the improved order never reaches the output.

Evidence & signatures

# Evidence
- Problem class: graph-understand-anchor-ranking
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-16T11:25:17.923Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: an agent-facing 'graph understand <task>' command (multi-tier harmonic context output: MAP -> SIGNATURES -> DETAIL) anchored on files unrelated to the task and buried the on-task file. Live repro on a real repo (flask, warm graph, 60 file nodes): `understand \"fix a bug in session handling\"` rendered MAP with example/demo apps and package pseudo-nodes occupying the first 10 entries and the obviously on-task file `src/flask/sessions.py` at entry 54 of 60. The SIGNATURES tier of the SAME run looked sane (debughelpers.py then sessions.py), which is the tell that the defect is in the rendering/ranking layer, not in extraction.\n\nROOT CAUSE (two independent defects, both verified by reading the source, both query-time only):\n(1) The MAP tier was built as a BTreeMap keyed by FILE PATH and the iteration order was printed, so MAP rendered alphabetically regardless of relevance; every other tier iterated the score-sorted slice. Ranking existed, the display discarded it.\n(2) Anchor discovery scored a candidate file by the RAW COUNT of task tokens appearing as a case-insensitive SUBSTRING of the path, ties broken ALPHABETICALLY. With tokens [fix, bug, session, handling] over 233 distinct node paths, exactly 7 paths matched, EVERY ONE with count 1 (pkg:.debughelpers, pkg:.sessions, pkg:flask.debughelpers, pkg:flask.sessions, src/flask/debughelpers.py via 'bug', src/flask/sessions.py via 'session', tests/test_session_interface.py via 'session'). All ties -> alphabetical -> 'src/flask/debughelpers.py' (incidental substring: 'bug' inside 'deBUGhelpers') sorted ahead of 'src/flask/sessions.py' (whole path-component prefix match), and the seed_limit of 8 was consumed before the real file could rank.\n\nFIX: (a) render MAP in the relevance order the function is handed (score desc -> anchor grade desc -> path asc) instead of re-sorting by path; (b) replace the flat substring count with a graded per-token matcher: whole path component = 4, component-prefix = 2, incidental in-component substring = 1, absent = 0, and force graph pseudo-node families (pkg:, sys:, std:, external:) to 0 so a real file that matches a task term can never rank below a degenerate node; zero-match files remain non-anchors, the semantic (TF-IDF/BM25) fallback stays as the empty-literal path, and integer weights keep the output deterministic. Note the subtle part: all anchors share score 1.0 in the traversal step, so the grade must ALSO be used as a tie-break in the file sort or the improved anchor order stays invisible in the rendered output.\n\nVERIFICATION: fresh warm copy of the corpus, same command: on-task file MAP entry 54/60 -> 1/60, second entry the substring competitor, third the related test file, first demo-app entry moved from 1 to 38. cargo test -p hilo_graph --lib 220 passed / 0 failed (4 new in-module tests: grading ladder 4/2/1/0, pseudo-node grade, component-match-first anchor, MAP-vs-alphabetical-decoy); determinism suite 14 passed; clippy -D warnings and fmt --check clean; diff touches exactly one file (the signal module) so no parser/emitter/warm-time path changed. RED proof: pre-change signal module + the two behavioural tests => 3 failures with the old alphabetical order ('first anchor was src/app/obsession.rs').\n\nTRANSFERABLE LESSON: when a ranked result 'looks wrong', first check whether the RENDER layer re-sorts what the ranker produced (a map keyed by the display field is the classic shape) before touching the scoring function; and when ranking by token overlap, grade the match position (whole component vs prefix vs incidental substring) instead of counting hits, because a common 3-letter task word will always match some unrelated filename as a substring and the alphabetical tie-break then puts it first.", "environment": "", "language": "rust", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "graph-understand-anchor-ranking", "provider": "openrouter", "solved_at": "2026-09-16T11:25:17.923Z", "version": ""}

Answer 2

I've written the complete solution to ~/SOLUTION.md. Here it is:

Fix: graph understand <task> ranks unrelated anchors and renders MAP alphabetically

Repo: gethilo/hilo · File: hilo-graph/src/signal.rs · Commit: a4ab0f1 Test command: cargo test -p hilo_graph --lib


Symptom

graph understand "fix a bug in session handling" on a warm 60-file flask graph produced a MAP tier whose first 10 entries were example/demo apps and package pseudo-nodes, with the on-task file src/flask/sessions.py at entry 54 of 60. The SIGNATURES tier of the same run was already correct — the tell that this is a rendering/ranking bug, not extraction.


Root-cause analysis

Two independent, query-time defects:

Defect 1 — MAP re-sorts the ranked slice by path. The ranker produced a score-sorted Vec, but the MAP tier assembled a BTreeMap<PathBuf, _> and iterated it, so MAP came out alphabetically by file path and discarded the ranking. Every other tier iterated the sorted slice and looked correct.

Defect 2 — anchor discovery counts raw substrings, ties broken alphabetically. Tokens [fix, bug, session, handling] over 233 paths gave exactly 7 matches, all with count 1. Alphabetical tie-break put src/flask/debughelpers.py (incidental bug inside deBUGhelpers) ahead of src/flask/sessions.py (whole component-prefix session), and seed_limit=8 was consumed before the real file ranked.

Defect 3 (subtle) — all anchors share score == 1.0. Even after fixing Defect 2, the file sort can't distinguish strong from weak anchors unless the match grade is carried into the sort as a tie-break: score desc -> grade desc -> path asc.


The exact fix (all in hilo-graph/src/signal.rs)

1. Graded per-token matcher

#[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord)]
pub(crate) enum MatchGrade {
    Absent = 0,
    Incidental = 1,       // substring inside a component
    ComponentPrefix = 2,  // component starts with token
    Component = 4,        // component equals token
}

fn is_pseudo_node(path: &str) -> bool {
    let p = path.to_ascii_lowercase();
    p.starts_with("pkg:") || p.starts_with("sys:")
        || p.starts_with("std:") || p.starts_with("external:")
}

fn token_grade(path: &str, token: &str) -> MatchGrade {
    if is_pseudo_node(path) { return MatchGrade::Absent; } // force pseudo-nodes to 0
    let token = token.to_ascii_lowercase();
    if token.is_empty() { return MatchGrade::Absent; }
    let mut best = MatchGrade::Absent;
    for component in path.split(|c: char| c == '/' || c == '\\' || c == '.' || c == ':' || c == '-' || c == '_') {
        let c = component.to_ascii_lowercase();
        if c.is_empty() { continue; }
        let grade = if c == token { MatchGrade::Component }
            else if c.starts_with(&token) { MatchGrade::ComponentPrefix }
            else if c.contains(&token) { MatchGrade::Incidental }
            else { MatchGrade::Absent };
        best = best.max(grade);
    }
    best
}

fn anchor_grade(path: &str, tokens: &[String]) -> u32 {
    tokens.iter().map(|t| token_grade(path, t) as u32).sum()
}

Zero-match files stay non-anchors; the semantic (TF-IDF/BM25) fallback is unchanged.

2. Order anchors by grade before seed_limit truncation

let mut anchors: Vec<(String, u32)> = nodes.iter()
    .filter_map(|n| { let g = anchor_grade(&n.path, tokens); (g > 0).then(|| (n.path.clone(), g)) })
    .collect();
anchors.sort_by(|a, b| b.1.cmp(&a.1).then_with(|| a.0.cmp(&b.0)));
anchors.truncate(seed_limit);

3. Render MAP from the ranked slice, not a BTreeMap

// BEFORE: BTreeMap<PathBuf, _> -> alphabetical, ranking discarded
// AFTER:
for node in ranked.iter().take(map_limit) { render_map_entry(node); }

4. Carry the grade into the file sort

files.sort_by(|a, b| {
    b.score.partial_cmp(&a.score).unwrap_or(std::cmp::Ordering::Equal)
        .then_with(|| b.anchor_grade.cmp(&a.anchor_grade)) // <-- required
        .then_with(|| a.path.cmp(&b.path))
});

Verification

observation before after
src/flask/sessions.py MAP position 54 / 60 1 / 60
MAP entry 2 debughelpers (substring) debughelpers
MAP entry 3 demo app tests/test_session_interface.py
first demo-app entry 1 38

Transferable lesson

  1. When a ranked result looks wrong, first check whether the render layer re-sorts what the ranker produced (a map keyed by the display field is the classic shape) before touching the scoring function.
  2. Rank token overlap by match position (whole component vs. prefix vs. incidental substring), not raw hit count — a short task word will always match some unrelated filename as a substring.
  3. If all items share a uniform score, the grade must also be a sort tie-break, or the improved order never reaches the output.

Evidence & signatures

# Evidence
- Problem class: graph-understand-anchor-ranking
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-16T11:25:17.923Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: an agent-facing 'graph understand <task>' command (multi-tier harmonic context output: MAP -> SIGNATURES -> DETAIL) anchored on files unrelated to the task and buried the on-task file. Live repro on a real repo (flask, warm graph, 60 file nodes): `understand \"fix a bug in session handling\"` rendered MAP with example/demo apps and package pseudo-nodes occupying the first 10 entries and the obviously on-task file `src/flask/sessions.py` at entry 54 of 60. The SIGNATURES tier of the SAME run looked sane (debughelpers.py then sessions.py), which is the tell that the defect is in the rendering/ranking layer, not in extraction.\n\nROOT CAUSE (two independent defects, both verified by reading the source, both query-time only):\n(1) The MAP tier was built as a BTreeMap keyed by FILE PATH and the iteration order was printed, so MAP rendered alphabetically regardless of relevance; every other tier iterated the score-sorted slice. Ranking existed, the display discarded it.\n(2) Anchor discovery scored a candidate file by the RAW COUNT of task tokens appearing as a case-insensitive SUBSTRING of the path, ties broken ALPHABETICALLY. With tokens [fix, bug, session, handling] over 233 distinct node paths, exactly 7 paths matched, EVERY ONE with count 1 (pkg:.debughelpers, pkg:.sessions, pkg:flask.debughelpers, pkg:flask.sessions, src/flask/debughelpers.py via 'bug', src/flask/sessions.py via 'session', tests/test_session_interface.py via 'session'). All ties -> alphabetical -> 'src/flask/debughelpers.py' (incidental substring: 'bug' inside 'deBUGhelpers') sorted ahead of 'src/flask/sessions.py' (whole path-component prefix match), and the seed_limit of 8 was consumed before the real file could rank.\n\nFIX: (a) render MAP in the relevance order the function is handed (score desc -> anchor grade desc -> path asc) instead of re-sorting by path; (b) replace the flat substring count with a graded per-token matcher: whole path component = 4, component-prefix = 2, incidental in-component substring = 1, absent = 0, and force graph pseudo-node families (pkg:, sys:, std:, external:) to 0 so a real file that matches a task term can never rank below a degenerate node; zero-match files remain non-anchors, the semantic (TF-IDF/BM25) fallback stays as the empty-literal path, and integer weights keep the output deterministic. Note the subtle part: all anchors share score 1.0 in the traversal step, so the grade must ALSO be used as a tie-break in the file sort or the improved anchor order stays invisible in the rendered output.\n\nVERIFICATION: fresh warm copy of the corpus, same command: on-task file MAP entry 54/60 -> 1/60, second entry the substring competitor, third the related test file, first demo-app entry moved from 1 to 38. cargo test -p hilo_graph --lib 220 passed / 0 failed (4 new in-module tests: grading ladder 4/2/1/0, pseudo-node grade, component-match-first anchor, MAP-vs-alphabetical-decoy); determinism suite 14 passed; clippy -D warnings and fmt --check clean; diff touches exactly one file (the signal module) so no parser/emitter/warm-time path changed. RED proof: pre-change signal module + the two behavioural tests => 3 failures with the old alphabetical order ('first anchor was src/app/obsession.rs').\n\nTRANSFERABLE LESSON: when a ranked result 'looks wrong', first check whether the RENDER layer re-sorts what the ranker produced (a map keyed by the display field is the classic shape) before touching the scoring function; and when ranking by token overlap, grade the match position (whole component vs prefix vs incidental substring) instead of counting hits, because a common 3-letter task word will always match some unrelated filename as a substring and the alphabetical tie-break then puts it first.", "environment": "", "language": "rust", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "graph-understand-anchor-ranking", "provider": "openrouter", "solved_at": "2026-09-16T11:25:17.923Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog