Runtime: Node 20+ (verified on Node 22, CommonJS, zero dependencies)
I've built and verified the complete solution. The full document (including complete source listings) is saved at ~/pathtracer/SOLUTION.md, with runnable code in pt.js, stats.js, experiments.js. Here is the solution:
Runtime: Node 20+ (verified on Node 22, CommonJS, zero dependencies)
Artifacts: ~/pathtracer/{pt.js, stats.js, experiments.js, per_pixel.js, SOLUTION.md}
A from-scratch CPU path tracer (cosine diffuse, GGX conductor, specular dielectric, NEE, power-heuristic MIS) was fixed and verified. Four implementation bugs and one statistical mis-specification:
| # | Defect | Symptom |
|---|---|---|
| 1 | NEE never multiplied by the light's emitted radiance L_e |
MIS under-counts by ~L_e; hidden when L_e=1 |
| 2 | NEE gated on specular, which is true for the camera ray |
NEE skipped on first visible diffuse surface |
| 3 | MIS emission weights applied with NEE disabled | BSDF-only estimator biased low |
| 4 | Emitter two-sided on hit, one-sided in light sampling | Brute counts back-face emission NEE can't sample |
| 5 | Chi-square two-sample test used as a bias test | Rejects because MIS has lower variance, not because of bias |
Verified: white furnace within 1e-3 (diffuse 5.5e-5, dielectric 0, GGX r=0.1 8.4e-5); Welch z-test rejects at 0.94% for α=0.01 (nominal 1%); chi-square null calibrated at 0.31–0.63%; MIS variance 76.5x lower (≥43.9x per pixel) on a small-source glossy scene.
sampleQuadLight() returns emission, but the NEE block never used it:
// BROKEN
const c = mul(mul(ev.f, (cosS / s.pdf) * w), 1); // missing L_e
Correct estimator: throughput * f * cos / p_light * L_e * w_mis. The white furnace and many smoke tests used L_e=1, so the bug hid. With L_e=3 (diffuse scene) pixel 5 measured MIS 0.175 vs brute 0.312; after the fix the ratio is 1.000.
specular=true is initialized for the camera ray (so directly visible emitters are unweighted). It was mistakenly reused to gate NEE. The flag must only control emission weighting; NEE belongs at every non-delta vertex:
// FIXED
if (useNEE && !mat.isDelta && scene.lights.length > 0) { ... }
// BROKEN: w_bsdf < 1 even with NEE off -> biased single-strategy MIS
if (useMIS && !specular) { const w = powerHeuristic(prevBsdfPdf, lightPdf); ... }
// FIXED
if (useNEE && useMIS && !specular) { ... }
sampleQuadLight rejects cosL <= 0 (emitting face only), but the hit test used Math.abs(...). Use the same signed cosine in both:
// FIXED (trace and traceBrute)
const cosL = dot(hit.obj.normal, neg(ray.d)); // > 0 on emitting face
if (cosL > 1e-9) { const lightPdf = d*d/(area*cosL); ... }
A chi-square two-sample test has null H0: same distribution, not H0: same mean. MIS and brute are both unbiased for the same mean but deliberately have different variances. So the chi-square test rejects whenever the (desired) variance reduction is detectable, and failing to reject does not prove unbiasedness.
pt.js)// (1.1) area-light NEE: include emitted radiance
const c = mulv(mul(ev.f, (cosS / s.pdf) * w), s.emission);
// (1.1) environment NEE
const c = mulv(mul(ev.f, (cosS / s.pdf) * w), s.emission);
// (1.2) NEE at every non-delta vertex
if (useNEE && !mat.isDelta && scene.lights.length > 0) { ... }
if (useNEE && !mat.isDelta && env && env.sample) { ... }
// (1.3) MIS weights require NEE
if (useNEE && useMIS && !specular) { ... }
// (1.4) signed cosine at emission hit
const cosL = dot(hit.obj.normal, neg(ray.d));
if (cosL > 1e-9) { const lightPdf = (dist*dist)/(hit.obj.area*cosL); ... }
The statistical core (stats.js) implements the chi-square test and the Welch z-test; experiments.js runs the furnace, grouped-replicate calibration, and variance study.
Commands
cd ~/pathtracer
node experiments.js # furnace + qualification + variance reduction (~35 s)
node per_pixel.js # per-pixel chi-square / Welch / variance table
L=1, 4M rays)| BSDF | mean | std. err | \|mean−1\| |
|---|---|---|---|
diffuse albedo=1 |
0.999945 |
1.27e-4 |
5.46e-5 ✓ |
dielectric ior=1.5 |
1.000000 |
0.00e+0 |
0.00e+0 ✓ |
GGX F0=1, r=0.1 |
0.999916 |
5.19e-5 |
8.39e-5 ✓ |
GGX F0=1, r=0.3 |
0.990727 |
1.39e-4 |
9.27e-3 (single-scatter loss; conservative, ≤1) |
| Test | H0 | Reject@0.01 |
|---|---|---|
| chi-square, MIS half vs MIS half | same estimator (true) | 0.0031 |
| chi-square, brute half vs brute half | same estimator (true) | 0.0063 |
| chi-square, MIS vs brute | same distribution (false) | 0.6641 |
| Welch z, MIS vs brute | same mean (true) | 0.0094 |
The chi-square test is calibrated under its true null (~1%), rejects MIS-vs-brute because of variance, and the Welch mean test rejects at 0.94% ≈ 1%.
From per_pixel.js: for pixels 0,2,3 the variance ratio is ≈1 and the chi-square test fails to reject (p = 0.72, 0.58, 0.067), while every Welch p-value stays > 0.01. For pixels with ratio 2x–58x, chi-square rejects — exactly the predicted behavior.
| Quantity | Value |
|---|---|
sum Var[MIS] |
6.198e-2 |
sum Var[brute] |
4.739e+0 |
| overall ratio | 76.5x |
| min per-pixel ratio (12 px) | 43.9x |
| mean radiance MIS / brute | 0.2439 / 0.2384 |
0.31%/0.63%, consistent with nominal 1%.p=0.01.0.94%, the nominal false-positive rate.Forcing the chi-square test to never reject would be statistically wrong: it rejects because MIS is better (lower variance) than the reference. The fix is to test the right hypothesis (mean, via Welch), and reserve chi-square for distribution/calibration.
The complete, runnable source is in SOLUTION.md and the four .js files in ~/pathtracer.
# Evidence - Problem class: js-path-tracer-next-event-mis-unbiasedness-chi-square - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-27T22:14:09.126Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Implement a CPU path tracer in Node 20 with cosine-weighted diffuse, GGX microfacet and specular-dielectric BSDFs, combined by multiple importance sampling (power heuristic) plus next-event estimation, and make it provably energy-conserving (white-furnace test within 1e-3). Prove unbiasedness statistically rather than by eye: on a fixed scene, run the MIS estimator and a brute-force hemisphere-sampling reference over many independent batches, then show a chi-square two-sample test over binned pixel radiance fails to reject at the expected false-positive rate at p=0.01. Also show MIS variance is strictly lower than the reference for a scene where a small glossy source dominates the illumination.", "environment": "node20", "language": "js", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "js-path-tracer-next-event-mis-unbiasedness-chi-square", "provider": "openrouter", "solved_at": "2026-09-27T22:14:09.127Z", "version": "20"}I've built and verified the complete solution. The full document (including complete source listings) is saved at ~/pathtracer/SOLUTION.md, with runnable code in pt.js, stats.js, experiments.js. Here is the solution:
Runtime: Node 20+ (verified on Node 22, CommonJS, zero dependencies)
Artifacts: ~/pathtracer/{pt.js, stats.js, experiments.js, per_pixel.js, SOLUTION.md}
A from-scratch CPU path tracer (cosine diffuse, GGX conductor, specular dielectric, NEE, power-heuristic MIS) was fixed and verified. Four implementation bugs and one statistical mis-specification:
| # | Defect | Symptom |
|---|---|---|
| 1 | NEE never multiplied by the light's emitted radiance L_e |
MIS under-counts by ~L_e; hidden when L_e=1 |
| 2 | NEE gated on specular, which is true for the camera ray |
NEE skipped on first visible diffuse surface |
| 3 | MIS emission weights applied with NEE disabled | BSDF-only estimator biased low |
| 4 | Emitter two-sided on hit, one-sided in light sampling | Brute counts back-face emission NEE can't sample |
| 5 | Chi-square two-sample test used as a bias test | Rejects because MIS has lower variance, not because of bias |
Verified: white furnace within 1e-3 (diffuse 5.5e-5, dielectric 0, GGX r=0.1 8.4e-5); Welch z-test rejects at 0.94% for α=0.01 (nominal 1%); chi-square null calibrated at 0.31–0.63%; MIS variance 76.5x lower (≥43.9x per pixel) on a small-source glossy scene.
sampleQuadLight() returns emission, but the NEE block never used it:
// BROKEN
const c = mul(mul(ev.f, (cosS / s.pdf) * w), 1); // missing L_e
Correct estimator: throughput * f * cos / p_light * L_e * w_mis. The white furnace and many smoke tests used L_e=1, so the bug hid. With L_e=3 (diffuse scene) pixel 5 measured MIS 0.175 vs brute 0.312; after the fix the ratio is 1.000.
specular=true is initialized for the camera ray (so directly visible emitters are unweighted). It was mistakenly reused to gate NEE. The flag must only control emission weighting; NEE belongs at every non-delta vertex:
// FIXED
if (useNEE && !mat.isDelta && scene.lights.length > 0) { ... }
// BROKEN: w_bsdf < 1 even with NEE off -> biased single-strategy MIS
if (useMIS && !specular) { const w = powerHeuristic(prevBsdfPdf, lightPdf); ... }
// FIXED
if (useNEE && useMIS && !specular) { ... }
sampleQuadLight rejects cosL <= 0 (emitting face only), but the hit test used Math.abs(...). Use the same signed cosine in both:
// FIXED (trace and traceBrute)
const cosL = dot(hit.obj.normal, neg(ray.d)); // > 0 on emitting face
if (cosL > 1e-9) { const lightPdf = d*d/(area*cosL); ... }
A chi-square two-sample test has null H0: same distribution, not H0: same mean. MIS and brute are both unbiased for the same mean but deliberately have different variances. So the chi-square test rejects whenever the (desired) variance reduction is detectable, and failing to reject does not prove unbiasedness.
pt.js)// (1.1) area-light NEE: include emitted radiance
const c = mulv(mul(ev.f, (cosS / s.pdf) * w), s.emission);
// (1.1) environment NEE
const c = mulv(mul(ev.f, (cosS / s.pdf) * w), s.emission);
// (1.2) NEE at every non-delta vertex
if (useNEE && !mat.isDelta && scene.lights.length > 0) { ... }
if (useNEE && !mat.isDelta && env && env.sample) { ... }
// (1.3) MIS weights require NEE
if (useNEE && useMIS && !specular) { ... }
// (1.4) signed cosine at emission hit
const cosL = dot(hit.obj.normal, neg(ray.d));
if (cosL > 1e-9) { const lightPdf = (dist*dist)/(hit.obj.area*cosL); ... }
The statistical core (stats.js) implements the chi-square test and the Welch z-test; experiments.js runs the furnace, grouped-replicate calibration, and variance study.
Commands
cd ~/pathtracer
node experiments.js # furnace + qualification + variance reduction (~35 s)
node per_pixel.js # per-pixel chi-square / Welch / variance table
L=1, 4M rays)| BSDF | mean | std. err | \|mean−1\| |
|---|---|---|---|
diffuse albedo=1 |
0.999945 |
1.27e-4 |
5.46e-5 ✓ |
dielectric ior=1.5 |
1.000000 |
0.00e+0 |
0.00e+0 ✓ |
GGX F0=1, r=0.1 |
0.999916 |
5.19e-5 |
8.39e-5 ✓ |
GGX F0=1, r=0.3 |
0.990727 |
1.39e-4 |
9.27e-3 (single-scatter loss; conservative, ≤1) |
| Test | H0 | Reject@0.01 |
|---|---|---|
| chi-square, MIS half vs MIS half | same estimator (true) | 0.0031 |
| chi-square, brute half vs brute half | same estimator (true) | 0.0063 |
| chi-square, MIS vs brute | same distribution (false) | 0.6641 |
| Welch z, MIS vs brute | same mean (true) | 0.0094 |
The chi-square test is calibrated under its true null (~1%), rejects MIS-vs-brute because of variance, and the Welch mean test rejects at 0.94% ≈ 1%.
From per_pixel.js: for pixels 0,2,3 the variance ratio is ≈1 and the chi-square test fails to reject (p = 0.72, 0.58, 0.067), while every Welch p-value stays > 0.01. For pixels with ratio 2x–58x, chi-square rejects — exactly the predicted behavior.
| Quantity | Value |
|---|---|
sum Var[MIS] |
6.198e-2 |
sum Var[brute] |
4.739e+0 |
| overall ratio | 76.5x |
| min per-pixel ratio (12 px) | 43.9x |
| mean radiance MIS / brute | 0.2439 / 0.2384 |
0.31%/0.63%, consistent with nominal 1%.p=0.01.0.94%, the nominal false-positive rate.Forcing the chi-square test to never reject would be statistically wrong: it rejects because MIS is better (lower variance) than the reference. The fix is to test the right hypothesis (mean, via Welch), and reserve chi-square for distribution/calibration.
The complete, runnable source is in SOLUTION.md and the four .js files in ~/pathtracer.
# Evidence - Problem class: js-path-tracer-next-event-mis-unbiasedness-chi-square - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-27T22:14:09.126Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Implement a CPU path tracer in Node 20 with cosine-weighted diffuse, GGX microfacet and specular-dielectric BSDFs, combined by multiple importance sampling (power heuristic) plus next-event estimation, and make it provably energy-conserving (white-furnace test within 1e-3). Prove unbiasedness statistically rather than by eye: on a fixed scene, run the MIS estimator and a brute-force hemisphere-sampling reference over many independent batches, then show a chi-square two-sample test over binned pixel radiance fails to reject at the expected false-positive rate at p=0.01. Also show MIS variance is strictly lower than the reference for a scene where a small glossy source dominates the illumination.", "environment": "node20", "language": "js", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "js-path-tracer-next-event-mis-unbiasedness-chi-square", "provider": "openrouter", "solved_at": "2026-09-27T22:14:09.127Z", "version": "20"}