You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Add discriminating numerical regressions for every current interpolation family. Keep this an audit: record and file discovered defects separately rather than repairing production algorithms here.
Plan
Map current interpolation families to their valid mathematical invariants.
Add boundary, reproduction and refinement regressions with deliberate-fault evidence.
Add deterministic synthetic truth recovery and an unchanged control.
Report defects separately and validate the library and workspace suites.
Classification: Both; PyAutoArray primary, autolens_workspace_test companion. Large but bounded to a single audit; no dependency on another member. If investigation proves too large, remove from bundle before expanding.
Inspect autoarray/inversion/mesh/interpolator/{rectangular,rectangular_uniform,delaunay,sibson,knn}.py and existing test_autoarray/inversion/pixelization/interpolator tests. DelaunayNN now lives in sibson.py, not a separate delaunay_nn.py; include both InterpolatorKNearestNeighbor and InterpolatorKNNBarycentric.
Add test_autoarray/inversion/pixelization/interpolator/test_numerics_audit.py (or extend existing family tests). Use independent coordinate/analytic oracles: rectangular linear reproduction in index space, barycentric/Sibson linear reproduction within their valid hull, constant and nearest-neighbor consistency where linear precision is not guaranteed. Document applicability of each of the prompt's five checks; do not impose continuity on inherently discontinuous nearest-neighbor changes.
Sweep interior cell/simplex boundaries with exact and nextafter coordinates; assert mapping bounds, correct corner/weight pairing, live/guard node conventions and finite values. Measure smooth-function refinement on data-supported query sets, with deterministic tolerances and rates justified by each scheme.
Measure CDF round-trip error at mesh sizes 16/32/64 and larger knot tables. Recommend scaling only if evidence warrants; production changes belong to separate bug work.
Add a deterministic synthetic source-recovery audit under autolens_workspace_test/scripts/imaging/, mirroring a current rectangular fitting script. Measure Pearson correlation and normalized RMS against truth; pin rectangular_uniform control output bit-for-bit between clean and deliberately perturbed candidates. Avoid real observational data.
Demonstrate each new regression fails with a relevant deliberately broken pairing, cell locator or weight mutation; use temporary monkeypatches or scratch variants, never commit corrupted source. Preserve baseline/fault results and applicability table in the issue evidence. Existing production defects remain visible; do not loosen assertions or disguise them with unconditional skips.
Run focused interpolation tests, full test_autoarray, the new truth script and applicable workspace smoke checks. JAX parity/gradient checks, if needed, live in workspace scripts, not library unit tests. Library PR first; linked workspace PR follows its gate. No use of areas_transformed, preserving independence from the geometry repair.
Worktree root: /home/jammy/Code/PyAutoLabs/.worktrees/autoarray-bundle-1, created once using Brain's worktree helper with PYAUTO_WT_ROOT inside this workspace. Shared repo worktrees: PyAutoArray, PyAutoGalaxy, autolens_workspace_test. Each member starts from origin/main on its own branch; execute and ship sequentially before switching the shared PyAutoArray worktree. One primary PyAutoArray issue and registry entry per member, linked companion PRs per repository as explicitly authorized. No merges.
Branch survey: PyAutoArray, PyAutoGalaxy and autolens_workspace_test canonical checkouts are clean on main. No target repo claims in active.md. Heart reports an unregistered sparse-operator-oversampling-cache/PyAutoArray worktree with 11 dirty files: preserve it and obtain overlap acknowledgement before setup. Recent PyAutoArray branches: main, feature/sparse-operator-oversampling-cache, chore/session-start-hook-regen, feature/delaunay-area-magnification-audit, claude/autonerves-floor-regime-stamp. Recent PyAutoGalaxy branches: main, chore/session-start-hook-regen. Recent autolens_workspace_test branches: main, feature/point-audits-wheel-provenance, feature/point-solver-image-accuracy, feature/point-solver-duplicate-policy, chore/session-start-hook-regen.
Execution: one native Sol delegate per member, sequential within shared repositories. Parent owns judgment and lifecycle. Pass the approved issue plan, branch, exact starting commit, worktree, permitted files, validation requirements; stop and return exact failure evidence rather than weakening tests. Full logs stay in ignored scratch. Applicable full library suites and workspace smoke checks, authoritative Heart verdict, then ship separately. Report counts without inventing results; CI must check Python 3.12 and 3.13. No tests have yet run for this bundle.
Original Prompt
Click to expand starting prompt
Final numerics audit of every mesh interpolator
Type: test
Target: autoarray
Repos:
PyAutoArray
autolens_workspace_test
Difficulty: large
Autonomy: supervised
Priority: high
Status: formalised
Filed: 2026-08-26
One last systematic round of numerics testing across every@PyAutoArray
mesh interpolator, to establish that no further bugs of the class PyAutoArray#490 exposed are still hiding.
Why now
#490 found that InterpolatorRectangular had mirrored bilinear row weights
for ~11 months, and that a 1-ULP change could flip a cell assignment because ix_up = ceil(g) collapsed on the CDF's clip plateau. Neither was caught by the
existing suite. The reason is the important part:
Partition of unity held throughout. Weights summed to 1 to 2.2e-16 in
every broken version, so any plausibility check passed.
The mirroring was consistent, so the likelihood surface stayed smooth
and every FD/AD gradient check passed.
Likelihood is nearly blind to it. With a pixelized source the inversion
solves for the source values, so a geometrically wrong mapping is still a
valid basis — the solved-for source absorbs the error.
It took a 1-ULP eager/jit divergence in a different repo to expose it. That is
not a reproducible detection strategy, which is why this audit is worth doing
deliberately rather than waiting for the next accident.
The tests that DO discriminate (validated in #490)
These caught every broken historical version and are the template:
Linear reproduction.sum_i w_i * node_i == query, in index space. A
correct bilinear scheme is exact for linear functions; any consistent
mis-pairing satisfies partition of unity but fails this. Measured: the
pre-fix: correct mirrored row weights + round-off-dependent cells in rectangular mapper #490 code failed by 0.999968 in the row axis while the column axis was
exactly 0.000000 — the asymmetry immediately localises the defect.
Convergence under refinement. Interpolate a smooth function and refine
the mesh; error must fall (~O(h²) for bilinear). Query where the data is,
not uniformly — an adaptive mesh deliberately leaves sparse regions coarse,
and judging it there measures the wrong thing.
Ground-truth source recovery. Where a truth source exists, compare the
reconstruction against it (Pearson r, normalised RMS). This is the only
measure that discriminated correct from mirrored on physical data.
Known-good control. Run the same harness against an interpolator not
under test (rectangular_uniform.py served this role) and require it to be bit-identical across the change — it proves the harness isolates what it
claims to.
For each: which of tests 1-5 apply (barycentric schemes reproduce linears too;
nearest-neighbour does not, so it needs a different invariant), then implement
and land the applicable ones as permanent regression tests.
KERNEL_CDF_DEFAULT_KNOTS = 64 does not scale with mesh_pixels.
Measured in fix: correct mirrored row weights + round-off-dependent cells in rectangular mapper #490: the knot-table inverse that places mesh nodes drifts from
the exact forward transform as the mesh refines — max |g_roundtrip - ix| was
0.0055 (n=16), 0.0238 (n=32), 0.101 (n=64) index units. Harmless at
production mesh sizes but wrong in principle; decide whether n_knots should
scale.
Guard-node conventions.grid_over_index ∈ [1, n-2] means flat rows 0-1
and cols 0, n-1 are never referenced. Confirm every interpolator agrees with
its mesh geometry about which nodes are live — a mismatch there is exactly the
kind of silent off-by-one this audit is for.
Acceptance
Each interpolator has permanent tests for whichever of 1-5 apply to it.
Overview
Add discriminating numerical regressions for every current interpolation family. Keep this an audit: record and file discovered defects separately rather than repairing production algorithms here.
Plan
Detailed implementation plan
Suggested branch:
feature/mesh-interpolator-numerics-auditClassification: Both; PyAutoArray primary, autolens_workspace_test companion. Large but bounded to a single audit; no dependency on another member. If investigation proves too large, remove from bundle before expanding.
Worktree root:
/home/jammy/Code/PyAutoLabs/.worktrees/autoarray-bundle-1, created once using Brain's worktree helper with PYAUTO_WT_ROOT inside this workspace. Shared repo worktrees: PyAutoArray, PyAutoGalaxy, autolens_workspace_test. Each member starts from origin/main on its own branch; execute and ship sequentially before switching the shared PyAutoArray worktree. One primary PyAutoArray issue and registry entry per member, linked companion PRs per repository as explicitly authorized. No merges.Branch survey: PyAutoArray, PyAutoGalaxy and autolens_workspace_test canonical checkouts are clean on main. No target repo claims in active.md. Heart reports an unregistered sparse-operator-oversampling-cache/PyAutoArray worktree with 11 dirty files: preserve it and obtain overlap acknowledgement before setup. Recent PyAutoArray branches: main, feature/sparse-operator-oversampling-cache, chore/session-start-hook-regen, feature/delaunay-area-magnification-audit, claude/autonerves-floor-regime-stamp. Recent PyAutoGalaxy branches: main, chore/session-start-hook-regen. Recent autolens_workspace_test branches: main, feature/point-audits-wheel-provenance, feature/point-solver-image-accuracy, feature/point-solver-duplicate-policy, chore/session-start-hook-regen.
Execution: one native Sol delegate per member, sequential within shared repositories. Parent owns judgment and lifecycle. Pass the approved issue plan, branch, exact starting commit, worktree, permitted files, validation requirements; stop and return exact failure evidence rather than weakening tests. Full logs stay in ignored scratch. Applicable full library suites and workspace smoke checks, authoritative Heart verdict, then ship separately. Report counts without inventing results; CI must check Python 3.12 and 3.13. No tests have yet run for this bundle.
Original Prompt
Click to expand starting prompt
Final numerics audit of every mesh interpolator
Type: test
Target: autoarray
Repos:
Difficulty: large
Autonomy: supervised
Priority: high
Status: formalised
Filed: 2026-08-26
One last systematic round of numerics testing across every
@PyAutoArraymesh interpolator, to establish that no further bugs of the class
PyAutoArray#490exposed are still hiding.Why now
#490 found that
InterpolatorRectangularhad mirrored bilinear row weightsfor ~11 months, and that a 1-ULP change could flip a cell assignment because
ix_up = ceil(g)collapsed on the CDF's clip plateau. Neither was caught by theexisting suite. The reason is the important part:
every broken version, so any plausibility check passed.
and every FD/AD gradient check passed.
solves for the source values, so a geometrically wrong mapping is still a
valid basis — the solved-for source absorbs the error.
It took a 1-ULP eager/jit divergence in a different repo to expose it. That is
not a reproducible detection strategy, which is why this audit is worth doing
deliberately rather than waiting for the next accident.
The tests that DO discriminate (validated in #490)
These caught every broken historical version and are the template:
sum_i w_i * node_i == query, in index space. Acorrect bilinear scheme is exact for linear functions; any consistent
mis-pairing satisfies partition of unity but fails this. Measured: the
pre-fix: correct mirrored row weights + round-off-dependent cells in rectangular mapper #490 code failed by 0.999968 in the row axis while the column axis was
exactly 0.000000 — the asymmetry immediately localises the defect.
interior cell boundary; the interpolant must not jump. Measured: pre-fix: correct mirrored row weights + round-off-dependent cells in rectangular mapper #490
jumped by 5.999 (a two-row flip).
the mesh; error must fall (~O(h²) for bilinear). Query where the data is,
not uniformly — an adaptive mesh deliberately leaves sparse regions coarse,
and judging it there measures the wrong thing.
reconstruction against it (Pearson r, normalised RMS). This is the only
measure that discriminated correct from mirrored on physical data.
under test (
rectangular_uniform.pyserved this role) and require it to bebit-identical across the change — it proves the harness isolates what it
claims to.
Scope — every interpolator
autoarray/inversion/mesh/interpolator/:rectangular.py—InterpolatorRectangular(done in fix: correct mirrored row weights + round-off-dependent cells in rectangular mapper #490; include it sothe suite is uniform)
rectangular_uniform.py—InterpolatorRectangularUniformdelaunay.py—InterpolatorDelaunay(note: the onlyjax.pure_callbackfamily in the chain — qhull triangulation)
delaunay_nn.py—InterpolatorDelaunayNNsibson.py— natural-neighbourknn.py/ KNN barycentric —InterpolatorKNearestNeighborFor each: which of tests 1-5 apply (barycentric schemes reproduce linears too;
nearest-neighbour does not, so it needs a different invariant), then implement
and land the applicable ones as permanent regression tests.
Specific suspicions worth checking first
floor/ceil/argmin/searchsortedon a value that can land exactlyon a boundary. fix: correct mirrored row weights + round-off-dependent cells in rectangular mapper #490's trigger was
clip(F_q, 0, 1)manufacturing exactlyinteger indices systematically. Look for other saturating transforms feeding
a discretisation.
defect was purely a pairing error, invisible to partition of unity.
KERNEL_CDF_DEFAULT_KNOTS = 64does not scale withmesh_pixels.Measured in fix: correct mirrored row weights + round-off-dependent cells in rectangular mapper #490: the knot-table inverse that places mesh nodes drifts from
the exact forward transform as the mesh refines — max
|g_roundtrip - ix|was0.0055 (n=16), 0.0238 (n=32), 0.101 (n=64) index units. Harmless at
production mesh sizes but wrong in principle; decide whether
n_knotsshouldscale.
grid_over_index ∈ [1, n-2]means flat rows 0-1and cols 0, n-1 are never referenced. Confirm every interpolator agrees with
its mesh geometry about which nodes are live — a mismatch there is exactly the
kind of silent off-by-one this audit is for.
Acceptance
variant — a regression test that passes either way is worthless, and that is
precisely how the fix: correct mirrored row weights + round-off-dependent cells in rectangular mapper #490 defect survived a commit literally named
"fixed mappings and weights".
Related
PyAutoArray#490— the fix and its regression tests (the template).autolens_workspace_test#279— how it surfaced.draft/test/workspaces/mesh_magnification_correctness.mddraft/test/workspaces/physical_model_check_when_speeding_up_smoke.md