Skip to content

test: Audit numerical correctness of every mesh interpolator #603

Description

@Jammy2211

Overview

Add discriminating numerical regressions for every current interpolation family. Keep this an audit: record and file discovered defects separately rather than repairing production algorithms here.

Plan

  • Map current interpolation families to their valid mathematical invariants.
  • Add boundary, reproduction and refinement regressions with deliberate-fault evidence.
  • Add deterministic synthetic truth recovery and an unchanged control.
  • Report defects separately and validate the library and workspace suites.
Detailed implementation plan

Suggested branch: feature/mesh-interpolator-numerics-audit

Classification: Both; PyAutoArray primary, autolens_workspace_test companion. Large but bounded to a single audit; no dependency on another member. If investigation proves too large, remove from bundle before expanding.

  1. Inspect autoarray/inversion/mesh/interpolator/{rectangular,rectangular_uniform,delaunay,sibson,knn}.py and existing test_autoarray/inversion/pixelization/interpolator tests. DelaunayNN now lives in sibson.py, not a separate delaunay_nn.py; include both InterpolatorKNearestNeighbor and InterpolatorKNNBarycentric.
  2. Add test_autoarray/inversion/pixelization/interpolator/test_numerics_audit.py (or extend existing family tests). Use independent coordinate/analytic oracles: rectangular linear reproduction in index space, barycentric/Sibson linear reproduction within their valid hull, constant and nearest-neighbor consistency where linear precision is not guaranteed. Document applicability of each of the prompt's five checks; do not impose continuity on inherently discontinuous nearest-neighbor changes.
  3. Sweep interior cell/simplex boundaries with exact and nextafter coordinates; assert mapping bounds, correct corner/weight pairing, live/guard node conventions and finite values. Measure smooth-function refinement on data-supported query sets, with deterministic tolerances and rates justified by each scheme.
  4. Measure CDF round-trip error at mesh sizes 16/32/64 and larger knot tables. Recommend scaling only if evidence warrants; production changes belong to separate bug work.
  5. Add a deterministic synthetic source-recovery audit under autolens_workspace_test/scripts/imaging/, mirroring a current rectangular fitting script. Measure Pearson correlation and normalized RMS against truth; pin rectangular_uniform control output bit-for-bit between clean and deliberately perturbed candidates. Avoid real observational data.
  6. Demonstrate each new regression fails with a relevant deliberately broken pairing, cell locator or weight mutation; use temporary monkeypatches or scratch variants, never commit corrupted source. Preserve baseline/fault results and applicability table in the issue evidence. Existing production defects remain visible; do not loosen assertions or disguise them with unconditional skips.
  7. Run focused interpolation tests, full test_autoarray, the new truth script and applicable workspace smoke checks. JAX parity/gradient checks, if needed, live in workspace scripts, not library unit tests. Library PR first; linked workspace PR follows its gate. No use of areas_transformed, preserving independence from the geometry repair.

Worktree root: /home/jammy/Code/PyAutoLabs/.worktrees/autoarray-bundle-1, created once using Brain's worktree helper with PYAUTO_WT_ROOT inside this workspace. Shared repo worktrees: PyAutoArray, PyAutoGalaxy, autolens_workspace_test. Each member starts from origin/main on its own branch; execute and ship sequentially before switching the shared PyAutoArray worktree. One primary PyAutoArray issue and registry entry per member, linked companion PRs per repository as explicitly authorized. No merges.

Branch survey: PyAutoArray, PyAutoGalaxy and autolens_workspace_test canonical checkouts are clean on main. No target repo claims in active.md. Heart reports an unregistered sparse-operator-oversampling-cache/PyAutoArray worktree with 11 dirty files: preserve it and obtain overlap acknowledgement before setup. Recent PyAutoArray branches: main, feature/sparse-operator-oversampling-cache, chore/session-start-hook-regen, feature/delaunay-area-magnification-audit, claude/autonerves-floor-regime-stamp. Recent PyAutoGalaxy branches: main, chore/session-start-hook-regen. Recent autolens_workspace_test branches: main, feature/point-audits-wheel-provenance, feature/point-solver-image-accuracy, feature/point-solver-duplicate-policy, chore/session-start-hook-regen.

Execution: one native Sol delegate per member, sequential within shared repositories. Parent owns judgment and lifecycle. Pass the approved issue plan, branch, exact starting commit, worktree, permitted files, validation requirements; stop and return exact failure evidence rather than weakening tests. Full logs stay in ignored scratch. Applicable full library suites and workspace smoke checks, authoritative Heart verdict, then ship separately. Report counts without inventing results; CI must check Python 3.12 and 3.13. No tests have yet run for this bundle.

Original Prompt

Click to expand starting prompt

Final numerics audit of every mesh interpolator

Type: test
Target: autoarray
Repos:

  • PyAutoArray
  • autolens_workspace_test
    Difficulty: large
    Autonomy: supervised
    Priority: high
    Status: formalised
    Filed: 2026-08-26

One last systematic round of numerics testing across every @PyAutoArray
mesh interpolator, to establish that no further bugs of the class
PyAutoArray#490 exposed are still hiding.

Why now

#490 found that InterpolatorRectangular had mirrored bilinear row weights
for ~11 months, and that a 1-ULP change could flip a cell assignment because
ix_up = ceil(g) collapsed on the CDF's clip plateau. Neither was caught by the
existing suite. The reason is the important part:

  • Partition of unity held throughout. Weights summed to 1 to 2.2e-16 in
    every broken version, so any plausibility check passed.
  • The mirroring was consistent, so the likelihood surface stayed smooth
    and every FD/AD gradient check passed.
  • Likelihood is nearly blind to it. With a pixelized source the inversion
    solves for the source values, so a geometrically wrong mapping is still a
    valid basis — the solved-for source absorbs the error.

It took a 1-ULP eager/jit divergence in a different repo to expose it. That is
not a reproducible detection strategy, which is why this audit is worth doing
deliberately rather than waiting for the next accident.

The tests that DO discriminate (validated in #490)

These caught every broken historical version and are the template:

  1. Linear reproduction. sum_i w_i * node_i == query, in index space. A
    correct bilinear scheme is exact for linear functions; any consistent
    mis-pairing satisfies partition of unity but fails this. Measured: the
    pre-fix: correct mirrored row weights + round-off-dependent cells in rectangular mapper #490 code failed by 0.999968 in the row axis while the column axis was
    exactly 0.000000 — the asymmetry immediately localises the defect.
  2. Continuity across integer boundaries. Sweep a coordinate across every
    interior cell boundary; the interpolant must not jump. Measured: pre-fix: correct mirrored row weights + round-off-dependent cells in rectangular mapper #490
    jumped by 5.999 (a two-row flip).
  3. Convergence under refinement. Interpolate a smooth function and refine
    the mesh; error must fall (~O(h²) for bilinear). Query where the data is,
    not uniformly — an adaptive mesh deliberately leaves sparse regions coarse,
    and judging it there measures the wrong thing.
  4. Ground-truth source recovery. Where a truth source exists, compare the
    reconstruction against it (Pearson r, normalised RMS). This is the only
    measure that discriminated correct from mirrored on physical data.
  5. Known-good control. Run the same harness against an interpolator not
    under test (rectangular_uniform.py served this role) and require it to be
    bit-identical across the change — it proves the harness isolates what it
    claims to.

Scope — every interpolator

autoarray/inversion/mesh/interpolator/:

  • rectangular.py — InterpolatorRectangular (done in fix: correct mirrored row weights + round-off-dependent cells in rectangular mapper #490; include it so
    the suite is uniform)
  • rectangular_uniform.py — InterpolatorRectangularUniform
  • delaunay.py — InterpolatorDelaunay (note: the only jax.pure_callback
    family in the chain — qhull triangulation)
  • delaunay_nn.py — InterpolatorDelaunayNN
  • sibson.py — natural-neighbour
  • knn.py / KNN barycentric — InterpolatorKNearestNeighbor

For each: which of tests 1-5 apply (barycentric schemes reproduce linears too;
nearest-neighbour does not, so it needs a different invariant), then implement
and land the applicable ones as permanent regression tests.

Specific suspicions worth checking first

Acceptance

  • Each interpolator has permanent tests for whichever of 1-5 apply to it.
  • Every new test is demonstrated to fail against a deliberately-broken
    variant — a regression test that passes either way is worthless, and that is
    precisely how the fix: correct mirrored row weights + round-off-dependent cells in rectangular mapper #490 defect survived a commit literally named
    "fixed mappings and weights".
  • Any bug found is filed separately; this task is the audit, not the fixes.

Related

  • PyAutoArray#490 — the fix and its regression tests (the template).
  • autolens_workspace_test#279 — how it surfaced.
  • draft/test/workspaces/mesh_magnification_correctness.md
  • draft/test/workspaces/physical_model_check_when_speeding_up_smoke.md

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions