Skip to content

feat(interferometer): array-free Interferometer.from_stream / from_sparse_terms (streaming phase 1, Discussion #13) - #593

Merged
Jammy2211 merged 1 commit into
mainfrom
feature/streaming-p1-array-free-dataset
Sep 30, 2026
Merged

Jammy2211 merged 1 commit into
mainfrom
feature/streaming-p1-array-free-dataset

Conversation

@Jammy2211

Copy link
Copy Markdown
Collaborator

Summary

Streaming phase 1 (Mind epic streaming-visibilities; phase 2 of https://github.com/orgs/PyAutoLabs/discussions/13, HRSAstro). Closes #592.

After #589 / PyAutoLabs/PyAutoGalaxy#637 / PyAutoLabs/PyAutoLens#757 the sparse likelihood no longer touches visibilities, but the process still held them: Interferometer.__init__ always built a TransformerNUFFT from uv_wavelengths (its own ~48 B/vis copy on top of the dataset's 48 B/vis), AbstractDataset.__init__ rejected noise_map=None, and apply_sparse_operator_from_chunks returned a dataset retaining every array. This PR adds the array-free dataset:

  • Interferometer.from_stream(chunks, real_space_mask, *, transformer_class=TransformerNUFFT, method, eps, chunk_size, chunk_k, use_jax, show_progress, batch_size) accumulates (uv_wavelengths, data, noise_map) chunks into SparseTerms and builds a dataset whose data, noise_map, uv_wavelengths and transformer are all None; from_sparse_terms(terms, real_space_mask) does the same from an existing record. is_array_free says which kind a dataset is; sparse_terms is stored (also by apply_sparse_operator_from_chunks).
  • SparseTerms carries provenance (shape_native, pixel_scales, origin, eps, transformer_class_name); __add__ refuses to sum terms from different geometries/accuracies; from_sparse_terms refuses a mask of the wrong shape.
  • AbstractInversionInterferometer.mask reads the dataset's real-space mask instead of transformer.real_space_mask (the one transformer read on the sparse likelihood path).
  • Every quantity that needs the visibilities raises a typed exc.DatasetException naming the missing input and from_stream (dataset: amplitudes, phases, uv_distances, dirty_image, dirty_noise_map, signal_to_noise_map, dirty_signal_to_noise_map, psf_precision_operator_from, apply_sparse_operator; fit: mask, transformer, residual_map, normalized_residual_map, chi_squared_map, signal_to_noise_map, chi_squared, every dirty_*); mapped_reconstructed_operated_data_dict raises exc.InversionException. log_evidence / figure_of_merit of a sparse inversion fit never reach any of them (spy-tested).
  • New dirty_image_natural (Re(Fᴴ w d)/Σw) and dirty_beam (Re(Fᴴ w)/Σw) on both dataset kinds (from the terms when present, else from the arrays; agree to rel 1e-12). dirty_image (unweighted adjoint) is unchanged.

Design decisions (recorded in the Mind epic ledger): transformer=None with honest gating rather than a stub transformer; natural-weighted images under new names so in-memory behaviour is untouched; provenance on SparseTerms. Phases 2–5 (fit + save/reload, visualizer, non-linear light profiles via the data-term identity, cubes/phase centre) follow one at a time.

Parity: rectangular-mesh sparse inversion on the array-free dataset vs the in-memory apply_sparse_operator dataset — fast_chi_squared, regularization and log-det terms, reconstruction, log_evidence all at rel 1e-8 (numpy and jax); on the 5e5-visibility witness log_evidence −1577804.7762316698 streamed vs −1577804.7762316696 in memory (rel 1.5e-16).

Memory witness (fresh process per N, NUFFT, 400×400 circular mask at 0.05"/pix = 125,676 pixels, 4096-vis chunks read from an npz one chunk at a time; baselines: import-only 171 MB, one-chunk stream 536 MB):

N_vis peak RSS over import would be held at 48 B/vis wall
5e5 821 MB 650 MB 24 MB 25 s
1e6 776 MB 605 MB 48 MB 49 s
2e6 831 MB 660 MB 96 MB 153 s
4e6 929 MB 758 MB 192 MB 454 s

Peak RSS is set by the image size (the NUFFT working set for a 125k-pixel mask), not by N_vis: it moves by ~100 MB over an 8× range where the held arrays would add 168 MB (and the transformer's copy as much again). The discussion's 36 MB figure is "over baseline" for pyuvimage's accumulator on its own image; the absolute level here is the per-chunk NUFFT working set. Not in this phase (by design, review finding #4): an AnalysisInterferometer search on an array-free dataset does not yet run end to end — save_attributes reads data.in_array and the aggregator reads the visibility HDUs; both are phase 2 (streaming_p2_fit_save_reload). Phase 1 supports the fit / inversion / log_evidence only. Two things not chased in this phase: the mild +100 MB drift at 4e6 and the super-linear wall time (2× N ≈ 3× time above 1e6 — most likely the witness's per-chunk npz reads, not the accumulator); both are noted on the epic ledger for phase 2.

Heart RED override (development only)

Heart verdict at ship: RED (2026-09-30T12:42Z), exact reason: release validation FAILED (stage integrate) — a release-integrate leg unrelated to this branch. The live human authorized the development-only override for issue #592 in-session on 2026-09-30 ("Authorize override for #592"): commit, push and this pending-release PR only. Branch gates passed before the ask: test_autoarray 1783 passed; test_autogalaxy/interferometer 43 and test_autolens/interferometer 29 passed unchanged; red-checks on the new tests; independent Codex (gpt-6-astra) review: FINDINGS (4) — #1 geometry guard in from_sparse_terms (pixel_scales/origin), #2 origin in SparseTerms.__add__, #3 provenance merge when one operand is unrecorded: all three fixed in-branch with red-checked tests; #4 (AnalysisInterferometer.save_attributes / aggregator cannot handle an array-free dataset) is phase 2 by design, documented (full record in the comment below). This PR does not claim to repair Heart; Heart remains RED for release purposes and merge needs its own explicit human command with every required check green.

API Changes

Additive plus one behaviour change. Added Interferometer.from_stream, Interferometer.from_sparse_terms, Interferometer.is_array_free, Interferometer.sparse_terms (ctor kwarg + attribute), Interferometer.dirty_image_natural, Interferometer.dirty_beam; SparseTerms gains five optional trailing provenance fields and __add__ raises on mismatch. Interferometer accepts uv_wavelengths=None / transformer_class=None (transformer becomes None); AbstractDataset accepts noise_map=None when data is also None; AbstractDataset.shape_slim is None on such a dataset. AbstractInversionInterferometer.mask now comes from the dataset (same value everywhere a transformer exists). Array/transformer-dependent properties raise exc.DatasetException / exc.InversionException instead of AttributeError when the input is absent. In-memory datasets are byte-identical in behaviour.
See full details below.

Test Plan

  • pytest test_autoarray — 1783 passed (1779 before the review fixes)
  • Regression, no edits: test_autogalaxy/interferometer 43 passed; test_autolens/interferometer 29 passed
  • New tests: construction (arrays/transformer None, terms + provenance + data_term set, mask-shape guard); parity vs in-memory at rel 1e-8 numpy + jax; typed exceptions on every dataset/fit property and on mapped_reconstructed_operated_data_dict; spy that log_evidence never touches visibility properties; provenance sum / mismatch raise; apply_sparse_operator_from_chunks retains sparse_terms; mask regression on Interferometer / DatasetInterface with and without transformer; dirty_image_natural / dirty_beam terms-vs-arrays agree (NUFFT + DFT)
  • Red check: mask reroute reverted → 3 fail (AttributeError: 'NoneType' object has no attribute 'real_space_mask'); one guard removed → typed-exception test fails; provenance check disabled → 2 fail
  • Witness script (table above)
  • CI green on unittest 3.12 / 3.13 / nojax
Full API Changes (for automation & release notes)

Added

  • Interferometer.from_stream(chunks, real_space_mask, *, transformer_class=TransformerNUFFT, method="nufft", eps=None, chunk_size=None, chunk_k=2048, use_jax=False, show_progress=False, batch_size=128) -> Interferometer — array-free dataset accumulated from (uv_wavelengths, data, noise_map) chunks
  • Interferometer.from_sparse_terms(terms, real_space_mask, *, batch_size=128) -> Interferometer
  • Interferometer(..., sparse_terms=None) kwarg and .sparse_terms attribute (also set by apply_sparse_operator_from_chunks)
  • Interferometer.is_array_free -> bool
  • Interferometer.dirty_image_natural -> Array2D, Interferometer.dirty_beam -> Array2D — natural-weighted, normalised; both dataset kinds
  • SparseTerms.shape_native, .pixel_scales, .origin, .eps, .transformer_class_name (optional trailing fields, filled by sparse_terms_from_chunks; __add__ keeps the recorded value when one operand is unrecorded)

Changed Signature

  • Interferometer(uv_wavelengths=None, transformer_class=None, ...) permitted (array-free); AbstractDataset(noise_map=None) permitted when data is None

Changed Behaviour

  • AbstractInversionInterferometer.mask — dataset's real_space_mask, else transformer's, else dataset.mask (was transformer only)
  • SparseTerms.__add__ — raises exc.InversionException when both sides record different shape_native / pixel_scales / origin / eps; Interferometer.from_sparse_terms raises exc.DatasetException when the mask's shape / pixel_scales / origin differ from the recorded ones
  • Interferometer properties needing absent arrays/transformer raise exc.DatasetException; InversionInterferometerSparse.mapped_reconstructed_operated_data_dict raises exc.InversionException without a transformer; aa.FitInterferometer maps / mask / transformer / chi_squared raise exc.DatasetException on an array-free dataset
  • AbstractDataset.shape_slim returns None when data is None

Migration

  • None for in-memory datasets. Code that must work on both kinds should branch on dataset.is_array_free.

Generated by the PyAutoLabs agent workflow.

🤖 Generated with Claude Code

https://claude.ai/code/session_01JZZksyZ8LTA4LLxoZjQMNF

…arse_terms (streaming phase 1, #592)

Build an Interferometer with no data, noise_map, uv_wavelengths or transformer
from accumulated SparseTerms, so a sparse (w-tilde) pixelized inversion and its
log_evidence run with peak memory set by the image size, not the number of
visibilities. SparseTerms carries provenance and refuses mismatched sums; the
inversion base reads its mask from the dataset; every array-dependent property
raises a typed DatasetException; dirty_image_natural / dirty_beam exist on both
dataset kinds. In-memory behaviour is unchanged.

Discussion https://github.com/orgs/PyAutoLabs/discussions/13, phase 2 of the
proposal; Mind epic streaming-visibilities, phase 1 of 5.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JZZksyZ8LTA4LLxoZjQMNF
@Jammy2211 Jammy2211 added the pending-release PR queued for the next release build label Sep 30, 2026
@Jammy2211

Copy link
Copy Markdown
Collaborator Author

Independent review (Codex gpt-6-astra) of this branch before PR-open

  1. P1 — Wrong geometry accepted by from_sparse_terms (in-diff). dataset.py:284 checks only native shape. Terms accumulated at pixel scale 0.5 can be loaded with an identically shaped mask at 1.0, or with a different origin. The grids then describe different coordinates while the cached dirty image and precision operator retain the original geometry, silently corrupting reconstruction and evidence. Changing which pixels are masked can also expose zeros previously inserted into dirty_image_native. The new mismatch test checks only different shapes.

  2. P1 — Addition accepts incompatible origins (pre-existing unsafe addition; incomplete in-diff guard). inversion_interferometer_util.py:1883 records origin but excludes it from validation. Accumulate two DFT records on identical masks with different origins: their dirty images sample different sky coordinates, yet addition sums corresponding pixels and labels the result with the left origin. The resulting data vector is incorrect. Different mask patterns with identical native shape and bounding extent likewise pass. Neither case is covered by the new provenance tests.

  3. P2 — Unknown left provenance erases known constraints (in-diff). inversion_interferometer_util.py:1914 copies provenance exclusively from the left. Let U have unknown scales, A record (0.5, 0.5), and B record (1.0, 1.0), with matching array shapes. (U + A) + B succeeds because the first addition discards A’s known scales; (A + U) + B rejects the same incompatible contributions. Allowing None to skip an individual comparison should not erase known constraints for subsequent additions. An isolated execution of the actual class confirmed the metadata loss. The new test exercises only known-left/unknown-right.

  4. P1 — Array-free datasets cannot complete the standard analysis/search lifecycle (new input exposes unchanged consumer assumptions). The constructor returns data=None at dataset.py:300, but both Galaxy analysis.py:253 and Lens analysis.py:357 unconditionally access data.in_array during saving, raising AttributeError even with plotting disabled. Enabled dataset visualization also dereferences data.in_grid; the aggregator requires visibility/noise/UV FITS extensions and cannot reload sparse terms. This affects even an otherwise supported pixelization-only fit. The new fit tests inject an already-built inversion into MockFitInterferometer, so they pass despite this lifecycle failure.

FINDINGS (4)

Disposition: #1, #2, #3 fixed in this branch (tests test__from_sparse_terms__mask_pixel_scales_and_origin_mismatch__raises, test__sparse_terms__add__origin_mismatch_raises, test__sparse_terms__add__unrecorded_left_operand_keeps_right_provenance, each red with its fix reverted). #4 is the save/reload lifecycle, deliberately phase 2 (streaming_p2_fit_save_reload); the PR body says so.

🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

pending-release PR queued for the next release build

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(interferometer): array-free Interferometer.from_stream / from_sparse_terms (streaming phase 1, Discussion #13)

1 participant