diff --git a/docs/README.md b/docs/README.md index e26712f..fa08956 100644 --- a/docs/README.md +++ b/docs/README.md @@ -49,8 +49,9 @@ Start with these documents when extending the engine: ## Analysis and visualization - [Analysis recipes](analysis/recipes.md) covers lazy Polars workflows for colony geometry, species, lineage, contact graphs, and signal fields. +- [Species and signal labels](models/channel-labels.md) describes native and SBML channel metadata. - [Viewer guide](../viewer/README.md) covers static scenes, interactive sessions, controls, development, and tests. -- [Scene format v2](formats/scene-v2.md) defines the data exchanged with visualization clients. +- [Scene format v3](formats/scene-v3.md) defines the data exchanged with visualization clients. - [Live viewer protocol v1](protocols/live-viewer-v1.md) defines the authenticated loopback protocol for interactive sessions. ## Execution environments @@ -66,7 +67,7 @@ The [testing and validation guide](development/validation.md) distinguishes comp ## Formats and protocols - [Run manifest v1](formats/run-manifest-v1.md) defines reproducible batch jobs and parameter sweeps. -- [Scene format v2](formats/scene-v2.md) defines data-only visualization frames. +- [Scene format v3](formats/scene-v3.md) defines data-only visualization frames. - [Live viewer protocol v1](protocols/live-viewer-v1.md) defines interactive viewer messages and authority boundaries. - [Checkpoint design](architecture/0004-checkpoints.md) defines restart state and schema migration. - [Analysis dataset design](architecture/0013-analysis-datasets.md) defines Parquet/Zarr schemas and provenance. diff --git a/docs/architecture/0004-checkpoints.md b/docs/architecture/0004-checkpoints.md index 90e1e6f..c8e5a61 100644 --- a/docs/architecture/0004-checkpoints.md +++ b/docs/architecture/0004-checkpoints.md @@ -21,7 +21,9 @@ The public checkpoint is UTF-8 JSON with the format identifier `microsimulator-c - producer, source-backend, and caller-supplied provenance; and - a SHA-256 digest of the canonical simulation payload. -Version 2 additionally records an optional validated signal-grid specification and its complete signal-major concentration field. Version 3 adds the typed coupled cell/grid rate plan. Version 4 adds an optional data-only controller payload with its own SHA-256 digest. Native checkpoints write a JSON `null` controller. A non-null controller cannot be silently discarded by `load_checkpoint`; callers use `load_checkpoint_bundle` and restore it with the matching controller. Version 5 records the signal integration kind and its iterative-solver parameters. Version 6 records whether each rod cell is fixed in mechanics. Version 7 adds an optional spatial affine source/loss field to the signal-grid specification. Writers emit only v7; readers explicitly migrate v1 through v6, using Forward Euler defaults for older signal grids, movable cells for checkpoints predating v6, and no affine field reaction for checkpoints predating v7. +Version 2 additionally records an optional validated signal-grid specification and its complete signal-major concentration field. Version 3 adds the typed coupled cell/grid rate plan. Version 4 adds an optional data-only controller payload with its own SHA-256 digest. Native checkpoints write a JSON `null` controller. A non-null controller cannot be silently discarded by `load_checkpoint`; callers use `load_checkpoint_bundle` and restore it with the matching controller. Version 5 records the signal integration kind and its iterative-solver parameters. Version 6 records whether each rod cell is fixed in mechanics. Version 7 adds an optional spatial affine source/loss field to the signal-grid specification. Version 8 adds boxes and cylinders to the constraint set. Version 9 adds required `channel_metadata` with ordered `species` and `signals` arrays and a separate `integrity.channel_metadata` SHA-256 digest using the same canonical JSON encoding as controller state. Entries are strings or null, and lengths must match the native channel counts. Writers emit only v9; readers explicitly migrate v1 through v8, using Forward Euler defaults for older signal grids, movable cells for checkpoints predating v6, no affine field reaction for checkpoints predating v7, and unnamed channel labels for checkpoints predating v9. Each old payload is verified with its original integrity rules before supplying defaults. `load_checkpoint_bundle` exposes labels without executing model code; `load_checkpoint` refuses named metadata it would otherwise discard. See the [channel authoring guide](../models/channel-labels.md). + +For v1–v8, absent metadata remains compact as `ChannelMetadata(species=None, signals=None)` in the returned bundle. Migration does not allocate null arrays from claimed native channel counts before native restoration validates the checkpoint. Scene export resolves these unspecified groups only after enforcing its separate 4096-channel presentation budget per group. Native counts are not capped by that scene budget, and v9 checkpoint metadata retains its explicit count-matched arrays. Files are written to a temporary sibling, flushed, and atomically replaced. Loading rejects duplicate JSON keys, non-finite numbers, unknown fields for the declared version, unsupported versions, oversized files, digest mismatches, and any state that fails native domain validation. No module is imported and no source text, callback, pickle opcode, or other executable representation is accepted. diff --git a/docs/formats/scene-v2.md b/docs/formats/scene-v2.md index 6530e27..97a916a 100644 --- a/docs/formats/scene-v2.md +++ b/docs/formats/scene-v2.md @@ -64,4 +64,4 @@ The scene preserves all channels. A viewer chooses a channel and slice as presen ## Compatibility -Writers always emit the current version. Readers accept version 2 exactly and fail closed on other versions until an explicit migration is defined. Backend conformance compares frame semantics while ignoring the expected backend identity fields. Pixel output is tested separately by the viewer. +Current writers emit [version 3](scene-v3.md). Current readers verify version-2 frames against this original schema and digest, enforce the presentation budget of 4096 species and 4096 signals independently, then supply unnamed channel metadata in memory. The budget applies even to empty colonies and prevents a tiny document's claimed count from causing unbounded label allocation. It does not change native simulation or checkpoint channel limits. Backend conformance compares frame semantics while ignoring the expected backend identity fields. Pixel output is tested separately by the viewer. diff --git a/docs/formats/scene-v3.md b/docs/formats/scene-v3.md new file mode 100644 index 0000000..66dd18c --- /dev/null +++ b/docs/formats/scene-v3.md @@ -0,0 +1,20 @@ +# MicroSimulator scene format v3 + +Version 3 retains the [version 2 envelope, geometry, constraints and grid representation](scene-v2.md) and adds one required field inside the integrity-protected frame: + +```json +"channel_metadata": { + "species": ["Green reporter", "Red reporter"], + "signals": ["Nutrient", null] +} +``` + +Both groups are required arrays. `species` has exactly `species_count` entries; `signals` has exactly `signal_grid.signal_count` entries, or zero when the grid is null. Each entry is a Unicode scalar string or null. Null is the canonical serialized representation of an unspecified slot; unspecified groups are expanded to null-filled arrays. Empty and whitespace-only strings are retained verbatim and use the same display fallback as null. Duplicate names are valid. Unknown metadata fields and invalid lengths or types are errors. + +Scene presentation has a channel-count budget of **4096 species and 4096 signals independently**, inclusive (`MAX_SCENE_CHANNELS` in Python and TypeScript). This budget applies to both v2 and v3, including empty colonies. Readers check each claimed count before expanding missing labels, copying channel data, or constructing viewer controls; an oversized count raises a scene-format error identifying the count and budget. Python applies the same limit to `SceneFrame` construction, `capture_scene`, parsing/loading, and encoding/saving. The existing encoded-size and grid-shape checks remain separate. This presentation budget does not limit native simulation counts or checkpoint restoration, and exporters reject oversized scenes rather than silently truncating channels. + +The entire frame, including channel metadata, is hashed with RFC 8785 canonical JSON and SHA-256. Digests detect corruption; they do not authenticate a publisher. Labels must be rendered as text, never interpreted as HTML or executable code. + +Readers accept versions 2 and 3. They verify a version-2 frame's original digest and exact version-2 keys first, then return the current in-memory representation with null-filled channel arrays. A version-2 file containing a `channel_metadata` field is invalid, even with a matching digest. Readers reject all other versions. Writers emit only version 3. + +Python `SceneFrame.channel_metadata` and TypeScript `SceneFrame.channelMetadata` expose the same ordered values. TypeScript `channelLabel(frame, "species" | "signals", index)` provides missing-label fallback and duplicate-name disambiguation. Indices identify channels; display names never identify settings or alter stored numerical values. See the [authoring guide](../models/channel-labels.md) for native, low-level and SBML examples. diff --git a/docs/models/channel-labels.md b/docs/models/channel-labels.md new file mode 100644 index 0000000..e0181d1 --- /dev/null +++ b/docs/models/channel-labels.md @@ -0,0 +1,82 @@ +# Species and signal labels + +Numerical channel indices determine rate-plan inputs, storage order, and viewer preferences. Labels only describe those indices. They are never inferred from Python variable names. + +## Native models + +Pass immutable `ChannelMetadata` to `NativeController` after configuring the simulation's species count and signal grid: + +```python +from microsimulator import ChannelMetadata, NativeController + +return NativeController( + simulation, + model_id="my-model", + model_version=1, + rng=context.rng, + channel_metadata=ChannelMetadata( + species=("Green reporter", "Red reporter"), + signals=("Nutrient", "Extracellular cue"), + ), +) +``` + +Each supplied tuple must contain exactly one entry per corresponding numerical channel. Use `None` for an unnamed entry, or omit a whole group to leave all its channels unnamed. The constructor validates counts immediately; the runner checks again before the first step, and exporters validate against the current state. Empty strings and whitespace-only strings display the same fallback as `None`, while retaining their exact supplied text in files. Labels must be Unicode scalar strings; unpaired surrogates are rejected. Presentation collapses and trims ASCII whitespace like an HTML option label before checking for duplicates; serialized metadata retains the original text. Duplicate names are valid and display their numerical indices for disambiguation. If a supplied name imitates one of those generated labels (for example, `GFP`, `GFP`, and `GFP [0]`), all labels in that species or signal group receive their indices so every displayed name remains distinct. Labels, including HTML-like strings, render as text. + +The complete [named-channel model](../../examples/named_channels.py) declares two species and two signals: + +```sh +uv run microsimulator run --model examples/named_channels.py --backend cpu --seed 17 --steps 2 --dt 0.01 --output named.json +uv run microsimulator view --model examples/named_channels.py --resume named.json --backend cpu --dt 0.01 +``` + +`NativeController.from_checkpoint` restores the persisted labels automatically. A custom controller may optionally expose a typed `channel_metadata: ChannelMetadata` attribute; this is not a required member of `SimulationController`. Its `resume` function must restore `checkpoint.channel_metadata`. The native model runner rejects a resumed model that changes the saved labels. Unnamed native models and legacy adapters need no changes. + +## Data-only export and low-level APIs + +Labels are stored in checkpoint version 9 independently of the controller payload, so recovering them never requires running model code. Scene version 3 carries the same ordered arrays. Use the bundle's labels explicitly when exporting or saving native state directly: + +```python +from microsimulator import capture_scene, load_checkpoint_bundle, save_checkpoint, save_scene + +bundle = load_checkpoint_bundle("named.json") +frame = capture_scene(bundle.simulation, channel_metadata=bundle.channel_metadata) +save_scene(frame, "named.scene.json") +save_checkpoint( + bundle.simulation, + "copy.json", + provenance=bundle.provenance, + controller=bundle.controller, + channel_metadata=bundle.channel_metadata, +) +``` + +`capture_scene` and `save_checkpoint` also accept `channel_metadata` for a bare `Simulation` when no controller is needed. A bare native simulation does not own Python presentation metadata. Therefore `load_checkpoint` refuses a file with non-null labels, just as it refuses a non-null controller payload: use `load_checkpoint_bundle` to avoid silently losing labels. Such a named, bare-native checkpoint is exported or continued through the bundle API; `run --resume` without a model retains its existing unnamed-only contract. Standard named models use the controller resume command above. + +Scenes support at most 4096 species and 4096 signals per frame, independently, including unnamed channels and empty colonies. `MAX_SCENE_CHANNELS` exposes this presentation budget; oversized export fails with `SceneError` before copying native state or expanding labels. Native simulation and checkpoint counts retain their existing semantics. When loading checkpoints predating v9, `CheckpointBundle.channel_metadata` keeps both unspecified groups as `None`, without allocating labels from native counts. Bounded scene export supplies the null-filled arrays; callers that explicitly need resolved metadata can use `.resolved(species_count, signal_count)`. + +Within a live dataset, channel choices stay keyed by kind and index. Renaming a channel does not change its concentration or select another channel. Frames, reset, and replay retain the same labels through the shared scene parser. Opening another dataset establishes a new presentation identity. + +## SBML labels + +`SBMLRateModel.channel_metadata` explicitly maps nonempty species names to labels and falls back to SBML species identifiers when names are missing. It preserves the imported species order: + +```python +from microsimulator import CellInit, NativeController, load_sbml + +rates = load_sbml("model.xml") +simulation = context.simulation(species_count=rates.species_count) +simulation.set_species_rate_plan(rates.rate_plan) +cell = CellInit() +cell.species = list(rates.initial_levels) +simulation.add_cell(cell) +return NativeController( + simulation, + model_id="my-sbml-model", + model_version=1, + rng=context.rng, + channel_metadata=rates.channel_metadata, +) +``` + +SBML species identifiers remain authoritative for compilation. If the simulation also has extracellular signals, declare those explicitly with `ChannelMetadata(species=rates.channel_metadata.species, signals=(...))`. The SBML importer does not infer extracellular signal identities. diff --git a/examples/named_channels.py b/examples/named_channels.py new file mode 100644 index 0000000..830bc8e --- /dev/null +++ b/examples/named_channels.py @@ -0,0 +1,43 @@ +"""Two intracellular reporters and two extracellular signals with explicit labels.""" + +from microsimulator import ( + CellInit, + ChannelMetadata, + CheckpointBundle, + GridShape, + ModelContext, + NativeController, + SignalGridSpec, + Vec3, +) + + +def build(context: ModelContext) -> NativeController: + simulation = context.simulation(species_count=2) + grid = SignalGridSpec() + grid.signal_count = 2 + shape = GridShape() + shape.x, shape.y, shape.z = 4, 4, 1 + grid.shape = shape + grid.spacing = Vec3(1.0, 1.0, 1.0) + grid.diffusion = [0.1, 0.2] + grid.advection = [Vec3(), Vec3()] + simulation.configure_signal_grid(grid, [0.25] * 16 + [0.75] * 16) + cell = CellInit() + cell.length = 2.0 + cell.species = [0.25, 0.75] + simulation.add_cell(cell) + return NativeController( + simulation, + model_id="named-channels", + model_version=1, + rng=context.rng, + channel_metadata=ChannelMetadata( + species=("Green reporter", "Red reporter"), + signals=("Nutrient", "Extracellular cue"), + ), + ) + + +def resume(context: ModelContext, checkpoint: CheckpointBundle) -> NativeController: + return NativeController.from_checkpoint(checkpoint, model_id="named-channels", model_version=1) diff --git a/python/src/microsimulator/__init__.py b/python/src/microsimulator/__init__.py index d41483a..96eef75 100644 --- a/python/src/microsimulator/__init__.py +++ b/python/src/microsimulator/__init__.py @@ -53,6 +53,7 @@ backend_available, backend_device_count, ) +from .channels import ChannelMetadata, ChannelMetadataError from .checkpoint import ( CHECKPOINT_FORMAT, CHECKPOINT_VERSION, @@ -120,6 +121,7 @@ from .sbml import SBMLImportError, SBMLRateModel, load_sbml, parse_sbml from .scene import ( MAX_SCENE_BYTES, + MAX_SCENE_CHANNELS, SCENE_FORMAT, SCENE_VERSION, SceneBackend, @@ -148,6 +150,7 @@ "MAX_LEGACY_EXAMPLE_MATRIX_BYTES", "MAX_RUN_MANIFEST_BYTES", "MAX_SCENE_BYTES", + "MAX_SCENE_CHANNELS", "RUN_MANIFEST_FORMAT", "RUN_MANIFEST_VERSION", "SCENE_FORMAT", @@ -163,6 +166,8 @@ "CellInit", "CellSnapshot", "CellUpdate", + "ChannelMetadata", + "ChannelMetadataError", "CheckpointBundle", "CheckpointError", "CheckpointSourceBackend", diff --git a/python/src/microsimulator/analysis.py b/python/src/microsimulator/analysis.py index bfa0808..b1edee7 100644 --- a/python/src/microsimulator/analysis.py +++ b/python/src/microsimulator/analysis.py @@ -469,7 +469,7 @@ def _load_sources( for index, value in enumerate(checkpoints): path = Path(value) bundle = load_checkpoint_bundle(path, backend=backend, device_index=device_index) - scene = capture_scene(bundle.simulation) + scene = capture_scene(bundle.simulation, channel_metadata=bundle.channel_metadata) if previous_time is not None and scene.time < previous_time: raise AnalysisError( f"checkpoint {path} has time {scene.time:.9g}, before prior time " diff --git a/python/src/microsimulator/channels.py b/python/src/microsimulator/channels.py new file mode 100644 index 0000000..08aece4 --- /dev/null +++ b/python/src/microsimulator/channels.py @@ -0,0 +1,85 @@ +"""Optional presentation labels; numerical channel indices remain authoritative.""" + +from __future__ import annotations + +from dataclasses import dataclass +from typing import TYPE_CHECKING, cast + +if TYPE_CHECKING: + from .checkpoint import JSONValue + + +class ChannelMetadataError(ValueError): + """Raised when labels cannot describe the declared numerical channels.""" + + +@dataclass(frozen=True, slots=True) +class ChannelMetadata: + """Ordered labels, with ``None`` for an unnamed group or individual channel. + + Explicit tuples must match channel counts exactly. Empty/whitespace-only + strings are treated as missing labels for display, but preserved on disk. + """ + + species: tuple[str | None, ...] | None = None + signals: tuple[str | None, ...] | None = None + + def __post_init__(self) -> None: + for kind in ("species", "signals"): + labels = cast(object, getattr(self, kind)) + if labels is None: + continue + if not isinstance(labels, tuple): + raise ChannelMetadataError(f"channel_metadata.{kind}: expected a tuple or None") + for index, label in enumerate(cast(tuple[object, ...], labels)): + if label is not None and not isinstance(label, str): + raise ChannelMetadataError( + f"channel_metadata.{kind}[{index}]: expected a string or None" + ) + if isinstance(label, str): + try: + label.encode("utf-8") + except UnicodeEncodeError as error: + raise ChannelMetadataError( + f"channel_metadata.{kind}[{index}]: invalid Unicode scalar value" + ) from error + + def resolved(self, species_count: int, signal_count: int) -> ChannelMetadata: + """Validate explicit counts and expand unspecified groups to null labels.""" + + groups: list[tuple[str | None, ...]] = [] + for kind, labels, count in ( + ("species", self.species, species_count), + ("signals", self.signals, signal_count), + ): + if labels is not None and len(labels) != count: + raise ChannelMetadataError( + f"channel_metadata.{kind}: expected {count} labels, got {len(labels)}" + ) + groups.append((None,) * count if labels is None else labels) + return ChannelMetadata(species=groups[0], signals=groups[1]) + + def to_json(self, species_count: int, signal_count: int) -> dict[str, JSONValue]: + labels = self.resolved(species_count, signal_count) + return {"species": list(labels.species or ()), "signals": list(labels.signals or ())} + + @classmethod + def from_json(cls, value: object, species_count: int, signal_count: int) -> ChannelMetadata: + """Decode the closed data-only representation shared by scenes/checkpoints.""" + + if not isinstance(value, dict) or set(cast(dict[object, object], value)) != { + "species", + "signals", + }: + raise ChannelMetadataError("channel_metadata: expected exactly species and signals") + data = cast(dict[str, object], value) + groups: list[tuple[str | None, ...]] = [] + for kind in ("species", "signals"): + values = data[kind] + if not isinstance(values, list): + raise ChannelMetadataError(f"channel_metadata.{kind}: expected an array") + groups.append(cast(tuple[str | None, ...], tuple(cast(list[object], values)))) + return cls(species=groups[0], signals=groups[1]).resolved(species_count, signal_count) + + +UNNAMED_CHANNELS = ChannelMetadata() diff --git a/python/src/microsimulator/checkpoint.py b/python/src/microsimulator/checkpoint.py index fa87c6c..ad29ccc 100644 --- a/python/src/microsimulator/checkpoint.py +++ b/python/src/microsimulator/checkpoint.py @@ -43,9 +43,10 @@ _SphereConstraint, _WorldStateCheckpoint, ) +from .channels import UNNAMED_CHANNELS, ChannelMetadata, ChannelMetadataError CHECKPOINT_FORMAT = "microsimulator-checkpoint" -CHECKPOINT_VERSION = 8 +CHECKPOINT_VERSION = 9 MAX_CHECKPOINT_BYTES = 1 << 30 _NATIVE_CHECKPOINT_VERSION = 4 @@ -83,6 +84,7 @@ class CheckpointBundle: provenance: dict[str, JSONValue] schema_version: int source_backend: CheckpointSourceBackend + channel_metadata: ChannelMetadata = UNNAMED_CHANNELS _RATE_OP_NAMES = { @@ -336,12 +338,17 @@ def save_checkpoint( *, provenance: Mapping[str, JSONValue] | None = None, controller: JSONValue = None, + channel_metadata: ChannelMetadata = UNNAMED_CHANNELS, ) -> None: """Atomically save a complete simulation checkpoint as validated JSON.""" checkpoint = simulation._checkpoint() checkpoint.validate() state = _simulation_to_json(checkpoint) + try: + labels = channel_metadata.to_json(simulation.species_count, simulation.signal_count) + except ChannelMetadataError as error: + raise CheckpointError(str(error)) from error digest = hashlib.sha256(_canonical_json(state)).hexdigest() controller_digest = hashlib.sha256(_canonical_json(controller)).hexdigest() backend = simulation.backend_info @@ -361,9 +368,11 @@ def save_checkpoint( "algorithm": "sha256", "simulation": digest, "controller": controller_digest, + "channel_metadata": hashlib.sha256(_canonical_json(labels)).hexdigest(), }, "simulation": state, "controller": controller, + "channel_metadata": labels, } try: encoded = ( @@ -704,9 +713,7 @@ def _signal_grid(value: object, path: str, schema_version: int) -> _SignalGridCh if schema_version >= 8: spec.obstacles = [ _integer(item, f"{path}.spec.obstacles[{index}]", 0, 1) - for index, item in enumerate( - _array(spec_data["obstacles"], f"{path}.spec.obstacles") - ) + for index, item in enumerate(_array(spec_data["obstacles"], f"{path}.spec.obstacles")) ] field_value = spec_data["velocity_field"] if field_value is not None: @@ -958,7 +965,7 @@ def load_checkpoint_bundle( if "version" not in root: _fail("$", "missing keys ['version']") schema_version = _integer(root["version"], "$.version", 0, _UINT32_MAX) - supported_versions = {1, 2, 3, 4, 5, 6, 7, CHECKPOINT_VERSION} + supported_versions = {1, 2, 3, 4, 5, 6, 7, 8, CHECKPOINT_VERSION} if schema_version not in supported_versions: _fail("$.version", f"unsupported checkpoint version {schema_version}") required = { @@ -972,6 +979,8 @@ def load_checkpoint_bundle( } if schema_version >= 4: required.add("controller") + if schema_version >= 9: + required.add("channel_metadata") _keys( root, "$", @@ -1007,6 +1016,8 @@ def load_checkpoint_bundle( integrity_keys = {"algorithm", "simulation"} if schema_version >= 4: integrity_keys.add("controller") + if schema_version >= 9: + integrity_keys.add("channel_metadata") _keys(integrity, "$.integrity", integrity_keys) if _string(integrity["algorithm"], "$.integrity.algorithm") != "sha256": _fail("$.integrity.algorithm", "unsupported integrity algorithm") @@ -1017,20 +1028,38 @@ def load_checkpoint_bundle( controller = cast(JSONValue, root["controller"]) if schema_version >= 4 else None if schema_version >= 4: - expected_controller_digest = _string( - integrity["controller"], "$.integrity.controller" - ) + expected_controller_digest = _string(integrity["controller"], "$.integrity.controller") actual_controller_digest = hashlib.sha256(_canonical_json(controller)).hexdigest() if not hmac.compare_digest(actual_controller_digest, expected_controller_digest): _fail("$.integrity.controller", "controller digest does not match") + if schema_version >= 9: + expected_labels_digest = _string( + integrity["channel_metadata"], "$.integrity.channel_metadata" + ) + actual_labels_digest = hashlib.sha256(_canonical_json(root["channel_metadata"])).hexdigest() + if not hmac.compare_digest(actual_labels_digest, expected_labels_digest): + _fail("$.integrity.channel_metadata", "channel metadata digest does not match") checkpoint = _native_checkpoint(root["simulation"], schema_version) + species_count = checkpoint.world.species_count + signal_count = checkpoint.signal_grid.spec.signal_count if checkpoint.signal_grid else 0 + try: + labels = ( + ChannelMetadata.from_json(root["channel_metadata"], species_count, signal_count) + if schema_version >= 9 + # Legacy files contain no labels. Preserve that omission compactly: + # claimed native counts must not allocate new presentation arrays. + else UNNAMED_CHANNELS + ) + except ChannelMetadataError as error: + raise CheckpointError(str(error)) from error return CheckpointBundle( simulation=Simulation(backend, checkpoint, device_index), controller=controller, provenance=provenance, schema_version=schema_version, source_backend=source_backend, + channel_metadata=labels, ) @@ -1047,4 +1076,12 @@ def load_checkpoint( raise CheckpointError( "checkpoint contains controller state; load it with load_checkpoint_bundle" ) + if any( + label is not None + for group in (bundle.channel_metadata.species, bundle.channel_metadata.signals) + for label in (group or ()) + ): + raise CheckpointError( + "checkpoint contains channel metadata; load it with load_checkpoint_bundle" + ) return bundle.simulation diff --git a/python/src/microsimulator/controller.py b/python/src/microsimulator/controller.py index 2e41d4e..99f25c3 100644 --- a/python/src/microsimulator/controller.py +++ b/python/src/microsimulator/controller.py @@ -18,6 +18,7 @@ MechanicsSolveResult, Simulation, ) +from .channels import UNNAMED_CHANNELS, ChannelMetadata from .checkpoint import CheckpointBundle, JSONValue _RANDOM_STATE_KIND = "python-random-mt19937" @@ -327,6 +328,7 @@ def __init__( mechanics: MechanicsConfig | None = None, state: Mapping[str, JSONValue] | None = None, completed_steps: int = 0, + channel_metadata: ChannelMetadata = UNNAMED_CHANNELS, ) -> None: self._model_id, self._model_version = _model_identity(model_id, model_version) if not isinstance(cast(object, simulation), Simulation): @@ -334,6 +336,9 @@ def __init__( if not isinstance(cast(object, rng), random.Random): raise TypeError("native controller requires an explicit random.Random stream") self.simulation = simulation + self.channel_metadata = channel_metadata.resolved( + simulation.species_count, simulation.signal_count + ) self._rng = rng self._regulate = regulate self._on_division = on_division @@ -537,6 +542,7 @@ def from_checkpoint( mechanics = None if mechanics_value is None else MechanicsConfig.from_json(mechanics_value) return cls( checkpoint.simulation, + channel_metadata=checkpoint.channel_metadata, model_id=model_id, model_version=model_version, rng=restore_random_state(value["random"]), diff --git a/python/src/microsimulator/runner.py b/python/src/microsimulator/runner.py index 7aecc0b..82fc344 100644 --- a/python/src/microsimulator/runner.py +++ b/python/src/microsimulator/runner.py @@ -18,6 +18,7 @@ Simulation, backend_available, ) +from .channels import ChannelMetadata, ChannelMetadataError from .checkpoint import CheckpointBundle, JSONValue, save_checkpoint from .controller import SimulationController @@ -102,6 +103,19 @@ def native_simulation(model: object) -> Simulation: return simulation +def model_channel_metadata(model: RunnableModel) -> ChannelMetadata: + """Read optional model labels without extending the required controller protocol.""" + + value = cast(object, getattr(model, "channel_metadata", ChannelMetadata())) + if not isinstance(value, ChannelMetadata): + raise BatchError("model channel_metadata must be ChannelMetadata") + native = native_simulation(model) + try: + return value.resolved(native.species_count, native.signal_count) + except ChannelMetadataError as error: + raise BatchError(str(error)) from error + + def controller_state(model: RunnableModel) -> JSONValue: """Capture optional data-only controller state for a runnable model.""" @@ -197,12 +211,11 @@ def run_simulation( raise BatchError("cell-count threshold must be a positive uint64 value") native = native_simulation(simulation) native.validate() + model_channel_metadata(simulation) destination = Path(output) periodic_steps = ( - tuple(range(checkpoint_every, steps + 1, checkpoint_every)) - if checkpoint_every > 0 - else () + tuple(range(checkpoint_every, steps + 1, checkpoint_every)) if checkpoint_every > 0 else () ) periodic_paths = tuple(_periodic_path(destination, step) for step in periodic_steps) destination.parent.mkdir(parents=True, exist_ok=True) @@ -256,6 +269,7 @@ def run_simulation( stop_cell_count=stop_cell_count, ), controller=controller_state(simulation), + channel_metadata=model_channel_metadata(simulation), ) written_periodic.append(periodic) if reached_cell_count: @@ -275,6 +289,7 @@ def run_simulation( stop_cell_count=stop_cell_count, ), controller=controller_state(simulation), + channel_metadata=model_channel_metadata(simulation), ) return RunSummary( completed_steps=completed_steps, @@ -327,9 +342,7 @@ def build_model( else: resume_value = module.__dict__.get("resume") if not callable(resume_value): - raise BatchError( - f"model {source_path} must define resume(context, checkpoint)" - ) + raise BatchError(f"model {source_path} must define resume(context, checkpoint)") resume = cast(Callable[[ModelContext, CheckpointBundle], object], resume_value) model_value = resume(context, checkpoint) entrypoint = "resume(context, checkpoint)" @@ -346,19 +359,22 @@ def build_model( if not isinstance(model_value, Simulation | SimulationController): raise BatchError( - f"model {source_path} {entrypoint} did not return a Simulation or " - "SimulationController" + f"model {source_path} {entrypoint} did not return a Simulation or SimulationController" ) simulation = native_simulation(model_value) if checkpoint is not None and simulation is not checkpoint.simulation: raise BatchError( - f"model {source_path} resume(context, checkpoint) did not use " - "checkpoint.simulation" + f"model {source_path} resume(context, checkpoint) did not use checkpoint.simulation" ) info = simulation.backend_info if info.kind != context.backend or info.device_index != context.device_index: raise BatchError("model returned a simulation on a different backend or device") simulation.validate() + labels = model_channel_metadata(model_value) + if checkpoint is not None and labels != checkpoint.channel_metadata.resolved( + simulation.species_count, simulation.signal_count + ): + raise BatchError("resumed model channel metadata differs from checkpoint") provenance: dict[str, JSONValue] = { "model": { "path": str(source_path), diff --git a/python/src/microsimulator/sbml.py b/python/src/microsimulator/sbml.py index f17bb5f..d48de97 100644 --- a/python/src/microsimulator/sbml.py +++ b/python/src/microsimulator/sbml.py @@ -15,6 +15,7 @@ RateOp, SpeciesRatePlan, ) +from .channels import ChannelMetadata _FLOAT32_MAX = 3.4028234663852886e38 _UINT32_MAX = (1 << 32) - 1 @@ -36,6 +37,17 @@ class SBMLRateModel: rate_plan: SpeciesRatePlan warnings: tuple[str, ...] + @property + def channel_metadata(self) -> ChannelMetadata: + """Use nonempty SBML names, falling back to stable SBML identifiers.""" + + return ChannelMetadata( + species=tuple( + name if name.strip() else identifier + for name, identifier in zip(self.species_names, self.species_ids, strict=True) + ) + ) + @property def species_count(self) -> int: return len(self.species_ids) @@ -263,8 +275,7 @@ def _local_parameters(kinetic_law: _KineticLaw, reaction_id: str) -> dict[str, f for index in range(kinetic_law.getNumLocalParameters()) ] parameters.extend( - (kinetic_law.getParameter(index), True) - for index in range(kinetic_law.getNumParameters()) + (kinetic_law.getParameter(index), True) for index in range(kinetic_law.getNumParameters()) ) for parameter, check_constant in parameters: identifier = parameter.getId() @@ -427,9 +438,7 @@ def _species_metadata( f"species {identifier!r} must declare exactly one initial concentration or amount" ) initial = ( - species.getInitialConcentration() - if has_concentration - else species.getInitialAmount() + species.getInitialConcentration() if has_concentration else species.getInitialAmount() ) initial = _finite_float32(initial, f"species {identifier!r} initial level") if initial < 0.0: @@ -444,9 +453,7 @@ def _species_metadata( def _stoichiometry(reference: _SpeciesReference, reaction_id: str) -> float: if reference.isSetStoichiometryMath() or not reference.getConstant(): raise SBMLImportError(f"reaction {reaction_id!r} uses dynamic stoichiometry") - value = _finite_float32( - reference.getStoichiometry(), f"reaction {reaction_id!r} stoichiometry" - ) + value = _finite_float32(reference.getStoichiometry(), f"reaction {reaction_id!r} stoichiometry") if value < 0.0: raise SBMLImportError(f"reaction {reaction_id!r} stoichiometry must be non-negative") return value @@ -531,9 +538,7 @@ def parse_sbml(source: str) -> SBMLRateModel: raise SBMLImportError("SBML source must be a nonempty string") libsbml = _libsbml() document = libsbml.readSBMLFromString(source) - if document.getLevel() > 0 and ( - document.getLevel() != 3 or document.getVersion() != 2 - ): + if document.getLevel() > 0 and (document.getLevel() != 3 or document.getVersion() != 2): raise SBMLImportError("SBML import currently requires Level 3 Version 2 Core") warnings = _validate_document(document, libsbml) model = document.getModel() diff --git a/python/src/microsimulator/scene.py b/python/src/microsimulator/scene.py index c3a4668..5485c9f 100644 --- a/python/src/microsimulator/scene.py +++ b/python/src/microsimulator/scene.py @@ -25,11 +25,14 @@ Simulation, Vec3, ) +from .channels import UNNAMED_CHANNELS, ChannelMetadata, ChannelMetadataError from .checkpoint import JSONValue SCENE_FORMAT = "microsimulator-scene" -SCENE_VERSION = 2 +SCENE_VERSION = 3 MAX_SCENE_BYTES = 1 << 30 +# Presentation resource budget, independent of native simulation channel counts. +MAX_SCENE_CHANNELS = 4096 _UINT32_MAX = (1 << 32) - 1 _UINT64_MAX = (1 << 64) - 1 @@ -161,6 +164,21 @@ class SceneFrame: cells: tuple[SceneCell, ...] constraints: SceneConstraints signal_grid: SceneSignalGrid | None + channel_metadata: ChannelMetadata = UNNAMED_CHANNELS + + def __post_init__(self) -> None: + _scene_channel_count(self.species_count, "$.frame.species_count") + if self.signal_grid is not None: + _scene_channel_count( + self.signal_grid.signal_count, "$.frame.signal_grid.signal_count", 1 + ) + object.__setattr__( + self, + "channel_metadata", + self.channel_metadata.resolved( + self.species_count, self.signal_grid.signal_count if self.signal_grid else 0 + ), + ) def _installed_version() -> str: @@ -181,9 +199,14 @@ def _capture_boundary(boundary: GridBoundary) -> SceneGridBoundary: ) -def capture_scene(simulation: Simulation) -> SceneFrame: +def capture_scene( + simulation: Simulation, *, channel_metadata: ChannelMetadata = UNNAMED_CHANNELS +) -> SceneFrame: """Capture a complete immutable presentation frame after a simulation step.""" + # Reject before copying native state or expanding omitted channel labels. + _scene_channel_count(simulation.species_count, "$.frame.species_count") + _scene_channel_count(simulation.signal_count, "$.frame.signal_grid.signal_count") checkpoint = simulation._checkpoint() checkpoint.validate() backend = simulation.backend_info @@ -276,6 +299,9 @@ def capture_scene(simulation: Simulation) -> SceneFrame: cells=cells, constraints=constraints, signal_grid=signal_grid, + channel_metadata=channel_metadata.resolved( + simulation.species_count, simulation.signal_count + ), ) _validate_frame(frame) return frame @@ -375,6 +401,9 @@ def _frame_to_json(frame: SceneFrame) -> dict[str, JSONValue]: "cells": cells, "constraints": constraints, "signal_grid": grid, + "channel_metadata": frame.channel_metadata.to_json( + frame.species_count, frame.signal_grid.signal_count if frame.signal_grid else 0 + ), } @@ -400,13 +429,16 @@ def dumps_scene(frame: SceneFrame) -> str: }, "frame": payload, } - return json.dumps( - document, - allow_nan=False, - ensure_ascii=False, - indent=2, - sort_keys=True, - ) + "\n" + return ( + json.dumps( + document, + allow_nan=False, + ensure_ascii=False, + indent=2, + sort_keys=True, + ) + + "\n" + ) def save_scene(frame: SceneFrame, path: str | os.PathLike[str]) -> None: @@ -442,6 +474,13 @@ def _fail(path: str, message: str) -> NoReturn: raise SceneError(f"{path}: {message}") +def _scene_channel_count(value: object, path: str, minimum: int = 0) -> int: + count = _integer(value, path, minimum, _UINT32_MAX) + if count > MAX_SCENE_CHANNELS: + _fail(path, f"exceeds scene presentation channel budget of {MAX_SCENE_CHANNELS} per group") + return count + + def _reject_constant(value: str) -> NoReturn: raise SceneError(f"scene contains non-finite JSON number {value}") @@ -713,7 +752,7 @@ def _signal_grid(value: object, path: str) -> SceneSignalGrid | None: path, {"signal_count", "shape", "origin", "spacing", "boundaries", "levels"}, ) - signal_count = _integer(data["signal_count"], f"{path}.signal_count", 1, _UINT32_MAX) + signal_count = _scene_channel_count(data["signal_count"], f"{path}.signal_count", 1) shape_values = _array(data["shape"], f"{path}.shape") if len(shape_values) != 3: _fail(f"{path}.shape", "expected exactly three dimensions") @@ -746,11 +785,25 @@ def _signal_grid(value: object, path: str) -> SceneSignalGrid | None: ) -def _frame(value: object, path: str) -> SceneFrame: +def _frame(value: object, path: str, schema_version: int) -> SceneFrame: data = _object(value, path) - _keys(data, path, {"time", "backend", "species_count", "cells", "constraints", "signal_grid"}) - species_count = _integer(data["species_count"], f"{path}.species_count", 0, _UINT32_MAX) + keys = {"time", "backend", "species_count", "cells", "constraints", "signal_grid"} + if schema_version >= 3: + keys.add("channel_metadata") + _keys(data, path, keys) + species_count = _scene_channel_count(data["species_count"], f"{path}.species_count") + signal_grid = _signal_grid(data["signal_grid"], f"{path}.signal_grid") + signal_count = signal_grid.signal_count if signal_grid else 0 + try: + labels = ( + ChannelMetadata.from_json(data["channel_metadata"], species_count, signal_count) + if schema_version >= 3 + else ChannelMetadata().resolved(species_count, signal_count) + ) + except ChannelMetadataError as error: + raise SceneError(str(error)) from error frame = SceneFrame( + channel_metadata=labels, time=_number(data["time"], f"{path}.time"), backend=_backend(data["backend"], f"{path}.backend"), species_count=species_count, @@ -759,7 +812,7 @@ def _frame(value: object, path: str) -> SceneFrame: for index, item in enumerate(_array(data["cells"], f"{path}.cells")) ), constraints=_constraints(data["constraints"], f"{path}.constraints"), - signal_grid=_signal_grid(data["signal_grid"], f"{path}.signal_grid"), + signal_grid=signal_grid, ) _validate_frame(frame) return frame @@ -776,6 +829,15 @@ def _validate_boundary(boundary: SceneGridBoundary, signal_count: int, path: str def _validate_frame(frame: SceneFrame) -> None: + _scene_channel_count(frame.species_count, "$.frame.species_count") + if frame.signal_grid is not None: + _scene_channel_count(frame.signal_grid.signal_count, "$.frame.signal_grid.signal_count", 1) + try: + frame.channel_metadata.resolved( + frame.species_count, frame.signal_grid.signal_count if frame.signal_grid else 0 + ) + except ChannelMetadataError as error: + raise SceneError(str(error)) from error _number(frame.time, "$.frame.time") if frame.time < 0.0: _fail("$.frame.time", "must be non-negative") @@ -936,7 +998,7 @@ def parse_scene(source: str | bytes) -> SceneFrame: if _string(root["format"], "$.format") not in (SCENE_FORMAT, "cellmodeller2-scene"): _fail("$.format", "not a MicroSimulator scene") schema_version = _integer(root["version"], "$.version", 0, _UINT32_MAX) - if schema_version != SCENE_VERSION: + if schema_version not in {2, SCENE_VERSION}: _fail("$.version", f"unsupported scene version {schema_version}") producer = _object(root["producer"], "$.producer") _keys(producer, "$.producer", {"name", "version"}) @@ -947,12 +1009,10 @@ def parse_scene(source: str | bytes) -> SceneFrame: if _string(integrity["algorithm"], "$.integrity.algorithm") != "sha256": _fail("$.integrity.algorithm", "unsupported integrity algorithm") expected_digest = _string(integrity["frame"], "$.integrity.frame") - actual_digest = hashlib.sha256( - _canonical_json(cast(JSONValue, root["frame"])) - ).hexdigest() + actual_digest = hashlib.sha256(_canonical_json(cast(JSONValue, root["frame"]))).hexdigest() if not hmac.compare_digest(actual_digest, expected_digest): _fail("$.integrity.frame", "frame digest does not match") - return _frame(root["frame"], "$.frame") + return _frame(root["frame"], "$.frame", schema_version) def load_scene(path: str | os.PathLike[str]) -> SceneFrame: diff --git a/python/src/microsimulator/viewer_server.py b/python/src/microsimulator/viewer_server.py index b0d6bf2..4e2599a 100644 --- a/python/src/microsimulator/viewer_server.py +++ b/python/src/microsimulator/viewer_server.py @@ -19,7 +19,7 @@ from aiohttp import WSMsgType, web from .checkpoint import JSONValue, save_checkpoint -from .runner import RunnableModel, controller_state, native_simulation +from .runner import RunnableModel, controller_state, model_channel_metadata, native_simulation from .scene import capture_scene, dumps_scene MAX_COMMAND_BYTES = 4096 @@ -104,6 +104,7 @@ def checkpoint_enabled(self) -> bool: def _build(self) -> tuple[RunnableModel, dict[str, JSONValue]]: model, provenance = self._factory() native_simulation(model).validate() + model_channel_metadata(model) return model, dict(provenance) def step(self, steps: int = 1) -> None: @@ -141,12 +142,20 @@ def checkpoint(self) -> Path: destination, provenance=provenance, controller=controller_state(self._model), + channel_metadata=model_channel_metadata(self._model), ) return destination def frame_message(self, *, playing: bool) -> dict[str, JSONValue]: native = native_simulation(self._model) - scene = cast(dict[str, JSONValue], json.loads(dumps_scene(capture_scene(native)))) + scene = cast( + dict[str, JSONValue], + json.loads( + dumps_scene( + capture_scene(native, channel_metadata=model_channel_metadata(self._model)) + ) + ), + ) return { "type": "frame", "revision": self._revision, diff --git a/python/tests/test_channels.py b/python/tests/test_channels.py new file mode 100644 index 0000000..3ab2851 --- /dev/null +++ b/python/tests/test_channels.py @@ -0,0 +1,224 @@ +from __future__ import annotations + +# ruff: noqa: RUF001 -- explicit Unicode-label coverage. +import hashlib +import json +from dataclasses import replace +from pathlib import Path +from typing import Any + +import pytest +import rfc8785 +from microsimulator import ( + MAX_SCENE_CHANNELS, + BackendKind, + ChannelMetadata, + ChannelMetadataError, + CheckpointError, + ModelContext, + NativeController, + SceneError, + Simulation, + build_model, + capture_scene, + dumps_scene, + load_checkpoint, + load_checkpoint_bundle, + load_scene, + parse_scene, + run_simulation, + save_checkpoint, + save_scene, +) +from microsimulator.runner import BatchError, model_channel_metadata +from microsimulator.viewer_server import LiveSession + +ROOT = Path(__file__).resolve().parents[2] +EXAMPLE = ROOT / "examples/named_channels.py" +LABELS = ChannelMetadata( + species=("Green reporter", "Red reporter"), signals=("Nutrient", "Extracellular cue") +) + + +def _build() -> tuple[NativeController, dict[str, Any]]: + model, provenance = build_model(EXAMPLE, ModelContext(BackendKind.CPU, 0, 17)) + assert isinstance(model, NativeController) + return model, provenance + + +def test_named_model_periodic_checkpoint_resume_and_standalone_export(tmp_path: Path) -> None: + model, provenance = _build() + output = tmp_path / "named.json" + summary = run_simulation( + model, steps=2, dt=0.01, output=output, checkpoint_every=1, provenance=provenance + ) + for path in (*summary.periodic_checkpoints, output): + assert load_checkpoint_bundle(path).channel_metadata == LABELS + bundle = load_checkpoint_bundle(output) + resumed, provenance = build_model( + EXAMPLE, ModelContext(BackendKind.CPU, 0, 17), checkpoint=bundle + ) + assert model_channel_metadata(resumed) == LABELS + resumed_output = tmp_path / "continued.json" + run_simulation(resumed, steps=1, dt=0.01, output=resumed_output, provenance=provenance) + continued = load_checkpoint_bundle(resumed_output) + assert continued.channel_metadata == LABELS + model.step(0.01) + assert continued.simulation.signal_levels == model.simulation.signal_levels + assert continued.simulation.cell(1).species == model.simulation.cell(1).species + # Only data and the bundle API are needed to export; no model import/execute. + destination = tmp_path / "named.scene.json" + save_scene( + capture_scene(bundle.simulation, channel_metadata=bundle.channel_metadata), destination + ) + assert load_scene(destination).channel_metadata == LABELS + + +def test_live_labels_survive_step_reset_and_checkpoint(tmp_path: Path) -> None: + session = LiveSession(_build, dt=0.01, checkpoint_output=tmp_path / "live.json") + for operation in (lambda: None, session.step, session.reset): + operation() + message = session.frame_message(playing=False) + frame = parse_scene(json.dumps(message["scene"])) + assert frame.channel_metadata == LABELS + assert load_checkpoint_bundle(session.checkpoint()).channel_metadata == LABELS + + +def test_metadata_counts_fail_before_stepping_or_writing(tmp_path: Path) -> None: + model, _ = _build() + model.channel_metadata = ChannelMetadata(species=("only one",)) + with pytest.raises(BatchError, match=r"species: expected 2 labels, got 1"): + run_simulation(model, steps=1, dt=0.1, output=tmp_path / "bad.json") + assert model.simulation.time == 0 + assert not (tmp_path / "bad.json").exists() + with pytest.raises(ChannelMetadataError, match=r"species: expected 2 labels"): + capture_scene(model.simulation, channel_metadata=model.channel_metadata) + with pytest.raises(CheckpointError, match=r"species: expected 2 labels"): + save_checkpoint( + model.simulation, tmp_path / "bad.json", channel_metadata=model.channel_metadata + ) + with pytest.raises(ChannelMetadataError, match=r"signals: expected 2 labels"): + ChannelMetadata(signals=()).resolved(2, 2) + + +def test_missing_duplicate_unicode_empty_and_markup_labels_roundtrip(tmp_path: Path) -> None: + model, _ = _build() + labels = ChannelMetadata(species=("α 🧪", "α 🧪"), signals=(None, " ")) + save_checkpoint(model.simulation, tmp_path / "labels.json", channel_metadata=labels) + bundle = load_checkpoint_bundle(tmp_path / "labels.json") + assert bundle.channel_metadata == labels + frame = capture_scene(bundle.simulation, channel_metadata=bundle.channel_metadata) + assert parse_scene(dumps_scene(frame)).channel_metadata == labels + with pytest.raises(CheckpointError, match="contains channel metadata"): + load_checkpoint(tmp_path / "labels.json") + with pytest.raises(ChannelMetadataError, match="invalid Unicode"): + ChannelMetadata(species=("\ud800",)) + with pytest.raises(ChannelMetadataError, match="expected a string"): + ChannelMetadata.from_json({"species": [7], "signals": []}, 1, 0) + + +def test_metadata_tampering_rejected_and_v8_migrates_only_after_verification( + tmp_path: Path, +) -> None: + model, _ = _build() + path = tmp_path / "labels.json" + save_checkpoint(model.simulation, path, channel_metadata=LABELS) + document = json.loads(path.read_text()) + document["channel_metadata"]["species"][0] = "tampered" + path.write_text(json.dumps(document)) + with pytest.raises(CheckpointError, match="channel metadata digest does not match"): + load_checkpoint_bundle(path) + document["version"] = 8 + del document["channel_metadata"] + del document["integrity"]["channel_metadata"] + path.write_text(json.dumps(document)) + bundle = load_checkpoint_bundle(path) + assert bundle.channel_metadata == ChannelMetadata() + assert capture_scene( + bundle.simulation, channel_metadata=bundle.channel_metadata + ).channel_metadata == ChannelMetadata().resolved(2, 2) + document["simulation"]["time"] = 999 + path.write_text(json.dumps(document)) + with pytest.raises(CheckpointError, match="state digest does not match"): + load_checkpoint_bundle(path) + + +def test_legacy_checkpoint_keeps_unspecified_labels_compact_before_native_restore( + tmp_path: Path, monkeypatch: pytest.MonkeyPatch +) -> None: + # Native state is allowed more channels than a presentation scene. No cells + # or label arrays are necessary to represent the legacy checkpoint. + simulation = Simulation(BackendKind.CPU, species_count=MAX_SCENE_CHANNELS + 1) + path = tmp_path / "legacy.json" + save_checkpoint(simulation, path) + document = json.loads(path.read_text()) + document["version"] = 8 + del document["channel_metadata"] + del document["integrity"]["channel_metadata"] + path.write_text(json.dumps(document)) + + def forbid_expansion( + self: ChannelMetadata, species_count: int, signal_count: int + ) -> ChannelMetadata: + raise AssertionError("legacy restore must not expand absent presentation labels") + + monkeypatch.setattr(ChannelMetadata, "resolved", forbid_expansion) + bundle = load_checkpoint_bundle(path) + assert bundle.simulation.species_count == MAX_SCENE_CHANNELS + 1 + assert bundle.channel_metadata.species is None + assert bundle.channel_metadata.signals is None + assert load_checkpoint(path).species_count == MAX_SCENE_CHANNELS + 1 + with pytest.raises(SceneError, match="scene presentation channel budget of 4096"): + capture_scene(bundle.simulation, channel_metadata=bundle.channel_metadata) + + # Untrusted input must still pass its original integrity validation first. + document["simulation"]["world"]["species_count"] = (1 << 32) - 1 + path.write_text(json.dumps(document)) + with pytest.raises(CheckpointError, match="state digest does not match"): + load_checkpoint_bundle(path) + + +def test_scene_v2_verifies_original_payload_and_v3_rejects_invalid_labels() -> None: + model, _ = _build() + document = json.loads(dumps_scene(capture_scene(model.simulation, channel_metadata=LABELS))) + document["version"] = 2 + del document["frame"]["channel_metadata"] + document["integrity"]["frame"] = hashlib.sha256(rfc8785.dumps(document["frame"])).hexdigest() + assert parse_scene(json.dumps(document)).channel_metadata == ChannelMetadata().resolved(2, 2) + document["frame"]["time"] = 999 + with pytest.raises(SceneError, match="frame digest does not match"): + parse_scene(json.dumps(document)) + document["version"] = 3 + document["frame"]["channel_metadata"] = {"species": [], "signals": [None, None]} + document["integrity"]["frame"] = hashlib.sha256(rfc8785.dumps(document["frame"])).hexdigest() + with pytest.raises(SceneError, match="species: expected 2 labels"): + parse_scene(json.dumps(document)) + + +def test_unnamed_native_model_and_closed_channel_schema() -> None: + simulation = Simulation(species_count=2) + assert model_channel_metadata(simulation) == ChannelMetadata(species=(None, None), signals=()) + with pytest.raises(ChannelMetadataError, match="exactly species and signals"): + ChannelMetadata.from_json({"species": [], "signals": [], "extra": []}, 0, 0) + + +def test_resumed_model_cannot_silently_replace_persisted_labels(tmp_path: Path) -> None: + model, provenance = _build() + path = tmp_path / "labels.json" + run_simulation(model, steps=0, dt=0.1, output=path, provenance=provenance) + bundle = load_checkpoint_bundle(path) + changed = replace( + bundle, + channel_metadata=ChannelMetadata( + species=("Changed", "Red reporter"), signals=LABELS.signals + ), + ) + # Standard native restore treats the checkpoint's labels as authoritative. + resumed, _ = build_model(EXAMPLE, ModelContext(BackendKind.CPU, 0, 17), checkpoint=changed) + assert model_channel_metadata(resumed) == changed.channel_metadata + + +def test_shared_python_typescript_v3_fixture() -> None: + frame = load_scene(ROOT / "viewer/tests/fixtures/channels-v3.scene.json") + assert frame.channel_metadata.species == ("α 🧪", "α 🧪") + assert frame.channel_metadata.signals == (None, " ") diff --git a/python/tests/test_checkpoint.py b/python/tests/test_checkpoint.py index 3ce896d..17f959f 100644 --- a/python/tests/test_checkpoint.py +++ b/python/tests/test_checkpoint.py @@ -263,6 +263,8 @@ def test_version_one_checkpoint_migrates_to_an_empty_signal_state(tmp_path: Path save_checkpoint(simulation, path) document = _document(path) document["version"] = 1 + del document["channel_metadata"] + del document["integrity"]["channel_metadata"] del document["controller"] del document["integrity"]["controller"] del document["simulation"]["signal_grid"] @@ -284,6 +286,8 @@ def test_version_two_checkpoint_migrates_without_a_coupled_plan(tmp_path: Path) save_checkpoint(simulation, path) document = _document(path) document["version"] = 2 + del document["channel_metadata"] + del document["integrity"]["channel_metadata"] del document["controller"] del document["integrity"]["controller"] del document["simulation"]["coupled_rate_plan"] @@ -307,6 +311,8 @@ def test_version_three_checkpoint_migrates_without_controller_state(tmp_path: Pa save_checkpoint(simulation, path) document = _document(path) document["version"] = 3 + del document["channel_metadata"] + del document["integrity"]["channel_metadata"] del document["controller"] del document["integrity"]["controller"] del document["simulation"]["signal_grid"]["spec"]["integration"] @@ -330,6 +336,8 @@ def test_version_four_signal_grid_migrates_to_forward_euler(tmp_path: Path) -> N save_checkpoint(simulation, path) document = _document(path) document["version"] = 4 + del document["channel_metadata"] + del document["integrity"]["channel_metadata"] del document["simulation"]["signal_grid"]["spec"]["integration"] del document["simulation"]["signal_grid"]["spec"]["solver"] _remove_affine_reaction(document) @@ -350,6 +358,8 @@ def test_version_five_cells_migrate_to_movable(tmp_path: Path) -> None: save_checkpoint(simulation, path) document = _document(path) document["version"] = 5 + del document["channel_metadata"] + del document["integrity"]["channel_metadata"] _remove_affine_reaction(document) _remove_fixed_fields(document) _remove_constraint_boxes(document) @@ -366,6 +376,8 @@ def test_version_six_signal_grid_migrates_without_affine_reactions(tmp_path: Pat save_checkpoint(simulation, path) document = _document(path) document["version"] = 6 + del document["channel_metadata"] + del document["integrity"]["channel_metadata"] _remove_affine_reaction(document) _remove_constraint_boxes(document) _remove_grid_obstacles(document) @@ -383,6 +395,8 @@ def test_version_seven_checkpoint_migrates_without_boxes(tmp_path: Path) -> None save_checkpoint(simulation, path) document = _document(path) document["version"] = 7 + del document["channel_metadata"] + del document["integrity"]["channel_metadata"] _remove_constraint_boxes(document) _remove_grid_obstacles(document) _rewrite_with_state_digest(path, document) diff --git a/python/tests/test_sbml.py b/python/tests/test_sbml.py index fc5aa56..7b8f3d3 100644 --- a/python/tests/test_sbml.py +++ b/python/tests/test_sbml.py @@ -187,3 +187,10 @@ def test_malformed_or_empty_sbml_fails_explicitly() -> None: ) with pytest.raises(SBMLImportError, match="Level 3 Version 2"): parse_sbml(level_two) + + +def test_sbml_channel_metadata_uses_names_and_identifier_fallback() -> None: + source = _model_xml().replace('name="substrate"', 'name=""') + model = parse_sbml(source) + assert model.channel_metadata.species == ("A", "product") + assert model.channel_metadata.signals is None diff --git a/python/tests/test_scene_channel_budget.py b/python/tests/test_scene_channel_budget.py new file mode 100644 index 0000000..b71d3b1 --- /dev/null +++ b/python/tests/test_scene_channel_budget.py @@ -0,0 +1,139 @@ +from __future__ import annotations + +import hashlib +import json +from dataclasses import replace +from pathlib import Path +from types import SimpleNamespace +from typing import cast + +import pytest +import rfc8785 +from microsimulator import ( + MAX_SCENE_CHANNELS, + ChannelMetadata, + SceneBackend, + SceneConstraints, + SceneError, + SceneFrame, + SceneGridBoundary, + SceneSignalGrid, + Simulation, + capture_scene, + dumps_scene, + load_scene, + parse_scene, + save_scene, +) + + +def _frame(species_count: int = 0, signal_count: int = 1) -> SceneFrame: + boundary = SceneGridBoundary("no_flux", ()) + grid = SceneSignalGrid( + signal_count, + (1, 1, 1), + (0.0, 0.0, 0.0), + (1.0, 1.0, 1.0), + boundary, + boundary, + boundary, + boundary, + boundary, + boundary, + (0.0,) * signal_count, + ) + return SceneFrame( + 0.0, + SceneBackend("cpu", "CPU reference", "host", 0, False), + species_count, + (), + SceneConstraints((), (), (), ()), + grid, + ) + + +def _forbid_expansion( + self: ChannelMetadata, species_count: int, signal_count: int +) -> ChannelMetadata: + raise AssertionError("scene count must be bounded before label expansion") + + +@pytest.mark.parametrize("version", [2, 3]) +def test_scene_channel_budget_boundary_roundtrips(version: int, tmp_path: Path) -> None: + # Each group independently accepts the inclusive limit, including a scene + # without cells that still needs species selector labels. + frame = _frame(MAX_SCENE_CHANNELS, MAX_SCENE_CHANNELS) + document = json.loads(dumps_scene(frame)) + document["version"] = version + if version == 2: + del document["frame"]["channel_metadata"] + document["integrity"]["frame"] = hashlib.sha256(rfc8785.dumps(document["frame"])).hexdigest() + encoded = json.dumps(document) + assert parse_scene(encoded) == frame + path = tmp_path / "boundary.scene.json" + path.write_text(encoded) + assert load_scene(path) == frame + save_scene(frame, path) + assert load_scene(path) == frame + assert frame.channel_metadata == ChannelMetadata( + species=(None,) * MAX_SCENE_CHANNELS, signals=(None,) * MAX_SCENE_CHANNELS + ) + + +@pytest.mark.parametrize("version", [2, 3]) +@pytest.mark.parametrize("group", ["species", "signals"]) +@pytest.mark.parametrize("count", [MAX_SCENE_CHANNELS + 1, (1 << 32) - 1]) +def test_scene_reader_rejects_tiny_oversized_claim_before_expansion( + version: int, group: str, count: int, tmp_path: Path, monkeypatch: pytest.MonkeyPatch +) -> None: + document = json.loads(dumps_scene(_frame())) + document["version"] = version + if version == 2: + del document["frame"]["channel_metadata"] + if group == "species": + document["frame"]["species_count"] = count + else: + document["frame"]["signal_grid"]["signal_count"] = count + document["integrity"]["frame"] = hashlib.sha256(rfc8785.dumps(document["frame"])).hexdigest() + encoded = json.dumps(document) + assert len(encoded) < 2000 + monkeypatch.setattr(ChannelMetadata, "resolved", _forbid_expansion) + with pytest.raises(SceneError, match="scene presentation channel budget of 4096"): + parse_scene(encoded) + path = tmp_path / "oversized.scene.json" + path.write_text(encoded) + with pytest.raises(SceneError, match="scene presentation channel budget of 4096"): + load_scene(path) + + +@pytest.mark.parametrize("group", ["species", "signals"]) +@pytest.mark.parametrize("count", [MAX_SCENE_CHANNELS + 1, (1 << 32) - 1]) +def test_construct_capture_and_write_bound_counts_before_copy_or_expansion( + group: str, count: int, tmp_path: Path, monkeypatch: pytest.MonkeyPatch +) -> None: + frame = _frame() + assert frame.signal_grid is not None + grid = ( + replace(frame.signal_grid, signal_count=count) if group == "signals" else frame.signal_grid + ) + species_count = count if group == "species" else 0 + monkeypatch.setattr(ChannelMetadata, "resolved", _forbid_expansion) + with pytest.raises(SceneError, match="scene presentation channel budget of 4096"): + replace(frame, species_count=species_count, signal_grid=grid) + + # No _checkpoint method: rejection must precede any native-state copy. + simulation = cast( + Simulation, SimpleNamespace(species_count=species_count, signal_count=grid.signal_count) + ) + with pytest.raises(SceneError, match="scene presentation channel budget of 4096"): + capture_scene(simulation) + + # Encoders independently validate even objects restored without __init__. + object.__setattr__(frame, "species_count", species_count) + object.__setattr__(frame, "signal_grid", grid) + with pytest.raises(SceneError, match="scene presentation channel budget of 4096"): + dumps_scene(frame) + path = tmp_path / "rejected.scene.json" + with pytest.raises(SceneError, match="scene presentation channel budget of 4096"): + save_scene(frame, path) + assert not path.exists() diff --git a/viewer/README.md b/viewer/README.md index 237d6cc..87ae74f 100644 --- a/viewer/README.md +++ b/viewer/README.md @@ -72,3 +72,7 @@ Opening a scene file, live session, or recording begins a new dataset. Call `Dat The ground reference grid is separate from the scientific signal lattice. Its square extent and origin come from the first frame's finite device geometry (boxes, spheres, and cylinders), or from the initial cell capsule bounds when no finite device exists. Infinite plane constraints are excluded. The extent is at least 10 scene distance units with 20 equal divisions; the grid plane is 0.01 units below the lesser of the initial lower Z bound and zero. An initially empty dataset uses a 10-unit grid centered on the world origin. These values remain fixed even when cells or device geometry appear later, the colony expands beyond the grid, or all cells disappear. Opening another dataset initializes a new reference grid; camera Fit never changes its geometry. `browser/reference-grid.mjs` verifies the reference grid and presentation lifecycle in Chromium against a running Vite server. It uses Playwright (`@playwright/test`) and its installed Chromium; a shared installation can be supplied through `MICROSIMULATOR_PLAYWRIGHT_MODULE` as an absolute module filename. Set `VIEWER_URL` if the server is not on `http://127.0.0.1:4320`, and `EVIDENCE_DIR` to choose the screenshot directory. The test observes renderer transforms through test-only request instrumentation and introduces no production debug interface. + +## Channel labels + +Model-defined species and signal names appear in channel selectors, the species legend, and cell inspection. Duplicate names include their channel indices; unnamed channels retain `Channel N`. Names are presentation text; indices continue to identify selected channels. Current readers accept scene v2 and v3, while writers emit v3. See the [authoring guide](../docs/models/channel-labels.md) and [scene v3 schema](../docs/formats/scene-v3.md). diff --git a/viewer/browser/channel-labels.mjs b/viewer/browser/channel-labels.mjs new file mode 100644 index 0000000..0cfedac --- /dev/null +++ b/viewer/browser/channel-labels.mjs @@ -0,0 +1,147 @@ +// Uses the shared external Playwright harness; no application debug API is shipped. +import assert from "node:assert/strict"; +import { createHash } from "node:crypto"; +import { mkdir, readFile } from "node:fs/promises"; +import canonicalize from "canonicalize"; + +const { chromium, expect } = await import( + process.env.MICROSIMULATOR_PLAYWRIGHT_MODULE ?? "@playwright/test" +); +const url = process.env.VIEWER_URL ?? "http://127.0.0.1:4318"; +const evidence = + process.env.EVIDENCE_DIR ?? "/tmp/microsimulator-channel-labels"; +await mkdir(evidence, { recursive: true }); +const browser = await chromium.launch({ headless: true }); +try { + const page = await browser.newPage({ + viewport: { width: 1440, height: 960 }, + }); + const errors = []; + page.on("pageerror", (error) => errors.push(error.message)); + await page.route("**/src/colony-viewer.ts*", async (route) => { + const response = await route.fetch(); + const source = await response.text(); + const marker = "this.onSelection = onSelection;"; + assert.equal(source.split(marker).length, 2); + await route.fulfill({ + response, + body: source.replace( + marker, + `${marker}\nglobalThis.__testViewer = this;`, + ), + }); + }); + let socket; + await page.routeWebSocket("**/api/v1/session?*", (connection) => { + socket = connection; + }); + await page.goto(`${url}/?token=fixture`); + await expect.poll(() => socket !== undefined).toBe(true); + const document = JSON.parse( + await readFile( + new URL("../tests/fixtures/channels-v3.scene.json", import.meta.url), + "utf8", + ), + ); + let revision = 0; + async function send() { + document.integrity.frame = createHash("sha256") + .update(canonicalize(document.frame)) + .digest("hex"); + socket.send( + JSON.stringify({ + type: "frame", + revision: revision++, + completed_steps: revision, + playing: false, + checkpoint_enabled: false, + scene: document, + }), + ); + } + await send(); + await expect(page.locator("#species-channel option")).toHaveText([ + "α 🧪 [0]", + "α 🧪 [1]", + ]); + await expect(page.locator("#signal-channel option")).toHaveText([ + "Channel 0", + "Channel 1", + ]); + await page.selectOption("#color-mode", "species"); + await page.selectOption("#species-channel", "1"); + await page.selectOption("#signal-channel", "1"); + await expect(page.locator("#legend-title")).toHaveText("α 🧪 [1]"); + await page.evaluate(() => globalThis.__testViewer.selectCell(0)); + await expect(page.locator("#species-values li span")).toHaveText([ + "α 🧪 [0]", + "α 🧪 [1]", + ]); + await expect( + page.locator("#species-values b, #species-channel b, #legend-title b"), + ).toHaveCount(0); + document.frame.channel_metadata = { + species: ["Green reporter", "Red reporter"], + signals: ["Nutrient", "Extracellular cue"], + }; + document.frame.time = 1; + await send(); + await expect(page.locator("#species-channel")).toHaveValue("1"); + await expect(page.locator("#signal-channel")).toHaveValue("1"); + await expect(page.locator("#species-channel option")).toHaveText([ + "Green reporter", + "Red reporter", + ]); + await expect(page.locator("#signal-channel option")).toHaveText([ + "Nutrient", + "Extracellular cue", + ]); + await expect(page.locator("#legend-title")).toHaveText("Red reporter"); + await expect(page.locator("#species-values li span")).toHaveText([ + "Green reporter", + "Red reporter", + ]); + document.frame.time = 0; + await send(); + await expect(page.locator("#time-chip")).toHaveText("t = 0"); + await expect(page.locator("#species-channel")).toHaveValue("1"); + // Literal labels may imitate automatically generated duplicate suffixes. + document.frame.species_count = 3; + for (const cell of document.frame.cells) cell.species.push(0.25); + document.frame.channel_metadata.species = ["GFP", "GFP", "GFP [0]"]; + await send(); + const indexed = ["GFP [0]", "GFP [1]", "GFP [0] [2]"]; + await expect(page.locator("#species-channel option")).toHaveText(indexed); + await expect(page.locator("#species-values li span")).toHaveText(indexed); + await expect(page.locator("#species-channel")).toHaveValue("1"); + await page.selectOption("#species-channel", "2"); + await expect(page.locator("#legend-title")).toHaveText("GFP [0] [2]"); + document.frame.channel_metadata.species = [ + " GFP\tname ", + "GFP name", + "GFP\t name [0]", + ]; + await send(); + const normalized = ["GFP name [0]", "GFP name [1]", "GFP name [0] [2]"]; + await expect(page.locator("#species-channel option")).toHaveText(normalized); + assert.deepEqual( + await page + .locator("#species-channel option") + .evaluateAll((options) => options.map((option) => option.label)), + normalized, + ); + await expect(page.locator("#species-channel")).toHaveValue("2"); + await expect(page.locator("#legend-title")).toHaveText("GFP name [0] [2]"); + await page.screenshot({ path: `${evidence}/named-channels.png` }); + assert.deepEqual(errors, []); + console.log( + JSON.stringify({ + status: "passed", + assertions: + "selector names, duplicate, whitespace and generated-suffix disambiguation, HTML text safety, Unicode, inspector, legend, index-stable renamed frames, reset time", + screenshot: `${evidence}/named-channels.png`, + }), + ); +} finally { + await browser.close(); +} diff --git a/viewer/src/color.ts b/viewer/src/color.ts index 22c0b16..f5b93f0 100644 --- a/viewer/src/color.ts +++ b/viewer/src/color.ts @@ -1,4 +1,4 @@ -import type { SceneCell, SceneFrame } from "./scene"; +import { channelLabel, type SceneCell, type SceneFrame } from "./scene"; export type RGB = readonly [number, number, number]; export type ColorMode = "cell-type" | "species" | "growth-rate" | "fixed"; @@ -128,7 +128,7 @@ export function mapCellColors( return scalarMapping( frame.cells, frame.cells.map((cell) => cell.species[config.speciesIndex] ?? 0), - `Species ${config.speciesIndex}`, + channelLabel(frame, "species", config.speciesIndex), ); } } diff --git a/viewer/src/main.ts b/viewer/src/main.ts index 28aa1d8..56adf18 100644 --- a/viewer/src/main.ts +++ b/viewer/src/main.ts @@ -12,6 +12,7 @@ import { } from "./live"; import { MAX_SCENE_BYTES, + channelLabel, parseScene, type SceneCell, type SceneFrame, @@ -105,14 +106,14 @@ function setStatus(message: string, kind: "info" | "error" = "info"): void { function options( select: HTMLSelectElement, count: number, - prefix: string, + label: (index: number) => string, selected: number, ): void { select.replaceChildren(); for (let index = 0; index < count; index += 1) { const option = document.createElement("option"); option.value = String(index); - option.textContent = `${prefix} ${index}`; + option.textContent = label(index); select.append(option); } select.value = String(selected); @@ -210,7 +211,10 @@ function updateSelection(cell: SceneCell | null): void { const item = document.createElement("li"); const label = document.createElement("span"); const encoded = document.createElement("code"); - label.textContent = `Channel ${index}`; + label.textContent = + frame === null + ? `Channel ${index}` + : channelLabel(frame, "species", index); encoded.textContent = formatNumber(value); item.append(label, encoded); speciesValues.append(item); @@ -246,7 +250,12 @@ function presentScene( next.signalGrid === null ? "None" : next.signalGrid.shape.join(" × "); colorMode.value = display.colorMode; - options(speciesChannel, next.speciesCount, "Channel", display.speciesChannel); + options( + speciesChannel, + next.speciesCount, + (index) => channelLabel(next, "species", index), + display.speciesChannel, + ); const speciesOption = colorMode.querySelector( 'option[value="species"]', ); @@ -261,7 +270,7 @@ function presentScene( options( signalChannel, next.signalGrid.signalCount, - "Channel", + (index) => channelLabel(next, "signals", index), display.signalChannel, ); signalRange.max = String( diff --git a/viewer/src/scene.ts b/viewer/src/scene.ts index 439e555..28dcc49 100644 --- a/viewer/src/scene.ts +++ b/viewer/src/scene.ts @@ -1,8 +1,10 @@ import canonicalize from "canonicalize"; export const SCENE_FORMAT = "microsimulator-scene"; -export const SCENE_VERSION = 2; +export const SCENE_VERSION = 3; export const MAX_SCENE_BYTES = 1 << 30; +// Presentation resource budget; this does not limit native engine counts. +export const MAX_SCENE_CHANNELS = 4096; const UINT32_MAX = 2 ** 32 - 1; const UINT64_MAX = (1n << 64n) - 1n; @@ -97,10 +99,57 @@ export interface SceneSignalGrid { readonly levels: readonly number[]; } +export type ChannelKind = "species" | "signals"; +export interface ChannelMetadata { + readonly species: readonly (string | null)[]; + readonly signals: readonly (string | null)[]; +} + +// Metadata arrays are readonly. Weak keys retain no discarded frame/history; +// resolving once also avoids rebuilding a whole group for every selector row. +const displayChannelLabels = new WeakMap< + readonly (string | null)[], + readonly string[] +>(); + +/** Presentation only. Numerical indices, never labels, identify channels. */ +export function channelLabel( + frame: SceneFrame, + kind: ChannelKind, + index: number, +): string { + const labels = frame.channelMetadata[kind]; + if (!Number.isInteger(index) || index < 0 || index >= labels.length) { + throw new RangeError(`${kind} channel ${index} is out of range`); + } + let resolved = displayChannelLabels.get(labels); + if (resolved === undefined) { + const names = labels.map((label, slot) => + label === null || label.trim() === "" + ? `Channel ${slot}` + : label.replace(/[\t\n\f\r ]+/g, " ").replace(/^ | $/g, ""), + ); + const counts = new Map(); + for (const name of names) counts.set(name, (counts.get(name) ?? 0) + 1); + resolved = names.map((name, slot) => + counts.get(name)! > 1 ? `${name} [${slot}]` : name, + ); + // A supplied name can imitate an automatically indexed duplicate, e.g. + // ["GFP", "GFP", "GFP [0]"]. In that case index the entire group: the + // distinct final indices guarantee uniqueness even for nested suffixes. + if (new Set(resolved).size !== resolved.length) { + resolved = names.map((name, slot) => `${name} [${slot}]`); + } + displayChannelLabels.set(labels, resolved); + } + return resolved[index]!; +} + export interface SceneFrame { readonly time: number; readonly backend: SceneBackend; readonly speciesCount: number; + readonly channelMetadata: ChannelMetadata; readonly cells: readonly SceneCell[]; readonly constraints: SceneConstraints; readonly signalGrid: SceneSignalGrid | null; @@ -478,11 +527,10 @@ function parseSignalGrid(value: unknown, path: string): SceneSignalGrid | null { "boundaries", "levels", ]); - const signalCount = integer( + const signalCount = sceneChannelCount( data.signal_count, `${path}.signal_count`, 1, - UINT32_MAX, ); const shapeValues = array(data.shape, `${path}.shape`); if (shapeValues.length !== 3) { @@ -551,7 +599,54 @@ function parseSignalGrid(value: unknown, path: string): SceneSignalGrid | null { }; } -function parseFrame(value: unknown, path: string): SceneFrame { +function parseChannelMetadata( + value: unknown, + path: string, + speciesCount: number, + signalCount: number, +): ChannelMetadata { + const data = record(value, path); + exactKeys(data, path, ["species", "signals"]); + function labels( + kind: ChannelKind, + count: number, + ): readonly (string | null)[] { + const values = array(data[kind], `${path}.${kind}`); + if (values.length !== count) { + return fail( + `${path}.${kind}`, + `expected ${count} labels, got ${values.length}`, + ); + } + return values.map((value, index) => { + if (value === null) return null; + const label = string(value, `${path}.${kind}[${index}]`); + if (/[\uD800-\uDFFF]/u.test(label)) + return fail( + `${path}.${kind}[${index}]`, + "invalid Unicode scalar value", + ); + return label; + }); + } + return { + species: labels("species", speciesCount), + signals: labels("signals", signalCount), + }; +} + +function sceneChannelCount(value: unknown, path: string, minimum = 0): number { + const count = integer(value, path, minimum, UINT32_MAX); + if (count > MAX_SCENE_CHANNELS) { + return fail( + path, + `exceeds scene presentation channel budget of ${MAX_SCENE_CHANNELS} per group`, + ); + } + return count; +} + +function parseFrame(value: unknown, path: string, version: number): SceneFrame { const data = record(value, path); exactKeys(data, path, [ "time", @@ -560,16 +655,15 @@ function parseFrame(value: unknown, path: string): SceneFrame { "cells", "constraints", "signal_grid", + ...(version >= 3 ? ["channel_metadata"] : []), ]); const time = number(data.time, `${path}.time`); if (time < 0) { return fail(`${path}.time`, "must be non-negative"); } - const speciesCount = integer( + const speciesCount = sceneChannelCount( data.species_count, `${path}.species_count`, - 0, - UINT32_MAX, ); const cells = array(data.cells, `${path}.cells`).map((item, index) => parseCell(item, `${path}.cells[${index}]`, speciesCount), @@ -587,13 +681,29 @@ function parseFrame(value: unknown, path: string): SceneFrame { } identifiers.add(cell.id); } + const signalGrid = parseSignalGrid(data.signal_grid, `${path}.signal_grid`); + const signalCount = signalGrid?.signalCount ?? 0; + // Digest verification has already completed against the unmodified source frame. + const channelMetadata = + version >= 3 + ? parseChannelMetadata( + data.channel_metadata, + `${path}.channel_metadata`, + speciesCount, + signalCount, + ) + : { + species: Array(speciesCount).fill(null), + signals: Array(signalCount).fill(null), + }; return { time, + channelMetadata, backend: parseBackend(data.backend, `${path}.backend`), speciesCount, cells, constraints: parseConstraints(data.constraints, `${path}.constraints`), - signalGrid: parseSignalGrid(data.signal_grid, `${path}.signal_grid`), + signalGrid, }; } @@ -639,7 +749,8 @@ export async function parseScene(source: string): Promise { ) { return fail("$.format", "not a MicroSimulator scene"); } - if (integer(root.version, "$.version", 0, UINT32_MAX) !== SCENE_VERSION) { + const version = integer(root.version, "$.version", 0, UINT32_MAX); + if (version !== 2 && version !== SCENE_VERSION) { return fail( "$.version", `unsupported scene version ${String(root.version)}`, @@ -672,5 +783,5 @@ export async function parseScene(source: string): Promise { if (actualDigest !== expectedDigest) { return fail("$.integrity.frame", "frame digest does not match"); } - return parseFrame(root.frame, "$.frame"); + return parseFrame(root.frame, "$.frame", version); } diff --git a/viewer/tests/channels.test.ts b/viewer/tests/channels.test.ts new file mode 100644 index 0000000..dc678ee --- /dev/null +++ b/viewer/tests/channels.test.ts @@ -0,0 +1,160 @@ +import canonicalize from "canonicalize"; +import { describe, expect, it } from "vitest"; + +import { mapCellColors } from "../src/color"; +import { channelLabel, parseScene } from "../src/scene"; +import pythonScene from "./fixtures/channels-v3.scene.json?raw"; + +async function sign(frame: unknown, version = 3): Promise { + const bytes = new TextEncoder().encode(canonicalize(frame)); + const digest = [ + ...new Uint8Array(await crypto.subtle.digest("SHA-256", bytes)), + ] + .map((byte) => byte.toString(16).padStart(2, "0")) + .join(""); + return JSON.stringify({ + format: "microsimulator-scene", + version, + producer: { name: "test", version: "1" }, + integrity: { algorithm: "sha256", frame: digest }, + frame, + }); +} + +interface Document { + version: number; + frame: { + time: number; + channel_metadata?: { species: unknown[]; signals: unknown[] }; + }; +} + +const document = (): Document => JSON.parse(pythonScene) as Document; + +describe("channel metadata", () => { + it("reads the Python-produced v3 fixture with Unicode, markup and missing labels", async () => { + const frame = await parseScene(pythonScene); + expect(frame.channelMetadata).toEqual({ + species: ["α 🧪", "α 🧪"], + signals: [null, " "], + }); + expect(channelLabel(frame, "species", 0)).toBe("α 🧪 [0]"); + expect(channelLabel(frame, "species", 1)).toBe("α 🧪 [1]"); + expect(channelLabel(frame, "signals", 0)).toBe("Channel 0"); + expect(channelLabel(frame, "signals", 1)).toBe("Channel 1"); + expect(() => channelLabel(frame, "species", 2)).toThrow(RangeError); + }); + + it("keeps numerical values/color selection independent of label text", async () => { + const frame = await parseScene(pythonScene); + const config = { mode: "species" as const, speciesIndex: 1 }; + const renamed = { + ...frame, + channelMetadata: { + species: ["Changed", "Chosen"], + signals: [null, null], + }, + }; + const before = mapCellColors(frame, config); + const after = mapCellColors(renamed, config); + expect(before.colors).toEqual(after.colors); + expect(after.title).toBe("Chosen"); + expect(renamed.cells).toEqual(frame.cells); + }); + + it("disambiguates explicit names that collide with missing-label fallbacks", async () => { + const frame = await parseScene(pythonScene); + const labels = { + ...frame, + channelMetadata: { species: [null, "Channel 0"], signals: [] }, + }; + expect(channelLabel(labels, "species", 0)).toBe("Channel 0 [0]"); + expect(channelLabel(labels, "species", 1)).toBe("Channel 0 [1]"); + }); + + it.each(["species", "signals"] as const)( + "keeps %s labels unique when names imitate generated disambiguation suffixes", + async (kind) => { + const frame = await parseScene(pythonScene); + for (const names of [ + ["GFP", "GFP", "GFP [0]"], + [null, "Channel 0", "Channel 0 [0]"], + ["GFP", "GFP", "GFP [0]", "GFP [0] [2]"], + ]) { + const renamed = { + ...frame, + channelMetadata: { ...frame.channelMetadata, [kind]: names }, + }; + const displayed = names.map((_, index) => + channelLabel(renamed, kind, index), + ); + expect(new Set(displayed).size).toBe(names.length); + expect(displayed[0]).toBe(`${names[0] ?? "Channel 0"} [0]`); + expect(displayed[2]).toBe(`${names[2]} [2]`); + } + }, + ); + + it.each(["species", "signals"] as const)( + "disambiguates %s names after browser whitespace normalization", + async (kind) => { + const frame = await parseScene(pythonScene); + const renamed = { + ...frame, + channelMetadata: { + ...frame.channelMetadata, + [kind]: [" GFP\tname ", "GFP name", "GFP\t name [0]"], + }, + }; + expect( + [0, 1, 2].map((index) => channelLabel(renamed, kind, index)), + ).toEqual(["GFP name [0]", "GFP name [1]", "GFP name [0] [2]"]); + }, + ); + + it("authenticates v2 without inserting metadata before digest verification", async () => { + const old = document(); + delete old.frame.channel_metadata; + const encoded = await sign(old.frame, 2); + const frame = await parseScene(encoded); + expect(frame.channelMetadata).toEqual({ + species: [null, null], + signals: [null, null], + }); + expect(channelLabel(frame, "species", 1)).toBe("Channel 1"); + const tampered = JSON.parse(encoded) as Document; + tampered.frame.time += 1; + await expect(parseScene(JSON.stringify(tampered))).rejects.toThrow( + "frame digest does not match", + ); + }); + + it("rejects label tampering even when channel counts are unchanged", async () => { + const tampered = document(); + tampered.frame.channel_metadata!.species[0] = "forged"; + await expect(parseScene(JSON.stringify(tampered))).rejects.toThrow( + "frame digest does not match", + ); + }); + + it("rejects invalid Unicode before integrity verification", async () => { + const invalid = document(); + invalid.frame.channel_metadata!.species[0] = "\ud800"; + await expect(parseScene(JSON.stringify(invalid))).rejects.toThrow( + "Lone surrogate", + ); + }); + + it.each([ + { species: [null], signals: [null, null] }, + { species: [null, 3], signals: [null, null] }, + { species: [null, null], signals: [null, null], extra: true }, + ])( + "rejects malformed metadata after valid integrity verification: %j", + async (metadata) => { + const invalid = document(); + invalid.frame.channel_metadata = metadata; + await expect(parseScene(await sign(invalid.frame))).rejects.toThrow(); + }, + ); +}); diff --git a/viewer/tests/color.test.ts b/viewer/tests/color.test.ts index 8da7817..5ec1f98 100644 --- a/viewer/tests/color.test.ts +++ b/viewer/tests/color.test.ts @@ -34,6 +34,7 @@ const frame: SceneFrame = { native: false, }, speciesCount: 2, + channelMetadata: { species: [null, null], signals: [] }, constraints: { planes: [], spheres: [], boxes: [], cylinders: [] }, cells: [ cell(0, -1, 0.5, [2, 8]), diff --git a/viewer/tests/fixtures/channels-v3.scene.json b/viewer/tests/fixtures/channels-v3.scene.json new file mode 100644 index 0000000..a74880e --- /dev/null +++ b/viewer/tests/fixtures/channels-v3.scene.json @@ -0,0 +1,85 @@ +{ + "format": "microsimulator-scene", + "frame": { + "backend": { + "device": "host", + "device_index": 0, + "kind": "cpu", + "name": "cpu-reference", + "native": true + }, + "cells": [ + { + "cell_type": 0, + "direction": [1.0, 0.0, 0.0], + "fixed": false, + "growth_rate": 1.0, + "id": "1", + "length": 2.0, + "parent_id": null, + "position": [0.0, 0.0, 0.0], + "radius": 0.5, + "slot": 0, + "species": [0.25, 0.75] + } + ], + "channel_metadata": { + "signals": [null, " "], + "species": ["α 🧪", "α 🧪"] + }, + "constraints": { + "boxes": [], + "cylinders": [], + "planes": [], + "spheres": [] + }, + "signal_grid": { + "boundaries": { + "x_lower": { + "kind": "no_flux", + "values": [] + }, + "x_upper": { + "kind": "no_flux", + "values": [] + }, + "y_lower": { + "kind": "no_flux", + "values": [] + }, + "y_upper": { + "kind": "no_flux", + "values": [] + }, + "z_lower": { + "kind": "no_flux", + "values": [] + }, + "z_upper": { + "kind": "no_flux", + "values": [] + } + }, + "levels": [ + 0.25, 0.25, 0.25, 0.25, 0.25, 0.25, 0.25, 0.25, 0.25, 0.25, 0.25, 0.25, + 0.25, 0.25, 0.25, 0.25, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, + 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75 + ], + "origin": [0.0, 0.0, 0.0], + "shape": [4, 4, 1], + "signal_count": 2, + "spacing": [1.0, 1.0, 1.0] + }, + "species_count": 2, + "time": 0.0 + }, + "integrity": { + "algorithm": "sha256", + "frame": "4ed41b5e8443c5da0da1843135647a0ed1d36b1681a2e3fc902199dd75dd3ef7" + }, + "producer": { + "name": "microsimulator", + "version": "0.1.0" + }, + "version": 3 +} diff --git a/viewer/tests/presentation-state.test.ts b/viewer/tests/presentation-state.test.ts index aca53de..299050e 100644 --- a/viewer/tests/presentation-state.test.ts +++ b/viewer/tests/presentation-state.test.ts @@ -28,6 +28,10 @@ const frame: SceneFrame = { native: true, }, speciesCount: 3, + channelMetadata: { + species: Array(3).fill(null), + signals: Array(3).fill(null), + }, cells: [], constraints: { planes: [], spheres: [], boxes: [], cylinders: [] }, signalGrid: grid, diff --git a/viewer/tests/reference-grid.test.ts b/viewer/tests/reference-grid.test.ts index d11c377..a92b59c 100644 --- a/viewer/tests/reference-grid.test.ts +++ b/viewer/tests/reference-grid.test.ts @@ -35,6 +35,10 @@ const frame: SceneFrame = { native: true, }, speciesCount: 0, + channelMetadata: { + species: Array(0).fill(null), + signals: Array(0).fill(null), + }, cells: [cell], constraints, signalGrid: null, diff --git a/viewer/tests/scene.test.ts b/viewer/tests/scene.test.ts index c5eda20..fef3c00 100644 --- a/viewer/tests/scene.test.ts +++ b/viewer/tests/scene.test.ts @@ -1,7 +1,7 @@ import canonicalize from "canonicalize"; import { describe, expect, it } from "vitest"; -import { parseScene, SceneFormatError } from "../src/scene"; +import { MAX_SCENE_CHANNELS, parseScene, SceneFormatError } from "../src/scene"; const PYTHON_SCENE = `{ "format": "microsimulator-scene", @@ -117,7 +117,101 @@ async function digest(value: unknown): Promise { .join(""); } +async function channelBudgetScene( + version: number, + speciesCount: number, + signalCount: number, + complete: boolean, +): Promise { + const boundary = { kind: "no_flux", values: [] }; + const frame = { + time: 0, + backend: { + kind: "cpu", + name: "CPU reference", + device: "host", + device_index: 0, + native: false, + }, + species_count: speciesCount, + cells: [], + constraints: { planes: [], spheres: [], boxes: [], cylinders: [] }, + signal_grid: { + signal_count: signalCount, + shape: [1, 1, 1], + origin: [0, 0, 0], + spacing: [1, 1, 1], + boundaries: { + x_lower: boundary, + x_upper: boundary, + y_lower: boundary, + y_upper: boundary, + z_lower: boundary, + z_upper: boundary, + }, + levels: Array(complete ? signalCount : 1).fill(0), + }, + ...(version === 3 + ? { + channel_metadata: { + species: Array(complete ? speciesCount : 0).fill(null), + signals: Array(complete ? signalCount : 0).fill(null), + }, + } + : {}), + }; + return JSON.stringify({ + format: "microsimulator-scene", + version, + producer: { name: "microsimulator", version: "0.1.0" }, + frame, + integrity: { algorithm: "sha256", frame: await digest(frame) }, + }); +} + describe("scene reader", () => { + it.each([2, 3])( + "accepts both channel groups at the inclusive scene budget in v%i", + async (version) => { + const frame = await parseScene( + await channelBudgetScene( + version, + MAX_SCENE_CHANNELS, + MAX_SCENE_CHANNELS, + true, + ), + ); + expect(frame.cells).toEqual([]); + expect(frame.channelMetadata.species).toEqual( + Array(MAX_SCENE_CHANNELS).fill(null), + ); + expect(frame.channelMetadata.signals).toEqual( + Array(MAX_SCENE_CHANNELS).fill(null), + ); + expect(frame.signalGrid?.levels).toHaveLength(MAX_SCENE_CHANNELS); + }, + ); + + it.each([2, 3])( + "rejects tiny signed oversized channel claims before allocating v%i labels", + async (version) => { + for (const count of [MAX_SCENE_CHANNELS + 1, 2 ** 32 - 1]) { + for (const group of ["species", "signals"] as const) { + const encoded = await channelBudgetScene( + version, + group === "species" ? count : 0, + group === "signals" ? count : 1, + false, + ); + expect(encoded.length).toBeLessThan(2000); + await expect(parseScene(encoded)).rejects.toThrow( + `${group === "species" ? "$.frame.species_count" : "$.frame.signal_grid.signal_count"}: exceeds scene presentation channel budget of 4096 per group`, + ); + } + } + }, + ); + it("verifies and reads a Python-authored RFC 8785 scene", async () => { const frame = await parseScene(PYTHON_SCENE); expect(frame.time).toBe(1);