load#
Reads compressed 4D-STEM data straight onto CUDA or Apple Metal and returns a
typed LoadResult accepted directly by Show4DSTEM. Public import:
from quantem.gpu.io import load
Reference#
- quantem.gpu.io.load(source: str | ~os.PathLike[str] | ~collections.abc.Sequence[str | ~os.PathLike[str]], *, dtype: str | type | ~numpy.dtype | None = None, backend: str = 'auto', dataset_path: str | None = None, scan_shape: tuple[int, int] | None = None, scan_order: ~typing.Literal['row-major', 'serpentine'] = 'row-major', scan_region: tuple[int, int, int, int] | ~collections.abc.Sequence[tuple[int, int, int, int]] | None = None, detector_region: tuple[int, int, int, int] | None = None, target_scan_region: tuple[int, int, int, int] | None = None, scan_shift_row_col: tuple[float, float] | ~collections.abc.Sequence[tuple[float, float]] | ~numpy.ndarray | None = None, scan_resample_dtype: type | ~numpy.dtype = <class 'numpy.float32'>, scan_indices: ~collections.abc.Sequence[int] | ~collections.abc.Sequence[tuple[int, int]] | ~numpy.ndarray | None = None, index_mode: str = 'scan', random_positions: int | None = None, seed: int | None = None, replace: bool = False, same_random_positions: bool = False, drift: ~collections.abc.Sequence | ~numpy.ndarray | None = None, det_bin: int = 1, apply_mask: bool = True, auto_narrow: bool = True, output: str = 'native', stack: bool = True, device: int | str | None = None, devices: list[int] | str | None = None, verbose: bool = True) LoadResult | list[LoadResult]#
Load one or more 4D-STEM sources through an accelerated backend.
All spatial arguments use
(row, col)order.dtypeis the only output-precision control; use"u8"only when the count range proves it lossless, and use"u16"or the native dtype for scientific workflows.backend="auto"selects CUDA or MPS and never selects CPU silently.output="native"preserves the backend-native payload; useoutput="torch"when the consumer expects a Torch tensor.- Parameters:
source – One master/data HDF5 path, a folder, or a list of master paths.
dtype – Requested output dtype, such as
"u8","u16","u32","f32","u4","native", or"auto".backend –
"auto","cuda","mps", or explicit reference"cpu".output –
"native"(default) returns the backend-native array or chunked source."torch"converts the loaded result to a Torch tensor, assembling chunked MPS data when necessary.scan_shape – Optional full scan shape as
(row, col).scan_region – Optional bounds as
(row_start, row_stop, col_start, col_stop).detector_region – Optional bounds as
(row_start, row_stop, col_start, col_stop).target_scan_region – Shared target crop and per-source row/column shifts for drift-aware multi-file loading.
scan_shift_row_col – Shared target crop and per-source row/column shifts for drift-aware multi-file loading.
drift – Optional dense per-source additive drift fields for a sparse ptychography batch, shaped
(n_files, scan_rows, scan_cols, 2)or(n_files, 2, scan_rows, scan_cols). The loader keeps raw diffraction patterns unchanged and records corrected floating-point probe positions inresult.metadata["drift_batch"]. Use withrandom_positions=orscan_indices=; this is intentionally separate from scan-space resampling.devices – CUDA devices for a multi-GPU load. With
stack=Truethe result may be sharded by device; withstack=Falseeach source is returned separately.
- Returns:
Data stays backend-resident. MPS list/folder loads return a common multi-frame detector object in
LoadResult.datawhile background decoding fills its dataset slots.- Return type:
LoadResult or list[LoadResult]
Tip
The default load keeps native detector sampling when the memory budget allows.
Use det_bin=2 or 4 only as an explicit preview or memory policy; pass a list
of file paths to stack several datasets behind a single “Dataset” slider.
Backend (CUDA / Apple Silicon)#
load detects the native GPU automatically: an NVIDIA box loads onto CUDA
and a Mac loads onto Apple Metal (MPS). Scientific loading does not silently
fall back to CPU.
from quantem.gpu.io import load
from quantem.widget import Show4DSTEM
result = load("scan_master.h5") # CUDA on a workstation, MPS on a MacBook
Show4DSTEM(result)
On a MacBook the read uses a raw-Metal path, so a memory-rich laptop can browse native 4D-STEM directly and smaller machines can use an explicit detector-bin preview. A conservative preview session is:
result = load("scan_master.h5", det_bin=8) # MPS, detector binned 8x -> small preview
Show4DSTEM(result)
MPS loads include a preflight memory guard. Before allocating Metal buffers, the
loader estimates the output footprint from HDF5 metadata and compares it to the
Mac’s recommended Metal working set. No-bin loads that exceed that budget fail
early with a specific det_bin recommendation instead of risking a frozen
laptop:
data = load("scan_master.h5", backend="mps", det_bin=4)
For the smallest browse workflows, combine detector binning with compact dtype:
result = load("scan_master.h5", backend="mps", det_bin=8, dtype="u8")
Show4DSTEM(result)
That path is intended for screening, layout, and export-to-browser review. Use uint16/full precision when detector counts are part of the scientific claim and the GPU memory budget is clean.
The same is one shell command - see the CLI: quantem show4dstem scan_master.h5 --bin 8.
Multi-File / 5D Browsing#
Use one public entry point for a time-series or tilt-series stack:
from quantem.gpu.io import load
from quantem.widget import Show4DSTEM
masters = [
"/data/session/file_001_master.h5",
"/data/session/file_002_master.h5",
"/data/session/file_003_master.h5",
]
r = load(masters, det_bin=1, dtype="u8")
w = Show4DSTEM(r)
The returned data has shape (n_files, scan_row, scan_col, det_row, det_col).
Show4DSTEM labels the extra axis as Dataset and uses the source filenames
as slider labels when they are available.
dtype="u8" is a browse contract, not a reconstruction contract. It routes to
the direct uint8 output path before stacking or sharding, so the loader avoids
materializing a full uint16 stack first. Counts above 255 clip to 255; use
dtype="u16" or omit dtype when detector counts are part of the scientific
claim.
For larger series or solver workflows that should stay sharded across GPUs, use:
parts = load(masters, det_bin=2, gpus=[0, 1], stack=False)
For full-detector browse-scale multi-GPU sessions, use the same public dtype vocabulary without detector binning when the memory budget allows:
parts = load(masters, det_bin=1, dtype="u8", devices=[0, 1])
The sharded result keeps one stack per GPU and records the file-to-device map in the metadata. It is disk-aware: when masters are split across independent NVMe mounts, the loader interleaves files by physical disk before assigning them to GPUs. When all files live on one disk, sharding still increases GPU capacity, but cold load remains disk-bound.
For an explicit logical series object:
series = load(
masters,
det_bin=1,
dtype="u8",
series_type="time",
series=[0.0, 1.0, 2.0],
units=["s", "pixels", "pixels", "mrad", "mrad"],
)
Memory rule of thumb for Sample-scale 512x512x192x192 data:
mode |
approximate resident size per file |
|---|---|
no bin, uint16 |
18-20 GiB |
|
4.5-5 GiB |
|
1.1-1.3 GiB |
|
about half of uint16 |
Use det_bin=4, dtype="u8" only for an explicit reduced preview on constrained
hardware. Use no-bin uint16 for count-preserving science, or no-bin
dtype="u8" for compact full-detector browsing when clipping has been accepted
or audited.
Scan-region loading for ROI workflows#
Use load(..., scan_region=...) when a reconstruction or denoise workflow
needs only a rectangular scan patch, not the full scan plane. This is different
from loading the full frame and slicing afterward: the loader reads only the
selected HDF5 detector-frame chunks, decompresses them on CUDA, and returns a
local patch.
from quantem.gpu.io import load
patch = load(
"scan_master.h5",
scan_region=(160, 293, 234, 367), # row_start, row_stop, col_start, col_stop
).data
print(patch.shape)
# (133, 133, 192, 192)
The returned LoadResult.data shape is
(region_rows, region_cols, detector_rows, detector_cols). Metadata records
both the original scan grid and the loaded patch:
result = load("scan_master.h5", scan_region=(160, 293, 234, 367))
print(result.metadata["full_scan_shape"]) # e.g. (512, 512)
print(result.metadata["scan_region"])
For drift-corrected time-series work, compute the source scan box from the shared specimen ROI plus a small halo, load that patch, then apply your existing subpixel sampler in local patch coordinates. Do not save the sampled patch as a new raw acquisition; drift is scan-position metadata, and detector counts stay physically unchanged.
Measured on a native-detector 5D-STEM ROI loader timing check
(10 x 128 x 128 x 192 x 192, CUDA, no detector binning):
Path |
Loader wall time |
Max loaded CuPy buffer |
|---|---|---|
|
|
|
|
|
|
The patch path is CUDA-only today and targets chunked 4D-STEM masters with one
detector frame per HDF5 chunk. Use load() for full-field browsing and for
Apple Metal/MPS until the region loader is ported there.