Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 0 additions & 2 deletions .github/workflows/static_analysis.yml
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,6 @@ jobs:
run: |
pip install black "black[jupyter]"
black --check src/
black --check tutorials/

isort:
runs-on: ubuntu-latest
Expand All @@ -50,7 +49,6 @@ jobs:
run: |
pip install isort
isort --check src/
isort --check tutorials/

mypy:
runs-on: ubuntu-latest
Expand Down
6 changes: 1 addition & 5 deletions .github/workflows/testing.yml
Original file line number Diff line number Diff line change
Expand Up @@ -53,15 +53,11 @@ jobs:
run: |
pytest .

- name: Test tutorials
run: |
jupyter nbconvert --to notebook --execute tutorials/*.ipynb --output-dir=/tmp --ExecutePreprocessor.timeout=300

- name: Test docs build
run: |
pip install ".[docs]"
cd docs
make clean
make html
cd ..
ls docs/build/html/index.html
ls docs/build/html/index.html
21 changes: 21 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,28 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

## [0.2.0] - 2026-09-28

### Changed
- `xp.get_backend()` now returns the active global backend, mirroring `xp.set_backend()`. Use the new `xp.get_array_backend(array)` to inspect an individual array. Calls to the former `xp.get_backend(array)` must be updated.
- Expanded the README and documentation with backend selection, array conversion, GPU controls, kernel adaptation, and Pyodide examples.

### Fixed
- GPU shape round-trip coverage now creates a zero-dimensional NumPy array for the scalar case instead of calling `.astype()` on a Python float.
- GPU tests now account for CuPy's retained split memory blocks and provide an array conversion method on the custom host-array test fixture.
- `set_backend()` and `use_backend()` now reject unsupported backend names with `ValueError` without changing the active selection.

### Added
- `xp.get_array_module(array)`: Return the array-api-compat module (`numpy`/`cupy`) matching a given array's own backend, regardless of the process-wide active backend. Mirrors `cupy.get_array_module` and works in Pyodide.
- `xp.get_rng(seed=None)`: Return a `numpy.random.Generator`/`cupy.random.Generator` matching the active backend, without having to branch on the backend yourself.
- `xp.device_count()`: Number of visible CUDA devices (`0` without a functional CuPy/CUDA install), independent of the currently active backend.
- `xp.set_device_for_rank(rank, devices_per_node=None)`: Convenience for one-MPI-rank-per-GPU codes; selects `rank % devices_per_node` (defaulting `devices_per_node` to `device_count()`) via `set_device()` and returns the chosen device id.
- `xp.memory_info()`: `(free, total)` bytes of memory on the active CUDA device, or `None` on the NumPy backend.
- `xp.free_memory()`: Release all free blocks held by CuPy's device and pinned-host memory pools (no-op on the NumPy backend).
- `xp.default_float_dtype()`: Return the active backend's `float64` dtype object, for pinning a portable float precision instead of the backend/platform-dependent `dtype=float`.
- `xp.stream()`: Context manager for a CUDA stream, to overlap host/device transfers with compute (no-op, yielding `None`, on the NumPy backend).
- `xp.pin_memory(array)`: Copy a host array into pinned (page-locked) CUDA host memory for faster transfers.
- `PyccelKernel(..., is_array=...)`: Extension point overriding the default `isinstance(value, np.ndarray)` check used to decide which host values returned by (or reachable from a declared output of) the wrapped kernel are converted back to the device -- for kernels that return/mutate a NumPy subclass or other custom host array type.
- Pyodide NumPy support documentation and CI that installs the built wheel in Pyodide's WebAssembly runtime and runs compiler-free tests for arrays, conversions, contexts, and Python kernels without CuPy or Pyccel imports.
- `test-compiled` extra for native Pyccel tests. The `test` extra is now compiler-free; `dev` continues to include compiled-test dependencies.
- `xp.same_backend(*arrays)`: Return `True` if all given arrays live on the same backend.
Expand Down
254 changes: 183 additions & 71 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,115 +1,227 @@
# CuNumpy

Simple wrapper for numpy and cupy. Replace `import numpy as np` with `import cunumpy as xp`.
CuNumpy lets a Python program use a NumPy-like API while choosing NumPy arrays
on the CPU or CuPy arrays on an NVIDIA GPU. In the simplest case, replace
`import numpy as np` with `import cunumpy as xp`; the array operations you
already know then run on the selected backend.

# Install
```python
import cunumpy as xp

```bash
pip install cunumpy
values = xp.arange(5, dtype=xp.float64)
print(values * 2)
print(xp.get_backend()) # 'numpy' by default
```

Example usage:
CuNumpy selects an array library for newly requested operations. It does not
move existing arrays just because the selected backend changes. This guide
covers backend selection, array movement, mixed CPU/GPU workflows, and the
helper APIs CuNumpy provides around NumPy and CuPy.

## Install

```bash
python -m pip install cunumpy
```
export ARRAY_BACKEND=cupy
```

NumPy and `array-api-compat` are installed as dependencies. To use a GPU,
install a CuPy package compatible with your CUDA environment as well. CuPy
installation depends on the CUDA version and platform; follow the CuPy
installation instructions for your system. CuNumpy does not install CUDA.

`array-api-compat` supplies NumPy and CuPy compatibility modules with more
consistent behavior for shared array operations. CuNumpy uses them internally;
your arrays remain ordinary NumPy or CuPy arrays. See [why CuNumpy uses
`array-api-compat`](docs/source/array-api-compat.md) for a plain-language
explanation and examples.

## Choose a backend

CuNumpy starts with NumPy unless `ARRAY_BACKEND=cupy` is set before import.
You can also choose at runtime:

```python
import cunumpy as xp

xp.set_backend("cupy")
print(xp.get_backend()) # 'cupy' if CuPy and CUDA are functional

values = xp.arange(5) # created by the active backend
```

The accepted backend names are `"numpy"` and `"cupy"`. If CuPy is requested
but unavailable or not functional, CuNumpy falls back to NumPy. Always check
`get_backend()` when the effective backend matters, such as when reporting
configuration or deciding whether GPU-specific work will happen.

arr = xp.array([1, 2])
Use `use_backend()` for a temporary selection. It restores the previous
selection when the block exits, including when an exception is raised:

print(f"{type(arr) = }")
print(f"{xp.__version__ = }")
```python
with xp.use_backend("numpy"):
cpu_values = xp.linspace(0, 1, 100)
assert xp.get_backend() == "numpy"

# Convert to NumPy
arr_np = xp.to_numpy(arr)
# The previous global backend is active again here.
```

# Convert to active backend
arr_xp = xp.to_cunumpy(arr)
The backend selection is process-wide shared state. Do not switch it
independently from multiple threads or async tasks; those changes can
interfere. A context manager is useful for sequential code, tests, and
notebooks.

# Inspect backend
print(f"{xp.get_backend(arr) = }")
print(f"{xp.is_gpu(arr) = }")
print(f"{xp.is_cpu(arr) = }")
## Understand the two backend questions

# Temporarily switch backend
with xp.use_backend("numpy"):
# This code runs on CPU even if ARRAY_BACKEND=cupy
arr_cpu = xp.zeros(100)
The active backend controls which library CuNumpy exposes through its NumPy
like operations. The array backend reports where one particular array lives.
These can differ: changing the active backend does not convert arrays that
already exist.

# Set backend globally
xp.set_backend("cupy")
```python
xp.set_backend("numpy")
cpu_values = xp.arange(3)

# Synchronize GPU operations (no-op on CPU)
xp.synchronize()
gpu_values = xp.to_cupy(cpu_values) # explicit transfer
print(xp.get_backend()) # 'numpy'
print(xp.get_array_backend(gpu_values)) # 'cupy'
```

Output:
Use `is_cpu(array)`, `is_gpu(array)`, or `get_array_backend(array)` when
dispatch should follow the array passed to a function. `get_array_module()`
returns the matching `array_api_compat` module, which is useful when writing
backend-generic functions:

```python
def vector_norm(values):
array_xp = xp.get_array_module(values)
return array_xp.sqrt(array_xp.sum(values * values))
```
type(arr) = <class 'cupy.ndarray'>
xp.__version__ = '0.1.4'
xp.get_backend(arr) = 'cupy'
xp.is_gpu(arr) = True
xp.is_cpu(arr) = False

## Move data between CPU and GPU

Transfers are explicit so it is clear when data crosses the CPU/GPU boundary:

```python
host = xp.to_numpy(gpu_values) # CuPy -> NumPy (host)
device = xp.to_cupy(host) # NumPy/array-like -> CuPy (device)
active = xp.to_cunumpy(host) # convert to the currently selected backend
```

# Pyodide
`to_numpy()` also accepts ordinary array-like values. `to_cupy()` raises
`ImportError` when CuPy or a functional CUDA runtime is unavailable.
`to_cunumpy()` is useful at API boundaries where the consumer expects the
currently selected backend. It does not change the original array.

cuNumPy supports Pyodide with the NumPy backend. In an initialized Pyodide
JavaScript runtime, install the package and run a Python-source kernel:
Avoid transferring data inside a tight loop. Keep intermediate arrays on one
backend and move only at boundaries such as file I/O, plotting, or a
CPU-only library call. For example:

```javascript
await pyodide.loadPackage("micropip");
await pyodide.runPythonAsync(`
import micropip
await micropip.install("cunumpy")
```python
with xp.use_backend("cupy"):
signal = xp.asarray(host_signal)
filtered = xp.fft.rfft(signal)
result = xp.to_numpy(filtered) # one transfer for a CPU-only consumer
```

import cunumpy as xp
xp.set_backend("numpy")
## Random numbers and dtypes

def scale(values, factor):
values[:] *= factor
return values
`get_rng(seed)` returns a random generator for the active backend. NumPy and
CuPy have similar generator APIs, though exact bit-for-bit sequences are not
guaranteed to match between libraries:

values = xp.array([1.0, 2.0, 3.0])
result = xp.PyccelKernel(scale)(values, 2.0)
assert result is values
print(xp.to_numpy(result)) # [2. 4. 6.]
`);
```python
rng = xp.get_rng(seed=42)
samples = rng.normal(size=1000)
```

NumPy is the default backend when `ARRAY_BACKEND` is unset. CuPy/CUDA and
native Pyccel compilation are not supported in Pyodide. Despite its name,
`PyccelKernel` accepts ordinary Python callables and neither imports Pyccel nor
compiles code. Applications must supply Python-source kernels with
Pyodide-compatible imports.
Use `default_float_dtype()` when code needs to explicitly request the active
backend's `float64` dtype rather than rely on Python scalar inference:

CI tests the built wheel in Pyodide's WebAssembly runtime under Node.js, including
array operations, conversions, mutation/aliasing, backend context restoration,
and Python kernels, while rejecting CuPy or Pyccel imports. This does not test
browser-specific integration such as page loading or workers. See the
[Pyodide guide](docs/source/pyodide.md) for details and local test commands.
```python
x = xp.asarray([1.0, 2.0], dtype=xp.default_float_dtype())
```

# Development tests
## GPU selection and memory helpers

```bash
pip install -e '.[test]'
pytest tests/portable
These helpers are useful for multi-GPU programs and for understanding CuPy's
memory behavior:

```python
print("visible GPUs:", xp.device_count())
xp.set_device(0) # selects CUDA device 0 when CuPy is active
print("memory (free, total):", xp.memory_info())
```

The `test` extra is compiler-free. For compiled-kernel tests, install
`'.[test-compiled]'` and a working native compiler, then run `pytest`.
The `dev` extra includes these compiled-test dependencies as before.
`set_device()` is a no-op on NumPy. `device_count()` checks visible CUDA
hardware even if the active backend is NumPy; it returns zero when CuPy/CUDA
cannot be used. `memory_info()` returns `(free_bytes, total_bytes)` on the
active CuPy device and `None` on NumPy. `set_device_for_rank(rank)` is a
round-robin convenience for MPI layouts where local ranks map contiguously to
GPUs. If your scheduler uses a different mapping, select the device directly.

# Build docs
CuPy caches released allocations in memory pools. This can make process-level
GPU memory appear occupied after arrays go out of scope. `free_memory()` asks
CuPy to release currently free cached blocks; it does not free memory still
referenced by live arrays.

`pin_memory(host_array)` makes a pinned host copy, which can improve transfer
throughput for workloads that explicitly manage asynchronous transfers.
`stream()` creates a non-blocking CuPy stream and yields it; it yields `None`
on NumPy. GPU work is asynchronous, so synchronize before reading results on
the host:

```python
with xp.stream():
device = xp.to_cupy(host)
transformed = xp.fft.fft(device)

xp.synchronize()
result = xp.to_numpy(transformed)
```
make html
cd ../
open docs/_build/html/index.html

## Use NumPy-only kernels with CuPy arrays

`PyccelKernel` adapts a callable that expects NumPy arrays. When conversion is
needed, CuNumpy copies CuPy inputs to the host, calls the wrapped function,
copies in-place output changes back to the device, and moves returned NumPy
arrays to CuPy. With NumPy inputs, the wrapper calls the function directly.
CuNumpy does not compile functions or import Pyccel for you.

```python
import cunumpy as xp


def scale_in_place(values, factor):
values[:] *= factor
return values


scale = xp.PyccelKernel(scale_in_place, outputs=(0,))

with xp.use_backend("cupy"):
values = xp.arange(5, dtype=xp.float64)
returned = scale(values, 3.0)
xp.synchronize()
```

By default every converted argument is copied back, because the wrapper cannot
know which arguments the kernel changed. `outputs=(0,)` declares that
positional argument 0 is written, avoiding unnecessary copy-back for
read-only inputs. For a keyword call, declare the keyword name, such as
`outputs=("out",)`. A wrong declaration can leave GPU output values stale.
The wrapper can also traverse arrays nested in lists, tuples, dictionaries,
and selected application objects; see the full [API reference](docs/source/api.md)
for `object_modules`, `is_array`, aliasing, and output declarations.

## Pyodide

CuNumpy supports the NumPy backend in Pyodide. It does not provide CuPy/CUDA
there. Ordinary Python callables can be wrapped with `PyccelKernel` without
compilation. See the [Pyodide guide](docs/source/pyodide.md) for a complete
installation example and compatibility notes.

## Documentation

The [user guide](docs/source/quickstart.md) explains common workflows. The
[API reference](docs/source/api.md) documents each helper and its behavior.
The [Pyodide guide](docs/source/pyodide.md) covers WebAssembly usage.
Loading
Loading