Download the PHP package lcmialichi/php-gpu-tensors without Composer
On this page you can find all versions of the php package lcmialichi/php-gpu-tensors. It is possible to download/install these versions without Composer. Possible dependencies are resolved automatically.
Download lcmialichi/php-gpu-tensors
More information about lcmialichi/php-gpu-tensors
Files in lcmialichi/php-gpu-tensors
Package php-gpu-tensors
Short Description Native PHP extension for NVIDIA GPU tensors, CUDA-accelerated computing, and machine-learning workloads.
License MIT
Informations about the package php-gpu-tensors
PHP GPU Tensors
GPU tensors, runtime-compiled CUDA kernels and fused tensor expressions for PHP.
No Python runtime required.
Run numerical workloads on an NVIDIA GPU without leaving PHP. This native extension gives PHP code GPU-resident tensors, common math operations, and a way to compile and launch your own CUDA kernels. For repeated work, you can optionally compile tensor expressions into replayable Fusion plans.
Status: beta · Linux · PHP 8.1–8.5, NTS and ZTS · NVIDIA GPU required.
0.1.0-beta.5 is the documented package baseline; this source checkout may
contain newer features. See Validation status for the
tested PHP, CUDA, and GPU combinations.
Is this a good fit?
Use PHP GPU Tensors when a PHP application needs numerical work on an NVIDIA GPU—for example, tensor calculations, custom CUDA operations, or a small machine-learning pipeline—and you want to keep that workflow in PHP.
It is a low-level GPU computing library, not a drop-in replacement for a full machine-learning framework. It requires Linux, an NVIDIA GPU, and the CUDA runtime; it does not fall back to CPU execution. Automatic differentiation is available for a documented subset of operations and is opt-in.
What you can do
- Work with GPU tensors. Create arrays and apply arithmetic, broadcasting, matrix multiplication, reductions, views, and Python-style slicing.
- Keep data on the device. Upload packed buffers or
.npyfiles, run a sequence of operations, and copy results back only when PHP needs them. - Write CUDA in PHP projects. Compile CUDA C++ at runtime with NVRTC and launch kernels synchronously or asynchronously.
- Reduce repeated work. Opt in to Fusion to compile a tensor-expression plan once and replay it. Compatible plans can use streams, asynchronous execution, and CUDA Graphs.
- Experiment with training. Use opt-in reverse-mode autograd and the end-to-end MLP classifier example as a starting point.
Contents: Results · GPU tensors · Data pipelines and autograd · Optimizers · Fusion · Streams and CUDA Graph · Training example · Validation · Contribute
Quick look
Tensor operations run on the GPU by default. Fusion is optional: use it when you want a reusable compiled plan for repeated work.
The CudaArray objects hold GPU data; toArray() is where this example copies
the result back to PHP. To explore the library, start with the
examples guide, which walks from basic tensors to custom
kernels, Fusion, and a complete training example.
For example, train the included classifier on the GPU:
Results
Measured on PHP 8.3 NTS with an NVIDIA GeForce MX570 A (4 GB, compute capability 8.6), driver 12.6 and CUDA runtime 12.3. The models are small, so these numbers mostly show how much per-step overhead Fusion removes; they are not peak GPU throughput and not a general-purpose GPU benchmark.
Fusion vs. eager execution on the same model and data, with identical metrics (accuracy 77.29%, ROC AUC 0.853, same confusion matrix in both modes):
| Mode | Time per step | Patches per second | Training time (1,280 steps) |
|---|---|---|---|
| Eager | 4.08 ms | 125,444 | 5.22 s |
| Fusion (compiled replay) | 0.66 ms | 776,251 | 0.84 s |
Workload: a small MLP (hidden size 256) on 32,768 training and 4,096 test patches of PatchCamelyon (CC0), using 480 handcrafted features per patch (RGB mean/std over an 8×8 grid and its central 4×4), batch size 512, 20 epochs, learning rate 0.02. This is a performance demonstration, not a clinical model. The planner turned the 68 captured nodes of the training step into 11 fused kernels plus 9 native boundaries (matmul and reductions).
The end-to-end training example is a complete multi-layer perceptron (MLP) for MNIST, Fashion-MNIST, or custom CSV datasets. It demonstrates tensor operations, runtime kernel compilation, AdamW/SGD, and PHP-based data orchestration. Its throughput is workload- and GPU-dependent; see the benchmark details above for the measured setup.
Full benchmark reports live in the benchmarks repository.
Install
Requirements
| Need | Details |
|---|---|
| GPU and driver | CUDA-capable NVIDIA GPU; the host driver provides libcuda.so.1 at runtime |
| Build toolchain | CUDA Toolkit (including NVRTC), C/C++ toolchain, make, autoconf |
| PHP | 8.1–8.5 development headers (phpize, php-config), NTS or ZTS |
| OS | Linux |
Building requires the toolkit; running requires the host driver.
From source
The build stays in cuda_build-<PHP major.minor>/modules/cuda.so; replace 8.3
above with the PHP version used by php-config. To install the extension and its
INI configuration instead, run ./compile.sh --install with permission to write
to your PHP extension/INI directories. To choose another PHP ABI:
Build options:
CUDA_HOMEif the toolkit is not at/usr/local/cuda.CUDA_ARCH=sm_86when cross-building.CUDA_USE_CUBLAS=no ./compile.shto use the extension's built-in CUDA matrix kernels instead of cuBLAS (cuBLAS is used when available for compatible larger matrix products).CUDA_USE_CUBLASLT=noto retain cuBLAS without cuBLASLt. Both default to enabled when their toolkit headers and libraries are available.CUDA_USE_CUDNN=auto(default),yes(require it), orno(disable it).CUDA_CUDNN_ROOT=/path/to/cudnnselects a prefix containinginclude/cudnn.handlib/libcudnn.soorlib64/libcudnn.so. Runtime libraries must also be discoverable by the dynamic loader. cuDNN is optional; CUB headers from the CUDA Toolkit are required.
CUDA compiler flags are tracked by the build, and changing backend flags or CUDA architecture rebuilds the affected CUDA objects after reconfiguration. Header dependencies are tracked too; a stale object must not silently retain an old backend configuration.
With PIE 🥧
The extension is published as
lcmialichi/php-gpu-tensors.
Use 0.1.0-beta.4 or newer for the Fusion APIs and the training example
described here; older releases may not include them. PIE builds the native
extension for the selected PHP installation; it does not install an NVIDIA driver
or CUDA Toolkit. Those must already be available on the system, and GPU execution
additionally requires a compatible NVIDIA driver and a visible GPU.
With Docker
Users with the NVIDIA Container Toolkit and a working host driver can build and test in the development image:
Running the tests
./run-tests.sh runs CPU-side C tests and the PHP test suite. Use --cpu-only to
run only the host-side C tests in build environments without a GPU; this does not
validate CUDA execution. --require-gpu fails immediately when no GPU is visible,
instead of treating skipped GPU tests as success.
Accelerated kernels and backend diagnostics
The backend additions and Cuda\NN below require 0.1.0-beta.5 or newer;
they are not included in the previously published 0.1.0-beta.4 package.
The execution path is hybrid: custom kernels/Fusion for elementwise work, CUB for large contiguous reductions, cuBLAS/cuBLASLt for eligible matrix products, and optional cuDNN for CNN inference. Small matrix products retain a lightweight CUDA kernel; larger incompatible layouts use shared-memory tiling. Padded/transposed BLAS-compatible views and batched broadcasts remain supported; large vector dot products use cuBLAS directly without the general GEMM size threshold.
Large global sum, mean, prod, argMin and argMax use parallel CUB
reductions, as do integral min/max. Non-last-axis reductions can use a
coalesced column kernel. Scratch storage is cached per device and stream, so
concurrent FusionGraph::runAsync() calls do not share writable workspace.
Arg reductions retain first-index tie behavior and ignore NaNs as before.
Floating min/max retain the original kernel and reduction order to preserve
their existing NaN semantics. Parallel floating sum/mean/product may change
low-order bits because the addition/multiplication order changes.
Matmul defaults to strict FP32. On compute capability 8.0+ with cuBLAS, callers may explicitly opt in to TF32 Tensor Core math:
The selected mode is per request/thread and resets to strict FP32 at request
shutdown. It affects eligible cuBLAS/cuBLASLt GEMMs only: dot products and
built-in fallback kernels retain their existing FP32 behavior, and tensor
dtypes are unchanged. Select the mode before compiling a Fusion graph that
contains matmul(); the graph captures the mode used during compilation.
cuBLASLt algorithms are cached by layout and precision with zero shared
workspace, allowing concurrent Fusion streams. Unavailable algorithms fall
back to cuBLAS; submission failures raise exceptions instead of silently
running another backend. This does not add bias/activation epilogues to Fusion.
Counters are per request/thread and reset on device reset. They count native
submissions, not completed GPU executions or CUDA Graph replays. The precision
field reports the selected cuBLAS matmul policy, not all tensor operations.
Use examples/10_kernel_benchmark.php for a
reproducible before/after benchmark with numerical checks, warm resident inputs,
configurable warmups/samples, latency distributions, memory-pool reuse workloads,
Fusion plan topology/statistics, raw samples and binary hashes:
This measures PHP/API wall-clock latency including output allocation and final
GPU synchronization, not isolated CUDA-event kernel time. Run without other
GPU workloads for useful comparisons; speedups depend on shape and hardware.
Its matmul cases cover square, skinny, MLP/classifier, transposed, strided,
batched, and broadcast layouts, and report the selected backend per case.
It also compares eager, Fusion, and CUDA Graph replay for elementwise chains and
reductions with fused work around a native boundary. The pool-reuse-* cases
repeat same-shape output allocation/destruction and report end-to-end reuse
latency; they are not allocator-only timings or internal pool counters. Adjust
sample counts and allocator churn with --warmups=N, --samples=N, and
--pool-iterations=N.
The benchmark defaults to strict FP32; pass --precision=tf32 to measure the
opt-in Tensor Core mode on supported GPUs.
CNN inference with optional cuDNN
Cuda\NN is a final, static inference API. All inputs must be initialized,
contiguous float32 tensors in NCHW format; filters use OIHW. Existing
tensor operations retain their dtype support. NN does not silently cast, pack
views, run on CPU or participate in Fusion capture.
NN::conv2d(input, weights, bias: null, stride: [1,1], padding: [0,0], dilation: [1,1], groups: 1)performs cross-correlation, with optional per-output channel bias. Groups must divide input and output channels; filter input channels must equal input channels divided by groups.NN::pool2d(input, window, stride: [2,2], padding: [0,0], mode: 'max')supports deterministic max pooling and'average'excluding padded cells. Padding must be smaller than the window.NN::softmax(input)uses accurate channel softmax.- Spatial output dimensions use floor division. Selectors are lists of exactly
two integers. Empty axes and tensors/output shapes exceeding
INT_MAXelements are rejected.
Convolution uses deterministic FP32 FMA algorithms with a cached shape plan and up to 32 MiB of reusable workspace. Calls complete on the default stream before returning. The initial integration uses cuDNN's fixed-function inference API, not the frontend graph API; it has been validated with cuDNN 8.9. Backpropagation, mixed precision, asynchronous CNN plans and convolution/activation graph fusion are not part of this API yet. Methods throw explicitly when cuDNN is unavailable.
GPU Tensors in PHP
CudaArray holds GPU storage; operations return GPU tensors. toArray()
transfers to the CPU and expands all values into PHP arrays. Operations include
arithmetic, broadcasting, comparisons, matmul() (including batches), shape
views, and sum(), mean(), min(), max(), prod(), argMax(), and
argMin() with an optional axis. PHP arithmetic operators also dispatch to tensor
methods. mean() reduces all values when called without an axis, or reduces one
dimension when given an axis. It returns float32 for float32 input and
float64 for float64, integer, and boolean input.
Eager elementwise operations and reductions enqueue work without a device-wide
wait after each kernel. Operations submitted on the default stream remain
ordered; host reads such as toArray(), toBuffer(), and item() wait for
their result. Operations that must validate GPU-resident values can still
synchronize for that validation.
More array operations
squeeze() removes size-one dimensions and unsqueeze() inserts one.
broadcastTo() expands a tensor as a zero-copy view where possible; expanded
dimensions use a zero stride. Elementwise maximum() and minimum() compare
two tensors or a tensor and a scalar, while clamp($min, $max) bounds values.
These operations follow the usual broadcasting rules.
all($axis) and any($axis) reduce boolean conditions and return boolean
tensors. var($axis, $correction) and std($axis, $correction) compute
variance and standard deviation; the default correction is zero. Calling
item() on a one-element tensor returns a PHP scalar and transfers that value
from the GPU.
gather($indices, $axis) selects values using an int32 index tensor with
matching rank and compatible dimensions outside the selected axis.
scatterAdd($indices, $updates, $axis) returns a new tensor and adds updates
at the indexed positions; duplicate indices accumulate. Gather supports the
available tensor dtypes, while scatter-add currently supports float32 and
float64. Both are eager-only and explicitly reject capture inside Fusion.
Other operations described here can participate in Fusion where their
underlying operation is supported by the captured plan.
Python-style slicing
CudaArray::slice() accepts a comma-separated expression, or one selector per
axis. Ranges have an exclusive stop, omitted bounds default to the axis limits,
and negative indices/bounds count from the end. Range bounds are clipped to the
axis size; individual indices outside the axis throw InvalidArgumentException.
Integers remove axes; null, ':', omitted axes and slice() with no arguments
select the full extent. Steps must be positive integers no larger than INT_MAX.
Zero/negative steps, ellipsis, new axes, index lists and executable expressions
are not supported. Expressions are parsed as integers/ranges, never evaluated as
PHP.
Views, empty ranges and legacy syntax
A fully indexed tensor has shape `[]` and size 1; `toArray()` and `toHost()->toArray()` represent it as a one-element array. Outside Fusion, slices are zero-copy views with shared storage: writes through a view affect its parent, and the parent remains alive while the view is retained. Host transfers pack strided views into contiguous host storage. Row assignments involving strided tensors stage the source on the host before writing, preserving overlapping-source correctness; this is not a GPU-only bulk scatter operation. Fusion captures slice index transformations without materializing during capture; returned compiled/scoped outputs use the existing independent-output storage contract. Empty ranges are supported, preserving the remaining shape: `[0, 4]` converts to `[]`, whereas `[3, 0]` converts to `[[], [], []]`. Elementwise operations, casts, transfers, reshape/transpose and matmul handle zero elements. Sum/product over an empty axis return 0/1; mean returns NaN. Min/max and arg reductions throw when an empty reduced axis would produce values, because no identity/index exists. An empty non-reduced output remains empty. Packed-buffer imports accept zero dimensions; materialized empty results support serialization. The legacy `__invoke()` and `[]` selection syntax remain unchanged, including their inclusive ranges. Do not interpret their bounds as the new exclusive `slice()` bounds.Data pipelines for machine learning
Use this extension to build GPU-accelerated numerical steps into PHP
applications: tensor arithmetic, matrix multiplication, broadcasting, reductions
such as mean(), and custom CUDA kernels. These primitives support
machine-learning data preparation, inference and small training loops while the
data remains in NVIDIA GPU memory.
This is a low-level GPU computing library, not a complete machine-learning framework. It provides opt-in reverse-mode automatic differentiation for floating-point tensors, including gradients built inside a Fusion capture.
Automatic differentiation
Gradient tracking is disabled by default. Mark leaf tensors with
requiresGrad(), build a scalar loss, and call backward(). Gradients
accumulate until zeroGrad() is called; detach() creates a shared-storage
view disconnected from the gradient history.
Non-scalar outputs require an explicit seed with the same shape and dtype:
$output->backward($seed). Scalar outputs use a seed of one by default.
Supported rules include arithmetic and broadcasting, 2D matmul, sum(),
mean(), max()/min() (ties share the gradient), where(), floating-point
casts, reshape(), transpose(), and the common exponential, logarithmic,
trigonometric, square-root and negation operations. Comparisons and where()
conditions do not receive gradients. Unsupported operations fail explicitly
during backward; batched matmul and slice/concat gradients are not implemented.
Backward can be part of a compiled Fusion callback. Mark the example inputs before compilation so the placeholders inherit gradient tracking, and return the gradients (or use them to compute functional parameter updates) as graph outputs:
Fusion still captures the operations into one execution plan; backward does not run kernels during graph construction. Optimizer updates remain functional: compute new parameters and state from the current values and gradients, then return them as outputs rather than mutating captured tensors.
Optimizers
Cuda\Optimizer provides SGD and AdamW updates without hiding tensor state.
That explicit state is useful in eager training and lets Fusion capture the
optimizer math along with the forward and backward operations:
step() returns new parameter tensors and new state; it does not mutate its
inputs. Pass a CudaArray learning rate when it must vary as a Fusion graph
input. weightDecayMask can disable decay for selected parameters, such as
biases. Parameters and gradients must have matching shapes and dtypes; optimizer
state preserves each parameter's float32 or float64 dtype. SGD state contains
a velocity tensor per parameter; AdamW state contains first and second moments
plus scalar beta powers. Pass the returned state to the next step.
For data already in packed row-major bytes, avoid creating individual PHP
scalars. fromFile() reads raw bytes, whereas fromNpy() parses NumPy's .npy
format:
The constructor validates rectangular numeric input while converting it, with
specialized packed-array loops for each dtype. Integer, float and boolean values
are accepted; array keys are ignored in iteration order. Ragged arrays,
nonnumeric values and excessive nesting throw Cuda\InvalidArgumentException;
empty rectangular arrays are supported. fromFlatArray() additionally verifies
that the flat element count matches the explicit shape.
toArray() preserves nested row-major PHP lists, using preallocated packed
arrays. Large contiguous imports use 4 MiB staging windows; PHP-array exports
use at most 32 MiB to avoid excessive CUDA copy calls, rather than a temporary
buffer the size of the tensor. Large or sparse
strided downloads are packed on the GPU before copying; small strided results
use a host gather. Scalars remain one-element arrays, and empty axes are
preserved. These optimizations do not remove the memory cost of one PHP value
per element: prefer toBuffer() when consuming binary data and keep intermediate
tensors on the GPU.
For reproducible constructor/download measurements, run
examples/09_transfer_benchmark.php --output=transfer.json with the extension
loaded. Add --large to include the 33,554,432-element case and allow enough PHP
memory (-d memory_limit=-1). Add --random for random GPU download inputs.
The runner prepares inputs outside timing and
reports materialization, destruction and end-to-end medians separately, plus
retained PHP heap bytes; it does not claim to isolate CUDA copy time.
HostArray is an alias of Cuda\ContiguousArray, a contiguous CPU tensor. Pass
pinned: true to its constructor or fromBuffer() for page-locked host storage
when repeated transfers justify the extra host memory. where() broadcasts its
three inputs; its mask treats nonzero values as true, and its two value tensors
must have the same dtype. .npy imports support C-order little-endian numeric and
boolean arrays; Fortran order and big-endian data are rejected.
Custom kernels
Register the CUDA kernel's argument metadata, compile to PTX, then launch with explicit grid and block dimensions:
launch() synchronizes; launchAsync() returns an operation ID for sync() or
wait(). Keep tensors alive until asynchronous work finishes. See
the JIT examples and
asynchronous execution.
Optional kernel fusion
Eager execution remains the default. Fusion is explicitly opt-in: existing code keeps using eager execution unless capture is enabled.
Fusion::run() captures tensor operations inside a callback and returns
materialized CudaArray outputs:
For repeated execution, compile a specialized plan once:
What to know at a glance:
- Fused: addition, subtraction, multiplication, division, unary operations,
comparisons,
where(), safeastype()conversions and reshape/transpose/slice index transformations. - Execution boundaries: reductions,
matmul()(currently float32) and powers run on existing kernels; the fused kernels run in dependency order around them. - Replay: the PHP callback is not invoked again after compilation, and input values and pointers can change between executions.
- Inspection:
getStats(),getPlan()andgetSource()show what the planner produced; for$a + $b * $cthe plan has one fused kernel and no intermediate data buffers.
Fusion reference: compile and replay, fusion rules, capture limits and caching
**Compile and replay.** Compilation invokes the callback once with metadata-only placeholders. Replay does not invoke PHP callback code again. Inputs must match the example shapes, dtypes and strides, and execution must use the compilation device and CUDA context. Input values and pointers can change between executions; keep the device/context alive until the graph is released. Tensors captured by a closure are retained by the graph; use callback parameters for replaceable inputs. **What fuses.** Addition, subtraction, multiplication, division, unary operations and comparison methods fuse into elementwise kernels, with broadcasting, strided/view inputs, scalar operands and dtype promotion. Each node converts to its own result dtype, preserving intermediate rounding/narrowing. Generated kernels use the eager backend's fast-math settings but disable cross-node FMA contraction. `where()`, safe explicit `astype()` conversions and reshape/transpose/slice index transformations also fuse. Safe casts work in eager execution too, using the same generated conversion kernel; unsafe narrowing retains the existing rejection policy. **Boundaries and planning.** Reductions, `matmul()` (currently float32) and powers are execution boundaries using existing kernels. All generated elementwise kernels in a plan are compiled together through the existing `Compiler` NVRTC infrastructure, then executed in dependency order around these boundaries. Expressions split at a weighted cost budget of 32 (math functions, index transforms and float64 have higher cost). Shared expensive expressions can be materialized instead of recomputed. Pure repeated binary/unary/cast nodes are deduplicated, and unreachable nodes are not included in compiled plans. Capture is limited to 512 nodes. **Outputs and diagnostics.** Callbacks can return tensors or nested arrays of tensors, preserving array keys. Up to four adjacent independent outputs with the same shape can share a kernel. Other outputs use separate kernels; fusion does not promise one kernel for an entire callback. `getPlan()` reports step kinds, output counts and the reason for each materialization, including zero-copy `view` steps for layouts consumed by compatible native operations. `getStats()` exposes planned fused kernels, boundary steps, intermediate buffer count, scratch reuse and successful replay count. **Capture limits.** During `run()`, CPU reads, legacy slicing (`__invoke()` and `[]`), serialization and operations not captured by the planner materialize their required inputs and continue eager execution; later elementwise operations can form a new segment. During `compile()`, reads of placeholder data are rejected rather than specializing on example values. Shape, stride and dtype queries do not execute kernels. Nested capture, Fiber switching, tensor mutation (including compound assignments), custom kernel launches and device changes/reset are prohibited during capture. Exceptions restore eager execution; escaped tensors from an aborted capture cannot be read. `run()` remains synchronous. **PTX cache.** Generated PTX is cached per PHP request/thread, with LRU eviction at 16 entries or 16 MiB. Keys include generated source (operations, constants, dtypes and layouts), compute capability and CUDA driver/runtime versions, not input pointers. `Fusion::getCacheStats()` reports hits, misses, compilations and evictions. `Fusion::clearCache()` drops PTX without invalidating existing graphs. `compile()` also avoids repeated planning and module loading during replay.Streams, asynchronous execution and CUDA Graph
Plans containing generated kernels, matmul and reductions run on a private nonblocking stream. Reduction descriptors are kernel parameters rather than shared global state, so concurrent replays cannot overwrite each other's shapes.
cudaGraph: true opts into a CUDA Graph executable for compatible plans. Generated
kernel-only plans update kernel parameters directly for new input/output pointers.
Plans with matmul or reduction boundaries capture the complete stream sequence,
then recapture it on replay and update the executable for current pointers. Those
plans warm up the native library/workspace path once before capture. Recapture adds
host work, so native-boundary graphs should be benchmarked against stream replay
for the target workload.
getStats()['backend'] is cuda-graph, stream or native.
| Plan contains | Stream and runAsync() |
CUDA Graph |
|---|---|---|
| Generated kernels only | ✅ | ✅ |
| Matmul / reductions | ✅ | ✅ |
| Power boundaries | ❌ (synchronous native executor) | ❌ |
Check cudaGraphCompatible and cudaGraphIncompatibility to distinguish graph
support from asyncCompatible. Power boundaries report the reason in
incompatibility and reject runAsync(). CUDA failures are reported rather than
hidden behind fallback.
Concurrency rules, scratch memory and profiling
Synchronous compiled replay keeps a reusable stream, readiness event and scratch workspace; async replays use separate scratch storage. Slots are planned by last consumer, including native consumers and aliased views. Returned outputs always have independent storage across replays. Scratch remains allocated until the graph is released and counts toward the configured GPU memory budget. Stream plans allow concurrent `runAsync()` calls. A CUDA Graph executable allows only one outstanding replay: call `wait()` (or release the execution) before replaying that graph again. A completion query alone does not release the execution's retained resources. Inputs and closure-captured tensors remain alive until completion is collected. While an execution is outstanding, tensor mutation, custom kernel launches and device changes/reset are blocked. Inputs must be ready before submission; independent custom async producer streams still require their existing synchronization contract. Repeated `wait()` calls return the same outputs. Destroying a pending `FusionExecution` synchronizes before releasing storage. Execution submission preallocates tensors and can incur allocation synchronization; asynchronous kernel submission does not imply a zero-blocking PHP call. Device shape/stride metadata is allocated lazily, only for kernels that need it. Generated kernels, reductions, unaries and matmul use compiled or host descriptors instead of uploading metadata for every result tensor. Optional synchronous replay phase timing: `profiledExecutions`, `bindTimeNs`, `executeTimeNs` and `collectTimeNs` accumulate successful synchronous `run()` calls only. Execution timing includes preparation, submission and waiting; it is not isolated GPU kernel time. Profiling is off by default. `tensorAllocations` counts replay-created data tensors, not pool cache misses or CUDA allocation calls. `workspaceBuffers` counts retained scratch slots; `bufferReuses` and `synchronizations` are cumulative replay counters. Run the focused comparison of eager, cached scoped, stream replay and CUDA Graph replay with: The example validates output bytes, reports cold/cache compilation costs and end-to-end timings, and does not assume CUDA Graph is faster for every workload.Real training with Fusion
examples/08_gpu_classifier.php trains an
MLP classifier on MNIST, Fashion-MNIST, or a custom CSV dataset without
requiring users to write CUDA source:
- Dataset parsing and normalization are handled in PHP.
- Packed
float32batches are uploaded and kept on the GPU. - Forward, backward, and optimizer elementwise operations are captured and replayed through Fusion; matrix multiplication and reductions use native kernels at plan boundaries.
- The example reports loss and evaluation metrics, including macro-F1 and a confusion matrix.
Defaults are MNIST, hidden layers of 512 and 256 units, 40 epochs, batch size
128, and AdamW. Initialization is deterministic. By default, the example saves
parameters after successful evaluation; pass --no-save to disable saving or
--load to evaluate a compatible saved model without training. Dataset and
model files are stored next to the script and are ignored by Git.
Use --profile to print Fusion plan timing and resource counters, or
--predict-index=N to run inference on one sample. Run the script with
--help to see the available dataset, optimizer, and training options.
API and limits
| API | Purpose |
|---|---|
Cuda\CudaArray |
GPU allocation, tensor math, reductions, views, imports, where() and opt-in autograd |
Cuda\HostArray / Cuda\ContiguousArray |
CPU storage, packed buffers, optional pinned memory and toGpu() |
Cuda\Optimizer |
Functional SGD/AdamW updates with explicit state, compatible with Fusion capture |
Cuda\Fusion / Cuda\FusionGraph |
Optional expression capture, compiled replay, PTX cache and plan diagnostics |
Cuda\FusionExecution |
Pending compatible execution, completion query and synchronized result collection |
Cuda\Compiler / Cuda\CompiledModule |
NVRTC compilation, cached PTX, synchronous and asynchronous kernels |
cuda_get_device_count() and other cuda_* functions |
Device selection, properties, memory and synchronization |
Cuda\Exception |
Base class for runtime, argument, allocation and compilation errors |
The annotated signatures are in class stubs and device function stubs; runnable examples live in examples.
Current limits:
- NVIDIA GPUs only, on Linux. GPU data has no CPU fallback.
- Reverse-mode automatic differentiation is opt-in and covers only the operations listed in Automatic differentiation. Batched matmul and slice/concat gradients are not implemented.
astype()supports safe dtype conversions; autograd tracks floating-point casts only.- The core API baseline is frozen at
0.1.0, and the current release is beta. - Compatible plans with matmul or reductions support asynchronous replay and CUDA Graphs. Power boundaries still require synchronous execution.
Validation status
The core API has a frozen 0.1.0 baseline; eager execution remains the default
and Fusion is explicitly opt-in. Build and CPU checks do not substitute for GPU
runtime validation on each PHP version and thread mode.
| PHP / mode | GPU | What was validated |
|---|---|---|
| 8.1, Docker development image | GeForce MX570 A | Current source: 45 tests passed; one optional cuDNN test skipped; MLP training smoke test completed |
| 8.3 NTS with cuDNN 8.9.7 | GeForce MX570 A | Beta.5: all 44 GPU PHPT tests passed, including CNN inference; CPU/API checks passed |
| 8.3 and 8.5 NTS without cuDNN | GeForce MX570 A | Beta.5: 43 GPU PHPT tests passed on each runtime; one optional cuDNN test skipped. CPU/API checks passed |
| 8.3 NTS without cuBLAS/cuDNN | GeForce MX570 A | Beta.5: 43 GPU PHPT tests passed; one optional cuDNN test skipped. CPU/API checks passed |
| 8.5 NTS and ZTS, 8.1 ZTS | RTX A2000 | Earlier tensor/JIT validation only; does not cover the latest Fusion changes |
| All ten PHP 8.1–8.5 × NTS/ZTS combinations | none | CI build/CPU/API matrix; GPU execution is not covered by CI |
Tested on another GPU, PHP version or CUDA version? Reports are very welcome:
please open an issue with
your PHP version, thread mode, GPU, driver and CUDA versions, and whether
./run-tests.sh --require-gpu passed.
Contribute
Contributions can be code, tests, documentation, runnable examples, or reports from another PHP/CUDA/GPU combination. Browse open issues, report a bug, or propose a feature. Start with the contribution guide; documentation, examples, and host-side tests are possible without an NVIDIA GPU.
For possible directions, see the roadmap. A feature does not need to be listed there to be worth discussing.
Benchmarks
The benchmark suite is maintained in the separate
PHP GPU Tensors Benchmarks repository.
It includes focused --matmul, --import and --fusion runs, JSON/HTML reports,
and a beta.4 full-suite report
covering 398 cases on PHP 8.3 NTS and an MX570 A, including Fusion replay and
slicing. Earlier results include a
published PHP 8.5 NTS vs ZTS comparison,
as well as a full-suite benchmark report
covering 368 cases across five workload groups with downloadable raw reports.
License
Licensed under the MIT License.