Skip to content
EN FR

WebGPU Tensor Execution

Documentation status: architecture — see Maturity and evidence.

WebGPU is an accelerator backend used by the current tensor runtime for selected numerical workloads. This page documents the public architecture rather than private implementation classes.

Verified capability families

The engine source and tests cover WebGPU paths for:

  • dense tensor operations;
  • resident tensor buffers;
  • matrix multiplication and transposed-matrix multiplication;
  • softmax, RMS normalization and element-wise activation flows;
  • rotary positional encoding (RoPE);
  • causal attention;
  • transformer feed-forward operations;
  • resident transformer blocks;
  • CSR/sparse operations;
  • quantized kernels and quantized transformer weights.

Execution policy

WebGPU use is not simply an on/off flag. The runtime contains policy logic that can consider workload thresholds and device-memory budget before selecting a GPU path. Some buffers can also be marked as resident or pinned to avoid inappropriate eviction.

operation
  -> capability check
  -> workload / memory policy
  -> WebGPU execution when admissible
  -> fallback path when available

Public contract

Do not make application semantics depend on WebGPU. A model should remain valid when the Runtime selects another compatible backend. WebGPU is therefore best treated as an execution capability that can be observed and tuned, not as a business-model primitive.

Maturity

The source tree includes dedicated WebGPU tests for tensor buffers, resident operations, transformer operations, transformer blocks, CSR operations, runtime policy and multiple quantized kernels. That supports documenting WebGPU as an implemented engine capability. Exact platform availability and public binding surface still depend on the distributed Runtime build.