Skip to content
EN FR

Tensor Compute Backends

Documentation status: architecture — see Maturity and evidence.

logiCells separates tensor semantics from the compute backend that executes numerical kernels. This lets the Runtime choose an implementation according to capability and availability without making the conceptual model depend on one accelerator API.

Capability-driven dispatch

The engine distinguishes three useful capability levels:

  • host F32 execution for ordinary floating-point operations;
  • resident-device execution for buffers that remain on an accelerator across several operations;
  • resident-transformer execution for transformer-oriented kernels and state that benefit from avoiding repeated host/device transfers.

A backend can report that an operation is unsupported. Dispatch can then continue to another backend or to a CPU implementation when the operation has a fallback path.

Why residency matters

Moving a tensor to an accelerator for every primitive can cost more than the operation itself. Resident buffers make the transfer boundary explicit:

host data
  -> upload once
  -> resident operations
  -> optional intermediate reuse
  -> download only when the host needs the result

This is especially important for transformer inference, where normalization, projections, attention, feed-forward layers and KV-cache updates form a long chain.

Current architecture

The current engine source contains registered accelerator paths for WebGPU and MLX, in addition to host/CPU execution paths. Availability remains platform- and build-dependent; applications should therefore reason in terms of capabilities rather than assume that a named accelerator is always present.

Developer rule

Treat the backend as an execution concern. Domain models, conceptual identity and published business contracts should not encode a GPU or accelerator choice.

See also: