Tensor Compute Backends
Documentation status: architecture — see Maturity and evidence.
logiCells separates tensor semantics from the compute backend that executes numerical kernels. This lets the Runtime choose an implementation according to capability and availability without making the conceptual model depend on one accelerator API.
Capability-driven dispatch
The engine distinguishes three useful capability levels:
- host F32 execution for ordinary floating-point operations;
- resident-device execution for buffers that remain on an accelerator across several operations;
- resident-transformer execution for transformer-oriented kernels and state that benefit from avoiding repeated host/device transfers.
A backend can report that an operation is unsupported. Dispatch can then continue to another backend or to a CPU implementation when the operation has a fallback path.
Why residency matters
Moving a tensor to an accelerator for every primitive can cost more than the operation itself. Resident buffers make the transfer boundary explicit:
host data
-> upload once
-> resident operations
-> optional intermediate reuse
-> download only when the host needs the result
This is especially important for transformer inference, where normalization, projections, attention, feed-forward layers and KV-cache updates form a long chain.
Current architecture
The current engine source contains registered accelerator paths for WebGPU and MLX, in addition to host/CPU execution paths. Availability remains platform- and build-dependent; applications should therefore reason in terms of capabilities rather than assume that a named accelerator is always present.
Developer rule
Treat the backend as an execution concern. Domain models, conceptual identity and published business contracts should not encode a GPU or accelerator choice.
See also: