Skip to content
EN FR

Transformer Runtime

Documentation status: architecture — see Maturity and evidence.

The current logiCells tensor engine contains a transformer-oriented execution path rather than only generic tensor primitives. The public contract remains capability-oriented: applications should not depend on private runtime class names.

Verified building blocks

The implementation and test suites cover the main inference building blocks:

  • RMS normalization;
  • Q/K/V projections;
  • rotary positional encoding (RoPE);
  • causal attention;
  • grouped-query attention (GQA) paths;
  • residual connections;
  • SiLU/SwiGLU-style feed-forward computation;
  • output normalization and logits;
  • resident buffers;
  • KV-cache append and single-token decode flows.

The source also includes model-runtime tests that compare token-by-token decoding with full-sequence execution and tests that produce logits from floating-point and quantized weights.

Prefill and decode

Transformer inference naturally separates into two phases:

prompt tokens
  -> prefill
  -> populate model state / KV cache
  -> decode one token at a time
  -> append new K/V state

Keeping transformer state resident can substantially reduce transfer overhead during repeated decode steps.

Backend independence

The tensor backend registry distinguishes ordinary host execution, resident-device execution and resident-transformer capability. WebGPU and MLX execution paths exist in the current engine, but a public model should depend on the requested capability rather than a backend name.

Quantized weights

The current engine includes tested transformer paths for several quantized representations. See Weight quantization. Quantization is a deployment/execution choice and does not alter conceptual identity.

Conceptual integration

Conceptual structures can provide entities, relations, masks or features that are projected into numerical tensors. Transformer results remain learned numerical projections. When those results affect conceptual reasoning, reattach them to stable conceptual identity and keep symbolic validity separate from learned ranking or estimation.

Status

This page is an architecture reference for an implemented engine subsystem. It does not certify that every transformer kernel, weight format or accelerator is exposed through every public SDK version.