Weight Quantization
Documentation status: architecture — see Maturity and evidence.
The current tensor runtime contains quantized-weight execution paths intended to reduce model memory and bandwidth while keeping the surrounding tensor and transformer semantics stable.
What is verified in the engine
The source tree contains implementations and tests for several quantization families, including Q4, Q8, GGUF Q4_K and additional low-bit packing paths. Transformer tests exercise both floating-point and quantized weight flows, including resident accelerator variants.
This is evidence that quantized execution is part of the current engine architecture. It is not a promise that every format is supported on every backend or exposed through every public binding.
Runtime boundary
A public application should normally express:
- the model or tensor operation it wants to execute;
- an acceptable precision or model artifact when that is part of the deployment contract;
- performance and memory constraints.
The Runtime chooses or validates the compatible storage/backend path.
Accuracy and compatibility
Quantization changes numerical representation. Validate at least:
- output quality on representative data;
- memory reduction;
- latency and throughput on the target backend;
- model-format compatibility;
- deterministic fallback behavior when a requested accelerated format is unavailable.
Do not treat a quantized kernel name as a domain-level concept.