Skip to content
EN FR

LLM Service Contract

Documentation status: architecture — see Maturity and evidence.

The LLM service contract is the stable application-facing boundary. A provider may implement it with llama.cpp, the native logiCells runtime, a remote endpoint or another inference engine.

Core lifecycle

A service exposes the following semantic operations:

configure model
    -> start service
    -> submit request(s)
    -> stream generated text
    -> complete / cancel
    -> stop service
    -> optionally unload model

The current contract includes model configuration, Start, Stop, explicit model unload, cancellation of current work, request submission, and basic state/diagnostic queries.

Request contract

A request carries at least:

  • a caller-provided or generated request identifier;
  • the prompt;
  • a maximum generated-token budget;
  • a token callback for streaming output;
  • a completion callback;
  • a callback-dispatch policy.

An empty prompt and a non-positive token budget are rejected before the request enters the backend queue.

Provider independence

Provider-specific objects must not cross this boundary. In particular, public bindings should not expose llama.cpp model/context/sampler handles or implementation scheduler objects.

The provider registry should identify the llama.cpp implementation as:

logicells.llm.llama.cpp

This lets higher-level code select the backend without coupling the application model to its implementation.

Contract evolution

The current contract is intentionally small. The code-evolution proposal supplied with this documentation recommends adding per-request cancellation/state, structured generation options and a structured completion result while keeping this provider-neutral boundary.