Skip to content
EN FR

Run a Local Model with the llama.cpp Provider

Documentation status: roadmap — see Maturity and evidence.

This page documents the target public naming for the source rename. The underlying implementation already loads and generates through llama.cpp; the logiCells.LLM.LlamaCpp.* names are the canonical names to use once the rename is applied.

Prerequisites

You need:

  • a llama.cpp native library compatible with the bindings shipped for the target platform;
  • a local model file accepted by that llama.cpp build;
  • a model context size greater than zero;
  • the llama.cpp provider registered with the LLM service registry.

Provider identity

logicells.llm.llama.cpp

Target provider-specific factory

The provider-specific API should offer a convenience factory equivalent to:

service = CreateLlamaCppService(
    modelPath,
    contextSize = 8192,
    useMMap = true)

Higher-level code may instead resolve the same capability through the generic provider registry.

First request

The application flow is:

service.Start()
request = {
    prompt: "Explain the current conceptual context.",
    maxTokens: 128,
    onToken: stream token text,
    onDone: observe completion
}
requestId = service.Submit(request)

The provider tokenizes the prompt, decodes it into the llama.cpp context, samples generated tokens, converts token pieces to valid UTF-8 text, and streams text fragments to the caller.

Stop versus unload

Stop stops request scheduling and cancels current/pending work, but deliberately keeps the successfully loaded model/session resident. A later Start can therefore reuse the resident model.

UnloadModel is the explicit resource-release boundary and is only valid while the service is stopped. Reconfiguring a stopped service to a different model/context/mmap configuration also releases the previous resident session.

See llama.cpp provider lifecycle for the detailed state model.

Current limitations of the public contract

The current request contract exposes a prompt and maximum-token budget but not yet the complete llama.cpp sampling surface. Temperature, top-k/top-p, seed, stop sequences, chat templates and per-request sampler policies should be added through provider-neutral generation options rather than llama.cpp-specific parameters.