Continuum C++ API
Unified runtime for token + tensor execution
Loading...
Searching...
No Matches
continuum::backend::VllmShimBackend Class Reference

#include <vllm_shim.hpp>

Inheritance diagram for continuum::backend::VllmShimBackend:
continuum::backend::Backend

Classes

struct  PrefixState
 Prefix text and model a state handle refers to. More...
 

Public Member Functions

std::vector< std::uint8_t > export_state (const BackendState &state) const override
 
std::optional< BackendState > import_state (const std::vector< std::uint8_t > &bytes) override
 
std::size_t state_count () const
 Live state handles (bounded; oldest dropped first).
 
std::size_t rewarm_count () const
 Re-warm requests sent by import_state.
 
BackendCapabilities capabilities () const override
 
BackendRunResult run_with_cache (const ir::Node &node, const std::vector< continuum::Value > &inputs, const std::optional< BackendState > &prefix_state, std::int32_t remaining_tokens) override
 
- Public Member Functions inherited from continuum::backend::Backend
virtual ~Backend ()=default
 
virtual std::string tensor_backend_type () const
 

Static Public Member Functions

static std::string ExtractJsonString (const std::string &body, const std::string &key)
 
static std::int32_t ExtractJsonInt (const std::string &body, const std::string &key)
 First non-negative integer stored under key in body, or 0.
 

Detailed Description

Token backend for any server speaking the OpenAI /v1/completions wire format: vLLM, Ollama, llama.cpp server, and similar.

With VLLM_BASE_URL set, each call POSTs the full prompt to $VLLM_BASE_URL/v1/completions and returns the completion text as a string value. The model name is the node's model_id with a leading vllm/ stripped (vllm/gemma4 -> gemma4), else VLLM_MODEL. Without VLLM_BASE_URL it runs offline and returns deterministic token ids.

Prefix reuse: the server keeps the KV blocks (vLLM automatic prefix caching, Ollama / llama.cpp prompt cache) and skips recomputing a shared prefix; its usage.prompt_tokens_details.cached_tokens is reported as tokens_saved. Continuum's state handle records which prefix the server holds warm. It is portable (export_state / import_state), so it survives a checkpoint / resume, and with VLLM_REWARM_ON_IMPORT=1 an imported state re-warms the server with a 1-token request over the prefix.

Member Function Documentation

◆ capabilities()

BackendCapabilities continuum::backend::VllmShimBackend::capabilities ( ) const
overridevirtual

◆ export_state()

std::vector< std::uint8_t > continuum::backend::VllmShimBackend::export_state ( const BackendState &  state) const
overridevirtual

Reimplemented from continuum::backend::Backend.

◆ ExtractJsonInt()

static std::int32_t continuum::backend::VllmShimBackend::ExtractJsonInt ( const std::string &  body,
const std::string &  key 
)
static

First non-negative integer stored under key in body, or 0.

◆ ExtractJsonString()

static std::string continuum::backend::VllmShimBackend::ExtractJsonString ( const std::string &  body,
const std::string &  key 
)
static

Decode the first JSON string value stored under key in body, resolving escapes (including \uXXXX and surrogate pairs) to UTF-8. Returns an empty string when the key is absent.

◆ import_state()

std::optional< BackendState > continuum::backend::VllmShimBackend::import_state ( const std::vector< std::uint8_t > &  bytes)
overridevirtual

Reimplemented from continuum::backend::Backend.

◆ rewarm_count()

std::size_t continuum::backend::VllmShimBackend::rewarm_count ( ) const
inline

Re-warm requests sent by import_state.

◆ run_with_cache()

BackendRunResult continuum::backend::VllmShimBackend::run_with_cache ( const ir::Node &  node,
const std::vector< continuum::Value > &  inputs,
const std::optional< BackendState > &  prefix_state,
std::int32_t  remaining_tokens 
)
overridevirtual

◆ state_count()

std::size_t continuum::backend::VllmShimBackend::state_count ( ) const

Live state handles (bounded; oldest dropped first).


The documentation for this class was generated from the following file: