API reference¶
Two surfaces: the pure-Python frontend (continuum) for writing
programs, and the runtime classes re-exported from continuum._native.
Frontend¶
Public Python API surface for Continuum.
The execution engine lives in the compiled extension, reached through
continuum._native. This module exposes the ergonomic Python frontend:
the program() tracer, tunable Param values, and the
Optimizer.
- class DurableAgent(*args: Any, **kwargs: Any)[source]¶
Bases:
DurableAgentStep-sequenced agent run that can be checkpointed, resumed, and forked.
With
CONTINUUM_OTEL=1in the environment, each agent exports its reuse events through OpenTelemetry (seecontinuum.telemetry).- run_until_step(step_index: int, store: CheckpointStore | None = None, key: str | None = None) bytes[source]¶
Run through
step_indexand return the checkpoint bytes.With
storeandkey, the checkpoint is also written there.
- resume_from(checkpoint: bytes | None = None, *, store: CheckpointStore | None = None, key: str | None = None) list[Any][source]¶
Resume from checkpoint bytes, or read them from
store[key].
- class Optimizer(program, metric, *, lr_tensor=0.1, lr_text=1.0, seed=0)[source]¶
Bases:
objectUnified optimizer facade for all Continuum parameter kinds.
- class Param(kind: str, value: Any, metadata: dict[str, ~typing.Any]=<factory>)[source]¶
Bases:
objectTyped program parameter tracked by the Continuum optimizer.
- kind: str¶
- value: Any¶
- metadata: dict[str, Any]¶
- program(fn)[source]¶
Decorate a Python function as a Continuum program; traced to CIR on first call.
- class Param(kind: str, value: Any, metadata: dict[str, ~typing.Any]=<factory>)[source]¶
Bases:
objectTyped program parameter tracked by the Continuum optimizer.
- kind: str¶
- value: Any¶
- metadata: dict[str, Any]¶
- class TensorOpt(lr: float, rng: Random)[source]¶
Bases:
objectLocal optimizer for numeric tensor-like scalar parameters.
- class TextGradOpt(lr: float, rng: Random)[source]¶
Bases:
objectText mutation optimizer used for prompt-style parameters.
- class GEPAOpt(rng: Random)[source]¶
Bases:
objectDiscrete choice optimizer for categorical parameters.
- class BayesOpt(rng: Random)[source]¶
Bases:
objectSimple random-search optimizer over bounded continuous values.
- class Optimizer(program, metric, *, lr_tensor=0.1, lr_text=1.0, seed=0)[source]¶
Bases:
objectUnified optimizer facade for all Continuum parameter kinds.
- class Linear(in_features, out_features)[source]¶
Bases:
ModuleTiny reference linear layer used by examples and tests.
Checkpoint storage: pluggable object stores and an incremental checkpoint log.
A checkpoint (DurableAgent.run_until_step()) is plain bytes. This
module moves those bytes to durable storage so a run can resume on another
machine without the caller shipping them around.
Stores (CheckpointStore) are minimal key/value object stores:
LocalDirectoryStorewrites files under a directory (atomic renames).S3Storeusesboto3(pip install boto3).GCSStoreusesgoogle-cloud-storage.
Each store supports an atomic create-if-absent write, which is what makes concurrent writers safe.
The log (CheckpointLog) stores a stream of checkpoints for one run
on top of any store:
Checkpoints are content-addressed: the id is the SHA-256 of the full checkpoint bytes, and every load is verified against it.
A checkpoint committed with a
parentis stored as a delta (only the values and KV entries that changed, seecontinuum._native.checkpoint_delta()), with a full checkpoint everymax_chainlinks so reconstruction stays bounded.Each checkpoint has a small JSON manifest record linking it to its parent and delta base, so any step can be rebuilt, and
CheckpointLog.fork()records fork lineage (source checkpoint + edited node).Objects and records are write-once (create-if-absent) and never modified, so any number of workers can resume from and commit to one log without locks and without overwriting each other.
- exception CheckpointCorruptionError[source]¶
Bases:
RuntimeErrorA stored checkpoint did not rebuild to the bytes its id promises.
- class CheckpointLog(store: CheckpointStore, run_id: str, max_chain: int = 16)[source]¶
Bases:
objectIncremental, content-addressed checkpoint stream for one run.
- Parameters:
store – Where objects and manifest records live.
run_id – Namespace for this run inside the store.
max_chain – Longest run of deltas before a full checkpoint is stored.
- commit(checkpoint: bytes, parent: str | None = None, *, step: int | None = None, label: str | None = None, fork: dict[str, Any] | None = None) str[source]¶
Store
checkpointand return its id.With a
parent, the payload is a delta against the parent unless the chain is alreadymax_chainlong or the delta would not be smaller. Committing the same checkpoint twice (from any worker) is a no-op that returns the same id.
- fork(checkpoint_id: str, node_id: int, new_value: Any, *, label: str | None = None) str[source]¶
Commit a fork of
checkpoint_idwithnode_idset tonew_value.The new record’s
parentis the source checkpoint and itsforkfield names the edited node, so lineage survives in the manifest.
- record(checkpoint_id: str) CheckpointRecord[source]¶
- manifest() list[CheckpointRecord][source]¶
Every checkpoint record in the run, oldest first.
- lineage(checkpoint_id: str) list[CheckpointRecord][source]¶
Records from the root of
checkpoint_id’s history down to it.
- children(checkpoint_id: str) list[CheckpointRecord][source]¶
- heads() list[CheckpointRecord][source]¶
Checkpoints nothing has been committed on top of (branch tips).
- class CheckpointRecord(id: str, parent: str | None, kind: str, base: str | None, depth: int, object: str, size: int, full_size: int, executed_nodes: int, step: int | None, label: str | None, fork: dict[str, Any] | None, created: float)[source]¶
Bases:
objectManifest record for one checkpoint in a
CheckpointLog.- id: str¶
- parent: str | None¶
Checkpoint this one was derived from (previous step, or fork source).
- kind: str¶
"full"or"delta".
- base: str | None¶
the checkpoint it applies to (equals
parent).- Type:
For a delta
- depth: int¶
Deltas between this checkpoint and the nearest full one.
- object: str¶
Store key of the payload.
- size: int¶
Stored payload bytes.
- full_size: int¶
Bytes of the reconstructed full checkpoint.
- executed_nodes: int¶
- step: int | None¶
- label: str | None¶
- fork: dict[str, Any] | None¶
{"node_id": ...}when this checkpoint forksparentby editing a node.
- created: float¶
- classmethod from_json(raw: str | bytes) CheckpointRecord[source]¶
- class CheckpointStore[source]¶
Bases:
ABCMinimal object store for checkpoint bytes.
Keys are
/-separated relative names (no.., no leading/).- abstractmethod put(key: str, data: bytes) None[source]¶
Write
dataunderkey, replacing any existing object atomically.
- class GCSStore(bucket: str, prefix: str = '', client: Any = None)[source]¶
Bases:
CheckpointStoreGoogle Cloud Storage via
google-cloud-storage.- Parameters:
bucket – Bucket name.
prefix – Object-name prefix inside the bucket.
client – A
google.cloud.storage.Client; created when omitted.
put_if_absentuses a generation precondition (if_generation_match=0).- put(key: str, data: bytes) None[source]¶
Write
dataunderkey, replacing any existing object atomically.
- class LocalDirectoryStore(root: str | PathLike[str])[source]¶
Bases:
CheckpointStoreFiles under
root. Writes go to a temp file and are renamed into place, so readers never see a partial checkpoint.- put(key: str, data: bytes) None[source]¶
Write
dataunderkey, replacing any existing object atomically.
- class S3Store(bucket: str, prefix: str = '', client: Any = None)[source]¶
Bases:
CheckpointStoreAmazon S3 (or any S3-compatible service) via
boto3.- Parameters:
bucket – Bucket name.
prefix – Key prefix inside the bucket, e.g.
"continuum/ckpt".client – A
boto3S3 client; created withboto3.client("s3")when omitted.
put_if_absentuses S3 conditional writes (If-None-Match: *).- put(key: str, data: bytes) None[source]¶
Write
dataunderkey, replacing any existing object atomically.
OpenAI-compatible proxy that runs requests through Continuum’s reuse stack.
Point any OpenAI-format client at it and it gains Continuum’s caching with no code changes:
python -m continuum.proxy --upstream https://api.openai.com/v1 --port 8787
export OPENAI_BASE_URL=http://localhost:8787/v1
The OpenAI SDK, LangChain (ChatOpenAI), LlamaIndex (OpenAI), and plain
HTTP all work, because the proxy speaks the same wire format as the upstream.
Any OpenAI-compatible upstream works: OpenAI, Azure OpenAI’s /openai/v1,
vLLM, Ollama (http://localhost:11434/v1), LM Studio, …
What is reused. POST /v1/chat/completions and POST /v1/completions
requests become one Continuum TokenOp each: the messages (or prompt) and the
sampling parameters are the key, and the upstream call is the backend behind
the memo tier (exact repeats), the prefix-KV tier (shared prefixes, tracked as
metrics), and, if enabled, the semantic tier. A cache hit returns the stored
response without calling the upstream.
What is never cached. Requests that carry tools / functions or
tool-result messages, requests with n > 1, and requests sent with
Cache-Control: no-cache / no-store are forwarded untouched. Every other
path (/v1/models, embeddings, …) is passed straight through.
Streaming. "stream": true works both ways: a miss streams the
upstream’s server-sent events through as they arrive (and caches the assembled
response), and a hit is replayed as a server-sent-event stream.
Metrics. Each response carries x-continuum-cache (hit / miss /
bypass), x-continuum-served-by, and x-continuum-tokens-saved
headers. GET /metrics returns Prometheus text and GET /continuum/metrics
returns JSON: per-tier lookups and hits, requests, and tokens saved.
- class ContinuumProxy(config: ProxyConfig | None = None, host: str = '127.0.0.1', port: int = 8787)[source]¶
Bases:
objectThe proxy server.
serve_forever()blocks;start()runs it on a thread.- property address: str¶
- start() ContinuumProxy[source]¶
- class ProxyConfig(upstream: str = 'https://api.openai.com/v1', api_key: str | None = None, memo_entries: int = 4096, prefix_entries: int = 8192, semantic_threshold: float | None = None, semantic_entries: int = 2048, embedder: EmbeddingProvider | None = None, verifier: HitVerifier | None = None, timeout: float = 600.0)[source]¶
Bases:
objectProxy settings.
- upstream¶
Base URL of the OpenAI-compatible upstream, including
/v1.- Type:
str
- api_key¶
Used when the client sends no
Authorizationheader.- Type:
str | None
- memo_entries¶
Capacity of the exact-match tier.
- Type:
int
- semantic_threshold¶
Enables the semantic tier at this similarity when set. Off by default: see
benchmarks/reports/semantic-false-hits.mdbefore enabling it, and pass a realembedder.- Type:
float | None
- embedder¶
Embedding provider for the semantic tier (
WordLlamaEmbeddingProviderwith threshold 0.7 is the measured starting point).- Type:
EmbeddingProvider | None
- verifier¶
Hit verifier for semantic candidates;
Nonekeeps the engine default (LexicalNearMissVerifier).- Type:
HitVerifier | None
- timeout¶
Upstream timeout in seconds.
- Type:
float
- upstream: str = 'https://api.openai.com/v1'¶
- api_key: str | None = None¶
- memo_entries: int = 4096¶
- prefix_entries: int = 8192¶
- semantic_threshold: float | None = None¶
- semantic_entries: int = 2048¶
- embedder: EmbeddingProvider | None = None¶
- verifier: HitVerifier | None = None¶
- timeout: float = 600.0¶
OpenTelemetry export for per-tier reuse events.
The interpreter can report every reuse-tier lookup (memo, semantic, prefix KV,
layer KV, memory graph) and every node it executes as a
ReuseEvent. Nothing is emitted unless an observer
is attached, so telemetry costs nothing when off.
OpenTelemetryObserver turns those events into:
spans: one per executed node (
continuum.node), with one child span per tier lookup (continuum.reuse.<tier>), timed with the interpreter’s own clock readings;metrics:
continuum.reuse.lookups/continuum.reuse.hits/continuum.reuse.tokens_savedcounters and acontinuum.reuse.lookup.durationhistogram (ms), all by tier, plus acontinuum.node.executionscounter by node kind andserved_by.
Turn it on per session, or for every DurableAgent with an
environment variable:
from continuum.telemetry import instrument
instrument(session) # uses the global OTel providers
CONTINUUM_OTEL=1 python my_agent.py # DurableAgent instances self-instrument
Requires opentelemetry-api (and an SDK / exporter to send data anywhere):
pip install "continuum-ai[otel]".
- class CallbackObserver(*args: Any, **kwargs: Any)[source]¶
Bases:
ReuseObserverCall
fn(event_dict)for every event; exceptions are logged, not raised.
- class OpenTelemetryObserver(*args: Any, **kwargs: Any)[source]¶
Bases:
ReuseObserverExport reuse events as OpenTelemetry spans and metrics.
- Parameters:
tracer_provider – Defaults to the global
opentelemetry.traceprovider.meter_provider – Defaults to the global
opentelemetry.metricsprovider.
- auto_instrument(target: _Observable) OpenTelemetryObserver | None[source]¶
Instrument
targetifCONTINUUM_OTELis set; otherwise do nothing.
- event_to_dict(event: continuum._native.ReuseEvent) dict[str, Any][source]¶
Copy an event into a plain dict (events are only valid during the callback).
- instrument(target: _Observable, tracer_provider: Any = None, meter_provider: Any = None) OpenTelemetryObserver[source]¶
Attach an
OpenTelemetryObserverto aSessionorDurableAgent.
Embedding providers for the semantic cache and memory-graph tiers.
Every provider subclasses continuum._native.EmbeddingProvider, so it
can be handed straight to Session.set_embedding_provider. Each one reports
an identity(); the semantic cache stores it with every entry and only
compares vectors that share it, so switching embedders never produces a false
hit against vectors from another embedding space.
Three shapes cover the common cases:
CallableEmbeddingProviderwraps anytext -> vectorfunction, e.g. a local sentence-transformers model.OpenAICompatibleEmbeddingProvidercalls a hosted/v1/embeddingsendpoint (OpenAI, vLLM, Ollama, LM Studio, …).PrecomputedEmbeddingProviderserves vectors computed ahead of time, for fully reproducible runs and offline evaluation.WordLlamaEmbeddingProvideris a small semantic model that runs locally on CPU with no download (pip install "continuum-ai[semantic]"). It is the recommended starting point for the semantic tier.
Pair any of them with the semantic tier’s hit verifier (on by default, see
continuum.verifiers): similarity alone cannot tell a paraphrase from a
near-miss edit such as “enable” vs “disable”.
The built-in continuum._native.BruteForceEmbeddingProvider (character
n-gram hashing, identity continuum/char-ngram-v1:<dim>) stays the default
in examples: dependency-free and deterministic, but lexical rather than
semantic.
- class CallableEmbeddingProvider(*args: Any, **kwargs: Any)[source]¶
Bases:
EmbeddingProviderAdapt a
text -> vectorcallable (a local model, a cached lookup, …).- Parameters:
fn – Returns the embedding of one string.
dimension – Length of every vector
fnreturns; checked on each call.identity – Stable name of the embedding space, e.g.
"sentence-transformers/all-MiniLM-L6-v2". Change it whenever the vectors would change (new model, new normalization, …).
- class OpenAICompatibleEmbeddingProvider(*args: Any, **kwargs: Any)[source]¶
Bases:
EmbeddingProviderCall an OpenAI-format
POST {base_url}/v1/embeddingsendpoint.Works with OpenAI, vLLM (
--task embed), Ollama, and any server that speaks the same wire format.- Parameters:
base_url – Server root, e.g.
"http://localhost:11434".model – Embedding model name sent in the request.
dimension – Expected vector length.
Noneprobes the server once.api_key – Sent as a bearer token when given.
identity – Defaults to
"openai-compatible/<model>"; the host is left out on purpose, since one model yields one embedding space wherever it is served.timeout – Per-request timeout in seconds.
- class PrecomputedEmbeddingProvider(*args: Any, **kwargs: Any)[source]¶
Bases:
EmbeddingProviderServe vectors computed ahead of time, keyed by the exact prompt text.
- Parameters:
vectors – Prompt text to embedding. All vectors must share one length.
identity – Name of the embedder that produced
vectors.fallback – Provider for texts not in
vectors. It must report the same identity and dimension; without one, unknown text raisesKeyError.
- class WordLlamaEmbeddingProvider(*args: Any, **kwargs: Any)[source]¶
Bases:
EmbeddingProviderLocal semantic embeddings from WordLlama (static token embeddings distilled from an LLM’s input layer; ~16 MB, CPU, sub-millisecond).
The model files ship inside the
wordllamawheel, so this works with no network access. Requirespip install "continuum-ai[semantic]".- Parameters:
config – WordLlama model config (
"l2_supercat"ships in the wheel).dim – Embedding width (the bundled model is 256; smaller truncates).
Hit verifiers for the semantic cache tier.
Similarity search finds the cached prompt closest to a new one; a verifier then decides whether that cached answer is actually correct for it. This second stage is what keeps near-miss edits (“enable” vs “disable” two-factor auth, “Australia” vs “Austria”, “5” vs “50” euros) from being served the wrong answer: embedders score those pairs as high as true paraphrases.
LexicalNearMissVerifier(the default on everySemanticCacheIndex) rejects minimal edits: swapped numbers, swapped content words with everything else unchanged, flipped polarity or negation. Fast and deterministic, but it cannot tell a synonym swap (“cancel” vs “end”) from an antonym swap, so it also turns those down.LLMJudgeVerifierasks a chat model whether both prompts have the same answer. It understands synonyms and antonyms, at the cost of one model call per candidate hit (verdicts are cached).AllOfrequires every verifier in a chain to accept. Note that chaining the lexical verifier in front of a judge keeps the lexical verifier’s synonym rejections; use the judge alone when recall on rewordings matters.
Use one with SemanticCacheIndex.set_verifier(...); None disables
verification. Measurements: benchmarks/reports/semantic-false-hits.md.
- class AllOf(*args: Any, **kwargs: Any)[source]¶
Bases:
HitVerifierAccept a hit only if every verifier accepts it (evaluated in order).
- class LLMJudgeVerifier(*args: Any, **kwargs: Any)[source]¶
Bases:
HitVerifierVerify candidate hits with a chat model over an OpenAI-compatible API.
Works with OpenAI, vLLM, and Ollama (
base_url="http://localhost:11434"). Any failure (network, unexpected reply) rejects the hit: a miss is safe, a wrong answer is not.- Parameters:
base_url – Server root, e.g.
"http://localhost:11434".model – Chat model name, e.g.
"gemma4".api_key – Sent as a bearer token when given.
timeout – Per-request timeout in seconds.
prompt – Judge prompt with
{a}and{b}placeholders.
Runtime¶
These classes are defined by the C++ pybind bindings and re-exported from
continuum._native. Import them from there:
from continuum import DurableAgent
from continuum._native import Session, ReusePolicy
DurableAgent¶
- class continuum._native.DurableAgent¶
A step-sequenced agent run that can be checkpointed, resumed, and forked. See Durable execution.
- begin(prompts, model_id=None, max_tokens=32)¶
Build the graph for a list of step prompts. Returns the step count. Each prompt becomes a
PromptOpfeeding aTokenOp; every step after the first also receives the previous step’s output.model_iddefaults to"vllm/gemma4"whenVLLM_BASE_URLis set, else"fake/model".
- run_until_step(step_index)¶
Execute through
step_index(0-based) and return the checkpoint asbytes: the graph, every computed value, and the portable KV state.
- resume_from(checkpoint)¶
Deserialize
checkpointbytes into a fresh runtime and run to completion. Returns the list of node outputs. Deterministic: two resumes of one checkpoint return identical output.
- cache_size()¶
Number of warm KV entries currently held.
- step_outputs()¶
Each step’s generated output in step order (
Noneif not yet run): text on a live server, token ids on the fake backend.
- backend¶
"vllm"whenVLLM_BASE_URLis set, else"fake".
- static inspect(checkpoint)¶
Read a checkpoint without a runtime. Returns a dict with
executed_nodesandcheckpoint_bytes.
- static fork(checkpoint, node_id, new_value)¶
Return a new checkpoint with the value at
node_idreplaced. Resuming it replays completed steps unchanged and diverges only at the edited node and downstream.
- prompt_node_ids¶
List of
PromptOpnode ids, one per step, in order.
- step_node_ids¶
List of
TokenOp(generation) node ids, one per step, in order.
Session¶
- class continuum._native.Session(id, backend_registry, max_cache=...)¶
A reuse-aware execution context. Holds the cache index and the reuse policy across many
runcalls.- run(graph, inputs)¶
Execute
graphwith a mapping of input node id to value, applying the reuse stack to everyTokenOp.
- policy¶
The
ReusePolicyfor this session. Assignable.
- metrics()¶
A
ReuseMetricssnapshot: tokens sent, tokens saved, per-tier hit counts.
- reset_metrics()¶
- save_cache_metadata(path)¶
- load_cache_metadata(path)¶
Persist and reload the cache index across processes. See The reuse stack on cross-session persistence.
- generate(prompt_parts, model_id, max_tokens=128, temperature=0.0, op_name='generate')¶
Run one generation through the reuse stack: each string in
prompt_partsbecomes aPromptOpfeeding a singleTokenOp. Returns the output value and appends a step tometrics().
- cache_size()¶
- cache_stats()¶
Per-tier occupancy:
{tier: {"entries", "capacity", "bytes"}}for every attached tier. See “Eviction and memory bounds” indocs/design/cache.md.
- set_memo_table(table)¶
- set_semantic_cache(index)¶
- set_embedding_provider(provider)¶
- set_layer_cache(index)¶
- set_memory_graph(store)¶
Attach the backing store for each reuse tier. A tier with no store attached is skipped.
ReusePolicy¶
- class continuum._native.ReusePolicy¶
- static always()¶
Reuse whenever any tier matches.
- static never()¶
Bypass every tier. Every call reaches the backend.
- static threshold(min_len)¶
Reuse a prefix only when it is at least
min_lentokens.
- kind¶
- min_prefix_len¶
The threshold, when
kindisThresholdPrefixLen.
- class continuum._native.ReusePolicyKind¶
Enum:
Always,Never,ThresholdPrefixLen.
Reuse-tier stores¶
Backing stores for the tiers, attached to a Session with the set_*
methods above:
MemoTable/MemoKey: exact-repeat memoization.SemanticCacheIndex: embedding-matched near-duplicates. Needs anEmbeddingProvider:BruteForceEmbeddingProvideris bundled, andcontinuum.embeddingsadapts local models, hosted endpoints, and precomputed vectors. SubclassEmbeddingProviderfor anything else.LayerKVCacheIndex: warm attention-layer state.MemoryGraphStore: prior-run context recall.FutureCache: in-flight de-duplication of concurrent identical calls.
IR¶
Graph,Node,NodeKind: the dataflow IR. See The execution model.GraphBuilder: incremental graph construction.Interpreter: the executor.SessionandDurableAgentwrap it.BackendRegistry: registers backends and their priorities.
Benchmark entrypoints¶
run_session_benchmark , run_cold_start_benchmark , run_v11_benchmark ,
benchmark_azure_agent , benchmark_vllm_agent , and related functions
drive the numbers in The reuse stack and docs/benchmarks.md. They return
plain dicts of metrics. examples/01_reuse_stack.py calls them directly.