Your agents keep paying to compute the same thing twice.

Continuum runs LLM calls and tensor ops as one dataflow graph. It reuses work it has already done, and it can save a running job to bytes and finish it in another process.

pip install continuum-ai Docs Python C++
92.5%
fewer tokens on a 20-step agent run
0 ms
to answer a call it has seen before
same result
after a resume or a fork

01 · The problem

Most of an agent run is work it has already done.

The graph repeats itself. Continuum removes the repeats at the layer where the tokens are actually spent.

  1. 01The system prompt goes to the model on every step, unchanged.
  2. 02The same question gets asked and answered more than once, word for word or close to it.
  3. 03A crash at step 19 throws away steps 1 through 18.

02 · Reuse

Each call takes the cheapest path that answers it.

A call checks four caches in order and stops at the first hit. Only a cold call reaches the backend.

  1. Memo The same call, seen before. Return the stored answer. 0 ms · 0 tokens
  2. Semantic Different wording, same question. Matched by embedding. 1 lookup · 0 tokens
  3. Prefix KV Shared prompt prefix, already tokenized. Send only the new part. ~30 tokens sent
  4. Layer KV Warm attention state carried forward. Resume the decode. no prefill
  5. Backend Nothing above matched. One real request to the provider. full cost

A fifth tier, memory-graph recall, lets an agent pull facts from earlier runs.

03 · Measured

Numbers from a live Azure OpenAI backend.

92.5%

fewer tokens on a mixed 20-step agent run: shared prefixes, exact repeats, paraphrases, and cold queries. Four of the twenty calls never reached the backend.

Tokens sent, baseline100%
Tokens sent, Continuum7.5%
~99%
fewer tokens on a 3,000-character shared prefix. About 30 tokens sent per call.
5 / 5
exact repeat calls served from cache, 0 ms each.
5.4s to 3.7s
median response time on prefix hits. The network round-trip stays, so time drops less than token cost.
80%+
cache hit rate on the first run after restarting in a new session.

04 · Durable execution

Save a running job. Finish it somewhere else.

A checkpoint holds the graph, every computed value, and the KV cache. It is a byte string. Write it to a file, an object store, or a queue, and any process that can read those bytes can carry the run forward.

one process
from continuum._native import DurableAgent

agent = DurableAgent()
agent.begin(["pull the ticket", "reproduce the bug", "draft a fix", "open the PR"])

ckpt = agent.run_until_step(1)              # steps 1 and 2 run now
DurableAgent.inspect(ckpt)                 # {'executed_nodes': 2, 'checkpoint_bytes': 4812}

outputs = DurableAgent().resume_from(ckpt) # a fresh runtime finishes steps 3 and 4
Checkpoint to bytes Resume in a new process Fork from any past step Deterministic replay

05 · One IR

One graph for tokens and tensors.

LLM calls and tensor ops are operators in the same IR, run by the same cache-aware interpreter. Change providers without touching the graph.

Azure OpenAI hosted OpenAI hosted Anthropic hosted vLLM self-hosted libtorch in-process MLX apple silicon FakeLLM tests, deterministic

06 · Start

Install it and run the examples.

Every example uses the FakeLLM backend, so the output is deterministic and safe to run in CI. The import path stays continuum.

pip install continuum-ai
run the demos
PYTHONPATH=python python examples/01_reuse_stack.py       # every reuse tier, one run
PYTHONPATH=python python examples/02_durable_agent.py     # checkpoint, crash, resume
PYTHONPATH=python python examples/03_time_travel_fork.py  # rewind, edit, replay