Your agents keep paying to compute the same thing twice.
Continuum runs LLM calls and tensor ops as one dataflow graph. It reuses work it has already done, and it can save a running job to bytes and finish it in another process.
- 92.5%
- fewer tokens on a 20-step agent run
- 0 ms
- to answer a call it has seen before
- same result
- after a resume or a fork
01 · The problem
Most of an agent run is work it has already done.
The graph repeats itself. Continuum removes the repeats at the layer where the tokens are actually spent.
- 01The system prompt goes to the model on every step, unchanged.
- 02The same question gets asked and answered more than once, word for word or close to it.
- 03A crash at step 19 throws away steps 1 through 18.
02 · Reuse
Each call takes the cheapest path that answers it.
A call checks four caches in order and stops at the first hit. Only a cold call reaches the backend.
- Memo The same call, seen before. Return the stored answer. 0 ms · 0 tokens
- Semantic Different wording, same question. Matched by embedding. 1 lookup · 0 tokens
- Prefix KV Shared prompt prefix, already tokenized. Send only the new part. ~30 tokens sent
- Layer KV Warm attention state carried forward. Resume the decode. no prefill
- Backend Nothing above matched. One real request to the provider. full cost
A fifth tier, memory-graph recall, lets an agent pull facts from earlier runs.
03 · Measured
Numbers from a live Azure OpenAI backend.
fewer tokens on a mixed 20-step agent run: shared prefixes, exact repeats, paraphrases, and cold queries. Four of the twenty calls never reached the backend.
- ~99%
- fewer tokens on a 3,000-character shared prefix. About 30 tokens sent per call.
- 5 / 5
- exact repeat calls served from cache, 0 ms each.
- 5.4s to 3.7s
- median response time on prefix hits. The network round-trip stays, so time drops less than token cost.
- 80%+
- cache hit rate on the first run after restarting in a new session.
04 · Durable execution
Save a running job. Finish it somewhere else.
A checkpoint holds the graph, every computed value, and the KV cache. It is a byte string. Write it to a file, an object store, or a queue, and any process that can read those bytes can carry the run forward.
from continuum._native import DurableAgent agent = DurableAgent() agent.begin(["pull the ticket", "reproduce the bug", "draft a fix", "open the PR"]) ckpt = agent.run_until_step(1) # steps 1 and 2 run now DurableAgent.inspect(ckpt) # {'executed_nodes': 2, 'checkpoint_bytes': 4812} outputs = DurableAgent().resume_from(ckpt) # a fresh runtime finishes steps 3 and 4
# plan.py · a small box: run the cheap steps, then park the job from pathlib import Path from continuum._native import DurableAgent agent = DurableAgent() agent.begin(["pull the ticket", "reproduce the bug", "draft a fix", "open the PR"]) Path("job.ckpt").write_bytes(agent.run_until_step(1)) # steps 1-2, a few KB # work.py · a different process, a GPU box, an hour later from pathlib import Path from continuum._native import DurableAgent outputs = DurableAgent().resume_from(Path("job.ckpt").read_bytes()) Path("job.result").write_text(repr(outputs)) # steps 3-4 finish here # audit.py · a third process, anywhere, replays the same bytes from pathlib import Path from continuum._native import DurableAgent blob = Path("job.ckpt").read_bytes() assert DurableAgent().resume_from(blob) == DurableAgent().resume_from(blob)
No shared variable, no message bus. The checkpoint carries the KV cache
with it, so work.py resumes warm, and every reader gets the
same result.
from continuum._native import DurableAgent rec = DurableAgent() rec.begin(["summarize the bug", "find the module", "draft a fix", "write the changelog"]) ckpt = rec.run_until_step(1) # steps 1 and 2 executed step_4 = rec.prompt_node_ids[3] # the step-4 prompt, still pending real = DurableAgent().resume_from(ckpt) what_if = DurableAgent().resume_from( DurableAgent.fork(ckpt, step_4, "write a haiku instead"), ) # steps 1-3 replay bit for bit; only step 4 and its output diverge
05 · One IR
One graph for tokens and tensors.
LLM calls and tensor ops are operators in the same IR, run by the same cache-aware interpreter. Change providers without touching the graph.
06 · Start
Install it and run the examples.
Every example uses the FakeLLM backend, so the output is deterministic and safe to run in CI. The import path stays continuum.
PYTHONPATH=python python examples/01_reuse_stack.py # every reuse tier, one run PYTHONPATH=python python examples/02_durable_agent.py # checkpoint, crash, resume PYTHONPATH=python python examples/03_time_travel_fork.py # rewind, edit, replay