|
Continuum C++ API
Unified runtime for token + tensor execution
|
#include <kv_prefix_cache.hpp>
Classes | |
| struct | SnapshotEntry |
| struct | TrieNode |
Public Member Functions | |
| KVCacheIndex (std::size_t max_entries=8192) | |
| std::optional< std::pair< CacheEntry, std::int32_t > > | longest_prefix (const std::string &model_id, const DecodeParams &decode, const std::vector< std::int32_t > &tokens, const std::string &cache_namespace={}) const |
| void | insert (CacheEntry entry, const std::vector< std::int32_t > &token_prefix) |
| void | insert_unlocked (CacheEntry entry, const std::vector< std::int32_t > &token_prefix) |
| void | invalidate (void *backend_handle) |
| void | clear () |
| std::size_t | size () const |
| std::size_t | max_entries () const |
| Capacity in entries passed at construction. | |
| std::size_t | estimated_bytes () const |
| bool | save_metadata (const std::string &path) const |
| bool | load_metadata (const std::string &path) |
| std::vector< SnapshotEntry > | snapshot () const |
Prefix-KV tier: a token trie of reusable backend states.
Eviction: least-recently-used. Every trie depth covered by an insert holds its own entry, so size() counts per-depth entries. insert and a longest_prefix hit refresh recency; once size() exceeds max_entries the least recently used entry is removed and empty branches are compacted.
|
explicit |
| void continuum::runtime::KVCacheIndex::clear | ( | ) |
| std::size_t continuum::runtime::KVCacheIndex::estimated_bytes | ( | ) | const |
Approximate resident bytes of the index itself: trie nodes plus entry metadata. Backend-owned state behind each handle is not counted.
| void continuum::runtime::KVCacheIndex::insert | ( | CacheEntry | entry, |
| const std::vector< std::int32_t > & | token_prefix | ||
| ) |
| void continuum::runtime::KVCacheIndex::insert_unlocked | ( | CacheEntry | entry, |
| const std::vector< std::int32_t > & | token_prefix | ||
| ) |
| void continuum::runtime::KVCacheIndex::invalidate | ( | void * | backend_handle | ) |
| bool continuum::runtime::KVCacheIndex::load_metadata | ( | const std::string & | path | ) |
| std::optional< std::pair< CacheEntry, std::int32_t > > continuum::runtime::KVCacheIndex::longest_prefix | ( | const std::string & | model_id, |
| const DecodeParams & | decode, | ||
| const std::vector< std::int32_t > & | tokens, | ||
| const std::string & | cache_namespace = {} |
||
| ) | const |
|
inline |
Capacity in entries passed at construction.
| bool continuum::runtime::KVCacheIndex::save_metadata | ( | const std::string & | path | ) | const |
| std::size_t continuum::runtime::KVCacheIndex::size | ( | ) | const |
| std::vector< SnapshotEntry > continuum::runtime::KVCacheIndex::snapshot | ( | ) | const |