Is provider KV caching sufficient for agent swarms and long run agents?

I’m trying to build a side project in the inference space. I’ve been talking to a few inference engineers and startups and I’ve been hearing how annoying it is to not have manual control over the KV cache at times and just constantly being subject to the black box caching methods of their inference providers. It is particularly annoying for agent swarms when you want to fork agents from the same cached prefix or manually store a cache for a longer period for a future agent to hit later.

Im curious if this problem is consistent across multiple people and if there are any solutions for it that people know about.

Hi — yes, I think this is a real and consistent frustration, and it mostly comes down to one thing: with providers you’re renting cache space, not owning it.

The way automatic prefix caching works is that the provider looks for an exact token-prefix match with a request it has already processed. If it matches and the cached entry is still alive, you skip recomputation and pay less. What makes it annoying for agent work is that every important decision — how long an entry lives, when it gets evicted under load, and whether your exact prompt (including any tokens the provider itself prepends, like system prompts) actually matches — is invisible to you. So two identical agent forks can get different costs and latencies for no reason you can see.

A test that could prove me wrong here: send the same long prompt twice through your provider, about 30 minutes apart, and compare time-to-first-token. If the second one is consistently much faster, the cache is surviving and something else is biting you. If it’s not, eviction is the answer and no prompt tweaking will fix it.

If you want real control, the smallest useful step I know is to run your own inference server once, just as an experiment — vLLM and SGLang both do prefix caching explicitly, and you can see the cache hits in the logs. That tells you what “owning the cache” actually buys you before you commit to anything. LMCache is worth a look too if you want to persist the KV blocks to storage yourself, so a future agent can reload them hours later.