ChiChieh HuangFOUNDER · AI ENGINEER

Rethinking Prompt Caching Through ngrok's Article

Date2025.12.28
Length187 words
Reading~1 min
ChiChieh HuangFounder · AI engineer

Translated from the Chinese original · Read the original

LocatorVibe Coding Valley
2 wks
Overview

ngrok’s article on prompt caching is well worth reading. When building products, I used to just calculate how much ten thousand tokens cost. Reading it made me realize that what really costs money is recomputing the same opening prompt hundreds of times a day.

It explains very clearly what gets cached inside a transformer. It’s not some mysterious black magic: it stores the K/V results of attention, so the next time the same opening prompt comes in, attention doesn’t have to run over it again.

For applications with long system prompts, long rules or long instructions, if you’re willing to spend the time restructuring the prompt once so the reusable parts are saved, your token costs and latency can drop significantly. Anyone building products has probably felt that helplessness: the model seems to be able to do it, but the cloud bill is just high. We’ve used a similar approach before, so I really recommend giving it a try. Treat low-level mechanisms like this as part of product design, instead of leaving it to the cloud provider to guess on your behalf. https://ngrok.com/blog/prompt-caching

End of the trail

187 words, and you made it to the end.

Newsletter

Get the next essay by email.

One email when a new essay goes up, nothing else. Unsubscribe in one click.

ChiChieh Huang
Surveyor

ChiChieh Huang

I build generative AI products and write about them, first in Chinese. Lately I’ve been researching agent memory and testing the ideas in Cairn.