295,107 tokens would have been sent without the proxy in place. That is the number every reduction below is measured against.
98,369 tokens were sent. The difference is roughly 67 percent, and the report attributes it optimization by optimization.
Removing repeated text blocks from earlier turns, which is worth 55 to 80 percent on its own in conversation-heavy work.
Cold intent classification adds 100 to 200 ms, and a warm cache adds effectively nothing. Streaming is preserved, so responses still arrive token by token.