Horace He Calls Cached Token Counts Incredibly Dumb
Researchers note that visible dashboard numbers lead users to include cached inputs when reporting token usage.
Horace He posted on X that people discussing LLM token usage usually include cached input tokens and called the practice incredibly dumb. Hieu Pham replied that users read the most visible dashboard number rather than seeking a breakdown. Lucas Beyer joked that a "kv tokens" metric multiplying by layer count would make the numbers even bigger. Horace He followed up suggesting counting every decode step load of cached tokens. The posts discuss reported LLM token counts that include cached inputs.
I've realized when people are talking about how many tokens they use, they're usually including cached input tokens... which is incredibly dumb...
Combined views
56.9K
Horace He Calls Cached Token Counts Incredibly Dumb
Researchers note that visible dashboard numbers lead users to include cached inputs when reporting token usage.
Horace He posted on X that people discussing LLM token usage usually include cached input tokens and called the practice incredibly dumb. Hieu Pham replied that users read the most visible dashboard number rather than seeking a breakdown. Lucas Beyer joked that a "kv tokens" metric multiplying by layer count would make the numbers even bigger. Horace He followed up suggesting counting every decode step load of cached tokens. The posts discuss reported LLM token counts that include cached inputs.