P
Back to articles
News4 min read

Tracking token usage across OpenAI, Anthropic, and Gemini: every streaming gotcha I hit

Developers tracking LLM costs across OpenAI, Anthropic, and Gemini face hidden gotchas in streaming token reporting — from different cache conventions to split-event stream formats.

Source: DEV Community
Tracking token usage across OpenAI, Anthropic, and Gemini: every streaming gotcha I hit

Why Token Tracking Matters More Than Ever

As organizations deploy large language models across OpenAI, Anthropic, and Google Gemini, one deceptively simple task has become a source of recurring frustration: accurately tracking token usage and cost. What sounds like a straightforward API read operation turns out to be a minefield of inconsistent conventions, split-stream events, and silent accounting differences.

Spanlens, an open-source LLM observability tool that proxies requests to all three providers, recently documented the streaming gotchas that developers encounter when trying to normalize token usage data. The findings serve as a practical warning for anyone building cost tracking or observability infrastructure around LLM APIs.

The Core Problem: Three Providers, Three Philosophies

Each major LLM provider reports token usage in streaming responses, but the similarities end at the abstract concept. Where the data lives, how cached tokens are counted, and what the fields are called all differ — and getting any of it wrong means your cost numbers are silently incorrect.

Gotcha 1: Token Counts Live in Different Stream Locations

For non-streaming calls, every provider returns a clean usage object on the response body. Streaming is where the complexity multiplies. OpenAI places token usage in a final chunk, after all content, right before the [DONE] signal — but only if you explicitly request it with stream_options: { include_usage: true }. Miss that opt-in flag and you receive the entire streaming response with zero usage data.

Anthropic takes a different approach entirely, splitting usage across two distinct events. Input tokens arrive early in message_start, while output tokens arrive at the very end in message_delta. Any parser that listens for only one event ends up with a grossly incomplete picture.

For developers building parsers, this means maintaining separate mental models for each provider — OpenAI requires preserving the last chunk, while Anthropic demands stitching together the first and last events.

Gotcha 2: Cached Token Accounting Is Reversed

This is arguably the most dangerous pitfall because it quietly corrupts cost data without throwing errors. Both OpenAI and Anthropic support prompt caching and report cached token counts, but they do so with opposite conventions.

OpenAI's prompt_tokens already includes cached tokens — the cached count is a subset within it. If you want the uncached portion, you subtract. Anthropic's input_tokens, by contrast, represents only the uncached portion. The cached tokens are reported separately and must be added to get the true total.

The same caching concept, mathematically inverted between providers. Code that naively passes both providers through the same function produces cost numbers that are wrong by the size of the cache — and cache hits are precisely the high-volume calls where the error is largest. Silent financial data corruption is the worst category of bug.

Gotcha 3: Gemini Streams in Two Formats

OpenAI and Anthropic both use standard server-sent events (SSE) with data: prefix lines. Gemini adds another wrinkle: it supports SSE only when ?alt=sse is appended to the URL. Without that parameter, the default streamGenerateContent endpoint delivers a single giant JSON array streamed character by character.

A robust Gemini parser must therefore handle both shapes — try SSE first, then fall back to parsing the buffer as a JSON array, then fall back again to scanning line by line for partial chunks. The field names add another layer of mismatch: OpenAI uses prompt_tokens/completion_tokens, while Gemini uses promptTokenCount/candidatesTokenCount inside a usageMetadata object.

Gotcha 4: Requested Tier vs. Served Tier

All three providers can report a service tier (default, flex, priority), and the cost depends on it. The critical insight is that the tier in the response is the tier actually served, which may differ from what was requested. OpenAI can downgrade a priority request to default under load, and only the response reveals the downgrade. Basing cost calculations on the request tier rather than the response tier introduces another class of silent errors.

Best Practices for Multi-Provider Token Tracking

The key lesson from Spanlens' experience is to resist the temptation of a shared early abstraction. Developers should write one parser per provider, validate each against real streaming responses, and only then collapse them behind a common interface. The differences are not cosmetic — where the number lives, whether cache is included, and what the field is named are all provider-specific concerns.

Equally important is assertion discipline around cost-bearing numbers. Type errors surface immediately on the first development request. A token count that is off by the cache size ships silently and emerges as a billing discrepancy weeks later. Every token-related assertion is worth writing.

As organizations increasingly adopt multi-provider LLM strategies, tooling that transparently normalizes usage and cost data across providers becomes essential infrastructure. The gotchas documented here are the difference between cost tracking that inspires confidence and one that quietly misleads.

Related Articles

RAG Pipeline: The Uncle-Nephew Complete Learning Guide
News4 min

RAG Pipeline: The Uncle-Nephew Complete Learning Guide

Learn how Retrieval-Augmented Generation (RAG) pipelines work — from vector retrieval and prompt augmentation to answer generation — and how to implement them effectively in production.

AI & Research
AI Agents in Group Chat: Octo Solves Multi-Agent Coordination
News3 min

AI Agents in Group Chat: Octo Solves Multi-Agent Coordination

Mininglamp Technology's Octo open-source platform reimagines AI agent orchestration by making agents first-class participants in group chat conversations, solving the session island coordination crisis.

AI & Research
After spat with Chinese gov't, Meta cuts AI Manus off from its internal systems and is 'sunsetting' platform, report claims — Beijing-ordered breakup of $2 billion AI deal begins
News4 min

After spat with Chinese gov't, Meta cuts AI Manus off from its internal systems and is 'sunsetting' platform, report claims — Beijing-ordered breakup of $2 billion AI deal begins

Meta has locked Manus AI out of its systems and is sunsetting the agentic AI platform after China's NDRC ordered the $2 billion acquisition unwound. Founders are now racing to raise $1 billion for a buyback.

AI & Research
LUCID: Learning Embodiment-Agnostic Intent Models from Unstructured Human Videos for Scalable Dexterous Robot Skill Acquisition
News7 min

LUCID: Learning Embodiment-Agnostic Intent Models from Unstructured Human Videos for Scalable Dexterous Robot Skill Acquisition

Researchers from Carnegie Mellon University have introduced LUCID, a two-stage framework that enables robots to learn dexterous manipulation skills from unstructured human videos at internet scale. The system separates intent prediction (what should happen next in a scene) from embodiment-specific motor control, allowing the same intent model to work across different robot platforms — from dexterous hands to parallel-jaw grippers. Evaluated on five real-world tasks including stirring, wiping, binning, push-T, and cable routing, LUCID achieved zero-shot transfer to novel scenes and objects using only internet video as training data.

AI & Research