The one-line answer
1 million tokens ≈ 750,000 words ≈ 1,500 single-spaced pages ≈ about 10 average novels. The math comes from OpenAI's widely cited rule of thumb: 1 token covers roughly 0.75 words of English text (equivalently, ~4 characters).
1M tokens in units you can picture
Estimates round to the nearest useful figure. Exact token counts vary by tokenizer (OpenAI's tiktoken, Anthropic's tokenizer, Gemini's SentencePiece), by language (non-English text often tokenizes 2–3× denser), and by content type (code and JSON pack more tokens per character than prose).
What does 1 million tokens cost?
Prices below are per 1M tokens (USD), so this is literally the per-million-token sticker price for each model — input vs. output.
| Model | 1M input | 1M output |
|---|---|---|
| Gemini 1.5 Flash | $0.075 | $0.30 |
| GPT-4o mini | $0.15 | $0.60 |
| Gemini 1.5 Pro | $1.25 | $5.00 |
| GPT-4o | $2.50 | $10.00 |
| Claude 3.5 Sonnet | $3.00 | $15.00 |
| Claude Opus 4.8 | $5.00 | $25.00 |
Read this way: feeding all 10 novels into GPT-4o as input costs $2.50; asking it to generate 10 novels back costs $10. Same amount of text, 4× the price — because output tokens are compute-heavier than input tokens.
Where the 0.75 words / token rule comes from
Modern LLMs use byte-pair encoding (BPE) or SentencePiece tokenizers. Common English words like "the", "and", "of" are usually a single token. Longer or rarer words split into pieces — "tokenization" typically breaks into "token" + "ization". Averaged over natural English prose, this comes out to ~0.75 words per token, or ~4 characters per token. That is the heuristic OpenAI publishes and what this site's calculator uses.
Content types that drift from the average:
- Code and JSON tokenize denser — braces, quotes, and identifiers eat tokens fast. Budget ~30% more tokens than the same character count of prose.
- Non-English languages (especially CJK, Arabic, Hindi) can be 2–3× denser per character. A 1,000-word Japanese document may consume more tokens than a 1,000-word English one.
- System prompts repeat on every call. A 2,000-token system prompt hit 500,000 times a month = 1 billion tokens before the user even types.
Budgeting rule of thumb
When you see a monthly cost quote like "$50 / month on GPT-4o", translate it back to human units:
- $50 on GPT-4o input ≈ 20M tokens ≈ 15M words ≈ ~200 novels of context read per month.
- $50 on GPT-4o output ≈ 5M tokens ≈ 3.75M words ≈ ~50 novels generated per month.
- $50 on Gemini 1.5 Flash input ≈ 666M tokens ≈ ~6,600 novels — two orders of magnitude more text for the same dollar.
Try it on your own text
Paste any prompt into the calculator and it will show the token count and monthly cost across every major model — 100% client-side, nothing uploaded.
Open the AI Cost Calculator