Skip to main content
For Developers4 min read

Understanding Token Limits in AI Models: GPT, Claude, and Gemini

Learn what tokens are in AI language models, how they affect costs, and strategies to optimise your token usage for GPT, Claude, and Gemini.

Northern Codes teamUpdated

What Are Tokens in AI Language Models?

When you interact with AI models like GPT, Claude, or Gemini, your text isn't processed as whole words. Instead, it's broken down into smaller units called tokens. Understanding tokens is essential for developers and businesses using AI APIs, as tokens directly impact both functionality and costs.

A token can be as short as one character or as long as one word. On average:

  • English text: 1 token ≈ 4 characters or ≈ 0.75 words

  • Code: Token count varies significantly based on syntax

  • Other languages: May use more tokens per word

Why Token Limits Matter

1. Context Window Limits

Each AI model has a maximum context window: the total number of tokens it can process in a single request (including both input and output). These are current models, checked on the providers' own pages on 27 September 2026:

Model

Context Window

GPT-6 Astra (OpenAI)

1.05 million tokens

GPT-6 Sol (OpenAI)

1.05 million tokens

Claude Opus 5.5 (Anthropic)

1 million tokens

Claude Sonnet 5 (Anthropic)

1 million tokens

Gemini 3.8 Flash (Google)

1,048,576 tokens

Exceeding these limits means your request will fail or be truncated.

2. Cost Implications

AI APIs charge per token, with different rates for input and output tokens. Standard prices checked on 27 September 2026, for example:

  • GPT-6 Astra: $10 per 1M input tokens, $50 per 1M output tokens

  • Claude Opus 5.5: $4 per 1M input tokens, $20 per 1M output tokens

  • Gemini 3.8 Flash: $0.75 per 1M input tokens, $3.75 per 1M output tokens (rising to $1.50 and $7.50 on 1 January 2027)

A single lengthy conversation can cost several dollars if not managed carefully.

How Tokenization Works

Different models use different tokenization algorithms:

OpenAI (GPT models)

Uses Byte Pair Encoding (BPE) via the tiktoken library. Common words become single tokens, while rare words are split into subwords.

Example: "tokenization" → ["token", "ization"] (2 tokens)

Anthropic (Claude)

Uses a similar BPE approach but with a different vocabulary, resulting in slightly different token counts for the same text.

Google (Gemini)

Uses SentencePiece tokenization, which can handle multiple languages more uniformly.

Strategies to Optimize Token Usage

1. Be Concise in Prompts

Remove unnecessary words and redundant instructions. Instead of "Could you please help me by writing a function that...", use "Write a function that...".

2. Use System Messages Wisely

System messages are included in every request. Keep them brief but effective.

3. Implement Conversation Summarization

For long conversations, periodically summarize earlier exchanges instead of including the full history.

4. Choose the Right Model

Don't use a provider's largest model for simple tasks that a smaller one can handle. GPT-6 Luna costs $0.10 per 1M input tokens against $10 for GPT-6 Astra, a hundredth of the price. OpenAI shuts down its older GPT-4 and GPT-3.5 Turbo models on 23 October 2026.

5. Truncate Strategically

When context is too long, remove middle portions rather than recent context—models often perform better with the beginning and end intact.

Counting Tokens Before API Calls

Always count tokens before making API calls to:

  • Prevent request failures from exceeding limits

  • Estimate costs accurately

  • Optimize prompt engineering

Our Token Counter estimates the tokens in your text and what they would cost with current models from OpenAI, Anthropic, Google, Mistral and DeepSeek. It gives an estimate, not the exact count each model's own tokenizer would give.

Common Tokenization Pitfalls

Whitespace Matters

"Hello World" (with space) and "HelloWorld" (no space) produce different token counts.

Code is Token-Heavy

A 100-line JavaScript function might use 500+ tokens due to syntax characters, variable names, and structure.

Non-English Text

Languages with non-Latin scripts (Chinese, Arabic, Japanese) typically use more tokens per word.

Practical Example

Let's tokenize a simple prompt:

Text: "Explain quantum computing in simple terms."

That is 6 words and 42 characters, so the rules of thumb above give about 8 to 11 tokens (6 words ÷ 0.75 = 8; 42 characters ÷ 4 = 10.5). Our Token Counter estimates 10. Each model's own tokenizer gives a slightly different exact count.

If the model responds with 500 words (about 667 tokens: 500 ÷ 0.75), your total usage is about 677 tokens.

Try Our Free Token Counter

Stop guessing about token counts and costs. Our Token Counter tool analyses your text instantly, showing estimated token counts for multiple models and estimated API costs.

Whether you're building a chatbot, processing documents, or fine-tuning prompts, knowing your token usage is essential for efficient AI development.

Sources

We checked these models, limits and prices on the providers' own pages on 27 September 2026. They change often, so check again before you rely on them:

Try the Token Counter

Put this knowledge into practice with our free tool.

Open Tool

Tags

AItokensGPT-4ClaudeGeminiAPILLM

Related Articles