Route LLM API calls through Mandarin Chinese to reduce token usage by 30–60%.
Chinese characters encode more information per token than English words. A BPE tokenizer (used by GPT, Claude, etc.) assigns roughly 1–3 tokens per Chinese character, but each character can represent a full word or concept. The same semantic content in Mandarin uses significantly fewer tokens.
mandiproxy is a local HTTP proxy that:
- Intercepts your LLM API call
- Translates message content to Mandarin before sending
- Receives the Mandarin response
- Translates back to English
- Returns it to your tool — transparently
pip install mandiproxy
# Start the proxy
mandiproxy start
# In another terminal, point your tool at the proxy:
ANTHROPIC_BASE_URL=http://localhost:8765 claude # Claude Code
OPENAI_BASE_URL=http://localhost:8765/v1 codex # OpenAI Codex CLI# Test a single prompt and see token savings
mandiproxy chat "Explain how neural networks work" --show-zh
# View session metrics
mandiproxy metrics| Variable | Default | Description |
|---|---|---|
MANDIPROXY_UPSTREAM_OPENAI |
https://api.openai.com/v1 |
Upstream OpenAI-compatible URL |
MANDIPROXY_UPSTREAM_ANTHROPIC |
https://api.anthropic.com/v1 |
Upstream Anthropic URL |
MANDIPROXY_TRANSLATION_MODEL |
gpt-4o-mini / claude-haiku-* |
Model used for translation |
Works with any tool that supports a custom base URL:
# Claude Code
export ANTHROPIC_BASE_URL=http://localhost:8765
claude
# OpenAI Codex
export OPENAI_BASE_URL=http://localhost:8765/v1
codexEvery response includes a mandiproxy_metrics field:
{
"mandiproxy_metrics": {
"input_savings_pct": 42.3,
"output_savings_pct": 38.1,
"total_savings_pct": 40.5,
"session_total_savings_pct": 39.2
}
}