Skip to content

US Region: DeepSeek and GLM

Call the US Region DeepSeek and GLM models through OpenAI Chat Completions, OpenAI Responses, or Anthropic Messages.

ItemValue
Base URLhttps://api.hop-base.com/v1
OpenAI Chat CompletionsPOST /v1/chat/completions
OpenAI ResponsesPOST /v1/responses
Anthropic MessagesPOST /v1/messages

The US Region models are their own set of model IDs, all ending in -us. Every model accepts all three protocols with the same key and the same model ID, so pick the protocol your client speaks.

curl https://api.hop-base.com/v1/chat/completions \
  -H "Authorization: Bearer sk-your-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4.1-flash-us",
    "messages": [{ "role": "user", "content": "Summarize this quarter'\''s risks in five bullets." }]
  }'

Models

Model IDContext windowMax output tokensImage input
deepseek-v4.1-flash-us1M393,216Yes
deepseek-v4-pro-us1M393,216No
deepseek-v4-flash-us1M393,216No
glm-5.3-us1M131,072No
glm-5.2-us1M131,072No

The DeepSeek models are in the DeepSeek US Region group and the GLM models in the GLM US Region group. Add the group to your key under API Keys in the console; a key without it gets a 404. In the key dialog, both groups are listed under US Region.

deepseek-v4-pro-us is the 2026-08-13 release and deepseek-v4-flash-us the 2026-07-31 release. Models without image input do not see images: the request succeeds and the model answers that no image was attached.

These IDs are separate from the DeepSeek and GLM-5.3 models on other pages. Limits, pricing, and protocol support differ, so follow this page for the -us models.

Choose a protocol

ClientProtocolEndpoint
OpenAI SDK, Cherry Studio, Cursor, DifyOpenAI Chat Completions/v1/chat/completions
Codex CLIOpenAI Responses/v1/responses
Claude Code, Anthropic SDKAnthropic Messages/v1/messages

Authenticate with Authorization: Bearer sk-your-key. On /v1/messages the x-api-key header also works. The model field in every response echoes the ID you requested.

Claude Code

Point Claude Code at HopBase and pick a -us model. Claude Code does not know these IDs, so also tell it the real context window; otherwise it compacts the conversation at 200K tokens.

export ANTHROPIC_BASE_URL="https://api.hop-base.com"
export ANTHROPIC_AUTH_TOKEN="sk-your-key"
export ANTHROPIC_MODEL="deepseek-v4.1-flash-us"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="deepseek-v4-flash-us"
export CLAUDE_CODE_MAX_CONTEXT_TOKENS=1000000
claude

ANTHROPIC_DEFAULT_HAIKU_MODEL is the model Claude Code uses for background tasks such as titles. For one-click setup and the settings.json form, see Claude Code.

Codex CLI

Add a provider to ~/.codex/config.toml. Set wire_api to responses.

model = "deepseek-v4.1-flash-us"
model_provider = "hopbase"

[model_providers.hopbase]
name = "HopBase"
base_url = "https://api.hop-base.com/v1"
env_key = "HOPBASE_API_KEY"
wire_api = "responses"

Then export HOPBASE_API_KEY and run codex. More options are in Codex CLI.

Thinking and tool calling

All five models think by default. Chat Completions returns the thinking trace in reasoning_content, Anthropic Messages in thinking blocks, and Responses in reasoning items. Output token counts already include thinking.

  • glm-5.3-us always thinks. Turning thinking off returns a 400.
  • On Anthropic Messages, output_config.effort for the GLM models is mapped to low, high, or max: medium becomes high and xhigh becomes max. DeepSeek receives the value unchanged.
  • Function calling works end to end: single and parallel calls, and tool results fed back in later turns.
  • The GLM models ignore a forced tool choice (a named tool, required, or any) and may answer in text instead. The DeepSeek models follow it.
400 The value of the enable_thinking parameter is restricted to True.

Caching and billing

Prefix caching is automatic; there is nothing to switch on. Cache hits are reported in usage.prompt_tokens_details.cached_tokens (Chat Completions), usage.input_tokens_details.cached_tokens (Responses), and usage.cache_read_input_tokens (Anthropic Messages).

cache_control breakpoints in Anthropic requests are removed before the request goes out, so explicit caching is not available; automatic caching still applies. In tests on 2026-10-02, Claude Code turns after the first read 94%–99% of the prompt from cache.

Prices follow the official US-region price, which differs from the other DeepSeek and GLM pages. The rates for your key are in the signed-in model catalog.

  • The three DeepSeek models follow a daily peak/off-peak schedule in Beijing time (UTC+8). Peak runs 08:00–22:00 every day, weekends included.
  • Off-peak unit prices for input, cached input, and output are all halved.
  • A request is priced by the moment it finishes, and carries one unit price.
  • The GLM models have no peak/off-peak schedule.

When you reconcile Anthropic Messages, input_tokens excludes the cached portion and matches Input tokens in your usage record. cache_read_input_tokens matches Cached input tokens.

Limits and errors

CaseResult
GLM max_tokens above 131,072400
DeepSeek max_tokens above 393,216Not rejected
DeepSeek n above 1, or logprobs400
GLM temperature of 2.0 or higher400
Capacity momentarily exhausted429
400 Range of max_tokens should be [1, 131072]
400 Invalid n value (currently only n = 1 is supported)
400 Temperature should be in [0.0, 2.0)
429 Upstream accounts are currently rate limited, please retry later

On Anthropic Messages, the GLM max_tokens error reads Range of max_completion_tokens should be [1, 131072]. A 429 is not billed. Retry with exponential backoff: after a burst, a model can stay unavailable for up to a minute.

Next steps