AI Key Supported Models

Supported Models for AI Keys

AI Keys is OpenClacky's deeply optimized model service, exclusively available for OpenClacky client use. We provide a proxy layer with smart routing and automatic caching — no third-party service configuration required.

Overview

Series Models Highlights
Turbo V4.1 Flash, V4 Pro High cost-efficiency, Auto context caching
GLM / Kimi / MiniMax GLM 5.3, Kimi K3, MiniMax M2.7 Reasoning models, Auto prefix caching (GLM 5.3 and MiniMax M2.7 text-only; Kimi K3 accepts images)

All prices include a 5% service fee — transparent billing, no hidden costs.


Turbo Series

Turbo models offer fast responses at extremely low cost, perfect for high-frequency calls and budget-sensitive workloads.

Flash-series price revision (in sync with DeepSeek's official 2026-09-10 adjustment)

Starting September 10, 2026 (12:00 Beijing time), DeepSeek revised the flash-series pricing: off-peak rates are $0.15/M input (cache miss), $0.003/M input (cache hit), and $0.60/M output; peak hours are 2x the off-peak price.

dsk-deepseek-v4-pro is unaffected by this revision — it stays available on its own model with unchanged billing (see the table below).

AI Key requests are billed by the actual time slot. Combined with our 99% cache hit rate, real-world spend on long conversations and Agent workflows drops even further.

Models & Pricing

Pricing (from Sep 10)

Prices are in USD per 1,000 tokens. Peak: 09:00-12:00 & 14:00-18:00 Beijing time; all other hours — and Beijing weekends — are off-peak.

Model Alias Tier Input Cache Read Output
dsk-deepseek-v4-pro Peak $0.001386 $0.0000462 $0.004158
dsk-deepseek-v4-pro Off-peak $0.000693 $0.0000231 $0.002079
dsk-deepseek-flash Peak $0.000315 $0.0000063 $0.00126
dsk-deepseek-flash Off-peak $0.0001575 $0.00000315 $0.00063

dsk-deepseek-v4-pro keeps serving DeepSeek V4 Pro upstream. The alias and your existing keys are unchanged, and its billing is not affected by the flash-series revision.

dsk-deepseek-v4-flash and dsk-deepseek-v4-flash-vision-exp have been retired — use dsk-deepseek-flash instead.

Choosing a Model

Use Case Recommended Model
Complex tasks, deep reasoning dsk-deepseek-v4-pro
Fast, cost-effective responses dsk-deepseek-flash

GLM / Kimi / MiniMax Series

Zhipu GLM 5.3, Moonshot Kimi K3 and MiniMax M2.7. All three are reasoning models — they think before answering — well suited to Chinese-language reasoning, long context, and long-horizon coding.

🎉 Launch Promo — 5% off: oc-glm-5.3, oc-kimi-k3 and oc-minimax-m2.7 are now live. All three are 5% off during the launch period.

The tables below show standard prices (5% service fee included). Actual billing is table price × 0.95. End-of-promo will be announced separately.

Models and Pricing

Prices are in USD per 1,000 tokens (including 5% service fee, promo discount not included).

Model Alias Input Output Cache Read Cache Write
oc-glm-5.3 $0.0011991 $0.0042 $0.0003 $0.0011991
oc-kimi-k3 $0.00315 $0.01575 $0.000315 $0.00315
oc-minimax-m2.7 $0.000315 $0.00126 $0.000063 $0.000315

All three models emit reasoning output: they generate reasoning tokens before the answer, billed as output. If max_tokens is too low the response can be truncated mid-thought and come back with an empty body — allow generous headroom.

All three support automatic prefix caching, with no explicit cache markers needed (unlike Claude): a repeated prefix hits from the second request onward.

Cache writes are not billed separately — the table lists them at the input rate because the upstream produces zero cache-write tokens.

Image input: oc-kimi-k3 accepts images; oc-glm-5.3 and oc-minimax-m2.7 are text-only — the former's upstream rejects any content block other than text, so sending an image directly fails, while the latter's upstream silently drops it. Clients handle both the same way: the image is downgraded to a local file reference and routed through the OCR sidecar, so the conversation keeps working — the model just sees a text description.

Protocol

These three models are served over the OpenAI-compatible spec (POST /chat/completions), unlike the Claude family which goes through Bedrock Converse (POST /model/{model}/converse). As a result:

  • Any client version that already supports custom OpenAI-compatible models can use them as-is — no client upgrade required;
  • If you configure them as a custom model, point the base URL at the gateway and use the OpenAI spec paths;
  • Requesting any of them over the Converse spec returns a clear error telling you to use /chat/completions instead.

Choosing a Model

Use Case Recommended Model
Chinese reasoning, cost-sensitive long context oc-glm-5.3
Long-horizon coding, end-to-end knowledge work oc-kimi-k3
Agent workflows, engineering delivery and office documents oc-minimax-m2.7

Why Choose AI Keys

Blazing Fast, Direct Connection

Direct connection to official APIs for fast, reliable responses. No multi-layer proxy forwarding — latency minimized to the absolute lowest.

Official Pricing, Transparent Billing

Same pricing as official APIs, no hidden fees — pay only for what you use. All prices include a 5% service fee, with clear and transparent billing.

Industry-Leading Caching, Massive Savings

Our proxy intelligently caches repeated prompt segments, achieving up to 99% cache hit rates and reducing overall costs by up to 50% compared to similar services.

When your request hits the cache:

  • Cache Read — pay only ~10% of the input price
  • Cache Write — tokens written to cache are billed at ~125% of input price, retained for 5 minutes

Turbo series models feature automatic context caching — cached tokens are automatically billed at a lower rate with no extra steps.


Usage

Create a Key

Generate an API key from your AI Keys dashboard. Keys use the format clacky-xxxx... and support quota limits, expiration dates, and usage tracking.

Configure Your Client

In the OpenClacky client, simply select OpenClacky as the Provider and enter your AI Key. That's it — no need to manually configure Base URL, API Type, or other parameters.

# Select OpenClacky in your client
# No additional configuration needed
provider: openclacky
api_key: clacky-your-key-here

Rate Limits & Quotas

  • Each API key can have optional usage quotas (daily, weekly, or monthly) set from your dashboard
  • The proxy enforces authentication and quota checks on every request
  • For high-volume needs, contact us about custom limits

Need help choosing a model? Visit the FAQ or contact support.