Supported Models for AI Keys
AI Keys is OpenClacky's deeply optimized model service, exclusively available for OpenClacky client use. We provide a proxy layer with smart routing and automatic caching — no third-party service configuration required.
Overview
| Series | Models | Highlights |
|---|---|---|
| Turbo | V4.1 Flash, V4 Pro | High cost-efficiency, Auto context caching |
| GLM / Kimi / MiniMax | GLM 5.3, Kimi K3, MiniMax M2.7 | Reasoning models, Auto prefix caching (GLM 5.3 and MiniMax M2.7 text-only; Kimi K3 accepts images) |
All prices include a 5% service fee — transparent billing, no hidden costs.
Turbo Series
Turbo models offer fast responses at extremely low cost, perfect for high-frequency calls and budget-sensitive workloads.
Flash-series price revision (in sync with DeepSeek's official 2026-09-10 adjustment)
Starting September 10, 2026 (12:00 Beijing time), DeepSeek revised the flash-series pricing: off-peak rates are $0.15/M input (cache miss), $0.003/M input (cache hit), and $0.60/M output; peak hours are 2x the off-peak price.
dsk-deepseek-v4-prois unaffected by this revision — it stays available on its own model with unchanged billing (see the table below).AI Key requests are billed by the actual time slot. Combined with our 99% cache hit rate, real-world spend on long conversations and Agent workflows drops even further.
Models & Pricing
Pricing (from Sep 10)
Prices are in USD per 1,000 tokens. Peak: 09:00-12:00 & 14:00-18:00 Beijing time; all other hours — and Beijing weekends — are off-peak.
| Model Alias | Tier | Input | Cache Read | Output |
|---|---|---|---|---|
dsk-deepseek-v4-pro |
Peak | $0.001386 | $0.0000462 | $0.004158 |
dsk-deepseek-v4-pro |
Off-peak | $0.000693 | $0.0000231 | $0.002079 |
dsk-deepseek-flash |
Peak | $0.000315 | $0.0000063 | $0.00126 |
dsk-deepseek-flash |
Off-peak | $0.0001575 | $0.00000315 | $0.00063 |
dsk-deepseek-v4-prokeeps serving DeepSeek V4 Pro upstream. The alias and your existing keys are unchanged, and its billing is not affected by the flash-series revision.
dsk-deepseek-v4-flashanddsk-deepseek-v4-flash-vision-exphave been retired — usedsk-deepseek-flashinstead.
Choosing a Model
| Use Case | Recommended Model |
|---|---|
| Complex tasks, deep reasoning | dsk-deepseek-v4-pro |
| Fast, cost-effective responses | dsk-deepseek-flash |
GLM / Kimi / MiniMax Series
Zhipu GLM 5.3, Moonshot Kimi K3 and MiniMax M2.7. All three are reasoning models — they think before answering — well suited to Chinese-language reasoning, long context, and long-horizon coding.
🎉 Launch Promo — 5% off:
oc-glm-5.3,oc-kimi-k3andoc-minimax-m2.7are now live. All three are 5% off during the launch period.The tables below show standard prices (5% service fee included). Actual billing is table price × 0.95. End-of-promo will be announced separately.
Models and Pricing
Prices are in USD per 1,000 tokens (including 5% service fee, promo discount not included).
| Model Alias | Input | Output | Cache Read | Cache Write |
|---|---|---|---|---|
oc-glm-5.3 |
$0.0011991 | $0.0042 | $0.0003 | $0.0011991 |
oc-kimi-k3 |
$0.00315 | $0.01575 | $0.000315 | $0.00315 |
oc-minimax-m2.7 |
$0.000315 | $0.00126 | $0.000063 | $0.000315 |
All three models emit reasoning output: they generate reasoning tokens before the answer, billed as output. If
max_tokensis too low the response can be truncated mid-thought and come back with an empty body — allow generous headroom.All three support automatic prefix caching, with no explicit cache markers needed (unlike Claude): a repeated prefix hits from the second request onward.
Cache writes are not billed separately — the table lists them at the input rate because the upstream produces zero cache-write tokens.
Image input:
oc-kimi-k3accepts images;oc-glm-5.3andoc-minimax-m2.7are text-only — the former's upstream rejects any content block other than text, so sending an image directly fails, while the latter's upstream silently drops it. Clients handle both the same way: the image is downgraded to a local file reference and routed through the OCR sidecar, so the conversation keeps working — the model just sees a text description.
Protocol
These three models are served over the OpenAI-compatible spec (POST /chat/completions), unlike the Claude family which goes through Bedrock Converse (POST /model/{model}/converse). As a result:
- Any client version that already supports custom OpenAI-compatible models can use them as-is — no client upgrade required;
- If you configure them as a custom model, point the base URL at the gateway and use the OpenAI spec paths;
- Requesting any of them over the Converse spec returns a clear error telling you to use
/chat/completionsinstead.
Choosing a Model
| Use Case | Recommended Model |
|---|---|
| Chinese reasoning, cost-sensitive long context | oc-glm-5.3 |
| Long-horizon coding, end-to-end knowledge work | oc-kimi-k3 |
| Agent workflows, engineering delivery and office documents | oc-minimax-m2.7 |
Why Choose AI Keys
Blazing Fast, Direct Connection
Direct connection to official APIs for fast, reliable responses. No multi-layer proxy forwarding — latency minimized to the absolute lowest.
Official Pricing, Transparent Billing
Same pricing as official APIs, no hidden fees — pay only for what you use. All prices include a 5% service fee, with clear and transparent billing.
Industry-Leading Caching, Massive Savings
Our proxy intelligently caches repeated prompt segments, achieving up to 99% cache hit rates and reducing overall costs by up to 50% compared to similar services.
When your request hits the cache:
- Cache Read — pay only ~10% of the input price
- Cache Write — tokens written to cache are billed at ~125% of input price, retained for 5 minutes
Turbo series models feature automatic context caching — cached tokens are automatically billed at a lower rate with no extra steps.
Usage
Create a Key
Generate an API key from your AI Keys dashboard. Keys use the format clacky-xxxx... and support quota limits, expiration dates, and usage tracking.
Configure Your Client
In the OpenClacky client, simply select OpenClacky as the Provider and enter your AI Key. That's it — no need to manually configure Base URL, API Type, or other parameters.
# Select OpenClacky in your client
# No additional configuration needed
provider: openclacky
api_key: clacky-your-key-here
Rate Limits & Quotas
- Each API key can have optional usage quotas (daily, weekly, or monthly) set from your dashboard
- The proxy enforces authentication and quota checks on every request
- For high-volume needs, contact us about custom limits
Need help choosing a model? Visit the FAQ or contact support.