// server models

Server Models

The tiers a hosted app can run on SkillSafe's servers, billed per token in credits — $1 = 10,000 credits — at the provider's list price, passed through with no markup; SkillSafe charges separately for the computation. The tiers come first; below them, every model grouped by what it takes in and gives out. The models that run in the browser are free.

Turn your skill into an app

Latest model of each tier

Provider list price · USD per 1M tokens

GPT-6 Astra

gpt-astra gpt-6-astra
OpenAI
$10.00 in / 1M tokens · list
$50.00 out / 1M tokens · list

GPT-6.1 Sol

gpt-sol gpt-6.1-sol
OpenAI
$2.00 in / 1M tokens · list
$10.00 out / 1M tokens · list

GPT-5.6 Terra

gpt-terra gpt-5.6-terra
OpenAI
$2.00 in / 1M tokens · list
$12.00 out / 1M tokens · list

GPT-6 Luna

gpt-luna gpt-6-luna
OpenAI
$0.10 in / 1M tokens · list
$0.50 out / 1M tokens · list

Gemma 4 26B (Workers AI)

gemma-fast @cf/google/gemma-4-26b-a4b-it
Workers AI
$0.10 in / 1M tokens · list
$0.30 out / 1M tokens · list

GPT Image 2.5 Flare

gpt-image gpt-image-2.5-flare
OpenAI
~$0.060 list, per 1024×1024 image

FLUX.2 Klein 9B (Workers AI)

flux-klein @cf/black-forest-labs/flux-2-klein-9b
Workers AI
~$0.015 list, per 1024×1024 image

Deepgram Aura 2 English (Workers AI)

aura @cf/deepgram/aura-2-en
Workers AI
$0.030 list, per 1k characters

Whisper Large v3 Turbo (Workers AI)

whisper @cf/openai/whisper-large-v3-turbo
Workers AI
$0.000513 list, per audio minute

Model cost is pass-through: these are the providers' published list prices, unchanged — see OpenAI pricing and Workers AI pricing. What SkillSafe charges for is the computation and API handling on its own servers: a 10% platform fee on each run's base — the model's list cost, with no per-job overhead. A creator's markup is a separate cut of the same base — they keep 100% of it. Only the latest model of each tier is shown; configure the alias and your app follows upgrades. Tiers that cannot run on this deployment are not shown.

Every model, by what goes in and what comes out

52 models · pinned ids

Text Text 22 models

Chat completion: the app's prompt plus the run input go in, prose comes back.

Model Provider in / out per 1M tokens · list Caps
Granite 4.0 H Micro (Workers AI) @cf/ibm-granite/granite-4.0-h-micro Workers AI $0.017 / $0.11 100k in · 4k out · 120s
Llama 3.2 1B (Workers AI) @cf/meta/llama-3.2-1b-instruct Workers AI $0.027 / $0.20 55k in · 4k out · 120s
Llama 3.1 8B fp8-fast (Workers AI, legacy) @cf/meta/llama-3.1-8b-instruct-fp8-fast Workers AI $0.045 / $0.38 27k in · 4k out · 120s
Llama 3.2 3B (Workers AI) @cf/meta/llama-3.2-3b-instruct Workers AI $0.051 / $0.34 75k in · 4k out · 120s
Qwen3 30B-A3B (Workers AI, thinks) @cf/qwen/qwen3-30b-a3b-fp8 Workers AI $0.051 / $0.34 28k in · 4k out · 120s
GLM-4.7 Flash (Workers AI) @cf/zai-org/glm-4.7-flash Workers AI $0.061 / $0.40 100k in · 4k out · 120s
Mistral 7B v0.1 (Workers AI, legacy, 2.8k context) @cf/mistral/mistral-7b-instruct-v0.1 Workers AI $0.11 / $0.19 2k in · 1k out · 120s
Llama 3.1 8B fp8 (Workers AI) @cf/meta/llama-3.1-8b-instruct-fp8 Workers AI $0.15 / $0.29 27k in · 4k out · 120s
GPT-OSS 20B (Workers AI, reasoning) @cf/openai/gpt-oss-20b Workers AI $0.20 / $0.30 100k in · 4k out · 120s
Llama 3.1 8B (Workers AI, legacy, 8k context) @cf/meta/llama-3.1-8b-instruct Workers AI $0.28 / $0.83 3k in · 4k out · 120s
Llama 3.3 70B fp8-fast (Workers AI) @cf/meta/llama-3.3-70b-instruct-fp8-fast Workers AI $0.29 / $2.25 19k in · 4k out · 120s
Llama 3.1 70B fp8-fast (Workers AI, legacy) @cf/meta/llama-3.1-70b-instruct-fp8-fast Workers AI $0.29 / $2.25 19k in · 4k out · 120s
GPT-OSS 120B (Workers AI, reasoning) @cf/openai/gpt-oss-120b Workers AI $0.35 / $0.75 100k in · 4k out · 120s
Gemma SEA-LION v4 27B (Workers AI) @cf/aisingapore/gemma-sea-lion-v4-27b-it Workers AI $0.35 / $0.56 100k in · 4k out · 120s
DeepSeek V4 Flash (Workers AI) @cf/deepseek-ai/deepseek-v4-flash-0731 Workers AI $0.44 / $1.32 100k in · 4k out · 120s
DeepSeek R1 Distill Qwen 32B (Workers AI, reasoning) @cf/deepseek-ai/deepseek-r1-distill-qwen-32b Workers AI $0.50 / $4.88 75k in · 4k out · 120s
Nemotron 3 Super 120B (Workers AI) @cf/nvidia/nemotron-3-120b-a12b Workers AI $0.50 / $1.50 100k in · 4k out · 120s
Qwen2.5 Coder 32B (Workers AI) @cf/qwen/qwen2.5-coder-32b-instruct Workers AI $0.66 / $1.00 28k in · 4k out · 120s
QwQ 32B (Workers AI, reasoning) @cf/qwen/qwq-32b Workers AI $0.66 / $1.00 19k in · 4k out · 120s
DeepSeek V4 Pro (Workers AI) @cf/deepseek-ai/deepseek-v4-pro-0813 Workers AI $1.32 / $3.96 100k in · 4k out · 120s
GLM-5.2 (Workers AI) @cf/zai-org/glm-5.2 Workers AI $1.40 / $4.40 100k in · 4k out · 120s
GLM-5.3 (Workers AI, thinks) @cf/zai-org/glm-5.3 Workers AI $1.40 / $4.40 100k in · 4k out · 120s

Text + image Text 17 models

The same, and the model can also see images attached to the run with $files. Only models that named the colour of a test image are listed here.

Model Provider in / out per 1M tokens · list Caps
GPT-6 Luna gpt-6-luna gpt-luna OpenAI $0.10 / $0.50 100k in · 4k out · 60s
Gemma 4 26B (Workers AI) @cf/google/gemma-4-26b-a4b-it gemma-fast default Workers AI $0.10 / $0.30 32k in · 4k out · 120s
GLM-5.3 Flash (Workers AI, thinks) @cf/zai-org/glm-5.3-flash Workers AI $0.15 / $0.50 100k in · 4k out · 120s
GPT-5.6 Luna gpt-5.6-luna OpenAI $0.20 / $1.20 100k in · 4k out · 60s
GPT-5 mini (retiring 2026-12-11) gpt-5-mini OpenAI $0.25 / $2.00 100k in · 4k out · 60s
Llama 4 Scout 17B (Workers AI) @cf/meta/llama-4-scout-17b-16e-instruct Workers AI $0.27 / $0.85 100k in · 4k out · 120s
Mistral Small 3.1 24B (Workers AI) @cf/mistralai/mistral-small-3.1-24b-instruct Workers AI $0.35 / $0.56 100k in · 4k out · 120s
Qwen 3.8 27B (Workers AI) @cf/qwen/qwen3.8-27b Workers AI $0.45 / $3.20 100k in · 4k out · 120s
Kimi K2.5 (Workers AI, legacy) @cf/moonshotai/kimi-k2.5 Workers AI $0.60 / $3.00 100k in · 4k out · 120s
Kimi K2.6 (Workers AI) @cf/moonshotai/kimi-k2.6 Workers AI $0.95 / $4.00 100k in · 4k out · 120s
Kimi K2.7 Code (Workers AI, thinks) @cf/moonshotai/kimi-k2.7-code Workers AI $0.95 / $4.00 100k in · 4k out · 120s
GPT-5.1 gpt-5.1 OpenAI $1.25 / $10.00 100k in · 8k out · 120s
GPT-6 Sol gpt-6-sol OpenAI $2.00 / $10.00 100k in · 8k out · 120s
GPT-6.1 Sol gpt-6.1-sol gpt-sol OpenAI $2.00 / $10.00 100k in · 8k out · 120s
GPT-5.6 Terra gpt-5.6-terra gpt-terra OpenAI $2.00 / $12.00 100k in · 8k out · 120s
GPT-5.6 Sol gpt-5.6-sol OpenAI $4.00 / $20.00 100k in · 8k out · 180s
GPT-6 Astra gpt-6-astra gpt-astra OpenAI $10.00 / $50.00 100k in · 8k out · 180s

Text Image 8 models

One 1024×1024 image per run from a text prompt, billed per image.

Model Provider per image · list Caps
FLUX.2 Klein 4B (Workers AI) @cf/black-forest-labs/flux-2-klein-4b Workers AI ~$0.0011 1 image · 60s
FLUX.1 Schnell (Workers AI) @cf/black-forest-labs/flux-1-schnell Workers AI ~$0.0019 1 image · 60s
FLUX.2 Klein 9B (Workers AI) @cf/black-forest-labs/flux-2-klein-9b flux-klein Workers AI ~$0.015 1 image · 60s
Leonardo Phoenix 1.0 (Workers AI) @cf/leonardo/phoenix-1.0 Workers AI ~$0.034 1 image · 60s
Leonardo Lucid Origin (Workers AI) @cf/leonardo/lucid-origin Workers AI ~$0.036 1 image · 60s
GPT Image 2 gpt-image-2 OpenAI ~$0.060 1 image · 180s
GPT Image 2.5 Flare gpt-image-2.5-flare gpt-image OpenAI ~$0.060 1 image · 180s
GPT Image 2.5 Sunburst gpt-image-2.5-sunburst OpenAI ~$0.060 1 image · 180s

Text Speech 4 models

Text-to-speech: the run's text field (up to 2,000 characters) comes back as one MP3, billed per character.

Model Provider per 1k characters · list Caps
MeloTTS (Workers AI) @cf/myshell-ai/melotts Workers AI $0.000205 1 audio file · 60s
Deepgram Aura 1 (Workers AI) @cf/deepgram/aura-1 Workers AI $0.015 2,000 chars · 120s
Deepgram Aura 2 English (Workers AI) @cf/deepgram/aura-2-en aura Workers AI $0.030 2,000 chars · 120s
Deepgram Aura 2 Spanish (Workers AI) @cf/deepgram/aura-2-es Workers AI $0.030 2,000 chars · 120s

Audio Text 1 model

Speech-to-text: one audio file attached with $files comes back as a transcript, billed per minute of audio.

Model Provider per audio minute · list Caps
Whisper Large v3 Turbo (Workers AI) @cf/openai/whisper-large-v3-turbo whisper Workers AI $0.000513 1 audio file · 180s

Pin a concrete id and you own the upgrades; a tier alias moves with the catalogue. "Text + image" is claimed only for models that named the colour of a test image through the same request shape a run uses — a text-only model handed an image fails the run and refunds it. Models this deployment cannot run are not shown; GET /v1/models still catalogues them with available: false and carries the same input_types / output_type fields.

How server pricing works

  • Pay for what a run actually uses. Each run is charged on the tokens the model actually consumed, rounded up to whole credits (minimum 1 credit ≈ $0.0001). There is no fixed per-job overhead, so a short turn on a cheap model bills a handful of credits.
  • The model is pass-through; the fee pays for the servers. A run's base is the provider's list cost — charged exactly as the provider publishes it, no markup and no overhead. SkillSafe adds a 10% platform fee on that base for the computation and API handling it does on its own servers; the creator takes their markup — up to 100% — of the very same base, shown on each app's detail page. At the default 10% the two are equal, and nothing is deducted from the creator's side: they keep 100% of their markup.
  • Prompt caching is passed through too. OpenAI caches any prompt of 1,024 tokens or more on its own and reports how many input tokens it served from the cache. Those are billed at OpenAI's cached-input price, 10% of the input price. On GPT-5.6 and GPT-6, the first run that stores a prompt is billed at the cache-write price (125%) for the stored part. An app whose system prompt stays the same across runs therefore pays much less for input after the first run. GET /v1/models lists both prices per model.
  • BYOK apps bill the 1-credit minimum ($0.0001) per run. When a publisher brings their own provider key, their key pays the provider directly and the platform charges only the minimum billable unit — no 10% platform fee and no creator markup.
  • Spend is capped per job. Every model carries hard caps on input tokens, output tokens, and wall-clock time, so a single run can never overrun its hold.
  • Image models bill per image. An image-generation run produces one 1024×1024 image. The hold reserves the worst-case cost of that image and the run settles down to the provider's actual (usage-reported where available) cost — same base, same two cuts as text runs.
  • No model configured? Apps without a model run on the Workers AI default — cheap, keyless, and never silently billed at premium rates.