Model rankings

What models real apps run in production. Measured from every metered run across hosted SkillSafe apps over the last 30 days — 394 runs — not benchmarks. Updated several times a day.

# Model Provider Runs (30d) Apps Success
1 Gemma 4 26B (Workers AI) @cf/google/gemma-4-26b-a4b-it Workers AI 197 2 100.0%
2 Qwen 3.8 27B (Workers AI) @cf/qwen/qwen3.8-27b Workers AI 51 1 100.0%
3 GPT-6 Sol gpt-6-sol OpenAI 44 1 100.0%
4 FLUX.2 Klein 4B (Workers AI) @cf/black-forest-labs/flux-2-klein-4b Workers AI 34 2 79.4%
5 Deepgram Aura 2 English (Workers AI) @cf/deepgram/aura-2-en Workers AI 26 1 100.0%
6 workflow workflow Anthropic 12 5 91.7%
7 GPT Image 2.5 Flare gpt-image-2.5-flare OpenAI 11 1 54.5%
8 GPT-5.6 Terra gpt-5.6-terra OpenAI 8 4 100.0%
9 MeloTTS (Workers AI) @cf/myshell-ai/melotts Workers AI 6 1 100.0%
10 FLUX.1 Schnell (Workers AI) @cf/black-forest-labs/flux-1-schnell Workers AI 2 1 100.0%
11 Kimi K2.6 (Workers AI) @cf/moonshotai/kimi-k2.6 Workers AI 1 1 100.0%
12 FLUX.2 Klein 9B (Workers AI) @cf/black-forest-labs/flux-2-klein-9b Workers AI 1 1 100.0%
13 Llama 3.3 70B fp8-fast (Workers AI) @cf/meta/llama-3.3-70b-instruct-fp8-fast Workers AI 1 1 100.0%

Aggregates only — no per-app or per-user figures. Data as of 2026-10-08.

How this is measured

  • Production runs, not benchmarks. Every row is built from metered runs that real apps executed and real users paid for over a trailing 30-day window.
  • Runs count toward the model that actually executed them. An app's configured model gets the credit unless the user overrode it — model aliases resolve to the concrete model that ran.
  • Success rate is the share of terminal runs that completed rather than failed (timeouts and provider errors count as failures; in-flight runs are excluded).