LiveCloudflare ships Clef-omni for multimodal decisions at $0.15 per million input tokens
IndieFounder
LatestAIAgents LearningRadar
Explore
Discover
FoundersStoriesTrendingActivityProductsCommunity
Build
Build ExperimentsRoadmapsGuidesCompareAlternativesBusiness ModelsHow It WorksCalculatorsGlossaryTeardownsStartup CostsIndustry GuidesStartup MetricsChecklistsAdvanced Calculators
Topics
StartupsAISaaSTechnologyProductGrowthMarketingMoney
Browse all topics
Sign in
IndieFounder

Practical intelligence for independent founders building products, companies, and useful things.

The founder brief

Ideas worth building. Delivered weekly.

Join the newsletter

IndieFounder

Read, learn, discover, and build with a community of independent founders.

Independent by design

Explore

01
  • Latest
  • Learning
  • Guides
  • Products
  • Founders
  • Radar
  • Community
  • Topics

Publication

02
  • About
  • Editorial policy
  • Newsletter
  • Contact
  • Corrections

Legal

03
  • Privacy
  • Cookies
  • Disclaimer
  • Sitemap
  • RSS feed

漏 2026 IndieFounder

RSSGet the brief
AI

Cloudflare ships Clef-omni for multimodal decisions at $0.15 per million input tokens

Open-weight model takes audio, video, images and text in one call. Clef-flash drops to $0.038 with a shorter context.

Kirtesh AdmuteKirtesh Admute路11 Oct 2026, 11:44 am IST路5 min read路941 words
Cloudflare ships Clef-omni for multimodal decisions at $0.15 per million input tokens

On 9 October 2026 Cloudflare released Clef-omni, a decision model that scores fixed options across audio, video, images and text without generating free text. Hosted at $0.15 per million input tokens on Workers AI, with weights on Hugging Face.

Cloudflare added audio and video to its decision models on 9 October 2026. Clef-omni takes WAV or MP3 audio, MP4 or WebM video, images and text in a single API call and returns calibrated probabilities over a fixed schema. No free-form text comes back.

The model sits on a Qwen3-Omni-30B-A3B-Instruct mixture-of-experts backbone with about 3 billion active parameters. Cloudflare kept the comprehension path and dropped the text-to-speech output pieces. They froze the base weights, trained LoRA adapters, and post-trained with label-smoothed cross-entropy plus Brier score calibration. The result scores every allowed answer in one forward pass.

Hosted pricing is $0.15 per million input tokens on Workers AI. Output tokens are free because the model never generates them. Media converts to input tokens: roughly 780 tokens per minute of audio, and video at higher rates depending on resolution (up to around 15,400 tokens per minute at maximum, or about 8,600 frame tokens per minute at 480p plus the audio). A 21-second video clip with sound scores in about 1.5 seconds median according to Cloudflare. Text-only decisions land around 130 ms median; images around 150 ms.

Context on the hosted endpoint is 64K tokens. Weights are open on Hugging Face under Apache 2.0 as Cloudflare/clef-omni. You can self-host if you want the full training context or different serving.

This is the third model in the Clef family. Clef itself stays at $0.24 per million input tokens with a 64K hosted context. Clef-flash dropped from $0.09 to $0.038 per million input tokens. The price cut came with a shorter hosted context: 24K tokens instead of the previous 64K. Cloudflare said only 0.24 percent of requests exceeded 24K, so they traded the long tail for the lower rate. The Hugging Face weights for Clef-flash still support up to 256K if you run them yourself. Clef itself also got faster on the hosted side through serving changes, including a move to SGLang. Median latency on roughly 800-token inputs fell from 262 ms to 152 ms.

Decision models like these sit next to TypeSafe's Jev and Microsoft's Decision-1. They answer yes/no, multiple-choice, rating or rubric questions by scoring fixed options instead of writing prose. That removes one source of hallucination and keeps latency low because there is no token generation. Clef-omni is the first in the family that can look at a video of a machine running while listening to the audio and checking a photo of the label, all in one request.

Cloudflare published a small benchmark table. On CLINC150+OOS macro-F1, Clef-omni scored 97.7 against Clef's 97.43 and Clef-flash's 66.77. On BANKING77 it led at 94.8. On some agent and tool-use sets it trailed the text-only siblings. The numbers are Cloudflare's own. Independent replications are not yet public.

The API follows the same System One style as the earlier Clef models. You send a state description, optional media arrays, and a questions object that defines the allowed answers. The model returns scores for each option. It is Jev-API compatible, so swapping the model ID is the main change for teams already calling decision endpoints through AI Gateway.

For an indie team building an agent that has to classify support tickets, route invoices, or flag security events, the practical numbers matter more than the architecture. At $0.15 per million input tokens you can score a few thousand short decisions for a dollar. A 21-second video costs more because of the token conversion, but it still finishes in a second and a half instead of a multi-step pipeline of transcription plus vision plus a text model. If your volume is high and most calls are under 24K tokens, Clef-flash at $0.038 is the cheaper option; move the long ones to Clef or Clef-omni.

Self-hosting is straightforward if you already run open weights. The model card on Hugging Face includes load code. Cloudflare also pointed to updated SGLang launch commands for the earlier Clef models. No new weights were released for the speed improvements on base Clef; those were serving changes.

Limits are the usual ones for this category. You define the options; the model does not invent new ones. Calibration is claimed but not independently audited at scale yet. Multimodal latency grows with media length and resolution. The hosted Clef-flash context cut will surprise anyone who was planning on 64K for the cheap tier. And because these models do not emit text, any workflow that needs a transcript or a written summary still requires a separate model.

Cloudflare said the whole Clef line started from a Friday evening decision to enter the decision-model space, trained over a weekend, and shipped the first versions the following Thursday. Clef-omni followed a week later. That cadence is fast for a company that is not primarily an AI lab. Whether the quality holds under real traffic is what teams will measure next.

The models are live on Workers AI today. Docs and pricing are on the Cloudflare developer site. Weights are on Hugging Face. If you already route decisions through AI Gateway, changing the model string is the shortest path to a test.

For most agent builders the choice is simple: use Clef-flash for the hot path under 24K tokens, Clef-omni when the input is a video or a mixed media packet, and keep a text generation model for anything that has to produce words. The price and latency numbers make the switch low-risk to try on a single workflow before you move more traffic.

That is the concrete change from 9 October. One new model that reads what you see and hear, two price and speed adjustments on the existing ones, and open weights if you want to run them yourself.

Community

What do you think?

0 comments

React to this article

Comments

0/2000

Trending now

What readers are opening

See all
The Solo Founder Playbook: Bootstrapping a Micro-SaaS to $50K MRR with AI Agents

Startups

The Solo Founder Playbook: Bootstrapping a Micro-SaaS to $50K MRR with AI Agents

Next.js 16 & Turbopack: Building and Shipping Micro-SaaS at Lightning Speed

AI & Code

Next.js 16 & Turbopack: Building and Shipping Micro-SaaS at Lightning Speed

Escaping Tutorial Purgatory: How Indie Hackers Ship From Idea to Production in 7 Days

Startups

Escaping Tutorial Purgatory: How Indie Hackers Ship From Idea to Production in 7 Days

AICloudflaredecision modelsClef-omnimultimodal

Written by

Kirtesh Admute

Kirtesh Admute

Founder

Kirtesh Admute is the founder of IndieFounder, a platform for founders, builders, and people curious about technology. He writes about AI, startups, software, product building, and the lessons that come from building in public.

See an issue with this story?

Continue reading

More from IndieFounder

Article cover

AI

Microsoft-Decision-1 lands in Foundry at $0.042 per million input tokens

1 day ago 路 8 min read

Article cover

AI

Xiaomi Open-Sourced MiMo-V2.6: A Cloud-Cost Exit for Builders Tired of API Tax

2 weeks ago 路 8 min read

Article cover

AI

Best AI Tools for Indie Founders in 2026

1 week ago 路 6 min read

Next storyMicrosoft-Decision-1 lands in Foundry at $0.042 per million input tokensArchiveBrowse all articles

Newsletter

Get the next brief

Useful founder stories and product lessons, without the noise.

No spam. Just the useful stuff. Unsubscribe whenever you want.

Learn more