LiveThe SaaS Conversion Funnel: From Visitor to Paying Customer
IndieFounder
LatestAIAgents LearningRadar
Explore
Discover
FoundersStoriesTrendingActivityProductsCommunity
Build
Build ExperimentsRoadmapsGuidesCompareAlternativesBusiness ModelsHow It WorksCalculatorsGlossaryTeardownsStartup CostsIndustry GuidesStartup MetricsChecklistsAdvanced Calculators
Topics
StartupsAISaaSTechnologyProductGrowthMarketingMoney
Browse all topics
Sign in
IndieFounder

Practical intelligence for independent founders building products, companies, and useful things.

The founder brief

Ideas worth building. Delivered weekly.

Join the newsletter

IndieFounder

Read, learn, discover, and build with a community of independent founders.

Independent by design

Explore

01
  • Latest
  • Learning
  • Guides
  • Products
  • Founders
  • Radar
  • Community
  • Topics

Publication

02
  • About
  • Editorial policy
  • Newsletter
  • Contact
  • Corrections

Legal

03
  • Privacy
  • Cookies
  • Disclaimer
  • Sitemap
  • RSS feed

漏 2026 IndieFounder

RSSGet the brief
AI

Mellum2.1 is live. Use it as the sub-agent, not the owner of the ticket.

JetBrains shipped Mellum2.1 on 8 October: a 12B MoE with 2.5B active parameters, Apache 2.0, and a SWE-bench Verified jump from 2 to 47. Qwen3.5-9B still leads the hard agent benches.

Kirtesh AdmuteKirtesh Admute路9 Oct 2026, 11:46 am IST路7 min read路1,384 words
Mellum2.1 is live. Use it as the sub-agent, not the owner of the ticket.

Mellum2.1 keeps the June architecture and spends the release on reinforcement learning in real repos. It is a self-hosted coding worker with a 131k context, not a hosted API and not the top of Terminal-Bench.

JetBrains put Mellum2.1 on Hugging Face on 8 October 2026. Same 12B mixture-of-experts shell as Mellum2 from June. Same 2.5B active parameters. Apache 2.0. The change is post-training.

I read the launch as a sub-agent bill, not a replacement for the model that owns the hard ticket. There is no hosted API in the post. You pay for the GPU you already have, or you do not run it.

What actually shipped

The checkpoint is Mellum2.1-12B-A2.5B-Thinking, in the Mellum2.1 collection. The blog post is by Bulat Salimzianov, dated 8 October.

Architecture did not move. 28 layers. 64 experts. A router turns on 8 of them per token. Grouped-query attention with 32 query heads and 4 key-value heads. Three of every four layers use a 1,024-token sliding window. Context is 131,072 tokens. Vocabulary is 98,304. Weights ship in bfloat16. Text only. No image input.

JetBrains says the model can explore a repository, edit files, and check its own changes. That is the claim missing from Mellum2. The June model was built for routing, short answers, and fast sub-agents. It was not good inside a repo.

GGUF builds for llama.cpp, Ollama, and LM Studio are listed as coming soon. The multi-token prediction head for speculative decoding in vLLM is in the same bucket. Do not plan a laptop rollout on files the vendor has not finished posting. The bf16 weights are the thing that is live.

Recommended sampling on the Thinking checkpoint is temperature 0.6, top_p 0.95, top_k 20. Serving shape from the prior Mellum2 recipe still applies: vLLM with a Qwen3-style reasoning parser, and Hermes-style tool parsing if you want tool calls. Confirm the flags on the 2.1 card before you copy a June command into production.

How they trained the jump

Almost all of the 2.1 work is reinforcement learning. JetBrains moved RL from a short final stage to the main part of training. They added tasks in math, competitive programming, science, tool use, and software engineering. Open RL sets were filtered for broken tests, unverifiable answers, and tasks that were too easy or impossible.

For the software side they ran the model in real repositories with a shell and file tools. Reward is tests passing. They say they stood up thousands of environments and launched millions of sandboxes. A sandbox reward teaches the model to satisfy the tests it can see. It does not teach it your migration policy.

The scores, and who owns them

JetBrains ran Mellum2.1, Mellum2, Qwen3.5-9B, and Gemma 4 E4B through one pipeline in thinking mode. These are JetBrains numbers. Qwen's own card does not match the Qwen row in that table. Treat it as a same-harness comparison, not a public leaderboard.

On that shared pipeline, as reported from the model card:

  • LiveCodeBench v6: Mellum2.1 at 82.0, Mellum2 at 69.4, Qwen3.5-9B at 75.4, Gemma 4 E4B at 69.4.
  • HumanEval+: 91.5 for Mellum2.1. MBPP+: 79.4.
  • BFCL v4: 62.3, ahead of Qwen3.5-9B at 58.5.
  • SWE-bench Verified: 2.0 on Mellum2, 47.0 on Mellum2.1. Qwen3.5-9B still at 50.0.
  • SWE-bench Pro: 0.0 to 28.0. Qwen3.5-9B at 38.0.
  • Terminal-Bench 2.1: 0.6 to 17.4. Qwen3.5-9B at 21.7.
  • GPQA Diamond: 64.6 versus 77.8 for Qwen3.5-9B. AIME 25/26: 83.3 versus 86.7.
  • HarmBench: 21.5 down to 8.5, lower is better.

Agentic runs used the open-source Pi harness at v0.73.1 with a 114k-token context. That is under the 131,072 limit. If your agent dumps the whole repo into context, you are outside the published setup.

The Verified jump, 2 to 47, is the useful number. It still loses the hard agent benches to a dense 9B model with more active parameters. 2.5B active is the point. You are buying throughput, not the top of SWE-bench Pro.

Speed, and what it costs

Post-training did not touch the architecture, so raw speed matches Mellum2. JetBrains says multi-token prediction makes a single request about 1.6 times faster. Under heavy load on one NVIDIA H200, they say Mellum2.1 serves almost twice as many tokens as Qwen3.5-9B.

There is no list price. Apache 2.0 means you can run it in a product without a seat fee. The bill is the card. A prior vLLM recipe for Mellum2 Thinking put bf16 around 29 GB and called a single H200, H100, or A100 enough. I would budget a 40 GB card for the full weights plus KV cache on a long agent trace. A 24 GB card is a maybe, not a plan, until the GGUF builds land and you have measured them.

Compare that with a hosted small model. Claude Haiku 5.5, shipped 7 October, is $0.10 per million input tokens and $0.50 per million output under 100k, then five times that over the line. Mellum2.1 is the other side of that choice. Haiku is a metered API. Mellum is a local worker with a 131k window and no per-token invoice. If the agent makes thousands of short tool calls a day on code that cannot leave the building, price the local model. If the job is a rare hard debug, still call the bigger one.

Where I would put it

Three jobs fit the card JetBrains wrote.

First, the inner loop of a coding agent. Identify why a test failed, draft a patch, run the check. Keep the planner on a stronger model. Pass a narrow file set and the failing test, not the monorepo.

Second, tool-call fan-out. BFCL v4 at 62.3 is the best of the four models in JetBrains' table. A router that only picks a function and fills arguments does not need GPQA Diamond at 77.

Third, the private box. If the repo cannot go to a vendor API, this is one of the few coding-agent checkpoints with a real SWE-bench Verified number and a license a lawyer can read in an afternoon. Apache 2.0 is not a compliance program. It does remove the objection that you cannot ship the weights.

I would not put it on knowledge questions, long research, or anything that needs images. Text only. Terminal-Bench 2.1 at 17.4 means most hard terminal tasks still fail. A 17 percent pass rate is a demo, not an on-call agent.

Limits I would write into the runbook

Context is 131,072, but the agent eval used 114k. Set a hard cap under that if you want behavior close to the published score. Sliding window on most layers means early tokens are not fully attended. Do not assume the system prompt at token 200 is as visible at token 120,000 as it would be in a full-attention model.

Thinking mode spends tokens before the answer. MTP is supposed to claw some of that back, and the vLLM head is not in the 8 October post as a finished artifact. Measure tokens per second on your card with thinking on and off before you promise a latency number.

The model is rewarded for tests that pass. Give it a weak test and it will look successful. Gate the patch with your own suite, a type check, and a diff size limit. A model that can edit files can edit the wrong file.

JetBrains is also shipping Air, an agent system inside the IDEs, as an early access. Mellum2.1 is not that product. Do not assume Air is running this checkpoint unless the IDE build says so.

What I would do this week

Pull the Thinking checkpoint. Run one failing unit test from a repo you know, with the file tree trimmed. Compare against Qwen3.5-9B on the same harness, same temperature band, same timeout. If Mellum wins on wall time and ties on the patch, keep it as the sub-agent. If it loses the patch and only wins tokens per second, it is a router, not a coder.

Skip the GGUF path until JetBrains marks those builds done. Skip any pitch that calls 47 on SWE-bench Verified a sweep. Qwen3.5-9B is still at 50 on the same sheet, and both numbers belong to the lab that trained one of the models.

The useful release is narrower than the headline. A fast Apache 2.0 coding worker, 2.5B active, 131k context, finally able to touch a repo and check a test. That is a part you can put on your own metal. It is not the model you hand the whole ticket.

Community

What do you think?

0 comments

React to this article

Comments

0/2000

Trending now

What readers are opening

See all
The Solo Founder Playbook: Bootstrapping a Micro-SaaS to $50K MRR with AI Agents

Startups

The Solo Founder Playbook: Bootstrapping a Micro-SaaS to $50K MRR with AI Agents

Next.js 16 & Turbopack: Building and Shipping Micro-SaaS at Lightning Speed

AI & Code

Next.js 16 & Turbopack: Building and Shipping Micro-SaaS at Lightning Speed

Escaping Tutorial Purgatory: How Indie Hackers Ship From Idea to Production in 7 Days

Startups

Escaping Tutorial Purgatory: How Indie Hackers Ship From Idea to Production in 7 Days

JetBrainsMellumopen weightscoding agentsApache 2.0

Written by

Kirtesh Admute

Kirtesh Admute

Founder

Kirtesh Admute is the founder of IndieFounder, a platform for founders, builders, and people curious about technology. He writes about AI, startups, software, product building, and the lessons that come from building in public.

See an issue with this story?

Continue reading

More from IndieFounder

Article cover

AI

Mistral Large 4 is live in preview. Weights wait until month end.

2 days ago 路 7 min read

Article cover

AI

AI Coding Agent Test Gates: How to Stop Broken Code From Shipping

4 days ago 路 6 min read

Article cover

AI

AI Coding Agent Secret Protection: How to Stop Credential Leaks

4 days ago 路 6 min read

Next storyMistral Large 4 is live in preview. Weights wait until month end.ArchiveBrowse all articles

Newsletter

Get the next brief

Useful founder stories and product lessons, without the noise.

No spam. Just the useful stuff. Unsubscribe whenever you want.

Learn more