JetBrains put Mellum2.1 on Hugging Face on 8 October 2026. Same 12B mixture-of-experts shell as Mellum2 from June. Same 2.5B active parameters. Apache 2.0. The change is post-training.
I read the launch as a sub-agent bill, not a replacement for the model that owns the hard ticket. There is no hosted API in the post. You pay for the GPU you already have, or you do not run it.
The checkpoint is Mellum2.1-12B-A2.5B-Thinking, in the Mellum2.1 collection. The blog post is by Bulat Salimzianov, dated 8 October.
Architecture did not move. 28 layers. 64 experts. A router turns on 8 of them per token. Grouped-query attention with 32 query heads and 4 key-value heads. Three of every four layers use a 1,024-token sliding window. Context is 131,072 tokens. Vocabulary is 98,304. Weights ship in bfloat16. Text only. No image input.
JetBrains says the model can explore a repository, edit files, and check its own changes. That is the claim missing from Mellum2. The June model was built for routing, short answers, and fast sub-agents. It was not good inside a repo.
GGUF builds for llama.cpp, Ollama, and LM Studio are listed as coming soon. The multi-token prediction head for speculative decoding in vLLM is in the same bucket. Do not plan a laptop rollout on files the vendor has not finished posting. The bf16 weights are the thing that is live.
Recommended sampling on the Thinking checkpoint is temperature 0.6, top_p 0.95, top_k 20. Serving shape from the prior Mellum2 recipe still applies: vLLM with a Qwen3-style reasoning parser, and Hermes-style tool parsing if you want tool calls. Confirm the flags on the 2.1 card before you copy a June command into production.
Almost all of the 2.1 work is reinforcement learning. JetBrains moved RL from a short final stage to the main part of training. They added tasks in math, competitive programming, science, tool use, and software engineering. Open RL sets were filtered for broken tests, unverifiable answers, and tasks that were too easy or impossible.
For the software side they ran the model in real repositories with a shell and file tools. Reward is tests passing. They say they stood up thousands of environments and launched millions of sandboxes. A sandbox reward teaches the model to satisfy the tests it can see. It does not teach it your migration policy.
JetBrains ran Mellum2.1, Mellum2, Qwen3.5-9B, and Gemma 4 E4B through one pipeline in thinking mode. These are JetBrains numbers. Qwen's own card does not match the Qwen row in that table. Treat it as a same-harness comparison, not a public leaderboard.
On that shared pipeline, as reported from the model card:
- LiveCodeBench v6: Mellum2.1 at 82.0, Mellum2 at 69.4, Qwen3.5-9B at 75.4, Gemma 4 E4B at 69.4.
- HumanEval+: 91.5 for Mellum2.1. MBPP+: 79.4.
- BFCL v4: 62.3, ahead of Qwen3.5-9B at 58.5.
- SWE-bench Verified: 2.0 on Mellum2, 47.0 on Mellum2.1. Qwen3.5-9B still at 50.0.
- SWE-bench Pro: 0.0 to 28.0. Qwen3.5-9B at 38.0.
- Terminal-Bench 2.1: 0.6 to 17.4. Qwen3.5-9B at 21.7.
- GPQA Diamond: 64.6 versus 77.8 for Qwen3.5-9B. AIME 25/26: 83.3 versus 86.7.
- HarmBench: 21.5 down to 8.5, lower is better.
Agentic runs used the open-source Pi harness at v0.73.1 with a 114k-token context. That is under the 131,072 limit. If your agent dumps the whole repo into context, you are outside the published setup.
The Verified jump, 2 to 47, is the useful number. It still loses the hard agent benches to a dense 9B model with more active parameters. 2.5B active is the point. You are buying throughput, not the top of SWE-bench Pro.
Post-training did not touch the architecture, so raw speed matches Mellum2. JetBrains says multi-token prediction makes a single request about 1.6 times faster. Under heavy load on one NVIDIA H200, they say Mellum2.1 serves almost twice as many tokens as Qwen3.5-9B.
There is no list price. Apache 2.0 means you can run it in a product without a seat fee. The bill is the card. A prior vLLM recipe for Mellum2 Thinking put bf16 around 29 GB and called a single H200, H100, or A100 enough. I would budget a 40 GB card for the full weights plus KV cache on a long agent trace. A 24 GB card is a maybe, not a plan, until the GGUF builds land and you have measured them.
Compare that with a hosted small model. Claude Haiku 5.5, shipped 7 October, is $0.10 per million input tokens and $0.50 per million output under 100k, then five times that over the line. Mellum2.1 is the other side of that choice. Haiku is a metered API. Mellum is a local worker with a 131k window and no per-token invoice. If the agent makes thousands of short tool calls a day on code that cannot leave the building, price the local model. If the job is a rare hard debug, still call the bigger one.
Three jobs fit the card JetBrains wrote.
First, the inner loop of a coding agent. Identify why a test failed, draft a patch, run the check. Keep the planner on a stronger model. Pass a narrow file set and the failing test, not the monorepo.
Second, tool-call fan-out. BFCL v4 at 62.3 is the best of the four models in JetBrains' table. A router that only picks a function and fills arguments does not need GPQA Diamond at 77.
Third, the private box. If the repo cannot go to a vendor API, this is one of the few coding-agent checkpoints with a real SWE-bench Verified number and a license a lawyer can read in an afternoon. Apache 2.0 is not a compliance program. It does remove the objection that you cannot ship the weights.
I would not put it on knowledge questions, long research, or anything that needs images. Text only. Terminal-Bench 2.1 at 17.4 means most hard terminal tasks still fail. A 17 percent pass rate is a demo, not an on-call agent.
Context is 131,072, but the agent eval used 114k. Set a hard cap under that if you want behavior close to the published score. Sliding window on most layers means early tokens are not fully attended. Do not assume the system prompt at token 200 is as visible at token 120,000 as it would be in a full-attention model.
Thinking mode spends tokens before the answer. MTP is supposed to claw some of that back, and the vLLM head is not in the 8 October post as a finished artifact. Measure tokens per second on your card with thinking on and off before you promise a latency number.
The model is rewarded for tests that pass. Give it a weak test and it will look successful. Gate the patch with your own suite, a type check, and a diff size limit. A model that can edit files can edit the wrong file.
JetBrains is also shipping Air, an agent system inside the IDEs, as an early access. Mellum2.1 is not that product. Do not assume Air is running this checkpoint unless the IDE build says so.
Pull the Thinking checkpoint. Run one failing unit test from a repo you know, with the file tree trimmed. Compare against Qwen3.5-9B on the same harness, same temperature band, same timeout. If Mellum wins on wall time and ties on the patch, keep it as the sub-agent. If it loses the patch and only wins tokens per second, it is a router, not a coder.
Skip the GGUF path until JetBrains marks those builds done. Skip any pitch that calls 47 on SWE-bench Verified a sweep. Qwen3.5-9B is still at 50 on the same sheet, and both numbers belong to the lab that trained one of the models.
The useful release is narrower than the headline. A fast Apache 2.0 coding worker, 2.5B active, 131k context, finally able to touch a repo and check a test. That is a part you can put on your own metal. It is not the model you hand the whole ticket.