Independent testing shows the true capabilities of this model

#9
by Gogeta70 - opened

The Luke's Dev Lab channel on Youtube did an independent test of this model (I am not affiliated with them in any way):
https://www.youtube.com/watch?v=I_Ay10006QY

The testing shown in the video very clearly contradicts the claims that this model is "frontier class".
It struggles with recalling in-context information at context sizes larger than 131k.
It also struggles with many coding/development tasks, sometimes completely failing at them.

Compared to Qwen 3.8 27B, this model is nowhere near frontier class.

Institute of Foundation Models org
โ€ข
edited 4 days ago

Hi, @Gogeta70 ,

Thanks for sharing this. I skim through the testing quickly. This may have pointed out something that the community care and what we need to improve. Some improvements are coming too!

I would argue that this does not contradict the claim though. This is a 4B activated model, which if far fewer than Qwen-27B. On benchmarks that we test, it is also clear that this model beats ones with similar sizes. We will be quite surprised if we can be comparable to a dense 27B in terms of FLOPS, but we are indeed pushing towards it.

For example, among the small models on AA, MoVA is behind Qwen27B, better than G9V3-39BA5B. This is consistent with the finding in the video, but still showing this model as leading at its compute range.

Artificial Analysis Intelligence Index (1 Oct '26) - Small Models

(OK I didn't realize this AA figure is so long)

Well, let me just say that I very much appreciate free, open-weight models and when something is free, it leaves little room for complaint. So this isn't a complaint, but I think that claiming that this model is frontier class is a bit optimistic.

You are right that comparing an MoE model to a dense model isn't exactly an apples-to-apples comparison. However, as an end-user, when you read a claim that a model is frontier class, the assumption is about the model's capabilities, regardless of it's architecture. Right now, in terms of small-weight models that can be run on consumer hardware, Qwen 3.8 27B is the one to meet/beat when it comes to claims of frontier capability.

An MoE model can run inference much faster than a dense model, but that speed is only useful when it results in useful output. In the test video I linked, K2 Horizon output quite a lot of tokens on most of the agentic tasks it was given, which largely negates that speed advantage over dense models.

So on a technical level, K2 Horizon might be better than most or all of the other MoE models in the same size class, and that's nothing to sneeze at for sure. But from an end user's perspective, I can't honestly say that it rises to the level of frontier capability. It's definitely not a bad model, especially for one that is effectively a 4B parameter model.

All that said, thank you for all the effort you've put into creating this model and for sharing it freely.

Institute of Foundation Models org
โ€ข
edited 4 days ago

We don't take these as complaints, these are valuable feedbacks. Feedbacks from real users matter to us, and this is almost the only source we can learn some truth. So if you don't mind, please keep these coming.

I can see now the definition of frontier is different, your point is valid: you want the best model runnable on the consumer hardware. When we are saying frontier, we are thinking about the model size vs. performance. Our definition also matters, since if we keep pushing these, we can eventually provide users with better experience and cost/benefit return. But I take your point. There are rough edges on our side since we release a whole family, if you notice we are fixing them graudally.

I sincerely thank you for your feedback, our whole lab has watched the video you posted. On our wish list, we do try to be the best on consumer hardware. Challenge acceptted.

P.S.

I would like to point out a few things about evaluation from a researcher point of view too:

  1. The speed is a more involved issue, really depends on the underlying setups and inference techniques. For example, our UNO method that comes with these models can speed the models up 2-3 times. Again, I understand that I am wearing a researcher hat now, which may not resonate on the user side.
  2. Performance on certain tasks may differ, this often depends on what tasks we feed to the models. We didn't train the model for a lot of Blender tasks. These are the lessons we learned from this release.
Institute of Foundation Models org

As researchers and people who support true open source, we may be more sensitive to comments such as "over claimed" or "not honest". Since we are trying to be as open as possible. Btw, we are hosting a localllama AMA to try to collect more feedbacks. Welcome if you have some questions.

MoE models depend much more on harness and the manner in which they are used compared to dense models. I found this model to be a bit better than Gemma-4-26B-A4B (more up-to-date information mostly).
If what you want is one fire and forget prompt, then MoE is probably not for you (not in this class anyway), and there's no free lunch.

We don't take these as complaints, these are valuable feedbacks. Feedbacks from real users matter to us, and this is almost the only source we can learn some truth. So if you don't mind, please keep these coming.

I can see now the definition of frontier is different, your point is valid: you want the best model runnable on the consumer hardware. When we are saying frontier, we are thinking about the model size vs. performance. Our definition also matters, since if we keep pushing these, we can eventually provide users with better experience and cost/benefit return. But I take your point. There are rough edges on our side since we release a whole family, if you notice we are fixing them graudally.

I sincerely thank you for your feedback, our whole lab has watched the video you posted. On our wish list, we do try to be the best on consumer hardware. Challenge acceptted.

P.S.

I would like to point out a few things about evaluation from a researcher point of view too:

  1. The speed is a more involved issue, really depends on the underlying setups and inference techniques. For example, our UNO method that comes with these models can speed the models up 2-3 times. Again, I understand that I am wearing a researcher hat now, which may not resonate on the user side.
  2. Performance on certain tasks may differ, this often depends on what tasks we feed to the models. We didn't train the model for a lot of Blender tasks. These are the lessons we learned from this release.

Thanks for clarifying. I think this is a common problem as I have seen a lot of models claim to be "frontier class," but not actually have frontier-grade capability. If the term "frontier class" means different things to researchers and end users, that explains the confusion.

I think this 36B-A4B MoVA (MoE-like) model is great for people with low to mid-grade consumer hardware because it's amenable to being offloaded due to it's low active parameter count. It's not a model I'm personally interested in running because my own hardware can handle larger models (48Gb VRAM, 128Gb RAM), so my interest is actually more focused toward your 32B dense model. Unfortunately, based on your own reported benchmarks, it doesn't yet compete with Qwen 3.8 27B in capability.

I think the work you're doing here is very valuable and would love to see another model that can seriously compete with Qwen 3.8 27B (and their 120B model too). Qwen is one of the only organization releasing frontier-grade models that are small enough to run on consumer hardware and I would really love to see more competition in this area. I wish you the best of luck in creating an amazing model and I'm looking forward to seeing your progress.

Institute of Foundation Models org

I can see your frustration, almost every model coming out claiming themselves as "frontier". We indeed reach the so called pareto frontier on a few capabilities that we care: basically the best cost perrformance tradeoff. So we actually use that word with confidence.

HSRuER_aoAAFsIq

While MoE indeed may need more effort to utilize, it is likely to be the future. In fact, among 32 and 36A4B, due to the compute differences, we priortize on 36A4B, we find that it is actually very close to the 32B (every other condition hold the same), so we believe we can provide a better offering.

Maybe the wording could have been a little bit different, currently it says:

Frontier-class results at 4B active parameters. On agentic and reasoning benchmarks it outscores open weight dense (approximately 30B model size) and MoE models up to 15ร— its size; and also performs competitively against closed frontier models (see Benchmark Results).

but an option would be:

Best in class results at 36B total parameters with 4B active.

I think the other claims cannot be sustantiated fully.

Institute of Foundation Models org

You can see how it is pretty hard to write it, since there is not a class for 36B/4B only. We draw some comparison with other models (some are moe with 15x size). Our technical report is coming out soon, it will have more precise languages and which models we are comparing to.

Sign up or log in to comment