MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis

(artificialanalysis.ai)

46 points | by theanonymousone 4 hours ago

5 comments

  • Gareth321 26 minutes ago
    OpenAI usage limits have been severely cut, and intelligence appears to be markedly declining, so I'm going to start trying these Chinese models seriously now. I don't mind if it takes longer. I just need the intelligence to predictably work the same way from day to day.
    • unsupp0rted 5 minutes ago
      Same- I pay $200/mo for Codex but whereas I used to get a week's work out of a weekly limit, now I get roughly 1~2 days.

      I've stopped using Astra entirely and remain on Sol orchestrating Luna Xhigh, but it's still not nearly a week's usage for a week's allotment.

      And even then, whenever a new model is about to come out, it feels like the model I'm using is being dumbed down substantially.

      I have no evidence for this and can have no evidence for this, but I can vote with my wallet regardless.

  • egeres 1 hour ago
    It feels suspicious that MiMo-V2.6 Pro gets 46 in de index while DeepSeek-V4.1 (https://artificialanalysis.ai/models/deepseek-v4-1-flash) gets 39. According to the appendix at the bottom of https://mimo.xiaomi.com/mimo-v2-6 the deepseek model sometimes surpasses mimo and it's not so far behind in capabilities. A week ago opus 5 appeared 1 points ahead of fable 5 despite fable being a much smarter model (now they corrected it)
  • dom96 1 hour ago
    It is an impressive model. Agreed on most that is written on this page, with the exception of it being fast. I ran it on my own LLM benchmark suite[1] and it is faster than DeepSeek but still much slower than leading models. But it's pricing is where it really shines.

    KillSwitch-Bench 1.0

      Claude Opus 5           66.9
      GPT-6 Astra             57.9
      Claude Fable 5.1        46.7
      MiMo-V2.6-Pro           38.8
      Muse Spark 1.3          36.5
    
    1 - https://bench.killswitch-lang.org/
  • kosolam 54 minutes ago
    Why sol is not in the comparison?
  • jampekka 2 hours ago
    "When evaluating the Intelligence Index, it generated 140M tokens, which is somewhat verbose in comparison to the median of 140M."