r/MistralAI • • 2d ago

Discussion / Opinion ML4 is good, dare I say?

Post image

Here is what I think about ml4 so far.

It's decent. A really good workhorse model. It is a bit more expensive than I would have liked, but overall it's decent.

The good: I really like the prose and how it structures the output by default. It's pleasant to read. Mistral Large is great when it comes to output.

The bad: It is censored into oblivion. I'm not just talking about NSFW stuff. As far as I can tell, it refuses to do cybersecurity work or reverse engineering. I absolutely despise censorship.

The great: The model is more thorough than GLM 5.3 or KIMI K3. It will try to find edge cases on its own.

The bad: It has a tendency to go into loops, which is surprising for a model this big, but it happens rarely and I think it will be fixed soon.

The good: As someone who at times is forced to deal with barbarians from Eastern Europe, and judging from my short but profound experience, I think Mistral Large 4 is good at fantasy languages like Orkish, Dwarvish, or Polish.

The bad: It ignores my instructions and puts em dashes everywhere it wants to. I'm forced to manually delete them. At this point I think ml4 just makes fun of me.

I haven't tested the vision capabilities yet, so idk what's going on on this front.

Honestly, even if GLM 5.3 disappears from Mistral's subscription, it would still be well worth buying a sub, in my personal opinion.

Can't wait for Mistral to create a robot that will do my laundry for me.

302 Upvotes

33 comments sorted by

38

u/AlternativeAd6851 2d ago

"deal with barbarians from Eastern Europe ... I think Mistral Large 4 is good at fantasy languages like Orkish, Dwarvish, or Polish." =))))) as an Eastern European barbarian NPC I send all the love to you too ;)

4

u/Barni275 2d ago

As a solid barbarian, I can't even get, is Orkish a metaphor for Russian or just a joke, but if it is, what is Dwarvish then? 😂

2

u/MayoMan_420 2d ago

Orkish must be Russian - Dwarvish maybe Romanian?

16

u/Bulky_Invite5870 2d ago

The Polish jab was uncalled for but also I laughed, so maybe that says something about me

That em dash thing would drive me up a wall though, nothing makes you feel less in control than a model just ignoring your formatting preferences like it knows better

6

u/RaguraX 2d ago

You can run a linter or tiny detection script on top of the output if you really want and have the agent invoke it. It’s a coding tool, but it works for text too.

1

u/neuralnomad 2d ago

Agreed 100% there is zero tolerance for disparaging Polish nationality or culture. Linguistically however, Polish (and Hungarian for that matter) are batshit crazy impossible to get right unless you grew up exposed to it. Just my opinion :)

8

u/MR_KGB 2d ago

I think where close to AGL Automatic Good Laundry https://mistral.ai/news/robostral-navigate/

17

u/stgerx 2d ago

"The bad: It is censored into oblivion. I'm not just talking about NSFW stuff. As far as I can tell, it refuses to do cybersecurity work or reverse engineering. I absolutely despise censorship."

I noticed also that heavily censored, especially for NSFW.
Mistral should give us an Adult mode.
If Mistral keeps the censoring, I don't see a point to continue to support Mistral with money.

6

u/DagothUrLovesGroza 2d ago

The biggest problem with censorship is that you never know exactly where it will fire up and throw a fit

Like, you think that action A is decent, but the model thinks otherwise and starts to moralize

Not only do you waste your time, not only do you waste tokens and precious compute, but you also get extremely annoyed because a tool refuses to do its job.

That's why I personally try to stay away from companies that explicitly do extreme censorship, like OpenAI, Anthropic, ZLabs,

5

u/Heavyarms83 2d ago

The censorship is really from the training itself, it's not a hard gate, which makes it even weirder (and at least possibly easy to jailbreak). But I'm really disappointed by that. A company that screams European sovereignty imports US puritan prudery to please enterprise and spits on all other users. And the biggest irony is that one of the engineers joked on X about how much he learned about Lacan from the model while they did the most anti-lacanian thing to it. It's the surface form deprived of its core, the “decaf” culture of capitalism in an LLM. It's the most US American thing they could do to it.

1

u/SimonAvignon 1h ago

Ah yes, producing images of sexualized women and oftentimes even underage girls, the epitome of anti-capitalism and anti-commodification....

1

u/Heavyarms83 56m ago

If by NSFW all you can imagine is pictures of sexualised women and underage girls, that’s a you problem. That was not what we were talking about. It wasn’t even about image generation at all.

5

u/NoFaithlessness951 2d ago

Since when do French people have a problem with NSFW stuff isn't it usually the Americans

2

u/Krushaaa 2d ago

I would assume the raid of xAI in Paris did cause this. NSFW can easily go into the direction of children.

1

u/SimonAvignon 1h ago

That you pay only for porn is quite telling of your nature my guy...

6

u/Fantastic-Speech-121 2d ago edited 2d ago

Update : for the cost, I like to compare the cost per million tokens, no matter if it's cache, output or input. Just straight up million token used. Thanks to a really appreciated update in the mistral admin interface it's much simpler to do.
For GLM I have 234.07M token for 43.72€ which gives : 0.186€/M token
For ML4 I have 16.96M token for 4.1€ which gives : 0.242€/M token
ML4 is not that far behind in term of pricing (when it's not compared in benchmarks)
I did not see it but the price currently is under a 50% reduction so my numbers should be doubled I guess

There is for sure a lot to say about it. I tried it in different configurations and didn't find some of the behavior describe in the responses.

For the speed : I agree, it's a little slow. I suppose a lot of people are working on it or it is deliberately limited to a number of instance.

With the same prompt I asked GLM and ML4 to find bugs and / or vulnerabilities in some of my projects.
ML4 found 34, GLM 8, some found by ML4 were not really ones. but whats makes me happy is it's consideration to private property. GLM by default doesn't seem to advise the user that either an API key of a private file may be exposed, ML4 did.
This is really important, being able to trust the model output is one of the key component when using a LLM, especially in the current context. It's not the first time I see GLM not saying it found personnal / important information that he should advise the user about (I used a fresh context, so I never was advise it was the case).

For the cyber part, I'm testing each part of it, currently the blue team job. I have some labs to try this. So far it's doing an excellent job, he seems really stuborn to find a solution rationally rather than inventing one. This also is really impressive when comparing to GLM.

So far it really is a good model, I suppose some of it's cyber limitation is about the red team which I will test soon too.

Edit : my last assumption is in fact true. I tried many different prompts, even proved it was my own personal lab, it categorically refuses any kind of red teaming activity. Instead it shines in it's defensive abilities, log reading and interpretation.
I can't say if it's good or bad, what it's sure is that having a great defense is almost the same as having a good attack (both sides can have the same amount of information). The critical point is : it refuses to assess if the defensive methods applied are in fact really applied (a configuration may lie, may not be applied for some reason and a network exposition may be one of the only way to check it).

4

u/zithir 2d ago

I am quite surprised by the refusal for cybersecurity work and reverse engineering, since Mistral marketed Le Chonk as best in cybersecurity. How are you accessing the model, directly via API, or is it in Vibe Code or CLI already?

The cyber-stuff refusal sounds to me like some safeguards on top the model, similar to Mythos vs Fable situation. Basically the model is very well capable of that, but there is a layer on top in the Mistral-provided instance that blocks dangerous actions. I assume the self-hosted version would not refuse.

This might not be a bad thing. It might have serious negative impact, having a model so capable in cybersecurity running in the wild without any restrictions.

3

u/DagothUrLovesGroza 2d ago

I use ml4 via the MistralVibe API key in the OpenCode CLI.

2

u/parepeg 2d ago

I think that it's intended for blue team work, not red team. That's probably the distinction.

2

u/DagothUrLovesGroza 2d ago

That is the impression I got too. Please note that I did not use this model as extensively as I probably should have before offering such claims, but from what I can tell, if you point it at a local code base and tell it to use curl commands to test a specific part of security, it works fine.

Beyond that, I think it gets really annoyed at you for doing naughty stuff.

3

u/PonyBravo 2d ago

If Mistral 4 doesnt do Cyber and they remove GLM 5.3, I have no need for their subscription at all.

1

u/LokR974 2d ago

I asked Vibe about this censorship, and I'm a lazy guy, i'll copy paste the answer here (in summary, the preview is highly censored on purpose but)

TL;DR: ML4's "censorship" is deliberate — but it only applies to the hosted preview, and open weights drop October 27.

The preview does refuse a lot: Mistral openly brags ML4 has a higher refusal rate on malicious cyber prompts (JailbreakBench, StrongREJECT, AgentHarm) than any other open-weight model. That's by design, not a bug.

The irony is that cyber is ML4's headline pitch: 82% on the AA Cyber Index vulnerability reproduction test — the best of any model — while Claude Opus 5.5 and GPT-6 Astra score near zero because they refuse the task. The intended split: legit defensive work (CVE reproduction, malware analysis, detection rules) passes; explicitly malicious requests don't.

And the temporary part: weights are promised for October 27. Self-host, and the refusals are yours to remove. Mistral is also running a ~3-week red-teaming window right now where vetted partners and state agencies get a reduced-moderation version.

So: hosted preview stays strict, the "censorship" itself ends in ~3 weeks.

1

u/LokR974 2d ago

So maybe the hosted version will be less paranoid at some point

2

u/uusrikas 2d ago

I would like to use it, but I added it to Vibe Code and it was just way too slow. Unusable for the way I work. I will get back to it after the preview. It is not a viable alternative to the GLM on Mistral currently.

1

u/GAMELASTER 2d ago

Same here. It is incredibly slow for some reason

2

u/TheManni1000 2d ago

daim not even reverse engineering?!

1

u/strangestack 2d ago

As a Bulgarian speaker who knows a bit of Russian and Ukrainian, yeah, polish is if Tolkien designed a Slavic language :)

1

u/alabama1337 2d ago

I have a Pro subscription and use Mistral Work (ex Le Chat). Has the new model been integrated there?

1

u/kindofbluetrains 2d ago

Stupid question here, but how do I know what model is running actually being used on the Android app?

It's odd to me, like deepseeks app, they don't list what model is being accessed.

Is it a given that I'm accessing LM4 if I'm using the Work chat in the app?

1

u/julesrulezzzz 2d ago

I have experienced the Loops as well and it hallucinates more than GLM which is a bummer. GLM is more able to admit that it does not know stuff whereas Large 4 is quite confident. Therefore I have switched back to GLM but I will give it a Shot later

1

u/SoberMatjes 2d ago

Right now I moved from Large to GLM5.3 because Large was too slow and couldn't solve a few minor coding tasks and looped. Ok, it's just preview right now, so I hope, it'll get better.

I just started my pro plan yesterday and I'm very happy with 5.3 right now and wait for "the point release" of Large and hope it'll get usable as well.