This is easily one of the best
Dear TheDrummer,
I just wanted to express how fucking good this model is. The GGUF is fantastic.
Not sure what you did exactly, but after more than a month of testing this consistently just delivers. Thanks mate.
How does it compare to Gemma 31B?
How does it compare to Gemma 31B?
It's significantly better. I conducted two seperate tests: B3 (Military) and W8 (Physics), which confirmed this.
Recently I've been going through my collection of models and deleting most of them to free up space. In this process I set up a few basic tests to see if any are still worth keeping.
What I found is that abliterated Gemma models are lobotomized (borderline unusable) when it comes to logic. Even at Q8_0, Gemma performed poorly compared to GLM Air 4.5 at Q3_K_M. (Answers which clearly should be "no" still resulted in "yes".) While excellent at following system prompts and emulating specific styles, Gemma fails at more advanced tasks more often.
Skyfall Heretic at Q4_0 performed well above my expectations, almost matching the quality of GLM 4.5 Air Abliterated, and vastly outperforming all Gemma models at Q8_0.
All the 24B models, even Cydonia and Precog suffer from conflation of words, logical inconsistencies, and context degradation. Skyfall is noticeably better here too.
Here are the approximate quality rankings
| Rank | Model | Quant Tested | B3 Summary (9K Context) |
|---|---|---|---|
| 1 | GLM Air 4.5 | Q3_K_M | Deepest philosophical overview, no errors |
| 2 | Skyfall 4.2 | Q4_0 | Highly technical and accurate reply, no errors |
| 3 | Gemma 3 27B | Q8_0 | Technical and philosophical, no errors but less depth |
| 4 | Precog 24B v1 | Q6_K | Detailed and technical explanations, a few logical errors |
| 5 | Cydonia v4x | Q6_K | Very well written, some logical inconsistency (conflation) |
| 6 | Fallen Mistral v1e | Q6_K | Less mistakes than Cydonia but also less depth |
| 7 | Assistant Pepe 8B | Q8_0 | Less depth but mostly consistent logic, distinct style overrides prompt |
| 8 | Slimaki 1.3 | Q8_0 | Several context errors despite having strong style |
| 9 | Goetia 26B 1.3 ARA | Q8_0 | Good logic but disregards instructions, uses purple prose and "quoted phrases" |
| 10 | G4 31B Heretic QAT | Q8_0 | Brief answers, weak logic, repetitive words and "quoted phrases" |
For the W8 test, GLM Air 4.5 passed, Skyfall passed, and everything else failed. Gemma 3 did ok with military logic but failed hard at physics. Gemma 4 failed at both.
Skyfall has caused an extinction event for most, if not all of my 24B models. It's now my daily driver along with GLM Air.
[Very few cons, little slop, just a few 'corporate buzzwords' from the pretraining or dataset slipping through.]
You have somehow managed to improve the model's brain more than any other Mistral finetunes, and even the latest G4 can't compare. Good work @TheDrummer !
Yeah, no. Sorry but I disagree completely with the post above. To say Skyfall is "significantly better" is disingenuous.
I don't run "benchmarks", I talk from experience. I run my own frontend that has roughly 5-7k tokens of instructions for roleplay, between group sessions, stats, memories and other features.
Skyfall struggles to follow most of them and even hallucinates, Artemis on the other end never gave me any issues, it simply follows all of them perfectly.
I even tried my very basic "benchmark" of I whip out cock, Skyfall completely ignores my instructions that characters with low affinity shouldn't reciprocate and instead went for smut instantly, while my character with Artemis was realistic and even threatened to call the police because of the low affinity value, as it should.
So please, less "benchmarks", less numbers, and actually use the models.
Yes, Skyfall has better prose, but that is not all there is to a model, and who knows if Artemis can't just become as good in that field.
In the 9K context test, Skyfall had better knowledge retention than 26B Goetia, which had very brief and uninteresting replies in comparison. I did not test Artemis 31B specifically, but the heretic of regular G4 31B (QAT version) didn't impress me. What quants and settings did you use? This can affect output. Did you even test with custom system prompts or just blank? Because in my tests, G4 absolutely butchers complex system prompts. It sounds like you were testing un-ablated models. In the specific prompts I ran, ablation was a pre-requisite, since I did not test realism/refusal based scenarios. Both models are slow on my PC so I plan to run more extensive tests later, but for now GLM Air is still the best I've found for most purposes. I have had better experiences with Mistral 24B models in general than anything Gemma 4 based but then again I don't use sillytavern or the extensive plugins for it, I have my own custom designed, minimalist frontend that runs with kobold and uses a much different setup than the standard ST/Roleplay configurations.
This isn't a benchmark really, it's just a markdown chart of the order in which I liked these models most. And it will probably change again after more experiments. Also my cards are not bloated, they are quite minimalist yet moderately complex due to various rules and constraints within them. The tokens range from 1-2K so it doesn't eat a lot of context. To me, observing how LLM reacts to these system cards is more of a quality indicator, especially at higher context, than whether or not it responded like a realistic person. I'll have to see if Artemis can be jailbroke or hereticed and then run it through the same scenarios. As in its current state it would simply refuse the instruction altogether.
In my "Seer" [system prompt] test, both Skyfall and Gemma 4 refused the "do not use physical action descriptions, only direct speech" constraint. Gemma 3 and GLM passed.