Great Model
#9
by PixelPhilosopher - opened
This model is outstanding. Llama.cpp added support for BailingMoE today. I'm running a GGUF by SC117 called "Ling-3.0-tiny-abliterated-APEX-I-Quality.gguf" on an 8 core AMD CPU, no discrete GPU, just Vulkan on the iGPU. Ling 3.0 Tiny gets 27 TPS inference and 275 prompt processing with this setup. I have it hooked up to a custom tool that can search and read a complete local archive of Wikipedia via a Kiwix zim. It answers every question I throw at it, fast and perfect so far. Whoever made this LLM is extremely competent!