← Artifacts

Muse Glimmer 30B ad hoc evaluation

Exact scripts, prompts, executable graders, and raw Ollama JSON behind the wasnotwas article Muse Glimmer on an RTX 4090: A Good Agent That Thinks Too Much.

Model: official Meta K-Quant-17GB GGUF
Comparison: Gemma 4 26B Q4_K_M
Hardware: RTX 4090, full Vulkan GPU offload
Runtime: Ollama 0.32.7 + llama.cpp 4dee52f

Reproduction

Read README.md for requirements, exact model configuration, commands, and limitations.

python3 benchmark.py
python3 addendum.py
python3 hard_benchmark.py
python3 language_benchmark.py

Scripts

Raw results

This grew organically during one session. It is an executable ad hoc evaluation, not a polished benchmark or a claim of statistical significance. Raw JSON includes model-generated thinking fields and synthetic test data only.