Exact scripts, prompts, executable graders, and raw Ollama JSON behind the wasnotwas article Muse Glimmer on an RTX 4090: A Good Agent That Thinks Too Much.
Read README.md for requirements, exact model configuration, commands, and limitations.
python3 benchmark.py python3 addendum.py python3 hard_benchmark.py python3 language_benchmark.py
This grew organically during one session. It is an executable ad hoc evaluation, not a polished benchmark or a claim of statistical significance. Raw JSON includes model-generated thinking fields and synthetic test data only.