Last time I wrote about the judge that scores my benchmark, I left one thing out. That post was about adding a model to the scoring. It was not about the fact that the model is the…
Last time I wrote about my benchmark (the code that kills your browser), I skipped the boring part. Sixteen models, six attempts each, one task – that is 96 HTML files somebo…
After Anthropic’s recent moves – higher prices, hard-trimmed limits – a question I keep hearing more and more comes back: “Codex or Claude Code?”. And…