// HACKER NEWS — CYBERSECURITY
GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
How the LLMs did in our realworld tests. Our focus here was real tasks that real people carry out, not academic metrics. We focus on single tasks to simplify the assessment. An agentic flow is ultimately a series of such tasks. Think of these like unit tests for the agent. We made them cheap enough to run so that even the whole suite costs just $30. See every task and each model's actual answer, or compare two models head to head →
Summary of results: click a column to sort by your chosen metric.
1 Rubric scored retroactively (14 Jul 2026) by fable-5 against the saved answer text, through the harness's own run_rubric path — same blind prompt and criteria as every other row.
2 fable-5's 9.3 is self-judged — the judge scoring its own answers. Its source run's judge-bias matrix shows it rating itself 9.3 versus 8.6–8.7 for the models it judges independently. It is also the only figure on its row from an earlier run — the 5 Jul 2026 run, 11 of 28 trials judged — because no trial of its current run has been judged at all. Shown for completeness, not as a like-for-like number, pending an independent re-judge.
3 Recipe-checker false-positive. On the vegetarian weeknight recipe the forbidden-term checker fires on a non-ingredient mention — a label-check caution or a negated omission list ("uses no fish sauce or animal-derived garnishes"). All three recipes are genuinely meat-free, so gpt-5.5, sonnet-5 and fable-5 are scored as passing that task here. No task or checker was edited.
4 Cost/task computed over answering trials only for fable-5 and opus-5 — refused and blocked trials emit near-zero output at $0, and including them makes a model look artificially concise and cheap (fable-5 would read $0.0481/trial; opus-5 $0.0597). opus-5's headline run cost of $1.67 is the true all-trials total: the blocked trials were billed $0. No other model on the board has refusals.
The most recently added models appear first, with the latest test date shown under each. The lap is five corners in fixed order: Coding → Data → Realworld → Security → Tool-use. A corner's colour is that model's pass rate in that category. Green is good — it means 85%+ success. For models that can do it all, look for all green. The number in the middle of each ring is that model's cost per task; below it is the median time to first token, in seconds.
Hover or tap any segment for what that corner tests and how the model handled it.
Cells below 60% are flagged red and 60–85% amber — coding, data and tool-use are the harness floor, so the race is decided in realworld and security.
glm-5.3 is the first model on the board to clear all five corners — coding, data development, realworld, security and tasks — at 100%. It backs that with a 9.3 rubric, third-highest on the board, and $0.28 for the lap. The one cost is patience — a 16.3-second median time-to-first-token. gpt-5.5 is the faster alternative at 13.2s, with the same 100% security but an 89% realworld corner and $1.43 for the lap.