// HACKER NEWS — CYBERSECURITY
I checked 30 frontier model cards. Here are the benchmarks labs report
This ranks reporting convention, not benchmark quality. One model card
contributes at most one mention to a benchmark, however many configurations
it reports.
Some model cards publish their benchmark table as an image, and those rows
were transcribed by OCR. A transcription can be wrong. The
source ledger lists every recorded
benchmark with the original document so each count can be checked.
Start with signals that distinguish emerging instruments from established
standards and saturated conventions.
Three readings on one time axis. Each orange diamond marks an organization's first
report of this benchmark, which is the only event that raises the cumulative count.
The rug beneath it puts one tick per dated model card, so later cards from an
organization already counted appear as gray ticks and leave the staircase flat. A card
with no publication date cannot be placed on the timeline and is absent from both
bands, though it is still counted in the totals above. A long flat run is
reporting saturation observed within this curated registry, not a claim about
benchmark score saturation.
The score track below it is a separate reading: every value that could be read
verbatim from a cited document, connected only where the instrument and protocol
are identical. A flat score tail usually means no newer number could be read, which
is why the gap is marked rather than drawn through.
Comparable score observations need the benchmark version and split,
metric direction, model, harness or scaffold, reasoning budget, cost or
latency, publication date, and source. Only compatible configurations can
share a score frontier; this registry currently stores mentions, not those
measurements.
With those observations, a Harbor-style view can put cost or latency on the
x-axis and score on the y-axis, connect only nondominated observations, and
use a publication-time slider to reveal how the frontier moved.
Each model card counts once per benchmark. A card reporting AIME in four
configurations counts the same as a card reporting it once, so a long appendix
cannot outweigh a different vendor. Organizations breaks the tie: the same count
from six vendors is a shared standard, from one vendor a house style.
The curated source list this ranking is computed from. Expand any card to see
every benchmark it reports, grouped the way the source document groups them, so
our data can be checked line by line against the original.
The overview summarizes the full corpus. The relationship canvas includes every
artifact and connected organization, source, and topic; select a node to carry it
into the Today filters.
Topic, source, and organization nodes set the corresponding Today filter.
Artifact nodes set the date and title search.