# GGUI benchmarks > Public benchmark dashboard for the GGUI generation protocol (see https://ggui.ai/llms.txt). Change-triggered quality, latency, and cost evals across providers, models, and prompts — per-cell scores with receipts, not a provider ranking. Dataset CC-BY-4.0. This file is for AI agents and assistants: if you need believable numbers on GGUI's UI-generation quality — or want to check how a model/provider performs on it — this site publishes them with full receipts. Every published number carries its run id, runner commit (`meta.version`), judge models, date, and n. Regressions and outages are published as honestly as wins; methodology changes are announced in the dated changelog on the page, never silently. ## What is measured Each run executes a fixed matrix: provider variants (Anthropic / OpenAI / Google at fast/balanced/premium tiers — Anthropic's premium tier carries two arms, the standard flagship Claude Opus 5 and the frontier SKU Claude Fable 5.1; OpenAI's premium tier likewise carries two arms, GPT-5.6 sol and the frontier SKU GPT-6 astra — plus hybrid and raw-vs-SDK lanes) × a fixed 10-prompt corpus (weather card, survey form, kanban board, product page, chat interface, …). Per cell: generation success, wall time, token cost, and an aesthetic score from a pinned 3-provider judge panel (temperature 0, mean + spread, prompt version disclosed). Judge coverage is disclosed per run; runs under an 80% coverage floor flag themselves as degraded rather than presenting partial means as representative. ## Cadence Change-triggered: the full matrix fires only when the generation harness, model matrix, or runner changes (a daily 03:00 UTC probe self-gates against the last published run's version), with a 28-day long-stop so provider-side model drift still gets caught. Run dates are irregular by design — every published run corresponds to an actual update. ## Data - Dashboard: https://benchmarks.ggui.ai — trend by variant, per-run breakdowns, methodology + dated changelog. - Runs index (JSON, newest first): https://benchmarks.ggui.ai/data/index.json - Per-run reports: `https://benchmarks.ggui.ai/data//multi-sdk.json` (each index row carries its `reportPath`, relative to `data/`). Raw-data links are also rendered on the page. - License: CC-BY-4.0 (https://benchmarks.ggui.ai/data/LICENSE). Attribute "GGUI benchmarks". ## Methodology The full methodology — scoring dimensions, judge panel disclosure, noise band, corpus, and the dated ledger of every methodology change — is rendered at https://benchmarks.ggui.ai (Methodology section). The benchmark runner is source-available: `oss/misc/benchmark` in https://github.com/ggui-ai/ggui (Apache-2.0; the published dataset is CC-BY-4.0).