Ringarc. Book free AI Audit

Labs · Open Model Benchmark

Vikas's Open Model Benchmark

How close are open-weight models to the frontier — on real work, not exam questions? Every time a notable open model ships, I run it through my private benchmark of ~163 realistic tasks and plot its frontier lag: the overall delta against four frozen frontier anchors, with a bootstrap 95% confidence interval. A model is "statistically at the frontier" when its interval includes zero.

Last updated 2026-08-13 · RSS

Frontier lag over time

One row per open-model release, oldest at the top · dot = overall lag, bar = 95% CI · the dashed line is the frontier (lag = 0). at frontier behind frontier

Current standings

Click a row to expand its 18-dimension scores. Read tiers, not second decimals — single run per model, the CIs are there for a reason.

Model Org Params (total/active) License Evaluated Frontier lag [95% CI] Verdict

The 16GB story — the deployment frontier

The other frontier I track: the best quality you can get for $0, on one consumer GPU. Current headline: Qwen3.6-35B-A3B lands statistically at the frontier (+0.005 [−0.01, +0.02]) on a single 16GB card, 66 tok/s, 10.9GB VRAM, thanks to expert offload. The story is in the write-up. Previous headline, still true: gpt-oss-20b ties GPT-5.4-mini on overall quality at the field's fastest time-to-first-token.

Quantization tax, measured: fp8 minus Q4 on identical weights across 160 paired tasks is +0.023 [0.00, +0.05] — zero on everyday tasks, concentrated in hard reasoning (+0.10) and repo work (+0.06). If your workload is everyday, Q4 is free lunch.

Test rig: RTX 5070 Ti 16GB + 128GB DDR5, Ollama / llama.cpp.

Field notes & serving traps

Things that bit me while measuring — the operational half of benchmarking that rarely gets written down.

Methodology show ▾

~163 private tasks across 18 dimensions of realistic work: coding, repo-level changes on two real codebases, everyday explanations, factuality (including hallucination baits), instruction following, long context, document/OCR/UI understanding, reasoning, and structured output — each split into a core band ("good enough for everyday use") and a hard band (frontier discrimination).

Four frontier models were run once and frozen as calibration anchors. Every new open model runs the full suite, is paired per task against the anchors on identical content, and gets an overall frontier-lag score with a bootstrap 95% confidence interval. Open-ended answers are graded criterion-by-criterion by a three-model judging panel (majority vote per criterion) that was audited for family bias before being seated.

Tasks stay private to avoid contamination — which is also why per-task details aren't published. Dimension-level scores are the finest granularity you'll see here.

Honest caveats

One person's benchmark, not an institution's. Single run per model (hosted APIs mostly reject sampling controls), so small deltas are noise — read tiers, not second decimals; the confidence intervals are there for a reason. ~163 tasks is small compared to public leaderboards; what it buys is that every task is private, realistic, and hand-validated. The suite measures single-shot work — it cannot yet see agentic multi-step tool use, which is exactly where several recent models claim their edge. Scores reflect a pinned serving configuration (provider, precision, reasoning effort), stated per model.

Changelog

2026-08-13 · Qwen3.8-Max weights land

Qwen3.8-Max's open weights arrived today as Qwen3.8-2.4T-A95B on Hugging Face, under a custom qwen3.8-max license. One honest footnote stays on its point in the chart: the released checkpoint is text-only with thinking mandatory, while the API model measured here on Aug 4 was multimodal with thinking off by default. The lag point describes the API sibling until the open checkpoint is measured in its own right. This update also catches the page up on four hosted intakes from last week, below.

2026-08-13 · Qwen3.6-35B-A3B: the first local model at the line, plus a correction

Qwen3.6-35B-A3B (Alibaba, 35B total / 3B active MoE, multimodal), running 4-bit on a single 16GB RTX 5070 Ti, lands at +0.005 [−0.01, +0.02]: statistically at the frontier, tied with Qwen3.8-Max for the best point estimate in the series, and the first model to get there running locally at $0. GLM-4.6V-Flash (Zhipu, 9B, MIT) lands at −0.175 with perfect image scores, which is the clearest evidence yet that the media dimensions have saturated and can no longer separate models. And a correction: Gemma 4 12B's July point (about −0.20) was measured with thinking off, a workaround for runaway thinking that silently became the measurement. Re-measured with thinking on and a sane cap, it lands at −0.045 [−0.08, −0.02], the best small local model on the board. The full story, including why the 35B MoE fits a 16GB card while the 27B dense model doesn't, is in the write-up.

2026-08-09 · GLM-5.2 and Inkling-Small

GLM-5.2 (Zhipu, 753B/40B MoE, MIT) lands at +0.004 [−0.01, +0.02], tied with the two Qwens for the best point estimate in the series, with the best repo score in the open field (0.98). The field calls it the top open model and this lag series agrees. Inkling-Small (Thinking Machines, 276B/12B, Apache-2.0) posts −0.002 [−0.02, +0.02], nominally ahead of the 975B big Inkling at 3.5x fewer parameters. Artificial Analysis scores the pair as tied, and this suite reaches the same verdict by a completely different method, a useful check on the instrument's resolution. Four open models now sit within ±0.005 of the frontier line.

2026-08-08 · DeepSeek V4 Pro and Flash, plus a scorer correction

DeepSeek V4 Flash (284B/13B, MIT) lands at −0.014 [−0.04, +0.01], statistically at the frontier at about $0.09 per Mtok in, the cheapest at-frontier result in the series. Its 1.6T sibling Pro lands at −0.021 [−0.05, 0.00] and shows no measurable overall gain over Flash at 4x the price (paired delta −0.006 [−0.03, +0.02]): Pro wins factuality, Flash wins long context and structured output. Both are pinned to Baidu's fp8 endpoint, because unpinned routing silently landed Flash on an fp4 host. Also on this date: three scorer regexes were corrected and every stored row re-scored. Repo scores rose about 0.02 field-wide, and earlier published figures are superseded.

2026-08-04 · Qwen3.8-Max

Qwen3.8-Max (Alibaba, 2.4T total / 95B active MoE, multimodal) posts the series' first positive lag point estimate: +0.005 [−0.01, +0.02] — statistically at the frontier. It's the first model clean on both historic failure modes at once: the hallucination baits and the full-stack repo band. Head-to-head with Kimi K3 it's statistically tied overall; Qwen wins repo and factuality, Kimi keeps the better everyday register. One asterisk: API-only at evaluation, with open weights announced for ~Aug 10 — the point stays flagged until weights land.

2026-07-27 · Inkling

Inkling (Thinking Machines, 975B/41B MoE, Apache-2.0) lands at −0.014 [−0.04, +0.01] — statistically at the frontier. It's the first open model to hold the full-stack repo band, which had been an open-model graveyard. Its weakness is the opposite of its coding strength: on the hallucination baits it invented confident plots for both fake films.

2026-07-22 · Inaugural cohort

The suite goes live: four frontier anchors (Opus 4.8, Sonnet 5, Haiku 4.5, GPT-5.4-mini) frozen in July as calibration reference, and four open models measured. Kimi K3 lands at −0.019 [−0.05, +0.01] — statistically at the frontier, with the best everyday-register answers in the open field and the only clean sweep of all 8 ethical-dilemma tasks ever recorded here. gpt-oss-20b becomes the efficiency headline: −0.078 overall, but tying GPT-5.4-mini at $0 on a 16GB GPU. Qwen3-Coder-30B and Gemma 4 12B trail at around −0.20, each with a distinct cliff (hard reasoning for both; repo work for Gemma).