Benchmarks · agency-ops

Agency-ops: historical evaluation results.

Historical agency-ops evaluation, not a claim of frontier parity or current production quality. Here’s exactly how it was measured — and where its limits are.

Other niches publish held-out validation loss; they have not been run through this frontier head-to-head yet.

Head-to-head quality

50 held-out items · reported n-gram overlap check, not proof of no contamination. blind panel of 2 model judges (Gemini 3.5 Flash + DeepSeek v3.1), absolute 1–10.

Fable 59.10
GPT-5.59.07
agency-ops — ours, MLX 8-bit · agency-ops only · Apple Silicon9.05
Qwen3.6-35B-A3B (base)8.01
GLM-5.2 · 744B6.88

Bars show historical mean model-judge scores (0–10). They do not establish frontier parity, statistical significance, or quality under the current shipped prompt.

What you get, by how you run it

Historical measurements cover different serving configurations and test sets. They are not interchangeable. The disputed hosted result is withheld until its provenance is resolved.

Run it viaHardwareQuality
MLX 8-bit · agency-opsApple Silicon · M-series, 24GB+ unified9.05 · historical run only
Q4_K_M GGUF · agency-opsNVIDIA GPU, CPU, or Mac7.97 · earlier 20-item run
Q4_K_M GGUF / hosted API · other 9 nichesNVIDIA GPU, CPU, Mac, or our APINo catalog-wide score · not yet measured

9.05 is a historical agency-ops MLX 8-bit result. The hosted score is withheld pending configuration reconciliation. Agency-ops measured 7.97 as Q4_K_M on an earlier 20-item run; the nine Q4_K_M catalog niches have no catalog-wide blind score yet.

How we measured it

The whole point is a result that survives scrutiny. Four rules make it honest:

Held-out test set

50 held-out questions with a reported n-gram overlap check against training and validation data. This check does not rule out every form of contamination.

Blind model judges

Every answer is graded 1–10 by two judges from different labs (Gemini 3.5 Flash + DeepSeek v3.1). Neither is a contestant, and the judge never sees which model wrote the answer.

Same prompt for everyone

Identical system prompt, temperature, and token budget for every model — ours and the frontier APIs alike. No home-field advantage.

Evaluation artifacts

The repository contains the harness, golden set, and judge prompts. Independent reproduction has not been established, and reruns may differ.

What we want to test next

Our hypothesis is that niche-specific training can make familiar tasks more useful and efficient. The next step is to validate the shipped prompt and artifact, test broader tasks, and measure regressions. This historical run does not establish those outcomes.

As of July 2026. These are historical results; new releases require fresh evaluation.