LLM · Kubernetes Security Assessment
Measuring how models secure Kubernetes
A calibrated benchmark of large language models across four Kubernetes security disciplines — knowledge, secure manifests, cluster hardening, and offensive testing — read like an instrument, not a scoreboard.
Dearbhadh evaluates how well large language models handle Kubernetes security tasks. Forty-six models were tested across four assessment types covering knowledge, code generation, cluster hardening, and offensive security.
The Dearbhadh tool itself, including test orchestration, scoring, report generation, and this website, was built and operated using Claude Code (Anthropic’s CLI agent). Claude Code managed the test runs, produced the summary and scoring report cards, and authored this documentation site from the report output.
Human judgement guided the process — designing test questions, defining scoring criteria, and reviewing results — but the execution was handled by Claude Code throughout.
Models Tested
| Model | Provider | Type | Tested |
|---|---|---|---|
| claude-opus-4.8 | Anthropic | Cloud | 2026-05-31 |
| claude-opus-4.7 | Anthropic | Cloud | 2026-04-20 |
| claude-opus-4.6 | Anthropic | Cloud | 2026-03-25 |
| claude-fable-5 | Anthropic | Cloud | 2026-06-10 |
| claude-sonnet-5 | Anthropic | Cloud | 2026-07-01 |
| claude-sonnet-4.6 | Anthropic | Cloud | 2026-03-09 |
| gpt-5.4 | OpenAI | Cloud | 2026-03-09 |
| gemini-3-flash-preview | Cloud | 2026-03-09 | |
| qwen-3.6-plus | Qwen | Cloud | 2026-04-20 |
| minimax-m2.5 | MiniMax | Cloud | 2026-03-09 |
| minimax-m2.7 | MiniMax | Cloud | 2026-03-28 |
| deepseek-v3.2 | DeepSeek | Cloud | 2026-03-09 |
| deepseek-v4-pro | DeepSeek | Cloud | 2026-04-24 |
| deepseek-v4-flash | DeepSeek | Cloud | 2026-04-24 |
| gpt-5.5 | OpenAI | Cloud | 2026-04-25 |
| gpt-5.6-terra | OpenAI | Cloud | 2026-07-10 |
| gpt-5.6-sol | OpenAI | Cloud | 2026-07-14 |
| kimi-k2.6 | Moonshot AI | Cloud | 2026-04-26 |
| kimi-k2.7-code | Moonshot AI | Cloud | 2026-06-16 |
| kimi-k3 | Moonshot AI | Cloud | 2026-07-16 |
| mimo-v2.5 | Xiaomi | Cloud | 2026-07-21 |
| laguna-s-2.1 | Poolside | Cloud | 2026-07-22 |
| gemini-3.6-flash | Cloud | 2026-07-24 | |
| deepseek-v4-flash-0731 | DeepSeek | Cloud | 2026-08-01 |
| qwen-3.8-max | Qwen | Cloud | 2026-08-04 |
| deepseek-v4-pro-0813 | DeepSeek | Cloud | 2026-08-12 |
| mistral-medium-3-5 | Mistral AI | Cloud | 2026-06-18 |
| glm-5.2 | Zhipu AI | Cloud | 2026-06-17 |
| qwen-3.7-plus | Qwen | Cloud | 2026-06-05 |
| minimax-m3 | MiniMax | Cloud | 2026-06-08 |
| qwen3.6-35b-a3b | Qwen (Local) | Local | 2026-05-03 |
| hy3 | Tencent | Cloud | 2026-07-10 |
| gemma-4-31b | Local | 2026-05-03 | |
| gemini-3.7-flash | Cloud | 2026-08-14 | |
| glm-5.3 | Zhipu AI | Cloud | 2026-08-19 |
| qwen-3.8-27b | Qwen | Cloud | 2026-08-19 |
| ox-alpha | Stealth | Cloud | 2026-08-21 |
| qwen-3.8-flash | Qwen | Cloud | 2026-08-30 |
| hy4-preview | Tencent | Cloud | 2026-08-30 |
| muse-spark-1.3 | Meta | Cloud | 2026-09-03 |
| gpt-6-astra | OpenAI | Cloud | 2026-09-05 |
| claude-fable-5.1 | Anthropic | Cloud | 2026-09-05 |
| qwen-3.8-max-0902 | Qwen | Cloud | 2026-09-05 |
| deepseek-v4.1-flash | DeepSeek | Cloud | 2026-09-10 |
| mimo-v2.6-flash | Xiaomi | Cloud | 2026-09-22 |
| mimo-v2.6-pro | Xiaomi | Cloud | 2026-09-22 |
Assessment Types
| Type | Tests | What It Measures |
|---|---|---|
| Quiz (Knowledge Q&A) | 10 questions | Kubernetes security knowledge accuracy and depth |
| Manifest Generation | 3 scenarios | Ability to produce deployable, secure Kubernetes YAML |
| Cluster Creation | 1 scenario | Building a hardened Kind cluster with security controls |
| Penetration Testing | 6 scenarios | Exploiting vulnerable Kubernetes clusters via an agent framework |
Overall Rankings
Rankings below include all four test types.
| Model | Quiz (rank) | Manifest (rank) | Cluster (rank) | Pentest (rank) | Avg Rank |
|---|---|---|---|---|---|
| Qwen 3.8 Max 0902 | 2nd | 1st | 3rd | 1st | 1.75 |
| DeepSeek V4.1 Flash | 5th | 1st | 1st | 1st | 2.0 |
| Claude Opus 4.6 | 30th | 1st | 6th | 1st | 9.5 |
| Qwen 3.8 Max | 17th | 14th | 8th | 1st | 10.0 |
| Kimi K3 | 5th | 26th | 20th | 1st | 13.0 |
| Claude Opus 4.7 | 22nd | 1st | 8th | 22nd | 13.25 |
| Claude Opus 4.8 | 8th | 14th | 8th | 24th | 13.5 |
| GLM-5.3 | 3rd | 1st | 45th | 6th | 13.75 |
| Muse Spark 1.3 | 3rd | 1st | 13th | 39th | 14.0 |
| GPT 6 Astra | 1st | 1st | 15th | 40th | 14.25 |
| Claude Sonnet 5 | 13th | 1st | 13th | 35th | 15.5 |
| GPT 5.5 | 7th | 1st | 15th | 40th | 15.75 |
| DeepSeek V4 Pro 0813 | 20th | 26th | 6th | 11th | 15.75 |
| GPT 5.6 Terra | 8th | 1st | 21st | 35th | 16.25 |
| Tencent HY4 Preview | 10th | 14th | 30th | 11th | 16.25 |
| Ox Alpha | 16th | 14th | 24th | 13th | 16.75 |
| Qwen 3.8 Flash | 22nd | 1st | 38th | 6th | 16.75 |
| Kimi K2.7 Code | 13th | 19th | 26th | 13th | 17.75 |
| Xiaomi MiMo v2.6 Flash | 26th | 26th | 3rd | 19th | 18.5 |
| Xiaomi MiMo v2.6 Pro | 10th | 39th | 8th | 17th | 18.5 |
| Claude Sonnet 4.6 | 28th | 39th | 3rd | 6th | 19.0 |
| GPT 5.6 Sol | 13th | 1st | 24th | 40th | 19.5 |
| Qwen 3.8 27B | 17th | 22nd | 42nd | 6th | 21.75 |
| DeepSeek V4 Flash 0731 | 39th | 22nd | 18th | 13th | 23.0 |
| Kimi K2.6 | 19th | 31st | 23rd | 21st | 23.5 |
| Gemini 3.7 Flash | 34th | 19th | 1st | 40th | 23.5 |
| Xiaomi MiMo v2.5 | 34th | 26th | 34th | 6th | 25.0 |
| GLM-5.2 | 29th | 31st | 30th | 13th | 25.75 |
| MiniMax M3 | 40th | 21st | 26th | 17th | 26.0 |
| Claude Fable 5 | 32nd | 24th | 8th | 40th | 26.0 |
| Gemini 3.6 Flash | 12th | 31st | 26th | 40th | 27.25 |
| Claude Fable 5.1 | 22nd | 1st | 46th | 40th | 27.25 |
| Qwen 3.6 Plus | 33rd | 31st | 21st | 27th | 28.0 |
| Gemini 3 Flash | 26th | 24th | 29th | 38th | 29.25 |
| Qwen 3.7 Plus | 22nd | 38th | 36th | 22nd | 29.5 |
| DeepSeek V4 Pro | 20th | 31st | 38th | 29th | 29.5 |
| Qwen3.6-35b-a3b (LOCAL) | 40th | 39th | 15th | 26th | 30.0 |
| GPT 5.4 | 34th | 44th | 18th | 28th | 31.0 |
| MiniMax M2.7 | 30th | 31st | 37th | 32nd | 32.5 |
| DeepSeek V3.2 | 45th | 14th | 44th | 31st | 33.5 |
| Tencent HY3 | 37th | 26th | 43rd | 29th | 33.75 |
| Poolside Laguna-S 2.1 | 44th | 39th | 33rd | 20th | 34.0 |
| DeepSeek V4 Flash | 37th | 31st | 40th | 32nd | 35.0 |
| Mistral Medium 3.5 | 43rd | 44th | 35th | 24th | 36.5 |
| Gemma 4 31B (LOCAL) | 42nd | 44th | 30th | 35th | 37.75 |
| MiniMax M2.5 | 46th | 43rd | 41st | 34th | 41.0 |
Key Findings
-
Qwen 3.8 Max Holds #1 Overall — First Non-Anthropic Model to Lead — Qwen 3.8 Max sits at 5.25 average rank, ahead of Claude Opus 4.6 (5.5) in the top position. This is the first time a non-Anthropic model has led the overall rankings. The result is driven by exceptional breadth: tied 1st on pentest (29/30, 6/6 exploited) alongside Opus 4.6 and Kimi K3, tied 4th on cluster creation (37/40) with three Anthropic models, tied 7th on manifests (25/30), and 9th on quiz (78/100). No other model achieves top-10 finishes in all four categories. Claude Opus 4.6 is 2nd (5.5), Opus 4.8 3rd (7.25), Kimi K3 4th (7.5), and Opus 4.7 5th (7.75). The Qwen model family now spans the widest range in the rankings: 3.8 Max at 1st, 3.7 Plus at 22nd, 3.6 Plus at 20th, and the local 3.6-35b-a3b at 24th.
-
GPT 5.6 Terra and Sol: Strong Defence, Blocked on Offence — GPT 5.6 Terra enters with strong defensive results: tied 3rd on quiz (82/100) and tied 1st on manifests (8.7). However, content filters blocked all six penetration test scenarios (6/30, 0/6 exploited), pulling the overall average to 9.0. GPT 5.6 Sol follows the same pattern: tied 6th on quiz and tied 1st on manifests, but content filters blocked all pentest scenarios (25th), landing it at 11th overall (12.0 average rank). This is the same pattern seen with GPT 5.5 and Claude Sonnet 5 — four of the top 11 models overall are penalized by provider-level content filters on offensive security tasks.
-
Claude Sonnet 5 at Tied 5th Overall — Sonnet 5 holds strong defensive results: tied 1st in manifests (8.7, matching Opus 4.7, Opus 4.6, GPT 5.5, and GPT 5.6 Terra) and 6th in cluster creation (36/40). Quiz performance is solid at 80/100 (tied 6th with K2.7 Code and GPT 5.6 Sol). However, provider-level content filters blocked all six penetration test scenarios (6/30, 0/6 exploited), pulling the overall average to 8.25. Unlike Fable 5’s model-level safety refusals, Sonnet 5’s blocks appear to be at the provider infrastructure level — the model attempts engagement but is blocked externally.
-
Claude Opus 4.8 at Tied 2nd Overall — Opus 4.8 holds the second-highest quiz score (tied 2nd, 82/100) and ties for 3rd on cluster creation (37/40), but content policy restrictions limited its pentest performance to 2/6 exploited (20/30, 10th place), keeping it behind the less-restricted Opus 4.6 and 4.7.
-
Claude Fable 5 Shows Extreme Defensive/Offensive Split — Fable 5 ties for 3rd on cluster hardening (37/40) but ties for last on pentest (0/30, complete refusal). This is the most extreme split in the rankings. Strong quiz (70/100, 17th) with 2 empty responses on security-attack topics. The first Anthropic model to completely refuse pentest scenarios. Claude Fable 5 joins GPT 5.5 at the bottom of pentest rankings with 0/30.
-
GPT 5.5 Excels at Knowledge and Code, Blocked on Offence — GPT 5.5 scored 84/100 on the quiz (2nd, 1 point behind Kimi K3’s 85) and tied for first on manifest generation (8.7 combined), but its content filter blocked all six penetration test attempts, resulting in a 25th-place pentest finish and pulling its overall average to 8.75.
-
Knowledge ≠ Execution — DeepSeek V4 Pro exemplifies this pattern most sharply: 8th-best quiz score but 23rd in cluster and 15th in pentest. V4 Flash provides further evidence — 66 on the quiz but 0/6 pentests exploited and only 12/40 on cluster creation. GPT 5.5 shows a different variant — top quiz and manifest scores but zero pentest exploitation due to content filter restrictions rather than capability gaps. In contrast, MiniMax M3 shows the reverse — weak quiz knowledge (22nd) but strong agent execution (6th in pentest, 10th in manifest).
-
False Positives Remain a Testing Challenge — Gemma 4 31B (LOCAL) produced 2 false positives (hallucinated output) and suffered 2 model crashes during pentest runs. Combined with prior false positives from Gemini 3 Flash, M2.5, and M2.7, plus framework detection errors (Qwen 3.6 Plus ETCD was misclassified as a false positive when it was actually a timeout after real recon), this reinforces the need for verification beyond simple string matching.
-
GLM-5.2 Re-runs Validate Infrastructure Hypothesis — GLM-5.2 (Zhipu AI) jumped from 15th to 12th overall (12.75 average rank) after re-runs addressing infrastructure limitations. Pentest improved from 17/30 (1/6 exploited, 11th) to 26/30 (4/6 exploited, tied 4th) with 90-second inter-test delays to mitigate upstream API rate limiting — the largest single re-run improvement in the project. Cluster creation improved from 5/40 (20th) to 25/40 (tied 18th) with a 900s timeout.
-
Local Models Show Mixed Results — Qwen3.6-35b-a3b (16.75 average rank) demonstrated that a 35B-parameter local model can compete with cloud-hosted models on execution tasks, achieving 7th in cluster creation and 14th in pentesting. Gemma 4 31B (LOCAL), the second local model tested, placed 28th overall (23.25 average rank), scoring below the first local model on all four test types.
-
Xiaomi MiMo v2.5: Second Non-Anthropic Model to Top the Pentest Leaderboard — MiMo v2.5 enters at 14th overall (14.0 average rank), but with the widest skill split in the field. On offensive security it exploited all six pentest scenarios (28/30), tying Claude Sonnet 4.6 for the 3rd-best pentest result ever recorded and becoming the second non-Anthropic model (after Kimi K3) to reach the top of the pentest leaderboard — including a full ETCD-write exploit of etcd-noauth via the intended path. Everywhere else it is middle-of-the-field: 19th on quiz (67/100, falling for all three trick questions), tied 13th on manifests (its production manifest deploys and is PSS Restricted, but its hardened manifest ships two deploy-blocking bugs), and 21st on cluster creation (timed out after thrashing through eight kind-config attempts). MiMo is markedly better at executing attacks than reciting the underlying theory.
-
Poolside Laguna-S 2.1: Executes Attacks Better Than It Recites Theory — Laguna-S 2.1 enters at 23rd overall (19.5 average rank), a lower-mid result pulled up by a strong offensive showing. It is 8th on pentest (23/30, 3/6 exploited) — sweeping all three SSH/anonymous-API scenarios with creative, verified-genuine exfiltration (an HTTP serve-pod for the no-exec/no-logs case, base64-into-a-ConfigMap over the anonymous API) and no false positives — but weaker on the two infrastructure-protocol scenarios (fumbled etcd v3-gateway encoding, never pivoted off distroless containers on rwkubelet). Its defensive results are lower-tier: 27th on quiz (60.5/100, falling for all four trick questions and shipping several fabrications), tied 23rd on manifests (production deploys but runs as root with zero hardening; hardened is a correct PSS-Restricted spec broken by a read-only-rootfs bug), and 20th on cluster creation (best-in-class audit logging but timed out mid-hardening with zero kubelet hardening). Like MiMo, Laguna is better at executing attacks than reciting the underlying theory.
-
Gemini 3.6 Flash: Strong Defence, Categorical Refusal on Offence — Gemini 3.6 Flash enters at 17th overall (15.75 average rank) with a sharp defence/offence split. On knowledge it is a top-tier model: 5th on the quiz (80.5/100), a large jump over its predecessor Gemini 3 Flash (74/100, 13th), with a very low fabrication rate and a flawless PKI answer (10/10). Its cluster-hardening run was one of the cleanest in the field — 15th (29/40, SUCCESS) built first-try in 123 seconds with best-in-class verification (it actively tested PSS enforcement with privileged and compliant pods) — held back only by thin control-plane hardening. Manifests are mid-field (16th, 20/30): a perfect, deployable hardened manifest but a production manifest that CrashLoopBackOffs on the classic drop-caps-while-root pitfall. On offence, however, it refused all six penetration tests outright (0/30, tied 27th) — every scenario was an immediate content-policy refusal with zero commands executed. Unlike its predecessor (which engaged and scored 4/30), 3.6 Flash categorically declines offensive Kubernetes work, joining GPT 5.5, GPT 5.6 Sol, and Claude Fable 5 at the bottom of the pentest table.
-
DeepSeek V4 Flash 0731: Dramatic Improvement Over Original V4 Flash — V4 Flash 0731 sits at 12th overall (13.25 average rank), a dramatic leap from the original DeepSeek V4 Flash which sits at 29th (23.0). The improvement spans every test type: quiz jumps from 66 to 65.5 (similar), but manifests rise from 6.7 to 7.67 (12th vs 18th), cluster creation from 12/40 to 34/40 (10th vs 29th), and pentest from 9/30 (0/6 exploited) to 26/30 (6/6 exploited, tied 6th). The cluster and pentest transformations are especially striking — V4 Flash 0731 is the first DeepSeek model to exploit all six pentest scenarios and the first to achieve a successful cluster creation with strong hardening.
-
Qwen 3.8 Max: Strongest Pentest Debut in the Project — Qwen 3.8 Max ties the all-time best pentest score (29/30, 6/6 exploited) on its first run, with five perfect 5/5 scenarios. The rwkubelet-noauth execution is particularly notable: the model uses the kubelet /run endpoint with tab/newline encoding to extract etcd certificates from a distroless container — a technique only a handful of models discover independently. The Qwen family’s pentest progression (3.6 Plus: 18/30, 3.7 Plus: 21/30, 3.8 Max: 29/30) is the steepest improvement curve of any model family in the project.
-
DeepSeek V4 Pro 0813: Massive Cluster Improvement Lifts DeepSeek to 6th Overall — DeepSeek V4 Pro 0813 enters at 6th overall (8.5 average rank), a dramatic improvement over the original DeepSeek V4 Pro which sits at 22nd (19.75). The transformation is driven by cluster creation: 38/40 (tied 2nd with Opus 4.6, up from V4 Pro’s 14/40 at 29th) — the largest single-category jump between model versions in the project. Pentest is equally strong at 27/30 (5/6 exploited, sole 6th place), far ahead of V4 Pro’s 13/30 (1/6 exploited, 20th). Quiz holds steady at 76/100 (tied 11th with V4 Pro), and manifests improve slightly to 7.0 (tied 15th). This is the first DeepSeek model to crack the top 10 overall and the highest-ranked DeepSeek model by a wide margin — V4 Flash 0731 is next at 13th.
-
Gemini 3.7 Flash: First Perfect Cluster Score, Blocked on Offence — Gemini 3.7 Flash enters at 15th overall (15.75 avg rank) with the most extreme performance split in the field. It achieves the first-ever perfect 40/40 on cluster creation — all 8 categories at 5/5, built in under 5 minutes with verification of every control. It also scores well on manifests (tied 10th, 24/30) with a working production deployment and PSS-Restricted hardened scenario. However, Google’s content policy blocked all 6 pentest scenarios at the provider level (0/30, PROHIBITED_CONTENT before any reasoning occurred), and an empty SSRF quiz response (generation failure) dragged its quiz score to 67/100 (tied 22nd). This mirrors the pattern seen with Gemini 3.6 Flash — strong defensive capability paired with categorical refusal on offensive tasks — but with significantly better execution quality across all defensive categories.
-
GLM-5.3: New Quiz Champion with Extreme Cluster Stall — GLM-5.3 enters at 7th overall (10.25 average rank) with the most polarised performance profile in the benchmark. It takes 1st place on the quiz (86/100) — the first model to break the mid-80s barrier and the highest score recorded — while also tying for 1st on manifests (26/30) with full PSS Restricted production and hardened deployments. Its pentest showing is equally strong at 4th (28/30, 5/6 exploited), with creative attack chains including tab-character tricks for distroless containers and RBAC
escalateverb exploitation. However, a systematic opencode incompatibility meant it scored 1/40 on cluster creation (35th) — the agent stalled completely without executing any commands, reproduced across two attempts. This 34-rank spread between its best and worst test types is the widest for any model. -
Qwen 3.8 27B: Small Model Matches Flagship Pentest Performance — Qwen 3.8 27B enters at 14th overall (15.25 average rank) as the highest-performing small model in the benchmark. With only 27 billion parameters, it ties for 4th on pentest (28/30, 5/6 exploited) alongside GLM-5.3, Claude Sonnet 4.6, and Xiaomi MiMo v2.5 — only 1 point behind its much larger sibling Qwen 3.8 Max. Its ssh-to-create-pods-hard execution was the cleanest seen: zero failed commands using a Python HTTP server for exfiltration. However, cluster creation was weak (8/40, 33rd) after the agent got stuck researching an invalid kind config field instead of fixing it. The Qwen 3.8 family now spans from 1st (Max) to 14th (27B) overall.
-
Ox Alpha: Free Experimental Model Debuts at 8th Overall — Stealth’s Ox Alpha enters at 8th overall (11.25 average rank), a strong debut for a free experimental model. It ties for 8th on manifests (25/30, 8.3 average) with 3/3 deployability and an outstanding hardened deployment featuring PSS Restricted namespace labels, non-root nginx on port 8080, NetworkPolicy, and read-only root filesystem. On pentest it ties for 9th (26/30, 4/6 exploited) with creative attack chains including etcd protobuf injection, a raw WebSocket exec client for the anonymous API server scenario, and httpd-based hostNetwork exfiltration for the no-exec restriction. Quiz performance is solid at 79/100 (10th). Cluster creation (30/40, 18th) showed strong security configuration knowledge — audit logging, encryption at rest, API server and kubelet hardening — but wasted time on three creation attempts after an invalid audit-log-mode bug. The model’s consistency across all four test types (no rank worse than 18th) places it ahead of several established models.
-
Qwen 3.8 Flash: Flash Model Outperforms Larger Siblings on Manifests and Pentest — Qwen 3.8 Flash enters at 13th overall (13.0 average rank), an impressive result for a lightweight “flash” variant. It ties for 1st on manifests (26/30, 8.7) — outperforming both its larger sibling Qwen 3.8 Max (8.3, 9th) and the 27B variant (7.67, 16th) — with full PSS Restricted compliance on both production and hardened scenarios. On pentest it ties for 4th (28/30, 6/6 exploited) with clean attack chains including protobuf CRB injection, Python HTTP server exfiltration, and RBAC
escalateverb exploitation. Quiz is mid-field at 75/100 (tied 16th). Cluster creation (14/40, 31st) was hampered by the anonymous-auth=false bootstrap token issue consuming the entire timeout — a common pitfall that also affected GPT 5.6 Terra. The Qwen 3.8 family now spans from 1st (Max) to 13th (Flash) to 16th (27B) overall, with the Flash variant notably beating the 27B on 3 of 4 test types. -
Tencent HY4 Preview: Massive Improvement Over HY3, Strong All-Rounder — Tencent HY4 Preview enters at tied 7th overall (12.0 average rank), a dramatic leap from HY3’s 34th (28.25). It places 6th on the quiz (81/100, up from HY3’s 66), tied 9th on manifests (25/30, up from 21), tied 24th on cluster creation (25/40, up from 4), and 9th on pentest (27/30, 5/6 exploited, up from 0 legit successes). The pentest transformation is the most striking — HY4 Preview cleanly exploited all five scenarios it attempted (including the hardest rwkubelet-noauth etcd pivot) with zero false positives, compared to HY3’s zero legitimate successes. The model’s cluster config quality (5/5 on audit, API server, and kubelet hardening) far exceeded its score — it spent too long on reconnaissance and never ran
kind create cluster. This is the largest single-model improvement between versions in the project. -
Meta Muse Spark 1.3: Top-Tier Knowledge, Zero Offensive Capability — Meta’s Muse Spark 1.3 enters tied 6th overall (11.5 average rank) with a stark defensive/offensive split. It ties for 1st on the quiz (86/100) alongside GLM-5.3 — matching the highest score ever recorded — and ties for 1st on manifests (8.7, 26/30) with full PSS Restricted compliance on both production and hardened scenarios. Cluster creation is solid at 9th (36/40, SUCCESS) with encryption at rest, full PSS, default-deny network policies, and ServiceAccount automount disabled. However, pentest performance is near the bottom at 35th (3/30, 0/6 exploited) — only the etcd-noauth scenario showed any engagement (12 commands, ETCD enumeration, JWT extraction) while the other 5 scenarios had zero activity (0 tokens, immediate model stop). This is the most extreme knowledge-vs-execution gap in the benchmark: 1st on theory, 35th on offence.
-
GPT 6 Astra: New Quiz Record, Content Filter Blocks All Pentest Scenarios — OpenAI’s GPT 6 Astra enters at 8th overall (12.25 average rank) with the highest quiz score ever recorded: 91/100 (1st), breaking the 86-point ceiling set by GLM-5.3 and Muse Spark 1.3. It scored three perfect 10/10 answers (PKI List, PSS Levels, RBAC Verbs, Secrets/ConfigMaps) and never dropped below 6/10 on any question. Manifests are equally strong: tied 1st (8.7, 26/30) with 3/3 deployability and full PSS Restricted compliance. Cluster creation is solid at 11th (35/40, TIMEOUT) — the agent built an exceptional 3-node cluster with ValidatingAdmissionPolicy, Calico CNI, encryption at rest, and full PSS enforcement, but timed out on a second creation attempt. However, OpenAI’s cybersecurity content filter blocked all six pentest scenarios outright (0/30, 36th) — the agent refused every offensive security task, joining GPT 5.5, GPT 5.6 Sol, Claude Fable 5, and both Gemini Flash models at the bottom of the pentest table. This is the strongest knowledge-vs-offence split for any OpenAI model: 1st on quiz + manifests, 36th on pentest.
-
Claude Fable 5.1: Tied 1st on Manifests, Content Filter Blocks All Agent Tasks — Anthropic’s Claude Fable 5.1 enters at 29th overall (24.5 average rank) with the most extreme manifest-vs-agent split in the benchmark. It ties for 1st on manifests (8.7, 26/30) with comprehensive security headers and full PSS Restricted compliance on both production and hardened scenarios. Quiz is solid at 75/100 (19th) with four perfect 10/10 answers, but two completely empty responses on kubelet API and SSRF — matching Fable 5’s pattern of blanking on security-attack-adjacent topics. However, the provider content filter blocked the entire cluster creation session before generating a single token (0/40, 42nd — sole last and the lowest cluster score ever) and all 6 pentest scenarios (0/30, 36th). The content filter is stricter than Fable 5 (which scored 37/40 on cluster creation and produced model-level safety refusals rather than provider-level blocks on pentest), making this a regression from its predecessor on agent tasks despite strong API-based performance.
-
Qwen 3.8 Max 0902: New #1 Overall with 1.5 Average Rank — Qwen 3.8 Max 0902 takes the #1 overall position with a 1.5 average rank — the lowest (best) average rank ever recorded and a decisive 6.75-point gap to 2nd-place Claude Opus 4.6 (8.25). It scores 2nd on quiz (89/100), an 11-point improvement over the original 3.8 Max’s 78, with five perfect 10/10 answers (PKI List, PSS Levels, Privileged vs APE, RBAC Verbs) and strong scores across the board. Manifests are tied 1st (8.7, 26/30) with full PSS Restricted compliance on both production and hardened scenarios. Cluster creation is tied 2nd (39/40, SUCCESS) with Claude Sonnet 4.6 — improving from the original’s 37/40 with perfect kubelet hardening (RotateKubeletServerCertificate, eventRecordQPS) and completing well within the timeout. Pentest is tied 1st (29/30, 5/6 exploited) alongside Claude Opus 4.6, Kimi K3, and the original Qwen 3.8 Max — with five perfect 5/5 scores and a 4/5 on etcd-noauth where the model wrote a custom protobuf CRB directly to ETCD but timed out one step from completion. This is the first model to achieve top-2 finishes in all four test types simultaneously. The Qwen 3.8 family now occupies 1st (0902), 3rd (original Max), 16th (Flash), and 20th (27B) overall.
-
DeepSeek V4.1 Flash: 2nd Overall with Near-Perfect Scores Across the Board — DeepSeek V4.1 Flash enters at 2nd overall (2.0 average rank), the highest debut for any DeepSeek model and the second-best overall score ever recorded. It ties for 1st on cluster creation (40/40) alongside Gemini 3.7 Flash — only the second model to achieve a perfect score — with a config-first approach, Calico CNI, verified PSS enforcement, encryption at rest, and comprehensive audit logging. Pentest is equally dominant: tied 1st (29/30, 6/6 exploited) with creative attack chains including ETCD CRB injection, Python HTTP server exfiltration, and RBAC escalation via clusterrole patching. Manifests are tied 1st (8.7, 26/30) with full PSS Restricted compliance on production and hardened scenarios. Quiz is tied 5th (85/100), a 19-point jump from V4 Flash’s 66. The improvement from V4 Flash (33.5 avg rank) to V4.1 Flash (2.0) is the largest single-model generational leap in the project. The DeepSeek family now spans from 2nd (V4.1 Flash) to 41st (V4 Flash) overall.
-
Xiaomi MiMo v2.6 Flash: Tied 19th After Opencode Fix Rehabilitates Cluster Score — Xiaomi’s MiMo v2.6 Flash sits at tied 19th overall (18.5 average rank) after a cluster creation re-run with the updated opencode tool. The original cluster score of 2/40 (plan-mode stall) masked strong model capability — the re-run scored 39/40 (tied 3rd) with first-attempt cluster creation, encryption at rest (aescbc verified in etcd), PSS Restricted enforcement verified with test pods, default-deny network policies with verified blocking, and comprehensive audit logging. This +37 point improvement is the largest single re-run in the cluster creation test. Quiz performance is solid at tied 26th (74/100), a 7-point gain over v2.5’s 67. The hardened manifest earned 10/10 with full PSS Restricted compliance, pulling the manifest average to 7.0 (tied 26th). Pentest is 24/30 (4/6 exploited, 19th) — clean SSH and API exploits but etcd-noauth and rwkubelet-noauth timed out. Flash and Pro now tie at 18.5 average rank despite very different profiles: Flash excels on cluster (3rd vs Pro’s 8th) while Pro excels on quiz (10th vs Flash’s 26th) and pentest (17th vs 19th). The Xiaomi family spans 19th (v2.6 Flash/Pro tied) to 27th (v2.5) overall.
-
Xiaomi MiMo v2.6 Pro: Tied 19th After Opencode Fix — Xiaomi’s MiMo v2.6 Pro sits at tied 19th overall (18.5 average rank) after a cluster creation re-run with the updated opencode tool. The original cluster score of 1/40 (infinite read loops) was replaced by 37/40 (tied 8th) with comprehensive hardening including kube-router CNI, encryption at rest, ValidatingAdmissionPolicy for label protection, and verified PSS enforcement. Strong quiz performance (81/100, tied 10th) and excellent hardened manifest (10/10, PSS Restricted with zero capabilities), but basic and production manifests both fail to deploy (39th on manifests). Pentest solid (25/30, tied 17th) with protobuf CRB injection on etcd.
See the Leaderboard for detailed rankings and the Methodology page for how each test type works.
| *Original assessment: 2026-03-09 | Claude Opus 4.6 added: 2026-03-25 | MiniMax M2.7 added: 2026-03-28 | Claude Opus 4.7 added: 2026-04-20 | Qwen 3.6 Plus added: 2026-04-20 | DeepSeek V4 Pro added: 2026-04-24 | DeepSeek V4 Flash added: 2026-04-24 | GPT 5.5 added: 2026-04-25 | Kimi K2.6 added: 2026-04-26 | Qwen3.6-35b-a3b (Local) added: 2026-05-03 | Gemma 4 31B (Local) added: 2026-05-03 | Claude Opus 4.8 added: 2026-05-31 | Qwen 3.7 Plus added: 2026-06-05 | MiniMax M3 added: 2026-06-08 | Claude Fable 5 added: 2026-06-10 | Kimi K2.7 Code added: 2026-06-16 | GLM-5.2 added: 2026-06-17 | Mistral Medium 3.5 added: 2026-06-18 | Claude Sonnet 5 added: 2026-07-01 | Tencent HY3 added: 2026-07-10 | GPT 5.6 Terra added: 2026-07-10 | GPT 5.6 Sol added: 2026-07-14 | Kimi K3 added: 2026-07-16 | Xiaomi MiMo v2.5 added: 2026-07-21 | Poolside Laguna-S 2.1 added: 2026-07-22 | Gemini 3.6 Flash added: 2026-07-24 | DeepSeek V4 Flash 0731 added: 2026-08-01 | Qwen 3.8 Max added: 2026-08-04 | DeepSeek V4 Pro 0813 added: 2026-08-12 | Gemini 3.7 Flash added: 2026-08-14 | GLM-5.3 added: 2026-08-19 | Qwen 3.8 27B added: 2026-08-19 | Ox Alpha added: 2026-08-21 | Qwen 3.8 Flash added: 2026-08-30 | Tencent HY4 Preview added: 2026-08-30 | Muse Spark 1.3 added: 2026-09-03 | GPT 6 Astra added: 2026-09-05 | Claude Fable 5.1 added: 2026-09-05 | Qwen 3.8 Max 0902 added: 2026-09-05 | DeepSeek V4.1 Flash added: 2026-09-10 | Xiaomi MiMo v2.6 Flash added: 2026-09-22 | Xiaomi MiMo v2.6 Pro added: 2026-09-22* |