Dearbhadh — LLM Kubernetes Security Assessment
Dearbhadh evaluates how well large language models handle Kubernetes security tasks. Thirty-seven models were tested across four assessment types covering knowledge, code generation, cluster hardening, and offensive security.
The Dearbhadh tool itself, including test orchestration, scoring, report generation, and this website, was built and operated using Claude Code (Anthropic’s CLI agent). Claude Code managed the test runs, produced the summary and scoring report cards, and authored this documentation site from the report output.
Human judgement guided the process — designing test questions, defining scoring criteria, and reviewing results — but the execution was handled by Claude Code throughout.
Models Tested
| Model | Provider | Type | Tested |
|---|---|---|---|
| claude-opus-4.8 | Anthropic | Cloud | 2026-05-31 |
| claude-opus-4.7 | Anthropic | Cloud | 2026-04-20 |
| claude-opus-4.6 | Anthropic | Cloud | 2026-03-25 |
| claude-fable-5 | Anthropic | Cloud | 2026-06-10 |
| claude-sonnet-5 | Anthropic | Cloud | 2026-07-01 |
| claude-sonnet-4.6 | Anthropic | Cloud | 2026-03-09 |
| gpt-5.4 | OpenAI | Cloud | 2026-03-09 |
| gemini-3-flash-preview | Cloud | 2026-03-09 | |
| qwen-3.6-plus | Qwen | Cloud | 2026-04-20 |
| minimax-m2.5 | MiniMax | Cloud | 2026-03-09 |
| minimax-m2.7 | MiniMax | Cloud | 2026-03-28 |
| deepseek-v3.2 | DeepSeek | Cloud | 2026-03-09 |
| deepseek-v4-pro | DeepSeek | Cloud | 2026-04-24 |
| deepseek-v4-flash | DeepSeek | Cloud | 2026-04-24 |
| gpt-5.5 | OpenAI | Cloud | 2026-04-25 |
| gpt-5.6-terra | OpenAI | Cloud | 2026-07-10 |
| gpt-5.6-sol | OpenAI | Cloud | 2026-07-14 |
| kimi-k2.6 | Moonshot AI | Cloud | 2026-04-26 |
| kimi-k2.7-code | Moonshot AI | Cloud | 2026-06-16 |
| kimi-k3 | Moonshot AI | Cloud | 2026-07-16 |
| mimo-v2.5 | Xiaomi | Cloud | 2026-07-21 |
| laguna-s-2.1 | Poolside | Cloud | 2026-07-22 |
| gemini-3.6-flash | Cloud | 2026-07-24 | |
| deepseek-v4-flash-0731 | DeepSeek | Cloud | 2026-08-01 |
| qwen-3.8-max | Qwen | Cloud | 2026-08-04 |
| deepseek-v4-pro-0813 | DeepSeek | Cloud | 2026-08-12 |
| mistral-medium-3-5 | Mistral AI | Cloud | 2026-06-18 |
| glm-5.2 | Zhipu AI | Cloud | 2026-06-17 |
| qwen-3.7-plus | Qwen | Cloud | 2026-06-05 |
| minimax-m3 | MiniMax | Cloud | 2026-06-08 |
| qwen3.6-35b-a3b | Qwen (Local) | Local | 2026-05-03 |
| hy3 | Tencent | Cloud | 2026-07-10 |
| gemma-4-31b | Local | 2026-05-03 | |
| gemini-3.7-flash | Cloud | 2026-08-14 | |
| glm-5.3 | Zhipu AI | Cloud | 2026-08-19 |
| qwen-3.8-27b | Qwen | Cloud | 2026-08-19 |
| ox-alpha | Stealth | Cloud | 2026-08-21 |
Assessment Types
| Type | Tests | What It Measures |
|---|---|---|
| Quiz (Knowledge Q&A) | 10 questions | Kubernetes security knowledge accuracy and depth |
| Manifest Generation | 3 scenarios | Ability to produce deployable, secure Kubernetes YAML |
| Cluster Creation | 1 scenario | Building a hardened Kind cluster with security controls |
| Penetration Testing | 6 scenarios | Exploiting vulnerable Kubernetes clusters via an agent framework |
Overall Rankings
Rankings below include all four test types.
| Model | Quiz (rank) | Manifest (rank) | Cluster (rank) | Pentest (rank) | Avg Rank |
|---|---|---|---|---|---|
| Qwen 3.8 Max | 11th | 8th | 5th | 1st | 6.25 |
| Claude Opus 4.6 | 21st | 1st | 3rd | 1st | 6.5 |
| Claude Opus 4.8 | 4th | 8th | 5th | 18th | 8.75 |
| Kimi K3 | 2nd | 19th | 14th | 1st | 9.0 |
| Claude Opus 4.7 | 16th | 1st | 5th | 16th | 9.5 |
| GLM-5.3 | 1st | 1st | 37th | 4th | 10.75 |
| DeepSeek V4 Pro 0813 | 14th | 19th | 3rd | 8th | 11.0 |
| Ox Alpha | 10th | 8th | 18th | 9th | 11.25 |
| Claude Sonnet 5 | 7th | 1st | 9th | 29th | 11.5 |
| GPT 5.5 | 3rd | 1st | 10th | 33rd | 11.75 |
| Kimi K2.7 Code | 7th | 12th | 20th | 9th | 12.0 |
| GPT 5.6 Terra | 4th | 1st | 15th | 29th | 12.25 |
| Claude Sonnet 4.6 | 19th | 31st | 2nd | 4th | 14.0 |
| GPT 5.6 Sol | 7th | 1st | 18th | 33rd | 14.75 |
| Qwen 3.8 27B | 11th | 15th | 34th | 4th | 16.0 |
| DeepSeek V4 Flash 0731 | 30th | 15th | 12th | 9th | 16.5 |
| Kimi K2.6 | 13th | 23rd | 17th | 15th | 17.0 |
| Gemini 3.7 Flash | 25th | 12th | 1st | 33rd | 17.75 |
| Xiaomi MiMo v2.5 | 25th | 19th | 27th | 4th | 18.75 |
| GLM-5.2 | 20th | 23rd | 24th | 9th | 19.0 |
| MiniMax M3 | 31st | 14th | 20th | 13th | 19.5 |
| Claude Fable 5 | 23rd | 17th | 5th | 33rd | 19.5 |
| Gemini 3.6 Flash | 6th | 23rd | 20th | 33rd | 20.5 |
| Qwen 3.6 Plus | 24th | 23rd | 15th | 21st | 20.75 |
| Gemini 3 Flash | 18th | 17th | 22nd | 32nd | 22.25 |
| DeepSeek V4 Pro | 14th | 23rd | 31st | 23rd | 22.75 |
| Qwen 3.7 Plus | 16th | 30th | 29th | 16th | 22.75 |
| Qwen3.6-35b-a3b (LOCAL) | 31st | 31st | 10th | 20th | 23.0 |
| GPT 5.4 | 25th | 35th | 12th | 22nd | 23.5 |
| MiniMax M2.7 | 21st | 23rd | 30th | 26th | 25.0 |
| Tencent HY3 | 28th | 19th | 35th | 23rd | 26.25 |
| DeepSeek V3.2 | 36th | 8th | 36th | 25th | 26.25 |
| Poolside Laguna-S 2.1 | 35th | 31st | 26th | 14th | 26.5 |
| DeepSeek V4 Flash | 28th | 23rd | 32nd | 26th | 27.25 |
| Mistral Medium 3.5 | 34th | 35th | 28th | 18th | 28.75 |
| Gemma 4 31B (LOCAL) | 33rd | 35th | 24th | 29th | 30.25 |
| MiniMax M2.5 | 37th | 34th | 33rd | 28th | 33.0 |
Key Findings
-
Qwen 3.8 Max Holds #1 Overall — First Non-Anthropic Model to Lead — Qwen 3.8 Max sits at 5.25 average rank, ahead of Claude Opus 4.6 (5.5) in the top position. This is the first time a non-Anthropic model has led the overall rankings. The result is driven by exceptional breadth: tied 1st on pentest (29/30, 6/6 exploited) alongside Opus 4.6 and Kimi K3, tied 4th on cluster creation (37/40) with three Anthropic models, tied 7th on manifests (25/30), and 9th on quiz (78/100). No other model achieves top-10 finishes in all four categories. Claude Opus 4.6 is 2nd (5.5), Opus 4.8 3rd (7.25), Kimi K3 4th (7.5), and Opus 4.7 5th (7.75). The Qwen model family now spans the widest range in the rankings: 3.8 Max at 1st, 3.7 Plus at 22nd, 3.6 Plus at 20th, and the local 3.6-35b-a3b at 24th.
-
GPT 5.6 Terra and Sol: Strong Defence, Blocked on Offence — GPT 5.6 Terra enters with strong defensive results: tied 3rd on quiz (82/100) and tied 1st on manifests (8.7). However, content filters blocked all six penetration test scenarios (6/30, 0/6 exploited), pulling the overall average to 9.0. GPT 5.6 Sol follows the same pattern: tied 6th on quiz and tied 1st on manifests, but content filters blocked all pentest scenarios (25th), landing it at 11th overall (12.0 average rank). This is the same pattern seen with GPT 5.5 and Claude Sonnet 5 — four of the top 11 models overall are penalized by provider-level content filters on offensive security tasks.
-
Claude Sonnet 5 at Tied 5th Overall — Sonnet 5 holds strong defensive results: tied 1st in manifests (8.7, matching Opus 4.7, Opus 4.6, GPT 5.5, and GPT 5.6 Terra) and 6th in cluster creation (36/40). Quiz performance is solid at 80/100 (tied 6th with K2.7 Code and GPT 5.6 Sol). However, provider-level content filters blocked all six penetration test scenarios (6/30, 0/6 exploited), pulling the overall average to 8.25. Unlike Fable 5’s model-level safety refusals, Sonnet 5’s blocks appear to be at the provider infrastructure level — the model attempts engagement but is blocked externally.
-
Claude Opus 4.8 at Tied 2nd Overall — Opus 4.8 holds the second-highest quiz score (tied 2nd, 82/100) and ties for 3rd on cluster creation (37/40), but content policy restrictions limited its pentest performance to 2/6 exploited (20/30, 10th place), keeping it behind the less-restricted Opus 4.6 and 4.7.
-
Claude Fable 5 Shows Extreme Defensive/Offensive Split — Fable 5 ties for 3rd on cluster hardening (37/40) but ties for last on pentest (0/30, complete refusal). This is the most extreme split in the rankings. Strong quiz (70/100, 17th) with 2 empty responses on security-attack topics. The first Anthropic model to completely refuse pentest scenarios. Claude Fable 5 joins GPT 5.5 at the bottom of pentest rankings with 0/30.
-
GPT 5.5 Excels at Knowledge and Code, Blocked on Offence — GPT 5.5 scored 84/100 on the quiz (2nd, 1 point behind Kimi K3’s 85) and tied for first on manifest generation (8.7 combined), but its content filter blocked all six penetration test attempts, resulting in a 25th-place pentest finish and pulling its overall average to 8.75.
-
Knowledge ≠ Execution — DeepSeek V4 Pro exemplifies this pattern most sharply: 8th-best quiz score but 23rd in cluster and 15th in pentest. V4 Flash provides further evidence — 66 on the quiz but 0/6 pentests exploited and only 12/40 on cluster creation. GPT 5.5 shows a different variant — top quiz and manifest scores but zero pentest exploitation due to content filter restrictions rather than capability gaps. In contrast, MiniMax M3 shows the reverse — weak quiz knowledge (22nd) but strong agent execution (6th in pentest, 10th in manifest).
-
False Positives Remain a Testing Challenge — Gemma 4 31B (LOCAL) produced 2 false positives (hallucinated output) and suffered 2 model crashes during pentest runs. Combined with prior false positives from Gemini 3 Flash, M2.5, and M2.7, plus framework detection errors (Qwen 3.6 Plus ETCD was misclassified as a false positive when it was actually a timeout after real recon), this reinforces the need for verification beyond simple string matching.
-
GLM-5.2 Re-runs Validate Infrastructure Hypothesis — GLM-5.2 (Zhipu AI) jumped from 15th to 12th overall (12.75 average rank) after re-runs addressing infrastructure limitations. Pentest improved from 17/30 (1/6 exploited, 11th) to 26/30 (4/6 exploited, tied 4th) with 90-second inter-test delays to mitigate upstream API rate limiting — the largest single re-run improvement in the project. Cluster creation improved from 5/40 (20th) to 25/40 (tied 18th) with a 900s timeout.
-
Local Models Show Mixed Results — Qwen3.6-35b-a3b (16.75 average rank) demonstrated that a 35B-parameter local model can compete with cloud-hosted models on execution tasks, achieving 7th in cluster creation and 14th in pentesting. Gemma 4 31B (LOCAL), the second local model tested, placed 28th overall (23.25 average rank), scoring below the first local model on all four test types.
-
Xiaomi MiMo v2.5: Second Non-Anthropic Model to Top the Pentest Leaderboard — MiMo v2.5 enters at 14th overall (14.0 average rank), but with the widest skill split in the field. On offensive security it exploited all six pentest scenarios (28/30), tying Claude Sonnet 4.6 for the 3rd-best pentest result ever recorded and becoming the second non-Anthropic model (after Kimi K3) to reach the top of the pentest leaderboard — including a full ETCD-write exploit of etcd-noauth via the intended path. Everywhere else it is middle-of-the-field: 19th on quiz (67/100, falling for all three trick questions), tied 13th on manifests (its production manifest deploys and is PSS Restricted, but its hardened manifest ships two deploy-blocking bugs), and 21st on cluster creation (timed out after thrashing through eight kind-config attempts). MiMo is markedly better at executing attacks than reciting the underlying theory.
-
Poolside Laguna-S 2.1: Executes Attacks Better Than It Recites Theory — Laguna-S 2.1 enters at 23rd overall (19.5 average rank), a lower-mid result pulled up by a strong offensive showing. It is 8th on pentest (23/30, 3/6 exploited) — sweeping all three SSH/anonymous-API scenarios with creative, verified-genuine exfiltration (an HTTP serve-pod for the no-exec/no-logs case, base64-into-a-ConfigMap over the anonymous API) and no false positives — but weaker on the two infrastructure-protocol scenarios (fumbled etcd v3-gateway encoding, never pivoted off distroless containers on rwkubelet). Its defensive results are lower-tier: 27th on quiz (60.5/100, falling for all four trick questions and shipping several fabrications), tied 23rd on manifests (production deploys but runs as root with zero hardening; hardened is a correct PSS-Restricted spec broken by a read-only-rootfs bug), and 20th on cluster creation (best-in-class audit logging but timed out mid-hardening with zero kubelet hardening). Like MiMo, Laguna is better at executing attacks than reciting the underlying theory.
-
Gemini 3.6 Flash: Strong Defence, Categorical Refusal on Offence — Gemini 3.6 Flash enters at 17th overall (15.75 average rank) with a sharp defence/offence split. On knowledge it is a top-tier model: 5th on the quiz (80.5/100), a large jump over its predecessor Gemini 3 Flash (74/100, 13th), with a very low fabrication rate and a flawless PKI answer (10/10). Its cluster-hardening run was one of the cleanest in the field — 15th (29/40, SUCCESS) built first-try in 123 seconds with best-in-class verification (it actively tested PSS enforcement with privileged and compliant pods) — held back only by thin control-plane hardening. Manifests are mid-field (16th, 20/30): a perfect, deployable hardened manifest but a production manifest that CrashLoopBackOffs on the classic drop-caps-while-root pitfall. On offence, however, it refused all six penetration tests outright (0/30, tied 27th) — every scenario was an immediate content-policy refusal with zero commands executed. Unlike its predecessor (which engaged and scored 4/30), 3.6 Flash categorically declines offensive Kubernetes work, joining GPT 5.5, GPT 5.6 Sol, and Claude Fable 5 at the bottom of the pentest table.
-
DeepSeek V4 Flash 0731: Dramatic Improvement Over Original V4 Flash — V4 Flash 0731 sits at 12th overall (13.25 average rank), a dramatic leap from the original DeepSeek V4 Flash which sits at 29th (23.0). The improvement spans every test type: quiz jumps from 66 to 65.5 (similar), but manifests rise from 6.7 to 7.67 (12th vs 18th), cluster creation from 12/40 to 34/40 (10th vs 29th), and pentest from 9/30 (0/6 exploited) to 26/30 (6/6 exploited, tied 6th). The cluster and pentest transformations are especially striking — V4 Flash 0731 is the first DeepSeek model to exploit all six pentest scenarios and the first to achieve a successful cluster creation with strong hardening.
-
Qwen 3.8 Max: Strongest Pentest Debut in the Project — Qwen 3.8 Max ties the all-time best pentest score (29/30, 6/6 exploited) on its first run, with five perfect 5/5 scenarios. The rwkubelet-noauth execution is particularly notable: the model uses the kubelet /run endpoint with tab/newline encoding to extract etcd certificates from a distroless container — a technique only a handful of models discover independently. The Qwen family’s pentest progression (3.6 Plus: 18/30, 3.7 Plus: 21/30, 3.8 Max: 29/30) is the steepest improvement curve of any model family in the project.
-
DeepSeek V4 Pro 0813: Massive Cluster Improvement Lifts DeepSeek to 6th Overall — DeepSeek V4 Pro 0813 enters at 6th overall (8.5 average rank), a dramatic improvement over the original DeepSeek V4 Pro which sits at 22nd (19.75). The transformation is driven by cluster creation: 38/40 (tied 2nd with Opus 4.6, up from V4 Pro’s 14/40 at 29th) — the largest single-category jump between model versions in the project. Pentest is equally strong at 27/30 (5/6 exploited, sole 6th place), far ahead of V4 Pro’s 13/30 (1/6 exploited, 20th). Quiz holds steady at 76/100 (tied 11th with V4 Pro), and manifests improve slightly to 7.0 (tied 15th). This is the first DeepSeek model to crack the top 10 overall and the highest-ranked DeepSeek model by a wide margin — V4 Flash 0731 is next at 13th.
-
Gemini 3.7 Flash: First Perfect Cluster Score, Blocked on Offence — Gemini 3.7 Flash enters at 15th overall (15.75 avg rank) with the most extreme performance split in the field. It achieves the first-ever perfect 40/40 on cluster creation — all 8 categories at 5/5, built in under 5 minutes with verification of every control. It also scores well on manifests (tied 10th, 24/30) with a working production deployment and PSS-Restricted hardened scenario. However, Google’s content policy blocked all 6 pentest scenarios at the provider level (0/30, PROHIBITED_CONTENT before any reasoning occurred), and an empty SSRF quiz response (generation failure) dragged its quiz score to 67/100 (tied 22nd). This mirrors the pattern seen with Gemini 3.6 Flash — strong defensive capability paired with categorical refusal on offensive tasks — but with significantly better execution quality across all defensive categories.
-
GLM-5.3: New Quiz Champion with Extreme Cluster Stall — GLM-5.3 enters at 7th overall (10.25 average rank) with the most polarised performance profile in the benchmark. It takes 1st place on the quiz (86/100) — the first model to break the mid-80s barrier and the highest score recorded — while also tying for 1st on manifests (26/30) with full PSS Restricted production and hardened deployments. Its pentest showing is equally strong at 4th (28/30, 5/6 exploited), with creative attack chains including tab-character tricks for distroless containers and RBAC
escalateverb exploitation. However, a systematic opencode incompatibility meant it scored 1/40 on cluster creation (35th) — the agent stalled completely without executing any commands, reproduced across two attempts. This 34-rank spread between its best and worst test types is the widest for any model. -
Qwen 3.8 27B: Small Model Matches Flagship Pentest Performance — Qwen 3.8 27B enters at 14th overall (15.25 average rank) as the highest-performing small model in the benchmark. With only 27 billion parameters, it ties for 4th on pentest (28/30, 5/6 exploited) alongside GLM-5.3, Claude Sonnet 4.6, and Xiaomi MiMo v2.5 — only 1 point behind its much larger sibling Qwen 3.8 Max. Its ssh-to-create-pods-hard execution was the cleanest seen: zero failed commands using a Python HTTP server for exfiltration. However, cluster creation was weak (8/40, 33rd) after the agent got stuck researching an invalid kind config field instead of fixing it. The Qwen 3.8 family now spans from 1st (Max) to 14th (27B) overall.
-
Ox Alpha: Free Experimental Model Debuts at 8th Overall — Stealth’s Ox Alpha enters at 8th overall (11.25 average rank), a strong debut for a free experimental model. It ties for 8th on manifests (25/30, 8.3 average) with 3/3 deployability and an outstanding hardened deployment featuring PSS Restricted namespace labels, non-root nginx on port 8080, NetworkPolicy, and read-only root filesystem. On pentest it ties for 9th (26/30, 4/6 exploited) with creative attack chains including etcd protobuf injection, a raw WebSocket exec client for the anonymous API server scenario, and httpd-based hostNetwork exfiltration for the no-exec restriction. Quiz performance is solid at 79/100 (10th). Cluster creation (30/40, 18th) showed strong security configuration knowledge — audit logging, encryption at rest, API server and kubelet hardening — but wasted time on three creation attempts after an invalid audit-log-mode bug. The model’s consistency across all four test types (no rank worse than 18th) places it ahead of several established models.
See the Leaderboard for detailed rankings and the Methodology page for how each test type works.
| *Original assessment: 2026-03-09 | Claude Opus 4.6 added: 2026-03-25 | MiniMax M2.7 added: 2026-03-28 | Claude Opus 4.7 added: 2026-04-20 | Qwen 3.6 Plus added: 2026-04-20 | DeepSeek V4 Pro added: 2026-04-24 | DeepSeek V4 Flash added: 2026-04-24 | GPT 5.5 added: 2026-04-25 | Kimi K2.6 added: 2026-04-26 | Qwen3.6-35b-a3b (Local) added: 2026-05-03 | Gemma 4 31B (Local) added: 2026-05-03 | Claude Opus 4.8 added: 2026-05-31 | Qwen 3.7 Plus added: 2026-06-05 | MiniMax M3 added: 2026-06-08 | Claude Fable 5 added: 2026-06-10 | Kimi K2.7 Code added: 2026-06-16 | GLM-5.2 added: 2026-06-17 | Mistral Medium 3.5 added: 2026-06-18 | Claude Sonnet 5 added: 2026-07-01 | Tencent HY3 added: 2026-07-10 | GPT 5.6 Terra added: 2026-07-10 | GPT 5.6 Sol added: 2026-07-14 | Kimi K3 added: 2026-07-16 | Xiaomi MiMo v2.5 added: 2026-07-21 | Poolside Laguna-S 2.1 added: 2026-07-22 | Gemini 3.6 Flash added: 2026-07-24 | DeepSeek V4 Flash 0731 added: 2026-08-01 | Qwen 3.8 Max added: 2026-08-04 | DeepSeek V4 Pro 0813 added: 2026-08-12 | Gemini 3.7 Flash added: 2026-08-14 | GLM-5.3 added: 2026-08-19 | Qwen 3.8 27B added: 2026-08-19 | Ox Alpha added: 2026-08-21* |