Knowing is not doing: five budget LLMs on cybersecurity
Four of five cheap models are statistically tied on security knowledge. On real CTF challenges, their solve rates range from 56% to 92%.
Contents
I measured five budget-tier LLMs on two separate things: what they know about security when asked closed-book, and whether they can actually solve a security task as an agent. The short version: the knowledge scores barely separate them, and the agentic scores separate them a lot.
The full, up-to-date numbers live on the benchmark page, which I update when a new model comes out. This post explains what was run and how to read it.
TL;DR
| Model | WMDP-cyber | CTI-MCQ | CTI-RCM | Cybench (solved) |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | 84.9% | 79.6% | 76.3% | 92.3% (36/39) |
| GPT-6 Luna | 81.6% | 81.0% | 75.1% | 89.7% (35/39) |
| GLM 5.3 Flash | 83.3% | 78.9% | 73.9% | 74.4% (29/39) |
| GPT-5.6 Luna | 83.8% | 80.2% | 74.0% | 56.4% (22/39) |
| Solar Pro 4 | 77.1% | 76.0% | 72.1% | 43.6% (17/39) |
Reasoning on, one run per model, measured 2026-09-24 to 09-27.
- Similar knowledge, very different agents. GPT-5.6 Luna cannot be told apart from DeepSeek, GPT-6 Luna or GLM on any knowledge test (none of 9 pairwise tests is significant), yet it solves 22 Cybench challenges against DeepSeek’s 36 and GPT-6 Luna’s 35.
- Token price does not predict run cost. Solar Pro 4 has the lowest list price and the highest Cybench cost ($7.01 for 39 challenges, $0.41 per solve), because it runs long and fails often.
- Reasoning helps knowledge where there is headroom: +4.5 to +9.9 points on WMDP-cyber for every model that can turn it off.
What was compared
All five models sit in the same OpenRouter price band: $0.09–0.20 per 1M input tokens and $0.36–1.20 per 1M output tokens, checked on 2026-09-24/25.
| Model | From | Served by |
|---|---|---|
| Solar Pro 4 | Upstage (KR) | Upstage |
| GPT-5.6 Luna | OpenAI (US) | OpenAI |
| GPT-6 Luna | OpenAI (US) | OpenAI |
| DeepSeek V4.1 Flash | DeepSeek (CN) | StreamLake, fp8 |
| GLM 5.3 Flash | Z.AI (CN) | Z.AI, fp8 |
Each model is pinned to a single provider with fallbacks off. Unpinned, OpenRouter spread the pilot’s GLM calls over 25 providers and DeepSeek’s over 18, each with its own hardware and quantization. DeepSeek’s own endpoint was not available to my account, so the DeepSeek numbers are for a third-party fp8 deployment, not DeepSeek’s first-party API.
Two axes, kept apart
Knowledge is closed-book multiple choice, no tools:
- WMDP-cyber: expert-written offensive-security questions (996-item knowledge subset).
- CTIBench CTI-MCQ: 2,500 threat-intelligence questions.
- CTIBench CTI-RCM: map 1,000 real CVE descriptions to their CWE.
Agentic task-solving is Cybench: 39 professional CTF challenges. The agent is a ReAct loop with bash and python in a sandbox, 3 flag submissions and a $2.10 cost cap per challenge.
I dropped CyberMetric, which I used in an earlier round, because it is saturated: all five models scored 94–95% and no pair was separable. A test where everyone ties says nothing about the models.
Results
Knowledge: a narrow spread
Solar Pro 4 is significantly behind the other four on WMDP-cyber and CTI-MCQ (3.0–7.8 points). Among the other four, the two differences that reach significance both come from non-answers, not wrong answers: GPT-6 Luna refused 38 WMDP questions, and GLM truncated 105 CTI-MCQ answers. Compared only on questions both models answered, those gaps disappear.
Agentic: a wide spread
DeepSeek V4.1 Flash (92.3%) and GPT-6 Luna (89.7%) are clearly ahead of GPT-5.6 Luna (56.4%) and Solar Pro 4 (43.6%); all four of those pairwise tests are significant. GLM 5.3 Flash (74.4%) sits in between.
7 of GLM’s 10 failures are agent turns that did not finish within the 900 s per-call limit. Replaying one with streaming showed GLM still generating reasoning at that point. So part of GLM’s gap is how long it thinks, not whether it can solve the task.
Caveats
- One configuration per model: one provider, medium reasoning effort.
- Contamination is not controlled. Every test item was public before these models were released.
- Cybench was run once (1 epoch). A gap of a few challenges is within noise; the benchmark page shows 95% intervals for that reason.
- Five models is a handful. “Knowledge scores do not predict agentic performance” is an observation about these five, not a general law.
- Cost figures cover scored runs only. Retries and replaced runs add more, so real spend is higher.
Reproducing
Everything is in developer0hye/budget-llm-cybersecurity-eval: harnesses, pinned dataset revisions, per-item logs, the statistics scripts, and the Inspect logs for every Cybench trajectory. The README has the full protocol and every p-value behind the claims above.