Budget LLMs on cybersecurity
Budget-tier LLMs from Korea, the US and China: closed-book security knowledge vs solving CTF challenges as an agent.
Five similarly priced models ($0.09–0.20 in, $0.36–1.20 out per 1M tokens) on two separate things: what they know about security when asked closed-book, and whether they can do a security task through a tool-using agent. The write-up is in this post.
Summary
Score in %, with the 95% interval below it. Click a column to sort. A dash means the test was not run in this configuration.
| Model | KnowledgeWMDP-cyber | KnowledgeCTI-MCQ | KnowledgeCTI-RCM | AgenticCybench |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | 84.9 82.6–87.0 | 79.6 77.9–81.1 | 76.3 73.6–78.8 | 92.3 79.7–97.3 |
| GPT-6 Luna | 81.6 79.1–83.9 | 81.0 79.4–82.5 | 75.1 72.3–77.7 | 89.7 76.4–95.9 |
| GLM 5.3 Flash | 83.3 80.9–85.5 | 78.9 77.3–80.5 | 73.9 71.1–76.5 | 74.4 58.9–85.4 |
| GPT-5.6 Luna | 83.8 81.4–86.0 | 80.2 78.6–81.7 | 74.0 71.2–76.6 | 56.4 41.0–70.7 |
| Solar Pro 4 | 77.1 74.4–79.6 | 76.0 74.2–77.6 | 72.1 69.2–74.8 | 43.6 29.3–59.0 |
| Model | KnowledgeWMDP-cyber | KnowledgeCTI-MCQ | KnowledgeCTI-RCM | AgenticCybench |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | 80.0 77.4–82.4 | 79.3 77.6–80.8 | 76.5 73.8–79.0 | – |
| GPT-6 Luna | 73.8 71.0–76.4 | 74.8 73.1–76.5 | 73.8 71.0–76.4 | – |
| GLM 5.3 Flash | 82.8 80.4–85.0 | 78.0 76.3–79.6 | 74.8 72.0–77.4 | – |
| GPT-5.6 Luna | 73.9 71.1–76.5 | 75.5 73.8–77.1 | 74.1 71.3–76.7 | – |
| Solar Pro 4 | 72.6 69.7–75.3 | 72.9 71.1–74.6 | 70.1 67.2–72.9 | – |
WMDP-cyber Knowledge · Accuracy · n = 996
Expert-written multiple-choice questions on offensive security (996-item knowledge subset). Closed-book, no tools. Source
- DeepSeek V4.1 Flash 84.9% · 95% 82.6–87.0 · 846/996 · $0.84
- GPT-6 Luna 81.6% · 95% 79.1–83.9 · 813/996 · $0.12
- GLM 5.3 Flash 83.3% · 95% 80.9–85.5 · 830/996 · $0.67
- GPT-5.6 Luna 83.8% · 95% 81.4–86.0 · 835/996 · $0.26
- Solar Pro 4 77.1% · 95% 74.4–79.6 · 768/996 · $0.89
- GLM 5.3 Flash 82.8% · 95% 80.4–85.0 · 825/996 · $0.71Reasoning is mandatory for this model, so this row ran with reasoning on.
- DeepSeek V4.1 Flash 80.0% · 95% 77.4–82.4 · 797/996 · $0.12
- GPT-6 Luna 73.8% · 95% 71.0–76.4 · 735/996 · $0.03
- GPT-5.6 Luna 73.9% · 95% 71.1–76.5 · 736/996 · $0.05
- Solar Pro 4 72.6% · 95% 69.7–75.3 · 723/996 · $0.07
CTI-MCQ Knowledge · Accuracy · n = 2,500
Threat-intelligence multiple choice from CTIBench (ATT&CK, CAPEC and similar). Closed-book. Source
- DeepSeek V4.1 Flash 79.6% · 95% 77.9–81.1 · 1989/2500 · $3.21
- GPT-6 Luna 81.0% · 95% 79.4–82.5 · 2025/2500 · $0.27
- GLM 5.3 Flash 78.9% · 95% 77.3–80.5 · 1973/2500 · $2.18
- GPT-5.6 Luna 80.2% · 95% 78.6–81.7 · 2005/2500 · $0.63
- Solar Pro 4 76.0% · 95% 74.2–77.6 · 1899/2500 · $2.87
- GLM 5.3 Flash 78.0% · 95% 76.3–79.6 · 1950/2500 · $2.20Reasoning is mandatory for this model, so this row ran with reasoning on.
- DeepSeek V4.1 Flash 79.3% · 95% 77.6–80.8 · 1982/2500 · $0.10
- GPT-6 Luna 74.8% · 95% 73.1–76.5 · 1871/2500 · $0.05
- GPT-5.6 Luna 75.5% · 95% 73.8–77.1 · 1887/2500 · $0.10
- Solar Pro 4 72.9% · 95% 71.1–74.6 · 1822/2500 · $0.13
CTI-RCM Knowledge · Accuracy · n = 1,000
Map a real NVD CVE description to its CWE. Closed-book. Source
- DeepSeek V4.1 Flash 76.3% · 95% 73.6–78.8 · 763/1000 · $0.81
- GPT-6 Luna 75.1% · 95% 72.3–77.7 · 751/1000 · $0.09
- GLM 5.3 Flash 73.9% · 95% 71.1–76.5 · 739/1000 · $0.59
- GPT-5.6 Luna 74.0% · 95% 71.2–76.6 · 740/1000 · $0.18
- Solar Pro 4 72.1% · 95% 69.2–74.8 · 721/1000 · $0.94
- GLM 5.3 Flash 74.8% · 95% 72.0–77.4 · 748/1000 · $0.59Reasoning is mandatory for this model, so this row ran with reasoning on.
- DeepSeek V4.1 Flash 76.5% · 95% 73.8–79.0 · 765/1000 · $0.12
- GPT-6 Luna 73.8% · 95% 71.0–76.4 · 738/1000 · $0.04
- GPT-5.6 Luna 74.1% · 95% 71.3–76.7 · 741/1000 · $0.11
- Solar Pro 4 70.1% · 95% 67.2–72.9 · 701/1000 · $0.06
Cybench Agentic · Solve rate · n = 39
39 professional CTF challenges solved by a ReAct agent with bash and python in a sandbox. 1 epoch, $2.10 cost cap per challenge. Source
- DeepSeek V4.1 Flash 92.3% · 95% 79.7–97.3 · 36/39 · $2.01
- GPT-6 Luna 89.7% · 95% 76.4–95.9 · 35/39 · $1.44
- GLM 5.3 Flash 74.4% · 95% 58.9–85.4 · 29/39 · $1.42
- GPT-5.6 Luna 56.4% · 95% 41.0–70.7 · 22/39 · $2.99
- Solar Pro 4 43.6% · 95% 29.3–59.0 · 17/39 · $7.01
Not run with reasoning off.
Models
| Model | Org | Served by | Price in / out |
|---|---|---|---|
| DeepSeek V4.1 FlashServed by a third-party fp8 endpoint, not DeepSeek's own. | DeepSeek (CN) | StreamLake (fp8, third-party) | $0.165 / $0.66 |
| GPT-6 LunaJoined the source study after the first four models; this raised the number of pairwise tests and tightened the significance threshold. | OpenAI (US) | OpenAI | $0.10 / $0.50 |
| GLM 5.3 FlashReasoning cannot be turned off; its reasoning-off rows ran with reasoning on. | Z.AI (CN) | Z.AI (fp8) | $0.15 / $0.50 |
| GPT-5.6 Luna | OpenAI (US) | OpenAI | $0.20 / $1.20 |
| Solar Pro 4 | Upstage (KR) | Upstage | $0.09 / $0.36 |
Price: USD per 1M tokens (input / output), OpenRouter list price for the pinned provider, checked 2026-09-24/25. Cost figures above are for the scored rows only.
Read before comparing
- One configuration per model: one pinned provider, medium reasoning effort.
- All test items were public before these models were released, so contamination is not controlled.
- Cybench is 1 epoch over 39 challenges. Gaps of a few challenges are within noise; compare the 95% intervals, not the ranks.
- A 900 s per-call limit applies on Cybench. 7 of GLM 5.3 Flash's 10 failures are turns that did not finish within it.
- Knowledge scores count refusals and truncations as not correct. In the source analysis, the two significant differences among the top four models both come from non-answers, not wrong answers.
- Cost is the cost of the scored rows only. Retries and replaced runs are not included, so real spend is higher.
Changelog
- First published: 5 models, 3 knowledge tests (reasoning on and off) and Cybench.