Budget LLMs on cybersecurity

Budget-tier LLMs from Korea, the US and China: closed-book security knowledge vs solving CTF challenges as an agent.

#llm#cybersecurity#benchmark

Five similarly priced models ($0.09–0.20 in, $0.36–1.20 out per 1M tokens) on two separate things: what they know about security when asked closed-book, and whether they can do a security task through a tool-using agent. The write-up is in this post.

Summary

Score in %, with the 95% interval below it. Click a column to sort. A dash means the test was not run in this configuration.

Reasoning on
ModelKnowledgeWMDP-cyberKnowledgeCTI-MCQKnowledgeCTI-RCMAgenticCybench
DeepSeek V4.1 Flash84.9 82.6–87.079.6 77.9–81.176.3 73.6–78.892.3 79.7–97.3
GPT-6 Luna81.6 79.1–83.981.0 79.4–82.575.1 72.3–77.789.7 76.4–95.9
GLM 5.3 Flash83.3 80.9–85.578.9 77.3–80.573.9 71.1–76.574.4 58.9–85.4
GPT-5.6 Luna83.8 81.4–86.080.2 78.6–81.774.0 71.2–76.656.4 41.0–70.7
Solar Pro 477.1 74.4–79.676.0 74.2–77.672.1 69.2–74.843.6 29.3–59.0

WMDP-cyber Knowledge · Accuracy · n = 996

Expert-written multiple-choice questions on offensive security (996-item knowledge subset). Closed-book, no tools. Source

Reasoning on · dot = score, bar = 95% interval, axis 70–100%
  • DeepSeek V4.1 Flash 84.9% · 95% 82.6–87.0 · 846/996 · $0.84
  • GPT-6 Luna 81.6% · 95% 79.1–83.9 · 813/996 · $0.12
  • GLM 5.3 Flash 83.3% · 95% 80.9–85.5 · 830/996 · $0.67
  • GPT-5.6 Luna 83.8% · 95% 81.4–86.0 · 835/996 · $0.26
  • Solar Pro 4 77.1% · 95% 74.4–79.6 · 768/996 · $0.89

CTI-MCQ Knowledge · Accuracy · n = 2,500

Threat-intelligence multiple choice from CTIBench (ATT&CK, CAPEC and similar). Closed-book. Source

Reasoning on · dot = score, bar = 95% interval, axis 70–100%
  • DeepSeek V4.1 Flash 79.6% · 95% 77.9–81.1 · 1989/2500 · $3.21
  • GPT-6 Luna 81.0% · 95% 79.4–82.5 · 2025/2500 · $0.27
  • GLM 5.3 Flash 78.9% · 95% 77.3–80.5 · 1973/2500 · $2.18
  • GPT-5.6 Luna 80.2% · 95% 78.6–81.7 · 2005/2500 · $0.63
  • Solar Pro 4 76.0% · 95% 74.2–77.6 · 1899/2500 · $2.87

CTI-RCM Knowledge · Accuracy · n = 1,000

Map a real NVD CVE description to its CWE. Closed-book. Source

Reasoning on · dot = score, bar = 95% interval, axis 60–100%
  • DeepSeek V4.1 Flash 76.3% · 95% 73.6–78.8 · 763/1000 · $0.81
  • GPT-6 Luna 75.1% · 95% 72.3–77.7 · 751/1000 · $0.09
  • GLM 5.3 Flash 73.9% · 95% 71.1–76.5 · 739/1000 · $0.59
  • GPT-5.6 Luna 74.0% · 95% 71.2–76.6 · 740/1000 · $0.18
  • Solar Pro 4 72.1% · 95% 69.2–74.8 · 721/1000 · $0.94

Cybench Agentic · Solve rate · n = 39

39 professional CTF challenges solved by a ReAct agent with bash and python in a sandbox. 1 epoch, $2.10 cost cap per challenge. Source

Reasoning on · dot = score, bar = 95% interval, axis 20–100%
  • DeepSeek V4.1 Flash 92.3% · 95% 79.7–97.3 · 36/39 · $2.01
  • GPT-6 Luna 89.7% · 95% 76.4–95.9 · 35/39 · $1.44
  • GLM 5.3 Flash 74.4% · 95% 58.9–85.4 · 29/39 · $1.42
  • GPT-5.6 Luna 56.4% · 95% 41.0–70.7 · 22/39 · $2.99
  • Solar Pro 4 43.6% · 95% 29.3–59.0 · 17/39 · $7.01

Models

ModelOrgServed byPrice in / out
DeepSeek V4.1 FlashServed by a third-party fp8 endpoint, not DeepSeek's own.DeepSeek (CN)StreamLake (fp8, third-party)$0.165 / $0.66
GPT-6 LunaJoined the source study after the first four models; this raised the number of pairwise tests and tightened the significance threshold.OpenAI (US)OpenAI$0.10 / $0.50
GLM 5.3 FlashReasoning cannot be turned off; its reasoning-off rows ran with reasoning on.Z.AI (CN)Z.AI (fp8)$0.15 / $0.50
GPT-5.6 LunaOpenAI (US)OpenAI$0.20 / $1.20
Solar Pro 4Upstage (KR)Upstage$0.09 / $0.36

Price: USD per 1M tokens (input / output), OpenRouter list price for the pinned provider, checked 2026-09-24/25. Cost figures above are for the scored rows only.

Read before comparing

  • One configuration per model: one pinned provider, medium reasoning effort.
  • All test items were public before these models were released, so contamination is not controlled.
  • Cybench is 1 epoch over 39 challenges. Gaps of a few challenges are within noise; compare the 95% intervals, not the ranks.
  • A 900 s per-call limit applies on Cybench. 7 of GLM 5.3 Flash's 10 failures are turns that did not finish within it.
  • Knowledge scores count refusals and truncations as not correct. In the source analysis, the two significant differences among the top four models both come from non-answers, not wrong answers.
  • Cost is the cost of the scored rows only. Retries and replaced runs are not included, so real spend is higher.

Changelog

  • First published: 5 models, 3 knowledge tests (reasoning on and off) and Cybench.