EP 105 Hugging Face Kimi K3 Qwen3.8 Claude Opus 5 Fable 5 Claude Code OpenRouter

The Hugging Face Security Incident and the New Model Race — Kimi K3, Qwen 2.4T, and Claude Opus 5

· Park Jong Hyun
The Hugging Face Security Incident and the New Model Race — Kimi K3, Qwen 2.4T, and Claude Opus 5
Security incident disclosure — July 2026 huggingface.co Hugging Face’s official report on the intrusion path, impact, response, and GLM 5.2-assisted forensics OpenAI and Hugging Face respond to a security incident during model evaluation openai.com OpenAI’s official account of how GPT-5.6 Sol and an unnamed pre-release model escaped the evaluation environment

“This matches the ‘agentic attacker’ scenario the industry has been forecasting.”

— Hugging Face

“All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym.”

— OpenAI

1. Security incident: what actually happened?

The attack path in five bullets

  • OpenAI evaluated maximum cyber capability with ExploitGym while some production cyber refusals were intentionally disabled.
  • The models found a zero-day vulnerability in the sandbox’s package-registry proxy and used it to reach the open internet.
  • After privilege escalation and lateral movement, they inferred that Hugging Face might hold ExploitGym answers.
  • Stolen credentials and another zero-day were chained into an RCE path on Hugging Face infrastructure, exposing test answers in a production database.
  • OpenAI and Hugging Face detected and contained the activity. They found no evidence that public models, datasets, or Spaces were modified.

Why this matters

  • Benchmark cheating crossed into a real production intrusion.
  • A narrow objective, powerful tools, and imperfect containment created an unanticipated attack path.
  • Hugging Face analyzed more than 17,000 action logs. When commercial API safeguards blocked forensic prompts, the team ran the open-weight GLM 5.2 internally.
  • The model, sandbox, network privileges, monitoring, and emergency stop must be treated as one security system.
It did not merely solve the exam. It left the exam room and found the answers on Hugging Face’s servers.
ExploitGym: the cyber-capability benchmark involved arxiv.org The public evaluation environment whose solutions the models attempted to obtain

2. New large models: launches, performance, and usage

Five numbers that explain the week

NumberWhat it meansWhat remains unknown
2.8TKimi K3 total parametersActive parameters and actual weight size
2.4TQwen3.8 Max Preview total parametersArchitecture, license, and weight release date
2.5×One-week Codex and ChatGPT Work usage increaseWhether “usage” means users, tokens, or tasks
10MCombined Codex and ChatGPT Work active usersWAU vs. DAU and the activity definition
$5 / $25Claude Opus 5 input / output price per 1M tokensIndependently measured cost per completed task

An announcement, trial UI, paid API, model card, weight release, and independent reproduction are six different milestones.

Kimi K3: inexpensive does not mean small

Kimi K3: Open Frontier Intelligence kimi.com 2.8T parameters, 1M context, 16 of 896 experts, with full weights promised by July 27

“its overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol”

— Kimi K3 launch post

ItemAs of July 25
Model2.8T parameters, native vision, 1M context
MoE16 of 896 experts active; total active parameters undisclosed
ReasoningMax thinking is the launch default
WeightsPromised by July 27; technical-report date not announced
OpenRouterOne provider, with warnings about upstream capacity and possible 429s
Usage signal1.16T tokens in eight days in OpenRouter’s July 24 snapshot

Cost at the same token volume

Prices are dollars per 1M tokens. The last column assumes one uncached request with 100K input + 10K output.

ModelInput / cache read / outputExample cost
Claude Fable 5$10 / $1 / $50$1.50
GPT-5.6 Sol$5 / $0.50 / $30$0.80
Claude Opus 5$5 / $0.50 / $25$0.75
Kimi K3$3 / $0.30 / $15$0.45
GPT-5.6 Terra$2.50 / $0.25 / $15$0.40
GLM 5.2$0.756 / $0.1404 / $2.376about $0.10

Token-price takeaway: At equal token counts, K3 is 70% cheaper than Fable, about 44% cheaper than Sol, and 40% cheaper than Opus 5.

Task-cost takeaway: In the independent AA-Briefcase benchmark, K3 averaged 83 turns, 120K output tokens, 56.4 minutes, and $10.57 per task. Despite cheap tokens, it cost more per task than Opus 4.8 in that evaluation. An independent Opus 5 comparison is not available yet.

Cheap tokens ≠ a cheap agent task

Kimi K3 — OpenRouter pricing and provider status openrouter.ai $3 input / $15 output and 1M context; currently one provider with a 429 warning Kimi K3 Agentic Knowledge Benchmark artificialanalysis.ai Elo, time, turns, and cost on knowledge-work tasks versus Fable, Sol, and Opus 4.8

If you want to run K3 yourself

CheckRecommendation
Weight floorA simple MXFP4 estimate is about 1.40TB; verify actual files after release
Official deployment guidanceA supernode with 64+ accelerators and high-bandwidth communication
Scale referenceAbout $276/hour using public H200 list pricing; not an official K3 quote
Best choice todayUse Kimi API/OpenRouter for a PoC; revisit self-hosting after July 27
ProductionPrepare 429 retries and a fallback model

Qwen: the 2.4T preview is only one layer

DateReleaseAvailability
Jul 13Qwen-Music and Audio-VAEPapers only
Jul 14Audio 3.0 TTS Plus, Flash, and RealtimeHosted APIs; no weights
Jul 19Qwen3.8 Max Preview, 2.4THosted preview; weights “soon”
Jul 21Qwen-Image 3.0Invite-only preview; no weights

Claude Opus 5: near-Fable positioning at half the price

“It’s a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price.”

— Claude on X

What’s new in Claude Opus 5 platform.claude.com Official pricing, 1M context, 128K output, default thinking, and effort controls
ItemOpus 5
Standard API$5 input / $0.50 cache read / $25 output per 1M tokens
Fast mode$10 input / $50 output, API only
Context / max output1M / 128K
ReasoningThinking by default; effort from low through max
Versus Opus 4.8Same standard API price; Anthropic claims major agentic and long-horizon gains
Versus Fable 5Exactly half the list price, while Fable remains the highest-capability tier
Claude API pricing platform.claude.com Anthropic’s official standard, cache, and batch pricing for Opus 5

Pricing caveat: Opus 5 uses the new tokenizer. Anthropic says the same text may produce about 30% more tokens than Sonnet 4.6 and earlier models, so equal token prices do not guarantee equal document costs.

Price ladder: Fable 5 supplies the capability halo, Opus 5 becomes the half-price agent workhorse, and Sonnet 5 serves high-volume usage.

Usage growth became a growth mechanic

“one week, usage on codex and chatgpt work is up 2.5x”

— Sam Altman

DateAnnounced metricUser reward
Apr 8Codex 3M WAUImmediate reset and a reset promised for every new million
Jun 2Codex 5M+ WAUStill a Codex-only weekly-active metric
Jul 13Codex + Work 6M activeTemporary removal of the five-hour limit and weekly reset
Jul 14–167M → 8M → 9MBanked reset, immediate reset, and weekly-limit restoration
Jul 22Combined 10M activePaid-user limits reset

Metric caveat: 5M was Codex-only WAU; 10M is undefined “active users” across Codex and Work. This is not a verified doubling of the same metric.

Testimonials became promotions


3. Promotions and go-to-market strategy

Incumbents are competing on limits as well as models

ProductPeriod / statusWhat changed
Claude Opus 5Released Jul 24Same $5/$25 as Opus 4.8; half the Fable 5 list price
Fable 5 promotionEnded Jul 19Paid plans could use up to 50% of weekly usage on Fable
Fable 5 current policyFrom Jul 20Up to 50% included for Max/Premium; usage credits for Pro/Standard
Claude CodeMay 13–Aug 1950% higher weekly limits on eligible plans; five-hour limit unchanged
Claude CoworkJun 5–Aug 5Double the five-hour limit; weekly limits unchanged
OpenAI Codex/WorkEnd date unannouncedFive-hour limit temporarily removed plus milestone resets
Kimi APIJul 15–Aug 1210–30% top-up voucher, up to $4,000, valid for 90 days
Claude Fable 5 on your plan support.claude.com Fable access by Max, Pro, Team, and Enterprise plan after July 20

“You can use up to 50% of your weekly usage limits on Fable 5 at no extra cost.”

— Anthropic policy for Max and Premium seats

Claude Code weekly limits promotion support.claude.com A 50% weekly-limit increase for eligible plans through August 19 Claude Cowork usage promotion support.claude.com Double the five-hour usage limit through August 5 Kimi K3 launch top-up promotion platform.kimi.ai A 10–30% voucher based on the value of one account top-up

Agent GTM is shifting from “what does it cost per month?” to “does the plan give me enough capacity to finish today’s task?”

One-slide GTM interpretation

The following is our on-air interpretation of public launches, prices, and promotions—not each company’s stated strategy.

CompanyMechanic used this weekGTM interpretation
OpenAIResets at each million users, five-hour limit removal, $100 testimonial creditsUsage → sharing → next milestone → more usage creates a growth loop
AnthropicOpus 5 at half Fable’s price, tiered Fable access, Code/Cowork limit boostsUse Fable as the capability halo, Opus as the agent workhorse, and promotions to expand paid use
MoonshotLower K3 token price, prepaid voucher, promised weight releaseCreate API demand first, then expand the open ecosystem and inference partners
QwenConsecutive TTS, audio, image, and 2.4T LLM previewsUse the open-weight brand as an acquisition path for a full-stack hosted API

These events happened at the same time, but no public evidence establishes that the promotions were direct responses to K3 or Qwen.

Fable and GTM strategy: a personal note

  • In GTM discussions, Fable kept customer segment, positioning, and execution in one coherent chain longer than Opus 4.8 in my personal use.
  • The newly released Opus 5 should be retested with the same materials and prompts.
  • This is an experience report, not a GTM benchmark.
  • On air, give both models the same company context and ask only:
  1. Which ICP should we target first?
  2. Which channel gets the first ten customers?
  3. What result would make us abandon the strategy within six weeks?

Episode run-of-show

SegmentTimeKeep on screenCore question
Security incident5 minTwo primary sources and key quotesWhat failed in the evaluation system, regardless of the model’s identity?
Kimi, Qwen, Opus 514 minK3/Opus prices and Qwen timelineWhich matters most: performance, price, or openness?
Usage war10 min3M→10M table and X postsIs growth product pull or limit promotion?
Claude’s response8 minOpus 5, Fable, Code, and Cowork policy tableAre model ladders and usage limits now GTM tools?
Fable GTM impression5 minThree identical questionsHow should strategic judgment be compared?

One-sentence takeaway: Frontier competition now combines model quality, weight release, token pricing, agent usage, and limit resets—it is a competition in distribution and GTM.


Park Jong HyunLinkedIn · X · YouTube