News Weekly
LV 10 XP
0% read
Your progress · 0/5 chapters
About 6 min total
ModelsISSUE #2 · STORY 3 OF 20Sep 21, 2026CONFIRMED

SpaceXAI's Grok 4.7 offers frontier coding at budget prices

SpaceXAI launched Grok 4.7 on Sept 21 at $2 per million input tokens, with an independent score of 46, just behind the leaders. The catch: it runs slow, spends many tokens, and its safety claims await third-party checks.

Illustration: an industrial engine core — a massive polished sphere of interlocking cogs — sits inside a vaulted machine hall, wrapped in concentric translucent shield rings and a luminous chain-and-padlock safe…

Read it your way

CHAPTER 1 · THE 60-SECOND VERSIONPicked for Explorers

A cheaper entry in the frontier race

On Monday Sept 21, SpaceXAI, the new home of xAI, released Grok 4.7 for coding and knowledge work. It was available the same day in Cursor, the Grok API, GitHub Copilot and other tools. Independent scores place it near the top, but behind two rivals.

Same price, higher scoreArtificial Analysis rates Grok 4.7 at 46, up from 44, with unchanged $2/$6 token pricing.
Not the outright leaderClaude Fable 5.1 and GPT-6 Astra lead the same index at 53, so this is a value story.
Everywhere on day oneIt shipped the same day in Cursor, Grok Build, the API, Copilot and router platforms.
Safety as a selling pointThe company claims a new safeguard stack, with biosafety and jailbreak numbers no outsider has re-run.
Finish this chapter for +15 XP
Flip the switch

Grok before and after the 4.7 launch

THE PRICESame $2/$6, index 46Grok 4.7 adds points without raising sticker price, though per-task cost runs higher.
THE BRANDNow SpaceXAIThe first frontier release after xAI folded into SpaceX Corp.
SAFETYSafety gets metricsCompany-published jailbreak and biosafety scores arrive, awaiting independent reruns.
Play with the numbers · +10 XP

Output tokens: Grok versus the premium leader

DRAG THE SLIDER
Grok 4.7
$12,000
about $6 per million output tokens
Claude Fable 5.1
$100,000
about $50 per million, five times Grok's rate
You'd keep
$88,000
every month

List prices from the research file as of Sept 21, 2026. Sticker savings can understate real spend: Grok 4.7 is slow and verbose, and Artificial Analysis measured a higher cost per benchmark task than GPT-6 Astra Max.

Your next move · as a Explorer

Judge models per task, not per token

1Pick five long coding or knowledge tasks you actually do
2Run Grok 4.7 at high and xhigh against your current model
3Log time, tokens, cost and pass rate, then decide

Switch your reading mode at the top to see a different next move.

Tap to open

Things to keep an eye on

Pop quiz · unlock the Bargain Hunter badge

Did it stick?

0/3
What score did Artificial Analysis give Grok 4.7 on its intelligence index?+20 XP
Why might the cheap token price understate real cost?+20 XP
What happened to the claim that Grok 4.7 has a 1 million token context window?+20 XP
Your call · +5 XP

Will Grok 4.7 become a default coding model at your organization this quarter?

Deep dive

The full research, labeled and sourced

CONFIRMED20 sources · 57 min
Story identity
FieldValue
Story IDS03
TitlexAI (SpaceXAI) releases Grok 4.7: "most powerful and safest AI" with new safeguard stack
OrganizationxAI / SpaceXAI (xAI folded into SpaceX Corp., now branded SpaceXAI)
CategoryModel Release
Event date2026-09-21
Announcement date2026-09-21
Article dates2026-09-21 (SiliconANGLE, AI Tech Daily, OfficeChai, GitHub changelog); 2026-09-22 (heise, OmniaKey, itbrief)
Window check2026-09-21 ∈ [2026-09-18, 2026-09-22] — eligible
Evidence statusCONFIRMED (official announcement fetched in full + Artificial Analysis page fetched in full + multiple independent outlets)

Corrections to the discovery record (important):

  • Context window is 500K tokens, not 1M. Artificial Analysis (fetched), OmniaKey and LLMReference all independently list a 500,000-token context window and a May 2026 knowledge cutoff for grok-4.7. The "1M-token context" figure in the discovery record is not supported and should be treated as an error. (Note: earlier Grok 4.3-list models had a 1000k window; Grok 4.6 and 4.7 both ship 500k.)
  • "Tied at the top of the AA Intelligence Index (46)" needs nuance. Grok 4.7 (xhigh) scores 46 on Artificial Analysis Intelligence Index v4.3.2, tying Xiaomi's MiMo-V2.6-Pro — but that tie is at the top of the open-weights field. The overall index is led by Claude Fable 5.1 and GPT-6 Astra at 53, with Claude Opus 5 at 51 (per heise and OrcaRouter reporting of AA data). Grok 4.7 sits in the top tier of proprietary models, not at the overall top.
  • "DeepSWE v1.1 71% (high effort) / 40% (low effort)" — only the 71.0% high-effort figure is published in the official table. No 40% low-effort DeepSWE figure appears in the announcement or any source found; the 40.4% figure in the same table belongs to Grok 4.6 on CursorBench 4.0. Treat "40% low-effort DeepSWE" as unverified/erroneous.
  • "~15% inference overhead" for the safeguard stack could not be verified. No source found in this research (including the official post and model card excerpts) publishes an overhead figure. The safeguard stack's outcomes are described (below); the overhead number is unverified.
  • "Head of Product R. Vifwala detailed the release in a launch thread" could not be verified — no trace of this person or thread surfaced in searches of official pages, model card, or coverage. Marked unverified (possible detail from the discovery pipeline's social capture).
  • "Most powerful and safest AI" is a paraphrase, not a verbatim quote. Official language: "SpaceXAI's most powerful model for coding and knowledge work" plus "best-calibrated safeguards to date" / "strongest model we've tested on refusals and jailbreak resistance." The stronger "world's most powerful and safest AI" framing is company marketing and Musk's pre-release "will exceed all current models" posturing, not an independently established fact.

✓

What happened?

🎓 For Explorer

On Monday September 21, 2026, SpaceXAI (the current identity of xAI after the startup was merged and folded into SpaceX Corp.) announced Grok 4.7 — its most powerful model to date for coding and knowledge work — available the same day in Cursor, Grok Build, the Grok API, third-party coding harnesses, model routers, and cloud platforms, and rolling out to GitHub Copilot (Pro, Pro+, Max, Business, Enterprise SKUs).

Key parameters (FACT unless labelled):

  • Pricing: $2/M input and $6/M output tokens — the same price tier as Grok 4.6; OpenRouter lists $1.60/$4.80 with a $0.40 cache-read rate. A "fast variant" serves at 2× output speed for 2× price. OmniaKey documents a premium tier above 200K prompt tokens ($4/$1/$12 per M input/cached/output). FACT (multiple independent spec trackers agree).
  • Model: new, larger base model vs Grok 4.6; a longer reinforcement-learning (RL) run on a harder mix of tasks weighted toward multi-hour problems; better self-verification; better long-context management; natively understands the Grok Bot harness. COMPANY CLAIM (architecture details not independently probed).
  • Reasoning: configurable effort (low / medium / high default / xhigh); reasoning model; text + image input, text output; model ID grok-4.7; knowledge cutoff May 2026; no batch API. FACT (spec trackers; AA confirms reasoning + modalities).
  • Company-published benchmarks (COMPANY CLAIM until replicated): CursorBench 4.0 46.3% (vs Grok 4.6 40.4%, GPT-5.6 Sol Max 41.7%, Fable 5.1 Max 51.8%); DeepSWE v1.1 71.0% (high effort; vs 4.6 65.2%, Sol 72.7%, Fable 70.0%); EEBench 64.0% (vs 53.0% / 39.4% / 56.4%); AA Briefcase v1.1 1,657 (vs 1,546 / 1,487 / 1,678); Terminal-Bench 4.0 37.6% (vs 20.3% / 37.3% / 57.9%); Harvey Legal Agent Benchmark 19.6% (vs 15.8% / 2.5% / 6.7%); HealthBench Professional 56.7% (vs 48.5% / 60.5% / 62.1%); GDPval 1,695 Elo (vs Fable 5.1 Max 1,735, GPT-6 Astra 1,542).
  • Safety (COMPANY CLAIM): an "entirely new safeguard stack" — the strongest model the company has tested on refusals and jailbreak resistance; 62.4% on LatchBio's biosafety benchmark; only 3.3% of risky dual-use prompts pass on HackerBench v0.3 while rarely blocking legitimate security work; invite-only red-team access for select cybersecurity partners. Outcomes published in the announcement; component design (e.g., focused-reasoning and vulnerability-scanning layers) is only described in the discovery record and community recaps — EARLY RESEARCH.
  • Independent evaluation (INDEPENDENTLY VERIFIED): Artificial Analysis measures Grok 4.7 (xhigh) at 46 on the Intelligence Index v4.3.2 (vs Grok 4.6 high at 44), 39.3 output tokens/s (notably slow; median 71), 0.88 s time-to-first-token (very competitive; median 3.84 s), ~240M output tokens on the index (very verbose; median 88M), $3.74 cost per index task, 500K context. On AA's Coding Agent Index, Grok 4.7 with Grok Build scores 56 — behind Claude Fable 5.1, GPT-6 Astra and Claude Opus 5 (via heise's independent readout of AA data).

Musk's pre-release commentary (August–September): claimed ~2.1T parameters and that Grok 4.7 "will exceed all current models"; in mid-September reportedly conceded parity-or-worse vs Opus 5.0 in some areas ("better in some ways, worse in others") and that the model needed "a few more days to cook." The 2.1T figure does not appear in the official announcement or model card — RUMOR/COMPANY CLAIM, unconfirmed.


Δ

What changed?

  • The frontier-for-less tier became real across multiple labs simultaneously. Grok 4.7 reaches top-tier intelligence (46 on AA) at $2/$6 per M — the same token price as Grok 4.6 — while GPT-6 Sol/Luna (announced the next day, Sep 22) undercut pricing and Xiaomi's open-weights MiMo-V2.6-Pro tied 46 at ~$0.43/$0.87 with a 99% cache discount. Price-per-intelligence-point has collapsed.
  • A default-on "safeguard stack" became a shipping feature of a frontier coding model — safety/refusal performance is now a headline differentiator, positioned as compatible with low refusal rates for legitimate cyber/security and biology work, with invite-only red-team access.
  • The organization identity hard-switched: this is the first major model release under the "SpaceXAI" brand (footer: "© 2026 SpaceXAI LLC"; the xAI startup was merged and folded into SpaceX Corp.). No more standalone "xAI" product surface.
  • Distribution widened immediately: same-day availability in Cursor, Grok Build, Grok API, GitHub Copilot (five SKUs), third-party harnesses, routers, and cloud platforms.
  • Launch cadence accelerated: Grok 4.5 (July) → Grok 4.6 (Aug 12) → Grok 4.7 (Sep 21), i.e., ~3 majors in ~10 weeks, with a TTS release the Friday before.

↔

Before → Change → After

🎓 For Explorer

Before (as of mid-September 2026):

  • Grok 4.6 (Aug 12) at $2/$6, 500K context, AA 44, CursorBench 4.0 40.4%; xAI still branded as xAI; launch cadence ~monthly; safety communicated as generic guardrail improvements; frontier index led by Fable 5.1 / GPT-6 Astra (53) with open weights at a distance (GLM-5.3 45, MiMo-V2.5-Pro 26 on old scale).

Change (Sep 21, 2026):

  • Grok 4.7 ships at the same $2/$6 price with +2 index points (46), +6 points CursorBench (46.3%), large gains on EEBench/Terminal-Bench, a native Grok Bot harness, an explicit new safeguard stack with published refusal/jailbreak metrics, and full multi-channel availability (Cursor, Grok Build, API, Copilot, routers).
  • First SpaceXAI-branded frontier release; safety published as benchmarkable claims (LatchBio 62.4%, HackerBench 3.3%).

After:

  • A three-front pricing war: SpaceXAI, OpenAI (GPT-6 Sol/Luna, Sep 22) and Xiaomi (open weights at $0.43/$0.87) all offer frontier-adjacent intelligence at commodity token prices; the "safest model" marketing axis is now quantified by vendors, not just regulators.
  • Agentic-coding price-performance, not raw benchmark supremacy, is the decision criterion; Fable 5.1 remains the performance leader on most official-table rows while costing 5× output tokens (FACT: $50/M output vs $6/M).

⚙

How it works

  • Base model + RL: a new, larger base model than Grok 4.6; a longer RL run on a harder task mix weighted toward problems that take hours — the model is trained to persist, verify its own work, and manage longer context. COMPANY CLAIM (no parameter counts, architecture, data mix, or compute details disclosed).
  • Grok Bot harness integration: trained to natively understand the Grok Bot agent harness (a team of always-on agents with their own computer that work inside tools/apps around the clock, introduced Aug 11, 2026), enabling multi-agent task splitting, parallelism, and cross-verification, per SiliconANGLE's summary of company statements.
  • Reasoning effort controls: low / medium / high (default) / xhigh; the 46 index score is at xhigh; speed/latency/verbosity tradeoffs are steep (39.3 t/s output; TTFT 0.88 s; very high token consumption). FACT (AA).
  • Safeguard stack: company describes an entirely new stack optimized for calibrated refusals and jailbreak resistance rather than maximal blocking: tops LatchBio biosafety (62.4%) while holding HackerBench v0.3 pass-through to 3.3% with low over-blocking of legitimate security work; red-team capabilities gated to invite-only cybersecurity partners. Outcomes COMPANY CLAIM; underlying components (focused reasoning, vulnerability scanning) reported by the discovery record/community — EARLY RESEARCH. No overhead figure published (~15% figure is unverified).
  • Serving: $2/$6 per M (below 200K prompt; $4/$1/$12 at/above 200K per OmniaKey); a fast variant at 2× output speed for 2× price; 75% cache-read discount (AA).
  • Capabilities: function calling, structured outputs, prompt caching, code execution, web search, X search; text + image in, text out; 500K context; no batch API. FACT (OmniaKey/LLMReference/AA agreement).

!

Why it matters

🎓 For Explorer
  1. Price-performance collapse at the frontier: Grok 4.7 delivers 46 AA-index intelligence at $3.74 per index task — roughly commodity pricing — in the same week GPT-6 Sol/Luna launched at ~50% under GPT-5.6 equivalents and Xiaomi's open-weights MiMo-V2.6-Pro tied 46 at ~$0.13 per task. "Frontier-class capability at budget prices" is now a three-company reality, which changes every developer/enterprise make-or-buy decision.
  2. The safety arms race became measurable and public: a frontier lab is now marketing quantified refusal/jailbreak results (LatchBio, HackerBench) and granting narrow red-team access. This sets a template competitors must answer — and regulators (e.g., California's EO N-9-26 kill-switch mandate, signed Sep 18; the federal AI Force EO, Sep 19) now have vendor-published safety baselines to point at.
  3. The SpaceXAI identity consolidates Musk's AI in SpaceX: grok is no longer an independent lab product; it is a SpaceX line with AI data centers, Tesla/X integration paths, and a very different strategic posture (per SiliconANGLE's corporate history).
  4. The Fable/GPT-6 distinction matters for a news weekly: Grok 4.7 is not the outright leader (Fable 5.1 and GPT-6 Astra top the index at 53; Fable leads most official-table rows) — the story is the price-performance tier, not a "new #1."

✦

What became possible?

🎓 For Explorer
  • Running frontier-adjacent agentic coding and knowledge work at ~$2/$6 per M tokens, with a same-day GitHub Copilot and Cursor integration path.
  • Deploying a reasoning model with a 500K context + native agent-harness understanding into long-horizon workflows (multi-hour coding, office/document work, engineering tasks) without paying Fable/Sol-tier output prices.
  • Comparing safety postures quantitatively across vendors for procurement and compliance (LatchBio/HackerBench-style evaluations become vendor talking points and enterprise checklists).
  • Invite-only cybersecurity defense research with Grok 4.7 red-team capabilities for select partners — a new channel between a model vendor and the security industry.

◎

Implications

Technical

  • Benchmark-lab independence problem: CursorBench 4.0, the flagship benchmark in the launch marketing, is built by Cursor — recently acquired by SpaceX (OfficeChai and SiliconANGLE both flag this). Self-adjacent benchmarking now requires independent replication (AA's coding-agent run: 56 points with Grok Build, behind Fable 5.1/GPT-6 Astra/Opus 5).
  • Token economy is the real spec: AA's measurements (81K output tokens per index task at xhigh vs ~38K for Grok 4.6 xhigh and ~27K for GPT-6 Astra Max; cost per index task $3.74 vs GPT-6 Astra Max $3.26) show the $2/$6 sticker understates effective cost because the model is very verbose and slow. Cost-per-task, not per-token, is the operative metric.
  • Speed/latency tradeoff: 39.3 t/s is below median (71); but TTFT 0.88 s is excellent — good for deliberative tasks, poor for interactive streaming.
  • Reasoning-effort surface: four effort levels multiply the effective cost/speed matrix; production tuning must pick effort per workload, not just per model.
  • 500K context + image input means long-horizon RAG and multimodal agent work are in scope, but the >200K prompt premium ($4/$1/$12 tier) creates a real bill-shock cliff at the 40% context mark.

Developer

  • Add a new default candidate for agentic coding: grok-4.7 at $2/$6 (or $1.60/$4.80 via OpenRouter) with high/xhigh effort for long-horizon tasks; evaluate per-task cost, not per-token price (verbosity is high).
  • Effort-level plumbing: expose effort (low/medium/high/xhigh) as a first-class param; default to high and measure.
  • Integration surface: available in VS Code/VS/Copilot CLI/cloud agent/app, JetBrains, Xcode, Eclipse via Copilot model picker; Cursor and Grok Build supported; API via Chat Completions + Responses APIs; function calling, structured outputs, code execution, web/X search, prompt caching; no batch API — plan for throughput.
  • Watch the context-price cliff: prompts ≥200K tokens hit $4/$1/$12 — budget caching carefully (75% cache-read discount helps long conversations).
  • Safety posture is now a product spec: for dual-use work (cyber, bio, security), Grok 4.7's calibrated refusal profile may change leak/allowance behavior vs prior models in tool-use flows — regression-test agent harnesses after switching.

Enterprise

  • Negotiation leverage: AAA-lab token prices are falling across the board (Grok $2/$6, GPT-6 Sol/Luna half of GPT-5.6-era, Xiaomi open weights near-free). Enterprises should re-bid API contracts and internal model-selection frameworks in the next 30 days.
  • Compliance timing: California EO N-9-26 (kill-switch safeguards, recommendations due Nov 16, 2026) and the federal AI Force EO raise the stakes on documented model-safety posture; vendor-published safety metrics (LatchBio/HackerBench) become input to vendor due-diligence.
  • Agentic coding pilots: Grok 4.7 + Grok Build at 56 on AA's Coding Agent Index (behind Fable 5.1/GPT-6 Astra/Opus 5) is viable for cost-sensitive long-horizon coding workloads, but retain a performance-first vendor for the highest-stakes work.
  • Procurement caution: verify "2× as fast, half the price" claims against measured per-task cost and speed (AA: slow output but excellent TTFT; verbose), and note benchmark-suite ownership conflicts (CursorBench is Cursor/SpaceX-owned).

Strategic

  • The frontier is now a pricing war with safety as a differentiator: OpenAI (Sol/Luna), SpaceXAI (Grok 4.7) and Xiaomi (MiMo-V2.6-Pro, MIT weights, published ~$2.62M RL bill) are squeezing per-task intelligence costs; differentiation shifts to safety posture, harness/agent integration, and speed.
  • xAI's absorption into SpaceX changes the chessboard: model releases now serve SpaceX's (and Musk's) vertical ambitions — space/engineering corpora, Tesla, X integration — rather than a standalone lab's roadmap. This is a structural, not just branding, shift.
  • Open weights closed the gap on the index (46 = 46): Xiaomi's open-weights tie at the top of the open field pressures closed labs' premium tiers and gives enterprises a credible open alternative.
  • Red-team access gating: vendor-controlled, invite-only red-team access is a nascent competitive moat and a governance question (who qualifies, under what terms?).

⚠

Risks & limitations

Risks
  • Unverified headline safety claims: LatchBio 62.4% and HackerBench 3.3% are self-published; no independent re-run yet. Over-blocking and under-blocking profiles in dual-use domains carry real-world stakes (bad actors probing cyber/bio knowledge).
  • Safety-on-default overhead: unknown inference overhead (~15% figure unverified); if real, it quietly raises per-task cost and could degrade latency in production.
  • Benchmark-suite independence risk: CursorBench being Cursor/SpaceX-owned weakens the lead claim (46.3% vs Fable 5.1's 51.8%, which still leads).
  • Bill shock: the >200K prompt premium and high verbosity (240M output tokens on AA's index; 81K/task at xhigh) can erase the sticker-price advantage — heise notes GPT-6 Astra Max's per-task cost ($3.26) is lower than Grok 4.7's ($3.74) despite far higher token prices.
  • Rebrand/legal churn: "SpaceXAI" branding on an AI model series inside a rocket company creates enterprise-contract and data-governance questions (where do prompts live? data-center portfolio per SiliconANGLE).
  • Rapid-cadence churn: 3 flagship majors in ~10 weeks hints at short-support lifecycles for each version; enterprises should pin model IDs and plan migrations.

Limitations
  • Not the overall #1: AA index 46 vs Fable 5.1 / GPT-6 Astra 53; Fable leads CursorBench, AA Briefcase, Terminal-Bench, HealthBench; GPT-5.6 Sol leads DeepSWE (72.7%) and EEBench (per official table GPT-6 Astra leads EEBench per SiliconANGLE).
  • Slow + verbose at xhigh: 39.3 t/s (well below median 71); highest token consumption in class.
  • Proprietary, closed weights: no parameter count published; the ~2.1T figure is Musk's claim, absent from the announcement and model card.
  • No batch API; limited provider count (3 providers on AA's tracker at launch; OmniaKey did not yet list a 4.7 route on Sep 22).
  • Safety metrics are narrow/self-run: single benchmarks, company-run, red-team access invite-only — not a formal third-party safety audit.
  • Context cliff: 500K window with a 200K pricing threshold; no disclosed long-context degradation measurements for Grok 4.7 yet.
  • Musk-parity admission: pre-release statements suggested only rough parity with Opus 5.0, "better in some ways, worse in others" — the marketing and the founder's own mid-September concession diverge.

?

Open questions

  1. What is the real inference overhead of the safeguard stack (the ~15% figure is unverified)? Does it vary by effort level?
  2. Will independent safety evaluations (LatchBio/HackerBench-style re-runs by third parties) reproduce 62.4% / 3.3%?
  3. Does the >200K token tier and high verbosity make Grok 4.7 more expensive per completed task than GPT-6 Astra/GPT-5.6 Sol in real agent workloads (AA says yes on the index today — will production data confirm)?
  4. What exactly is "the largest base model vs Grok 4.6"? Parameter count, MoE routing, training compute — none disclosed.
  5. What is the SpaceXAI product roadmap: Grok 4.8? Grok 5 (Musk teased AGI-level claims)? How does the TTS/voice line (Grok Voice Transcribe 2.0, Sep 18) factor in?
  6. How do enterprise data-governance terms change under SpaceXAI (data centers absorb xAI's portfolio per SiliconANGLE)?
  7. Will OpenAI's GPT-6 Sol/Luna at half price (Sep 22) pull the frontier-for-less price tier down further, eroding Grok 4.7's $2/$6 position within weeks?
  8. Does Xiaomi's open-weights 46 (MIT license) pull migration from Grok 4.7 for cost-sensitive deployments?

↗

What happens next?

🎓 For Explorer
  • Days: third-party benchmark re-runs (AA updates its leaderboard continuously; expect Coding Agent Index refresh), independent safety re-evaluation attempts, provider routes multiplying (reach via more cloud/router providers).
  • Weeks: OpenAI GPT-6 Sol/Luna pricing (Sep 22) sets the new low bar — watch for a SpaceXAI price response; Xiaomi MiMo-V2.6-Pro adoption in open-weights stacks; possible Grok 4.7-paired releases (voice, image) from SpaceXAI's recent cadence.
  • Months: Grok 4.8/Grok 5 signals (Musk teased AGI-level claims), Colossus data-center integration into SpaceXAI serving, enterprise adoption reports on the >200K pricing cliff, and (likely) NAIC/regulatory commentary citing vendor-published safety scores.

★

Editorial takeaway

🎓 For Explorer

Grok 4.7's real news is not "most powerful AI" — independent measurement puts it in the top proprietary tier (AA 46) behind Fable 5.1 and GPT-6 Astra — it is that frontier-adjacent intelligence is now a commodity sold by SpaceXAI at $2/$6, one day before OpenAI undercut its own pricing with GPT-6 Sol/Luna and a week after Xiaomi tied the score with MIT-licensed open weights. The second story is the safeguard stack: a frontier lab is now marketing quantified refusal/jailbreak safety (LatchBio 62.4%, HackerBench 3.3% pass-through) — unverified by third parties, but it turns safety into a spec sheet item, not a press-release paragraph. Cover it as: the price-performance frontier became a three-company race this week, Grok 4.7 is a strong but not dominant entry, and "safest AI" is a company claim awaiting independent replication.

Illustration: frame: a polished sphere of interlocking cogs at the center of a vaulted hall, wrapped in concentric translucent shield rings — an artistic impression of a newly proclaimed safeguard stack protecti…
⌘

Lab: COMPARE

Step 1 — Collect the public record (all figures cited in sources/S03.md)

Official table, xAI announcement (2026-09-21), Grok 4.7 (xhigh) vs Grok 4.6 (high) vs GPT-5.6 Sol Max vs Fable 5.1 Max:

ModelInput $/MOutput $/MCursorBench 4.0DeepSWE v1.1*Terminal-Bench 4.0HealthBench
Grok 4.7 (xhigh)$2.00$6.0046.3%71.0%37.6%56.7%
Grok 4.6 (high)$2.00$6.0040.4%65.2%20.3%48.5%
GPT-5.6 Sol Max$4.00$20.0041.7%72.7%37.3%60.5%
Fable 5.1 Max$10.00$50.0051.8%70.0%57.9%62.1%

Independent measurements, Artificial Analysis (v4.3.2):

  • Grok 4.7 (xhigh): Intelligence Index 46, output speed 39.3 t/s, TTFT 0.88 s, $3.74 cost per index task, 240M output tokens on the index (very verbose), 500K context.
  • Grok 4.6 (high): Index 44; Grok 4.6 (low) $0.48/task.
  • GPT-6 Astra Max: $3.26 per index task (heise readout) — the price-performance warning.
  • Xiaomi MiMo-V2.6-Pro (open weights): Index 46, $0.43/$0.87 per M, $0.13 per index task (OrcaRouter).

Vendor-adjacent number: CursorBench average cost of $4.69 per task for Grok 4.7, ahead of GPT-5.6 Sol and Fable 5.1 (SiliconANGLE; note CursorBench is Cursor/SpaceX-owned).

Step 2 — Reproduce the per-task economics
# cost_model_grok47.py — run with plain python3, no deps
# Verifies the "frontier in price-performance" claim from published numbers only.

rows = [
    # name, in$, out$, index, index_task_cost$, speed_tps
    ("Grok 4.7 (xhigh)",    2.00,  6.00, 46, 3.74, 39.3),
    ("Grok 4.6 (high)",     2.00,  6.00, 44, None, 60.0),  # per-task N/A from AA release page for high
    ("GPT-6 Astra Max",     None,  None, 53, 3.26, None),
    ("MiMo-V2.6-Pro (open)",0.43,  0.87, 46, 0.13, 134.3),
]

print(f"{'model':<22}{'idx':>4}{'$/task':>8}{'$/idx-pt':>9}{'tps':>7}")
for name, i, o, idx, task, tps in rows:
    per_pt = round(task / idx, 3) if task and idx else float("nan")
    print(f"{name:<22}{idx:>4}{str(task or '-') if task else '-':>8}"
          f"{per_pt:>9.3f}{str(tps) if tps else '-':>7}")

Expected output on the published numbers:

  • Grok 4.7: $0.0813 per index point ($3.74/46) and $10.13 per CursorBench point ($4.69/0.463).
  • MiMo-V2.6-Pro: $0.0028 per index point — 29x cheaper per intelligence point (open weights).
  • GPT-6 Astra Max: $0.0615 per index point — cheaper per index point than Grok 4.7 despite 10x+ token prices.
Step 3 — Verdict template (fill in your own verdict after running the script)
  1. "Frontier in price-performance" (vendor claim) — PARTIALLY CONFIRMED against Fable 5.1-class rivals (Fable costs ~$10–15/task at $50/M output), REFUTED against GPT-6 Astra Max on per-index-task cost ($3.26 vs $3.74), and overshadowed by open weights (MiMo $0.13/task).
  2. "Same price and speed as Grok 4.6" (official) — CONFIRMED on price; speed is close at similar effort (39.3 vs 53–60 t/s depending on effort), but verbosity jumped (240M index tokens vs 4.6's lower consumption) so per-task cost rose even though per-token price did not.
  3. "Twice as fast at half the price vs comparable models" (headline) — MARKETING FRAME: only true vs the premium max-effort tiers (Sol $4/$20, Fable $10/$50); false vs OpenAI's own budget tier (GPT-6 Sol/Luna, Sep 22) and Xiaomi open weights.
Step 4 — Optional one-command smoke test of availability (no key needed)
# Confirm the model exists in the wild via public docs index (read-only, no billing).
curl -s https://docs.x.ai/developers/models | grep -io "grok-4.7" | head -1
# Or check a public provider catalog: OpenRouter community page.
curl -s https://openrouter.ai/api/v1/models | grep -o "grok-4.7" | head -1

If the model IDs resolve, note that routing is already live; if not, the "available today" claim is at least partially aspirational (OmniaKey already observed a routing lag on Sep 22).


What this lab does NOT do (honest limits)

  • It does not benchmark the model — no API spend, no sandboxed runs.
  • It treats vendor table numbers as COMPANY CLAIM throughout (per evidence discipline).
  • Per-task figures from AA are index-specific; real production workloads will differ (prompt mix, caching, effort levels).

Result

A one-page, evidence-labelled cost model that shows Grok 4.7 is a genuine price-performance improvement over Grok 4.6 and Fable-tier rivals, but not a price-performance leader once GPT-6 Astra Max per-task cost and Xiaomi open weights enter the comparison. The lab doubles as a template for any future frontier-model release evaluation.

≡

Research sources

Primary Sources (2)
Primary
SpaceXAI — "Model Card: Grok 4.7" (official model card, PDF) - **URL:** https://media.x.ai/v1/website/4p7card-5eccc980.pdf** Existence and content of formal model card (Sep 21, 2026); safety-threshold evaluations of dual-use biology/chemistry knowledge and lab-protocol understanding; corroborates that the 2.1T parameter figure does not appear in official disclosures. — ** Primary documentation (partially read via search snippet and referenced content; cited for model-card existence and safety-evaluation framing). ---Date: ** 2026-09-21 (revision 2026-09-21)
URL unavailable
Primary
SpaceXAI — "Introducing Grok 4.7" (official announcement) - **URL:** https://x.ai/news/grok-4-7** Release fact (why: same-day announcement, Sep 21, 2026); SpaceXAI branding; "most powerful model for coding and knowledge work" and "best-calibrated safeguards to date" wording; pricing $2/$6 per M tokens + fast variant at 2x speed/2x price; company benchmark table (CursorBench 4.0 46.3%, DeepSWE v1.1 71.0%*, EEBench 64.0%, AA Briefcase 1,657, Terminal-Bench 4.0 37.6%, Harvey Legal 19.6%, HealthBench 56.7%); safety claims (LatchBio 62.4%, HackerBench v0.3 3.3% pass-through, invite-only red-team access); availability (Cursor, Grok Build, Grok API, harnesses, routers, cloud platforms); training description (larger base model, longer RL run, Grok Bot harness). — ** FACT baseline (release event) + COMPANY CLAIM (all benchmark and safety figures are vendor-published).Date: ** 2026-09-21
URL unavailable
Independent Sources (10)
Independent
Tesla North — "SpaceXAI Drops Grok 4.7, Its Smartest AI Yet for Writing Code" - **URL:** https://teslanorth.com/2026/09/21/grok-4-7-smartest-ai** "Ranks third for agentic coding on the Artificial Analysis Intelligence Index with a score of 46"; Musk comment "a strong combination of intelligence, speed & low cost"; launch-timeline context ("needed a few more days to cook"; Grok Voice Transcribe 2.0 last week). — ** Independent (Musk-ecosystem) reporting; confirms AA index rank framing and Musk quotes. ---Date: ** 2026-09-21
URL unavailable
Independent
AI Tech Daily — "xAI launches Grok 4.7 for coding and knowledge work" - **URL:** https://www.aitechdaily.com/xai-grok-4-7/** Recap of release; notes the 2.1T parameter figure (attributed to Decrypt's coverage of the release) does **not** appear in the SpaceXAI news post; benchmark table incl. Terminal-Bench and GDPval numbers (1,695 vs Fable 1,735); "price-performance, not a clean lead" framing. — ** Independent reporting; supports the RUMOR/COMPANY-CLAIM tagging of the 2.1T figure and the parity-not-supremacy framing.Date: ** 2026-09-21
URL unavailable
Independent
IT Brief (Australia) — "xAI launches Grok 4.7 for coding & knowledge work" - **URL:** https://itbrief.com.au/story/xai-launches-grok-4-7-for-coding-knowledge-work** Independent recap of release facts (pricing, availability, benchmark table incl. DeepSWE 71.0% high-effort, EEBench 64.0%); market framing: model releases increasingly framed around cost/speed/safeguard trade-offs, not overall leadership; immediate adoption test framing. — ** Independent reporting; corroborates official figures and the "mixed results / no clear leader" framing.Date: ** 2026-09-22
URL unavailable
Independent
OfficeChai — "SpaceXAI Launches Grok 4.7, Beats GPT-5.6 Sol And Fable 5.1 On Some Benchmarks" - **URL:** https://officechai.com/ai/grok-4-7-benchmarks/** Benchmark recaps (Grok 4.7 vs 4.6 vs Sol vs Fable 5.1); wins on DeepSWE, EEBench, Harvey Legal but Fable leads CursorBench, AA Briefcase, Terminal-Bench, HealthBench; GDPval Elo 1,695 (Grok 4.7) vs Fable 5.1 1,735 and GPT-6 Astra 1,542; caveats that numbers come from the company and CursorBench is built by Cursor (recently acquired by SpaceX); GPT-6 Astra launch context (earlier in September). — ** Independent reporting; useful adversarial caveats (benchmark ownership).Date: ** 2026-09-21
URL unavailable
Independent
heise online (English) — "Grok 4.7: Inexpensive API meets high token consumption" - **URL:** https://www.heise.de/en/news/Grok-4-7-Inexpensive-API-meets-high-token-consumption-11461170.html** Independent readout of AA data: 46 points at xhigh places SpaceXAI among top four labs but behind Claude Fable 5.1 and GPT-6 Astra; ~81,000 output tokens per index task at xhigh vs ~38,000 (Grok 4.6) and ~27,000 (GPT-6 Astra Max); cost per task $3.74 (Grok 4.7) vs $3.26 (GPT-6 Astra Max) — sticker price does not equal cost per task; AA Coding Agent Index: Grok 4.7 + Grok Build 56, behind Fable 5.1, GPT-6 Astra, Claude Opus 5; moderate gains over Grok 4.6 outside agentic knowledge work. — ** Independent technical analysis; strongest support for the "verbose/slow/effective-cost" interpretation.Date: ** 2026-09-22
URL unavailable
Independent
GitHub Changelog — "Grok 4.7 is now available in GitHub Copilot" - **URL:** https://github.blog/changelog/2026-09-21-grok-4-7-is-now-available-in-github-copilot/** Same-day distribution via GitHub Copilot (Pro, Pro+, Max, Business, Enterprise SKUs); model picker availability in VS Code, Visual Studio, Copilot CLI, cloud agent, Copilot app, JetBrains, Xcode, Eclipse; usage-based billing at provider list pricing; default model enablement and admin policy controls. — ** Independent (GitHub) platform confirmation of availability and developer-facing rollout mechanics.Date: ** 2026-09-21
URL unavailable
Independent
SiliconANGLE — "SpaceX launches Grok 4.7 with long-horizon processing, safety upgrades" - **URL:** https://siliconangle.com/2026/09/21/spacex-launches-grok-4-7-with-long-horizon-processing-safety-upgrades** SpaceX Corp. corporate-history framing (xAI merged into SpaceX; SpaceXAI as the new identity); CursorBench 4.0 developed by Cursor, "another of SpaceX's recent acquisitions"; average $4.69 cost per task on CursorBench ahead of GPT-5.6 Sol and Fable 5.1; LatchBio/HackerBench records claim; $2/$6 pricing and fast variant; Grok Bot harness multi-agent behavior; Grok Voice Transcribe 2.0 (Sep 18, Friday) context; EEBench behind GPT-6 Astra. — ** Independent reporting; enterprise-tech outlet; corroborates official claims with additional corporate context.Date: ** 2026-09-21 (updated 20:32 EDT)
URL unavailable
Independent
Artificial Analysis — "MiMo-V2.6-Pro vs Grok 4.7 (xhigh)" comparison page - **URL:** https://artificialanalysis.ai/models/comparisons/mimo-v2-6-pro-vs-grok-4-7** Existence of the direct comparison; both models score 46 on v4.3.2, i.e., the tie at the top of the open-weights field; Xiaomi is the other entrant at 46. — ** Independent evaluation; supports the corrected "tie at 46 = open-weights top, not overall top" framing.Date: ** consulted 2026-09-23
URL unavailable
Independent
Artificial Analysis — "Grok 4.7 vs Grok 4.6" release comparison - **URL:** https://artificialanalysis.ai/models/releases/comparisons/grok-4-7-vs-grok-4-6** Grok 4.7 (xhigh) 46 vs Grok 4.6 (high) 44 on the Intelligence Index; Grok 4.6 output speed 60 t/s vs Grok 4.7 (high) 53 t/s; cost per task Grok 4.6 (low) $0.48 vs Grok 4.7 (high) $2.73. — ** Independent evaluation; supports "modest index jump, steep verbosity/cost behavior" interpretation.Date: ** consulted 2026-09-23
URL unavailable
Independent
Artificial Analysis — "Grok 4.7 (xhigh)" model page - **URL:** https://artificialanalysis.ai/models/grok-4-7** INDEPENDENTLY VERIFIED Intelligence Index 46 (xhigh) on v4.3.2; output speed 39.3 t/s (median 71); TTFT 0.88s; $3.74 cost per Intelligence Index task; 240M output tokens (very verbose); 500K context window (corrects discovery's "1M" claim); $2/$6 pricing with 75% cache discount; text+image input, text output; reasoning model; proprietary; 3 API providers; released September 21, 2026. — ** Independent technical evaluation — the key counterweight to company claims.Date: ** 2026-09-21 (release data; page reflects ongoing independent testing as of research date)
URL unavailable
Secondary Sources (4)
Secondary
OrcaRouter — "Xiaomi MiMo-V2.6-Pro: 46 on the Open-Weights Index" - **URL:** https://www.orcarouter.ai/blog/xiaomi-mimo-v2-6-pro-release** Overall AA index tops: Claude Fable 5.1 53, GPT-6 Astra 53, Claude Opus 5 51; MiMo-V2.6-Pro serving specs (134.3 t/s, $0.13 per index task, 99% cache discount, MIT weights, 1M context); reinforces that 46 is the open-weights top while 53 leads the overall index. — ** Secondary industry analysis; supports index-tier context used in Section 14/11. ---Date: ** 2026-09-21
URL unavailable
Secondary
AI Weekly — "Xiaomi MiMo-V2.6-Pro ties Grok 4.7 atop open-weights index" - **URL:** https://aiweekly.co/alerts/xiaomi-mimo-v26-pro-ties-grok-47-atop-open-weights-index** The 46–46 tie framing: MiMo-V2.6-Pro (1.02T MoE, ~42B active, MIT license) debuted at the top of AA's Intelligence Index among open-weights systems at 46, tying Grok 4.7; DeepSeek V4.1 Pro at 36; Xiaomi's ~$2.62M RL bill and MIT release (weights, code, technical report, Sep 21, 2026). — ** Secondary industry analysis; supports the corrected "tie at top of the open-weights field" interpretation and the price-war context.Date: ** 2026-09-22
URL unavailable
Secondary
LLMReference — "Grok 4.7 — 500k context, multimodal" - **URL:** https://www.llmreference.com/model/grok-4.7** 500,000-token context window; knowledge cutoff 2026-05; provider routes (xAI Console $2/$6; OpenRouter $1.60/$4.80, cache read $0.40); capability set (coding, RAG, agents, long context, vision, JSON/tool use); no batch API; proprietary, commercial use conditional. — ** Secondary spec aggregation; corroborates pricing spread and spec sheet.Date: ** 2026-09-21 (last refreshed 2026-09-21)
URL unavailable
Secondary
OmniaKey — "Grok 4.7 Review: Benchmarks, API Price & 4.6 Upgrade" - **URL:** https://omniakey.com/blog/grok-4-7-review** API detail layer: model ID `grok-4.7`; 500K context; knowledge cutoff May 2026; reasoning effort low/medium/high(default)/xhigh; Responses API + Chat Completions; capabilities (function calling, structured outputs, web/X search, code execution); no batch API; pricing tiers: below 200K prompt tokens $2/$0.50/$6 per M; at/above 200K $4/$1/$12; OmniaKey did not yet list a 4.7 route on Sep 22 (routing nuance). — ** Secondary technical spec aggregation; corroborates AA's 500K context and adds the >200K pricing cliff.Date: ** 2026-09-22 (evidence checked Sep 22, 2026)
URL unavailable
Unverified Sources (4)
Unverified
NextBigFuture — "Elon Claims SpaceXAI Grok 4.7 Will Surpass All Current Models" - **URL:** https://www.nextbigfuture.com/2026/08/elon-claims-spacexai-grok-4-7-will-surpass-all-current-models.html** Musk's Aug 12, 2026 post verbatim ("Grok 4.7 will exceed all current models… SpaceX training corpus is so awesome & unique…"); Grok 4.6 context (launched Aug 2026, ~1.5T); token-efficiency claim. — ** Secondary blog quoting CEO claims — COMPANY CLAIM; shows the pre-launch supremacy rhetoric that the actual launch post conspicuously softens to "highly competitive in its class."Date: ** 2026-08-12
URL unavailable
Unverified
The Standard (HK) / AFP — "SpaceXAI to launch Grok 4.7 model in 10 days to outpace rivals" - **URL:** https://www.thestandard.com.hk/innovation/article/341654/SpaceXAI-to-launch-Grok-47-model-in-10-days-to-outpace-rivals** Musk's Sep 2, 2026 claims: ~10-day launch window (Sep 12 target), 2.1T parameters, "will surpass all existing models," SpaceX corpus uniqueness (real-world engineering). — ** Wire/aggregated reporting of company-founder statements — COMPANY CLAIM/RUMOR; supports the documented gap between Musk's projections and the actual Sep 21 ship date.Date: ** 2026-09-02
URL unavailable
Unverified
OrcaRouter — "Grok 4.7 Release Date: Musk Confirms Mid-September, and What xAI Has Actually Said" - **URL:** https://www.orcarouter.ai/blog/grok-4-7-release-date** Pre-release claim ledger: ~2.1T parameters (Musk, ~40% larger than 1.5T Grok 4.6); "Grok 4.7 will exceed all current models" (Musk, Aug 12/Sep 2); SpaceX supplemental-training corpus claim (Starlink telemetry, Raptor logs, internal engineering artifacts); the pattern of missed verbal deadlines; post-release note that the 2.1T figure is "claim, not spec." — ** RUMOR/COMPANY CLAIM ledger; labelled as founder claims without independent measurement; supports the "unverified 2.1T" and "founder-date unreliability" framing.Date: ** 2026-08-13 (tracked through 2026-09-22 updates)
URL unavailable
Unverified
Kie.ai — "Grok 4.7 Release: What Builders Should Know" (pre-release tracking) - **URL:** https://kie.ai/blog/grok-4-7-2-1t-signal-ai-builders** Pre-release rumour layer: 2.1T parameter figure (community + Musk-attributed); bot-error exposures of grok-4-7-0907; delay reports ("a few more days to cook"; possible RL failure diagnosis); Musk's roadmap comments (Grok 4.7 ≈ Claude Opus 5.0 parity, per attributed relay). — ** RUMOR/comms tracking only; explicitly labelled as community commentary without official spec-sheet confirmation.Date: ** 2026-09-11 (updated 2026-09-14, 2026-09-16, 2026-09-21)
URL unavailable