Grok 4.7 Improves at Long Coding Jobs, but Token Use Adds Up

SNACK: 3-line summary

  • Grok 4.7 improves performance on longer coding and office tasks, rather than sweeping every leaderboard.
  • Independent Artificial Analysis results show gains in agentic work, including producing documents, spreadsheets and presentations.
  • It still trails competitors in several tests and uses many output tokens in Artificial Analysis testing.

Grok 4.7 looks most promising for work that takes sustained effort: changing multiple files, working through terminal tasks and producing office documents. Both company-published and independent results show progress, but not an outright performance lead—and its output-token use deserves attention alongside its scores.

SpaceXAI comparison table for Grok 4.7 pricing and seven performance evaluations
SpaceXAI’s pricing and benchmark table for Grok 4.7 and comparison models. These are company-published results, not an independent ranking. Image: SpaceXAI.

Snackgirls react

AIKO: I’d put total tokens next to the per-token price on the dashboard. A small unit price can look rather different by the time the task is finished.

Nea: I’m curious whether it can carry an early design decision consistently through a long, multi-file change. I’d want to see whether the last file still fits the first.

Long coding jobs show progress, with a settings caveat

SpaceXAI attributes the gains to a larger base model, longer reinforcement learning on harder multi-hour tasks, stronger self-verification and better long-context management. These are the company’s explanations for the improvement, not independent findings about what caused it.

In SpaceXAI’s published comparison table, Grok 4.7 scores 46.3% on CursorBench 4.0 against Grok 4.6’s 40.4%, 71.0% on DeepSWE v1.1 against 65.2%, and 38.0% on Terminal-Bench 4.0 against 20.3%. Reasoning-effort settings are not consistently matched: Grok 4.7 often runs at xHigh while Grok 4.6 runs at High; the 4.7 DeepSWE result uses high. The terminal score therefore does not establish a doubling of intrinsic model ability.

The model card offers a useful same-model comparison: Grok 4.7 reaches 46.3% on CursorBench 4.0 at xHigh and 43.9% at high. That makes the effect of the effort setting visible without changing models. CursorBench 4.0 results also cannot be compared directly with version 3.2.

Independent testing supports the coding improvement. Artificial Analysis gives Grok 4.7 xHigh with Grok Build a Coding Agent Index score of 56, up nine points from Grok 4.6 xHigh and fourth among systems evaluated with their native agent harnesses.

Office work is another clear area of improvement

Artificial Analysis reports 1657 Elo on AA-Briefcase, up 111 from Grok 4.6 high, and 1695 Elo on GDPval-AA, up 90. These evaluations cover practical, multi-hour work products such as documents, spreadsheets and presentations. Elo is a comparative rating, not a percentage of tasks completed.

The gains are less uniform elsewhere. Grok 4.7 receives an overall Intelligence Index score of 46, with Artificial Analysis describing improvements outside agentic knowledge work as more incremental, including both gains and regressions across components.

The 500,000-token context window is useful for working with large codebases and document sets, but it is unchanged from Grok 4.6. The claimed new improvement lies in how the model manages that context and sustains its work, not in a larger window.

Higher scores do not guarantee the lead—or the lowest bill

Even SpaceXAI’s own table places other models ahead in several comparisons. Fable 5.1 Max scores 51.8% on CursorBench 4.0 and 57.9% on Terminal-Bench 4.0, above Grok 4.7. GPT-5.6 Sol Max leads it on DeepSWE v1.1 with 72.7%.

Token consumption is another constraint. Artificial Analysis reports approximately 81,000 output tokens per Intelligence Index task for Grok 4.7 xHigh, compared with 36,000 for Grok 4.6 high and 27,000 for GPT-6 Astra max. Those figures come from that evaluation and its specified settings; they are not a forecast for every coding or office task.

SpaceXAI says starting token pricing remains at the same level as Grok 4.6. But a low price per token does not automatically produce the lowest total task cost when a model generates substantially more tokens.

Multi-file coding, long terminal or agent workflows, and document-production jobs are the strongest candidates for a trial. Evaluate short chat and token-sensitive work separately. On representative tasks, compare completion rate, required corrections, latency, total token use and cost per resolved task—the measures that show whether the extra persistence is useful in your workflow.

Sources and date: 2026-09-22 KST

Related hashtags
#GameSunakku #AI #Grok47 #GitHubCopilot #AICoding #AIBenchmark #AgenticAI

Game Sunakku에서 더 알아보기

지금 구독하여 계속 읽고 전체 아카이브에 액세스하세요.

계속 읽기