This website uses cookies

Read our Privacy policy and Terms of use for more information.

I installed the Grok command-line tool and gave Grok 4.5 the same work I was running through Claude Code with Opus 4.8.

No benchmark charts. Two builds:

1. Create an interactive 3D solar system from a 200-line prompt

2. Recreate a CRM dashboard from a visual reference

The results made it impossible to give one model a permanent crown.

Round one: build a solar system

In my recorded run, Grok finished the first build in about three and a half minutes. Opus took roughly 35 minutes.

That is one test on my setup, with my prompt and that day's tools. It is not a universal speed guarantee.

Grok's result also looked better to me. The movement felt smoother, the camera controls were easier to use, and the simulation had more visual detail.

Grok's solar-system build focused on Saturn. Captured at 5:18.

It still had rough edges. Labels floated in distracting places, and some views felt chaotic. Grok won the round because it got to a stronger interactive result much faster.

Round two: recreate a CRM dashboard

For the second test, I gave both models an image of a dark analytics dashboard and asked them to rebuild the UI and UX.

This time the completion speeds were close. The design judgment separated them.

The Opus 4.8 CRM result. Captured at 8:28.

Opus produced the more polished version. The chart felt smoother, the cards were better balanced, the spacing was subtler, and the interface looked closer to a product I would keep refining.

The Grok 4.5 CRM result. Captured at 9:49.

Grok's version was good. Opus showed more taste.

The result

| Test | Grok 4.5 | Opus 4.8 | Winner in this run |

| --- | --- | --- | --- |

| Interactive solar system | Much faster, smoother, more detailed | Slow, functional, visually rougher | Grok 4.5 |

| CRM recreation | Strong and usable | More polished, better spacing and chart treatment | Opus 4.8 |

Grok was the better choice for speed and getting a complicated visual system moving. Opus was the better choice for refinement and product feel.

That is a more useful answer than "Which model is best?"

A pricing correction worth making

The video discussed Grok as dramatically cheaper. The current official API list prices support the direction, but the sticker-price gap alone is not 17x.

At the time of this article:

  • Grok 4.5 lists at $2 per million input tokens and $6 per million output tokens

  • Claude Opus 4.8 lists at $5 per million input tokens and $25 per million output tokens

That makes Grok 2.5x cheaper on input and a little over 4x cheaper on output before caching, long-context rules, tool charges, token efficiency, or platform markups.

SpaceXAI also reports that Grok 4.5 used about 4.2x fewer output tokens than Opus 4.8 on its SWE Bench Pro comparison. That is the developer's own benchmark, not a promise for your project.

The only cost number I trust for my business is the cost of completing my actual workload.

The comparison has already moved on

This test is a snapshot from July 9, 2026.

SpaceXAI released Grok 4.6 in August. Anthropic released Claude Opus 5 in July. That does not erase the test. It reinforces the point: model rankings move faster than most teams can rewrite their stack.

I would not hard-wire an entire business to one model unless the advantage is specific and measurable. Keep the prompts, tests, and acceptance criteria portable enough to run the same job somewhere else.

My take

Use Grok when speed, agentic execution, and cost are the priority. Use Opus when the work depends heavily on design taste, polish, and careful long-form collaboration.

Then test again next month.

The top models are close enough that workflow often matters more than brand loyalty. The best builder is the one who can evaluate the output, switch tools without drama, and keep shipping.

Tools from the video

Disclosure: Some links below are affiliate links. If you choose to buy through one of them, I may earn a commission at no extra cost to you.

Want to learn and build alongside other AI builders? Join us here.

Sources and notes

This is my result from two specific tests, not a controlled benchmark or a guarantee of speed, quality, cost, or business results. Model versions, pricing, limits, benchmarks, and availability change quickly.