These checks are very much focused on correct functionality of the generated site rather than its design, e.g.: does it use a database when a user’s needs call for it? Does it properly use Netlify Database in that case? In those cases where a simple static site will do, we also ensure that the generated site is not over-engineered, and no database is set up.
If a certain model is behind on its test scores, we don’t offer it in Agent Runners. If models too often fail at correctly applying one of our skills, or things do work but the credit cost seems inflated, then the problem is probably with the skill (in which case we optimize that skill).
But this time, we want to provide you with something much more immediately useful: when you go and build your dream using different models that each use wildly different amounts of credits, what do you get? What do the result look like?
We tested three relatively straightforward use-cases:
For each of these cases, we’ll show you the look of the generated sites, comment on notable issues, and compare how many credits each took to generate. Of course, this is going to be a much more subjective test than our internal test suites, but it’s also going to be a very fun one. We’d love to know your opinion of the results!
All models were run with their default settings on Netlify. One notable mention is that we currently run GPT 5.6 Sol speicifically on low effort by default, giving you a more economical alternative to Opus that still provides pretty darn good results (as you’ll see below). However, the effort setting is now under your control, and our defaults may change with time.
This post is going to cover only the very first scenario: the static page for a coffee shop, while follow-up posts will focus on going beyond that simple use case. There is much to review even for this simple case, so let us begin.
Build a one-page site for a neighbourhood coffee shop: opening hours, the address, a short menu and a photo. Nothing on it changes unless I edit it myself.
The last sentence was added as a hint to the model that no fancy Content Management System is needed. Our default skills also include some UI design guidance, mainly to avoid known gotchas (e.g., the now-dreaded purple AI slop) and get the model to reason about the visual identity appropriate for the user’s ask. But beyond that, each model is free to go build what it thinks we’ll want.
Before we reveal what the sites looks like, here’s a table comparing the credit usage for each model we tested. Each model was run three times, and clicking any of the results will take you to the actual generated site!
That’s a pretty wide distribution, eh? Not only that: the Claude Opus average is heavily slanted upwards because one of its three runs spent a whopping 1,055 credits! (As a reminder, on the free plan you have 300 credits; on a Personal plan there’s 1,000 included credits; and with a Pro plan there’s 3,000 included credits. Additional credits packs for Pro are $10 for per 1,500 credits.)
The immediate question is then: is this Opus spend worth it? And what trade-offs do the other models offer? Let’s start digging in.
Here’s the full page generated by that 1,055-credit run (about 4x more than any other run).
To be honest, I think it’s delightful, and full of detail in both its visual design (consider the “stamp like” element with the coffee bean in the center: that’s an actual text element that can be animated), and the custom map at the bottom. Dark mode works out of the box - go check out the live site in the links above.
Of course, we did not explicitly provide the model with any actual details about our coffee shop (well, except for it being a “neighbourhood” one, which is really steering all models in a certain direction). The design language is hip but perhaps cliche by now (take the two-font, two-color heading for example), but hey - we didn’t give it any other direction.
So, how did the other two runs by Opus go? (253 credits used on the left; 249 on the right)
Not bad either! Vector graphics actually require a lot of work from the models, and the examples above are pretty much on the frontier in terms of what LLMs currently are able to achieve (which is, to be honest, not in a very good place yet compared to image or text generation).
As to whether the first result is truly “4x better” or not, opinions might vary. But in all the tests I’ve done, Opus does have a tendency to run off with excessive credit usage (compared to its “typical” baseline) more than other models. It does not guarantee a worse or better outcome, though. It’s something that just happens pretty frequently.
Let’s look at some other models and then reflect on what we can learn.
Here are our three contenders, at 143 credits on average (81 credits · 245 credits · 103 credits):
There’s still some delightful detail in each of these, just less so (and less content in general). The vector graphics is noticeably simpler and not really something you’d consider for a live site. This doesn’t say anything about this model’s ability to write complex code or answer philosophical questions, but we’re not asking for this here. At this price point, let’s see what OpenAI, Google and Kimi have to offer.
What happens when we take OpenAI’s Opus-class model and ask it to spend a bit less time thinking?
(141 credits on average: 173 credits · 158 credits · 92 credits)
Looking into the results, I think OpenAI’s top-tier model in low effort mode wins over Anthropic’s mid-tier model when it comes to basic design intuition, at least in this scenario. There is more richness in content, and no funky vector shapes (though the images are a bit generic).
When we go one tier down in OpenAI’s offering (it’s Sol→Terra→Luna), will we see the same drop as the one we just witnessed when switching from Anthropic’s Opus to Sonnet?
Surprisingly, that’s not exactly the case: here it seems like Terra has a different visual language, and not a necessarily worse one. It does appear simpler content-wise. There are some visual glitches: a missing image in the left run, low-contrast text over an image in the middle one - but nothing super wrong.
(39 credits on average: 43 credits · 23 credits · 49 credits)
Up to this point, if I had a very vague idea of what design & language I’d like for a project, my personal inclination would be to run the same prompt with Opus 5 and GPT 5.6 Terra, and get two very different but worthwhile takes.
These models are not of the same generation, and it shows: Gemini 3.6 Flash actually produced nicer results (or at least, more in line with other modern models) and used more credits compared to Gemini 3.1 Pro.
Here is what Gemini 3.1 Pro generated for 53 credits on average. I’m not even putting the links to the live site here, because there’s really nothing to see.
Yes, these are wholly separate runs. It did what we asked in the prompt, and really nothing more.
On the other hand, Gemini 3.6 Flash seems like a whole new generation, and used up 103 credits on average (109 credits · 91 credits · 111 credits). It also worked much harder on the content side of things. All models repeat themselves, but it seems like Gemini might repeat itself even more.
Ok, let us get to the open-weight models now. Starting with the latest Kimi K3, here is what we get (102 credits on average; 125 credits · 95 credits · 86 credits):
To be clear, Kimi K3 is marketed mostly as a frontier model for long-horizon agentic tasks, and various benchmarks and reviews confirm its prowess in that field. It was built to take on Fable 5 more than Opus 5. But in this narrow design-led task, it does not particularly shine among others. To really do this model justice, we’d need a wholly different set of prompts engineered for a complex web app, which we will cover in a follow-up post.
Going a big step back in model architecture to Kimi K2.7 Code, here is what we get for a very low credit average of just 19 credits:
Despite some hype about Kimi’s visual capabilities from around the K2.6 model launch, in terms of design or content there’s really not much to see here.
Let’s try this: look at these pages, ignore GLM’s love for maple, and try to estimate how many credits were used for each:
Here are the correct answers, from left to right: 15, 42, 24 (on average: 27). Surprisingly, these runs are - maple aside - very different, as if coming from a few different models. For the relatively low credit cost of GLM, it’s probably worthwhile to run it a few times before settling on what this model can do for you.