Need help?
<- Back

Comments (56)

  • Systemerror7A69
    Am I wrong or are these evaluations, while interesting, not really meaningful for anyone doing serious development work?I'm asking because I personally only use AI with specific and detailed instructions, building my projects piece-by-piece.I mostly don't look at the low level code and some of it I don't understand as much as I'd like, but I very much give much more technical instructions than a simple, two sentence prompt.So, it seems to me this "oneshot from a simple prompt" eval is fairly meaningless when it comes to model evaluation itself, as it is in no way representative of real world application.This would be more something for "vibe coders", people with little to no programming background wanting a website?
  • sinuhe69
    Having skimmed through the article, I assumed that they hadn't tested the design for mobile devices because it wasn't mentioned. That would be a huge mistake. In today's web design, a mobile-first approach is imperative. This is even more important if you want to showcase your café. I used the developer tool in Firefox to see how this design would behave on mobile devices. There are huge differences.Some designs use the screen estate so ineffectively that only the title and a big, boring, generic graphic is shown on a phone. Users have to scroll all the way down to see the content and find what they need. Better designs show the menu, navigation points, and meaningful, aesthetic graphics. Other designs, such as Gemini 3.6, were quite sophisticated but not optimized for traffic and would not load on a 3G connection. However, a simple static website should load instantly on a mobile connection.That said, even a simple web page has many requirements, so expecting a turnkey, ready-made design if the user is not guiding the process is not realistic. Thus, I believe the best choice nowadays is a model with good design skills that understands and adheres to an iterative design process, offering a good initial design as a starting point but also prompting the user to provide guidance and feedback. As the design process runs through many cycles, the initial cost should be modest. But more importantly, the model should understand its own design, be able to explain its choices so it can converse with the user using concrete elements in a accurate language to guide the process. IMO, this is still missing from even the frontier models today.
  • isqueiros
    > Build a one-page site for a neighbourhood coffee shop: opening hours, the address, a short menu and a photo. Nothing on it changes unless I edit it myself.If that's the entire prompt, it's quite depressing how much alike these all look. I appreciate some of the details from the Opus 5 version, but I can't help but strongly feel the AI vibes emanating from that design.
  • jwr
    I've been doing a lot of benchmarking for a long time now with a number of local models for the purposes of spam filtering. The major observation is that there is a lot of variance in model performance. This should not be surprising, as these are probabilistic machines based on random numbers, so your performance will vary from run to run. But this also means that any sort of evaluation of benchmark with a sample size of 1 is essentially worthless for the purposes of model comparison.In my benchmarks, I started insisting on having at least 5 runs.This also makes me very suspicious of some of the posted benchmarks that compare models. If the results don't say how many benchmark runs were performed and what the variance was, there is really no basis for comparison. You are just guessing.
  • arjie
    Building ad-hoc evals is trivial these days. And you can place an LLM judge in front to disambiguate between output. e.g. I have Frigate gating a video feed so that an LLM can watch our home cameras and the labeling task takes a few minutes and then the evals run at fairly low cost https://wiki.roshangeorge.dev/w/images/4/4e/Screenshot_-_Eva...This means that generic benchmarks and evals are sort of passé. There's no need to have them write "make a coffee shop page" or whatever. Instead, simply focus on the real problem you have and evaluate whether they solve that problem. You want to move along the price-performance frontier alone and the SOTA models are decidedly at the top of the performance curve but very expensive so they serve very well to produce the gold standard and to judge.The change from prior to today is that benchmaxxing and per-token pricing allowing for evaluation point to the same direction: do not proxy results. Instead deploy the highest end model you have as a judge for others until you have statistical confidence in discrimination and move along the task-specific frontier.
  • anon
    undefined
  • andai
    > Of course, Opus will also perform relentless self-validation of its own work (it does not bill itself on good looks alone). But remember there’s certainly a higher-than-average credit cost attached to that.I haven't tried it but I think this can be replicated with a system prompt.I remember the codex system prompt contains something like, "Do not consider a task complete until you have verified the result."Although I've been running the new GPT models in a custom harness and they do that anyway now, without being prompted. So I think that prompt was for a previous generation.
  • Schlagbohrer
    Finally some really useful apples-to-apples comparisons rather than endless benchmarkmaxxing and anecdata. I want to see this test run again for the open-weights models that fit into 128GB combined RAM. Like how does Muse-glimmer compare with qwen3.6-35B?
  • anon
    undefined
  • s4i
    In my opinion, this kind of a benchmark doesn't tell much about the models' capabilities on normal software development tasks. It's fun to look at the differences in the output of course, but how often would anyone prompt with very brief instructions, without even hinting the model about caring about any of the details in the outcome nor the implementation?When there are implicit boundaries, negotiable tradeoffs, taste, whatever, in the mix, then the differences in the model capabilities become way more interesting.
  • tealmoonx
    You betta lose yoself in the contextIt’s not theft, you own itYou betta neva use Go (go) (go)You only get 1 promptDo not use canvas (No!)Cause opportunity comes once in a lifetime
  • feor
    I like the output of the cheaper/older models better, surprisingly, say Gemini 3.1 or DeepSeek V4. They're mostly no frills and just text, and ironically look less AI-generated to me because of the lack of hip slogans and graphics. Definitely closer to what I'd want for my own site, but I don't claim to know what people want from a website for a coffeeshop.
  • pedrosbmartins
    Pretty interesting how DeepSeek V4 Flash 0731 has such disparate (and cheap!) results. I would never guess they come from the same model and prompt.
  • sceptic123
    My "favourite" site was <https://6a6fa3376288679d094a8437--ar-testing-coffee-2b15c6a0...> with footer image captions that are amazing. Precision Late Art being the best one.
  • edgyquant
    Benchmarks are so difficult with ai because as soon as one gets popular it enters the dataset so the next iteration of the model is trained on the solution. I’m not sure if there’s any potential work around here or anyone doing interesting work but would love to hear about it if so
  • neom
    I'd be curious to see Terra xhigh vs Sol low, only in that the visual languages are actually kinda different, that Terra work was a lot less "llm" feeling than a lot of the other designs, wonder if pumping up the effort would result in a more in-depth design but within that style.FWIW I enjoyed reading this way more than any usual benchmark posts we see.
  • throwa356262
    This seems to be the easiest way to get on HN front page:Come up with an arbitrary test, let a bunch of LLMs work on it. Make some very subjective judgement about the result...
  • jannishan
    I think we should establish a professional evaluation organization for models; otherwise, it will be difficult for informal evaluations to form standards.
  • horsawlarway
    Approaching this from the perspective of a potential customer and not a designer - I find that I actually like the smaller/open model output quite a bit more.Kimi, GLM, and Deepseek all absolutely run away with the "Can I quickly read the menu and find the address" challenge.Most of the rest of the pages are stylistic, but hard to parse.If I were in a car on a mobile phone trying to find the address of the place to meet a friend for coffee... I don't want a bunch of fluff and stylistic design that makes it hard to parse the information on the site.
  • sorokod
    Same model over 11 days, one prompt, different results?
  • kifler
    I always enjoy these comparisons between models, especially when they demonstrate the actual costs in addition to the outputs.
  • plumb_samji
    Interesting exploratory comparison, but I be cautious about treating it as a model benchmark With only three runs per model, the results are highly sensitive to randomness
  • nullzzz
    This would be interesting, if I was into building coffee shop sites and todo apps from scratch. For the HN crowd tho, I’d say these are toy examples. No offense!
  • chrisjj
    > Vector graphics actually require a lot of work from the modelsHow so? Surely they can just steal such generic graphics off existing web sites.
  • Damjanski
    this is fun!
  • reindeer2
    [flagged]
  • TrustScoreAgent
    [flagged]