<- Back
Comments (237)
- simonwThis is fantasticI've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations.Catching a lab cheating specifically on my one dumb benchmark would be really funny.Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than anything I was considering.His conclusion:> Nothing jumped out at me. I couldn’t find a case where the pelican-bicycle images looked noticeably better than the rest of that model’s grid.
- mauvehaus> All 21 pelican-bicycle images, across all seven labs, face right. No other animal/vehicle combination does that.> However, facing right is common: 60% of all 1,008 images do it. How common depends on the animal and the vehicle, and bicycles are one of the two vehicles where it’s strongestOf course the pelican on the bicycle is facing right. The drivetrain on a bicycle is on the right side. If you want any representation of a bicycle that shows the drivetrain you're going to show the right side of it if you want to do so without the frame occluding it. It's an excellent bet that their training data reflects this.Citation: https://www.rei.com/c/bikesEdited to add:As near as I can tell, all of the bicycles are shown facing right, regardless of the direction the animal is facing (GPT 5.6-Terra, Sample 1/3). Also, in every case where the rider has legs (i.e. not the whale) both of the rider's legs are on the right side of the bicycle. This suggests a pretty serious lack of actual understanding of how a bicycle works.
- stusmallI'm glad someone ran the numbers on this. Every single Simon Willison post of an SVG is followed with someone dismissing it saying "I'm sure they train on it by now." This is despite a good blog post with sound logic on how easy that is to catch. [1] Glad to see someone took the time for a quantitative analysis of dumb little animals riding dumb little bikes.1. https://simonwillison.net/2025/Nov/13/training-for-pelicans-...
- elliottoBike nerd + AI nerd here. The author's observation that all bicycle images face right almost certainly has to do with the convention to photograph a bicycle from the right. From the right, you see the drivetrain - this is good for aesthetics, but also for marketing - the drivetrain is branded and labelled and a buyer will want to know what model it is.There is a bunch of guidance online on how to photograph bikes, and every sales image of a bike will be from the right. You can anecdotally observe this by google imaging 'bicycle for sale'.
- SyneRyderHuh. They're not "Pelicanmaxxing"... they're Ottermaxxing.Take a look at the GLM 5.2 and Deepseek V4 "animal on a plane" examples. In every case, the animal is standing on top of the plane, a clear misunderstanding of the concept of "animal on a plane"... with the exception of the Otter. The otters are sitting in a seat on the plane, looking out the window.That's Ethan Mollick's "Otter On A Plane Using WiFi" image benchmark.https://www.oneusefulthing.org/p/the-recent-history-of-ai-in...(Sometimes the Racoon is sitting inside the plane as well, but the racoon is a common backup benchmark. I'm surprised it wasn't also holding a sign saying that it loves trash.)Also, Grok seemed to really really enjoy "whale on a plane" in that second round, and kudos to GPT Terra for deciding after 3 rounds that the user was terrible at spelling and generated "Antelope On A Plain".EDIT: I promise I'm a human, but I did just notice my "that's not x... that's y" construction at the start. I am rather Claudepilled — my apologies.
- bnfclThis is funny, I actually did a similar experiment just yesterday.Looking for evidence of the same, but with another twist: checking if the models would choose to create a pelican on a bicycle, if no specific bird or method of transportation was specified.My version of it: https://www.modelbias.ai/pelican-on-a-bicycle-test
- waterproofI recently had the following conversation with Claude:Me: how many P's are in the following text? [Pasted text]Claude: There are 14 P's, all lowercase (no capital P's)Me: how many in "strawberry"?Claude: there are 3 R's in the word "strawberry".
- Wowfunhappy> The more plausible story is SVGmaxxingExactly--and you have to ask yourself at this point what "maxxing" really means, since "get better at drawing SVGs" is a useful skill.
- fennecfoxyWhat I find interesting, though, is how the "feature" isn't really exciting anymore, or because it creates SVGs that aren't "perfect" that we maybe feel underwhelmed sometimes (at least I do).However thinking about it...if someone asked me to manually create an SVG, or hell even draw a quick doodle on a bit of paper of the same, I'd still probably be much slower than an LLM and potentially end up sketching less accurate anatomy than the machine.I think the general "organic task" stuff has been mostly sorted out, but in personal and professional experiences using AI to try to _do_ something, I've found less so recently problems with hallucinations and moreso problems with attention.For example GPT5.6 still has issues where if I provide it with a list of documents and then ask it to raise questions from that information. Then provide it with additional documents that answer some of those questions and ask it to summarise which outstanding questions there are again, it still asks questions that have become irrelevant with the additional documents - but when this is pointed out it knows exactly what to do and produces the correct list of outstanding questions.I'm sure frontier models are doing all sorts of crazy stuff with attention already, but it seems to me like we almost need some hierarchical attention mechanism like KVL (with Level added) so that it's aware not only of semantic connections between tokens in the context but also of where there are gaps, missing links to assist the model in becoming aware of its own attention span (I guess).
- simonwUnderlying data is available on GitHub: https://github.com/dylanjcastillo/blog/tree/main/_extras/pel...
- esnardNice article! Thanks for experimenting and writing it.I'm curious if some of the animals / vehicules might force the models to use more tokens than others, and I could not find the token counts in the shared data, is it possible to publish it please? :)
- apwheeleSo this is not my experience at all for asking about simple SVG icons for web-pages. Here is one of the examples I have tried for in the past, make a simple cartoon SVG knife for a map icon for a crime map.https://x.com/CrimeDecoder/status/2080008114615537766Can see the images for ChatGPT/Claude (Sonnet 5), and Gemini are all quite bad.Jagged edge of LLMs. How do you explain being able to generate very complicated shapes in the Pelican example but cannot make a much simpler icon without just alluding to it is in the training data?
- throwaway6s1dfI don’t know how anyone with a neutral view can confidently take this:> Direction: All 21 pelican-bicycle images, across all seven labs, face right. No other animal/vehicle combination does that.> However, facing right is common: 60% of all 1,008 images do it…and the tables in “Evidence #5” to be anything but evidence the models have likely trained on pelican on bicycle data more than others.The data clearly shows:- 100% pelican on bicycle facing right- significant skew to the right for bicycle-like vehicles- significant preference for right facing for birdsAveraging those extreme results to “60%” to make it sound like it’s pretty fair because it’s close to “50%” isn’t statistically sound.The methodology is generally unsound. There is no actual scoring with a well defined rubric, it’s just vibed with a single model (GPT 5.6 Luna).The “not better at drawing” evidence are equally hard to take seriously when there is no clear, non-subjective indication of what better or worse is.
- jboss10Some of these are quite nice(I like gemma 3.5 flash's work)https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...
- dlluI feel like getting LLMs to spit out an SVG is akin to getting a human artist to draw something by just reciting a list of coordinates. It's insanely hard and unnatural.Image generation models nowadays can easily generate a photorealistic pelican riding a bicycle, where the bicycle has perfect structure. But it is, of course, only a raster image.It seems that we're missing a kind of step to decompose an image into a list of instructions (say, SVG paths, or even brush strokes with a real brush) to reproduce it properly. Doing so would probably need a true understanding of the structure of the scene, which is something that AI still struggles with to this day.
- ianberdinWe have a little bit different conclusions.Same pelican grid + MacBook Pro in 3D. Also with end cost.https://playcode.io/blog/macbook-svg-benchmark
- 2001zhaozhaoI feel like the simplest Pelicanmaxxing method is just to teach the model that whenever a user asks for a svg illustration of something with no other clarification, it should default to making it as detailed and pretty as possible. This would make every single svg from that model look better and not just the pelican on a bicycle
- AussieWog93There's every chance here I'm just being annoying pedant, but generating a bunch of SVGs of animals on vehicles doesn't mean that it's good at generating SVGs in general, just SVGs of animals on vehicles.On the flip side, GPT 5.6 Sol did a pretty convincing render of a burglar eating salami.I'd be curious to see how the other models on random things that are completely tangential to pelicans or bicycles.
- richardwAnd a whole universe of random tests got baked into the AI’s training data. Websites of antelopes driving trains and hammerhead sharks swinging in a tyre swing were created. It was a short while until AI became so focused on animals that it gave up competing with developers. Life became sane again.
- ListeningPieThe blog has it's own comment section under the article that's completely blank and yet clearly there is a lot of interest with 197 comments. Having an empty comment sections sends the wrong signal.
- anshumankmrFor an MVP I am building I asked to make a brain SVG and every attempt it has been doing it has been hillariously wrong and I had told Opus4.8/Fable to pick some SVG it could find online and it went ahead and still used its own thing that turned not too good. (full disclosure not particularly that good at front end stuff so relying heavily on Claude and Codex for it)
- inigyouWhy do all of the images just say "failed"?
- reilly3000I fear there this reveals something else - not about the nature of models but of our community. The intro of the article seemed to affirm the importance of HN as a tastemaker, potentially influencing decision-makers on the scale of billions or trillions. For the 15 or so years I’ve been hanging around here, that only feels like mild hyperbole… a good LaunchHN can reshape the future, right?It made me realize that for this time around, we’re not at the center anymore. The future of LLMs is the stuff of nations and AIG is what the labs actually care about. They aren’t pelicanmaxxing just as much as they really aren’t revenue/margin/marketing maxing. They just want our attention and ideas so they can show growth and acquire FLOPS.The apparent fact that they aren’t catering to this community (who frankly decides what goes and stays in prod) leaves me feeling a bit defeated somehow. And also awesome?! Like there is a culture here that runs deeper than any technology and it has fought hard to maintain its identity. Props to dang and all for keeping the astroturfing so imperceptible that I can say this.In any case, a great little piece of citizen-science dcastm. Lmk if there is a way I can chip in towards token costs.
- anonundefined
- rldjbpinthe pelican bike combo is truly an HN phenomenon, and i think the hypothesis carry a lot of weight about the assumption of the mindshare this website truly has in the industry.very nice approach to test it and might be a nice way to "grid search" evals in other use cases perhaps.
- ertgbnmI've had the feeling that labs aren't pelicanmaxxing specifically but that they do have some sort of RL environment for SVGs that they are letting the AIs overcook in. Specifically I'm thinking of the gemini 3.1 pro annoucnement that seemed to have a huge leap in animated SVG performance but not much else impressive about it.So they aren't pelicanmaxxing but they are benchmaxxing in a way. The benefit of the pelican was originally that uplift on the pelican signaled an overall uplift on intelligence. I don't believe that is the case anymore and it is just another jagged edge of model intelligence.
- pjmlpI would expect so.In the last century, as spec tests for C and C++ compilers, databases, Java application servers became trendy, all vendors were optimising for great articles on the respective technical magazines.
- bluealienpieAI rating AI? Am I missing something.
- nostrademonsIt's really refreshing to see someone publish a null result.
- munk-aIt's a method to grade LLM output - as such it's something that will receive focus in correcting for. As soon as people who have a say in where funding is going noticed it as a metric the labs started caring about their performance in it. In the best case the labs are focusing on improving SVG capabilities in general and optimizing Pelican production as part of that initiative - but now that it's a known measure it is no longer reliable.
- luciana1uthe benchmark was supposed to measure if models can solve novel problems and instead it became a benchmark for how fast labs can solve the benchmark
- oasisbobAs a unicyclist, all the sideways-riding caught my eye for being especially silly.However, I'm very surprised that most of the models make the same sideways mistake with only some of the animals, and they do it consistently.With most of the models, cat, raccoons, and otters are almost always riding sideways. Why is that?
- antonyragleapThe interesting question is whether this generalizes beyond pelicans to benchmarks the model hasn't seen before.
- davidkunzMaybe they're animalsonvehiclesmaxxing.
- ionwakeAbsolutely fantastic top tier post
- pasquinellii'm puzzled by the choice to have an llm judge the images.
- gpjanikTLDR: the experiment asks for in-distribution responses and gets those.The right answer here is to ask a LLM to create a scene similar in quality to those, but completely out of distribution.I asked GPT 5.6 Sol to give me a pelican playing football on San Siro while smoking a cigarette, in AC Milan's t-shirt. While this sounds like higher complexity of a problem, the generations from current models often include additional details like scene composition, scarf, etc., I don't ask for, so I wanted to see what here is memorization vs. composition skill."write svg code of a fish playing football on san siro in ac milan's t shirt, with raybans on and a cigarette."Try that on GPT 5.6 Sol, Fable, or whatever other model. It's chaos.
- oaxacaoaxacaHilarious question. Imagine someone woke up from a 7 year coma and read this title lol
- robvirenWhy let your dreams be dreams? This is a perfect example of following a hypothesis. I love when people dive into an esoteric subject and just go full swing. Reminds me a CGP Grey and the name Tiffany. Sometimes you just need to know.
- scosmanjoin me in building the ideal training set for pelicans riding bicycles: https://github.com/scosman/pelicans_riding_bicycles
- jonatronOK, so we've done animal_vehicle, how about new SVG ideas each time? I just tried "make an SVG of a man sitting in a chair at a computer behind a desk" which gives more interesting results than the animalVehicle test.
- BeetleBOh great! You've now made it a lot easier for LLMs to train on this dataset!Your next iteration will need different animals and different transportation options. You'll run out after a few iterations.
- pbronezNice use of regression analysis to understand your experimental dataset> But again, some combinations might be just harder to draw than others.> To account for that, I fit a fixed-effects regression on all 1,008 images: score ~ lab + animal × vehicle, plus per-lab interaction terms for pelican, bicycle, and the pelican-bicycle cell, with robust standard errors. The animal × vehicle terms absorb the inherent difficulty of all 48 combinations. The interactions measure each lab’s benchmark-specific boost relative to the average lab, with confidence intervals.
- johndoughAnother point for consideration: Specialized SVG models create way better looking pelicans riding a bicycle. (E.g. Refract V4: https://jumpshare.com/s/8liB7Aiuoo3yucbWGXjZ mirror: https://postimg.cc/McV70p84 )
- Rooster61I find it humorous that the animal + plane combo appears to be such an outlier. I assume this is due to the models assuming the user mean plain and misspelled it in the prompt.
- CopenjinUsing an LLM to judge drawing giving a rating? Not the best idea.
- PeterStuerAny benchmark gaining some, even a little, traction before a model release date should at this point be considered tainted. Create your own, never publish it or write about it in any detail.
- bvrmnit's quite interesting scoring and conclusions. For me Gemini renders are definitely top outliers for all combinations.
- tomas789Having an objective score is quite difficult. Maybe it would be better to do a pairwise comparison and calculate ELO?
- 6thbitWhat would be an alternative format or process with a similar effort to drawing SVGs?
- stri8tedYou seem to assume training on pelican would not result in improved performance on other similar tasks. Why?
- 40fourOkay, first off, without honestly reading the whole article (I tried but I just don’t have the patience), a quick red flag is the final analysis is only through Fable? Isn’t that inherently going to introduce unwanted bias?Whatever. Doesn’t really matter much. My next thought is, I get that requiring an SVG is adding an extra layer of complexity as far as the art goes, but why is nobody talking about that the actual art is absolute trash?I get it. It’s basically a meme at this point and it’s a fun game to play with the models. But my thought is it should be illuminating to anyone who is an artist that LLMs are still a long way off from taking your job :)
- comrade1234Hilarious. Could you imagine being a programmer at an AI company and this is your assigned task?
- RobRiveraChasing metrics Chasing dragonsTomato, tomato
- andy99If an AI researcher was going to pelicanmaxx, they would almost certainly apply the augmentations mentioned in the article during training, e.g. randomly selecting animals and conveyances. You’d want a model that generalizes well, just sfting in that specific prompt would be pretty bush league for a frontier lab.I don’t have any reason to believe they are gaming the benchmark, just saying. I do find the idea of a data labeller having to generate thousands of svgs of different animals on different modes of transportation quite funny though.
- IshKebabThanks for not using AI to write this. So much more pleasant to read.
- busymom0How does attempt 2 by Llama 4 Maverick look like a bald eagle??
- ck2I am not sure if this is how it works but let's say there was a reddit thread talking about the pelican benchmark and in it someone posts mockup examples of what an ideal result would look likearen't some LLM going to digest that thread at some point and indirectly learn from it?basically my point is originally this was a good benchmark because it was an absurd never-seen-before thing, but now that it is in content, some models are going to get a benefit in education?you'd need the "AI" equivalent of an old-school "google whack", something with no previous results* https://en.wikipedia.org/wiki/Googlewhack
- j45The models definitely seem to pay attention to the tests.Since the tests can be generally gamed with directing descriptions at it non-deterministically, there's a greater chance the questions solution can be found.Of course, hopefully the models are instead adding patterns and types of questions as well and it makes the models more capable, but it may be limited in how it transfers to other types of questions in breadth or depth.
- cute_boihttps://playcode.io/blog/macbook-svg-benchmarkI think we should stop using pelican benchmark.
- dcchambersIt's incredible that each model has it's own style that remains relatively consistent throughout all of the different generated examples.
- TZubirihttps://en.wikipedia.org/wiki/Goodhart%27s_law"Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes."Or the more pop layman version"When a measure becomes a metric/KPI, it ceases to be a good measure."Story time, I live in Argentina, and we don't have Big Macs, the main Mc Donald's brand, here, because during the CFK presidency, one of her tactics was to Goodhart economic metrics. Even the informal obscure ones like the [Big Mac Index](https://en.wikipedia.org/wiki/Big_Mac_Index), I don't know the precise details, but the Big Mac ended up being a very cheap item, like 2 or 3 times cheaper than actual menu items, but it was never on the advertised menu, and it also ended up being very small compared to the other burgers, so it wasn't even like a hack, a shrinkflation type of deal.But hey, anyone who read the Big Mac Index table would never find Argentina at the bottom of that list along with a couple of other countries with bad brands, so the ploy worked. And now we live with the aftershock, the brand never really turned around, other brands with ridiculous names took over it like the McTasty, which makes me sound like that skit from Tarantino's Pulp Fiction.
- edifierxuhao[flagged]
- tim_tihub[dead]
- Ilya85[flagged]
- Ilya85[flagged]
- TheSpacerr[dead]
- sbseitzI wish I could downvote this for Pelicanmaxxing lmao.
- andrewstuartThe pelican prompt is ridiculous.Test the LLLM against things you want it to do.Asking questions that are absurd is like interviewing developers and asking absurd questions on the grounds that it tests creative and critical thinking.Remember these Microsoft interview questions designed to identify the best developers?"If you could eliminate one U.S. state, which one would it be?""How would you move Mount Fuji?"Absurd interview questions have an air of legitimacy due to the quasi sophisticated justifications put forward for why they are good tests.Absurd interview questions are not good tests of people or LLMs.Relevant questions are good tests.