<- Back
Comments (75)
- simonwcurl https://api.openai.com/v1/decisions \ -H "Authorization: Bearer $(llm keys get openai)" \ -H "Content-Type: application/json" \ --data ' { "model": "gpt-6-luna", "input": [{ "role": "user", "content": [ {"type": "input_text", "text": "I am angry about the new product feature"} ] }], "questions": [{ "type": "predicate", "name": "complaint", "instructions": "Is this a complaint?" }, { "type": "predicate", "name": "compliment", "instructions": "Is this a compliment?" }] }' Returned: { "model": "gpt-6-luna", "answers": [ { "type": "predicate", "name": "complaint", "probability": 0.91 }, { "type": "predicate", "name": "compliment", "probability": 0.06 } ], "usage": { "input_tokens": 310, "input_tokens_details": { "cached_tokens": 0, "cache_write_tokens": 0 }, "output_tokens": 0, "output_tokens_details": { "reasoning_tokens": 0 }, "total_tokens": 310 } } That https://api.openai.com/v1/decisions endpoint is notable because usually when OpenAI define an endpoint like that it ends up as a defecto standard for other providers.(I turned this all into a new llm plugin: https://github.com/simonw/llm-openai-decisions)
- TSiegeThe response to Jev should be the nail in the coffin over whether or not the AI business is a commodity market.Out of no where Jev appeared as the next round of the price wars. Jev showed the value of System One models. A fast yes/no/confidence score not only is cheaper but also often all people want. Open source versions flood hugging face and now the big players are giving up a potentially big driver of output tokens to keep customers and race to the bottom price wise.If I were OpenAI or Anthropic I’d be racing to make their products as sticky as possible bc ppl will flock to what’s cheapest otherwise.
- elpakalHave there been any signals from Anthropic about matching this? We use AWS bedrock and just switched to Anthropic from OpenAI because of the ZDR guarantee. Would be great to not have to entertain switching back.
- TopfiRan my decisions evals (still rudimentary, less than 600 calls (UI component selection, chat charting, tag selection, PKM stuff)) on this via OpenRouter against Jev and Mercury Decide. Jev because it has replaced my mt0 efforts by sheer force of affordability (more importantly, the limits running on a MacBook Neo bring even after vocab pruning and quant insanity) and Mercury Decide because I do like dLLM efforts (and I'd like to use fewer model providers if possible).Preliminary of course, but seems to be slower than Jev and similar to Mercury Decides latency, though not in growing linearly with the amount of input (346ms p50 and 860ms p95, (Mercury Decide also had some extremes up to 1,3s that were around 800ms today, likely preview related, it scaled far more consistently with size)), less "confidence" concerning my ambiguous UI component and response shape specific tasks (have very specific use cases for these models which Luna often fails to meet at 0.6 and lower), lead to a few failed calls which neither competitor had (4 vs 0 for both) and measured more expensive than Jev to boot by a factor of 3,1 times on average (Mercury Decide pricing I think is still unknown so no numbers there).Basically slower, more expensive and less capable than Jev, roughly on par with Mercury Decide (provided, in my insane set of use cases and requirements that are a PKM focused Firefox fork with multiple infinite canvas using decision models to improve information synthesis from multiple sources).Seems a bit undercooked overall and I'd rather frontier-labs don't jump on bandwagons until they can offer something competitive in price, performance or both. In fairness, though, I have yet to test image input, maybe that makes all the difference. Also, again, mine is unlikely to reflect everyones use case, so interested in seeing others results.Didn't comment at the time, but having read up on Devday after the fact, there seems to have been a lot of that going around. Notion and GDocs, Jev, Muse, most seems to have been cloned from existing competitors (and despite infinite, ultrafast, ultra code tokens with unsandboxed Mega Astra not that amazing to boot).Prefer less announcements, but focused and at a higher quality. Considering ChatGPT Atlas (their Chromium based browser) and its insanely fast death, I'd be skeptical to put much into any of these even if they were in some way an improvement over what is out there. Maybe focus on a fresh pre-train and some sandboxing improvements.
- AM1010101It’s interesting they skipped caching. I could see wanting to ask follow up questions so having your first x tokens in cache would be interesting.Also if you have a long “system prompt” then caching would have saved a considerable amount on bulk data processing.There may well be a technical reason I don’t understand.
- agentdev001Note that, being that this is gpt-6-luna under the hood, this offers you 1m token input window, and multi-modal (image) input. In my testing so far, I'm seeing 160-175ms end to end. Worst 5% 285ms, worst so far was 743ms.
- AM1010101How well calibrated is it? Is 90% calibrated to be correct 9 times out of 10?I think Jev had put significant effort here and its not clear if luna will be well calibrated in this way.
- sidcoolJev really shook up the industry. This seems obvious in hindsight
- ashu1461If we compare this with using the older solution of writing a prompt to find out the answer of the classification- Cost : It is the same for both scenarios $0.10 per 1M tokens- Speed : decisions is 10x faster than responses API- Quality : I guess if we compare with luna which is a pretty good model it itself, both will be at parSo essentially it has to do more with speed vs any other factor.
- mrkn1If you rather run your decision model on your CPU, check gutsy [0][0] - https://news.ycombinator.com/item?id=49976996
- nicoJust going to drop this here: https://jeffyclassify.com/Open source classifier models you can run and train locally on CPU
- mritchie712it already supports image inputs, which was the first big gap I found in Jev.
- mohsen1Since it is fast and understand images, I wonder if it can play video games. I have a harness setup for the LLM play EA FC but even the fastest LLMs are too slow for it. I need to try this with Decisions API
- jasonjmcgheeDifferent APIs for different things reminds me of the early auto-complete vs instruction apis.Will this get folded into models / post training pipelines at some point and make them better at calibrated outputs?
- swader999I wonder why the decision routing isn't just integrated into all models in addition to this stand alone.
- stillatitOne difference between Decisions and Jev (for now) seems to be that Decisions can take image inputs, which is a pretty common need.
- nnoman7808Roblox
- MiroslavPokornyWhat value is there in knowing if a can has a dent ?
- ImustaskforhelpThis rather didn't take long for OAI to create*, I remember people giving opinions and discussions that it won't take too long and that openAI should do it[0], so looks like they were right.Interesting to see where all this leads us and if other major labs follow suitEdit: decisions voice looks really interesting as well[1][0]: https://news.ycombinator.com/item?id=49802161: OpenAI is well positioned to fast-follow Jev[1]: https://developers.openai.com/api/docs/guides/decisions-voic...
- waterTanuki> The tulip became a luxury item and many varieties were introduced. The varieties were classified and the most sought-after, prized tulips were the streaked tulips, especially yellow or white streaks on a red or purple background. These flame-like tulips were highly sought after. Interestingly, the streaks or “flames” of the tulip petals were caused by a virus. The virus is the tulip breaking virus, or tulip mosaic virus.Source: https://www.canr.msu.edu/news/tulip_mania_the_history_of_the...What's old is new.
- lab14How is the pricing vs Jev?
- esafakYou knew it was going to happen! Benchmarks or it didn't happen.
- OutOfHerev3.26.0 of the openai Python SDK covers its use. Those already using the SDK don't need to make explicit HTTP calls.
- peterson_lockCan we use this through subscription?
- hcentelles[dead]
- juanfranpaez[flagged]
- zane_shu[flagged]
- dvtI genuinely do not understand why anyone would pay OpenAI for this. Running something comparable to Jev is pretty trivial. The whole point of paying for ChatGPT is because OpenAI has a bunch of warehouses that can run a zillion-parameter model.Running a decision model is way easier and much cheaper. Are they really just trying to capitalize on the hype here? It feels like they really have absolutely zero moat.