Need help?
<- Back

Comments (32)

  • hbrn
    $40m in funding, 2 years in stealth.Performs on-par with SemIf which was built in a couple days and apparently uses raw Qwen, with no fine-tuning. SemIf runs in your freaking browser. Oh and Jev is twice as expensive?Is it surprising that Jev consistently thinks it's Qwen?I'm almost convinced that Jev is a scam. Take Qwen, fine tune it a little, tell investors it cost $10m, spend $1m on advertising, profit.
  • dmix
    You can spot vibecoded websites by how they include the prompt or commit-style comments into the literal interface, instead of communicating it via visual context (or simply excluding it)> Sort by any column; values the run could not produce always sort last. Hover a cost for how it was priced, a latency for the endpoint. Names link to each project.A designer would never write this, but an LLM just inserts it by making it small grey text next to the interface, just like it does with inane code comments.
  • faangguyindia
    Can you please add our model to benchmark? https://gambler-relay-us-west1.leo-fish.ts.net/demo
  • ks2048
    I was trying to figure out what exactly the tests here are. I guess I found some of the questions (here: https://github.com/fstandhartinger/jevbench/blob/main/datase...)e.g., "instructions": "Which intent does the user's message express?", "labels":["set_alarm", "play_music", "weather", "send_message", "turn_off_lights"], "state": "Play some Taylor Swift.", "expected": "play_music"
  • sean_pedersen
    Good project but this one also exists https://huggingface.co/spaces/multimodalart/jev-decision-ind... and the results do not seem to add up and also model sets are different... still needs time to mature likely
  • swyx
  • adityamishra241
    How do you handle task distribution and prevent the benchmark from favoring models that are tuned specifically to these 534 questions?
  • DylanMerigaud
    534 English decisions in one full run sounds substantial.
  • anon
    undefined
  • janalsncm
    It is strange that you put the BGE reranker in the list but not BART which is an actual zero shot classifier.
  • nzoschke
    https://is-it-ai-slop.app.mintapis.com/ is a fun tool. Is the source or methodology for that in the github repo? I couldn't find it immediately.We've been experimenting with Jev for classifying email, some thoughts here: https://housecat.com/blog/classifying-emailFlagging AI written email is a much requested feature too.
  • pushpendraw
    the slop detector giving 86% confidence on a keysmash is the real finding here, not the leaderboard score. confident and wrong is worse than an LLM that just hedges.
  • visarga
    [flagged]
  • kevinbaiv
    [flagged]
  • joserobles84
    [flagged]
  • anon
    undefined
  • rg1992
    [dead]