<- Back
Comments (19)
- AreibmanCould you say more about how caching works? One major advantage of sticking with a single model is saving money on cached input tokens. I'd imagine if you swap between a bunch of models, you may improve performance but cost would would balloon out of control
- swthbhtVery cool. Does your gateway decide effort levels as well? Or just models?
- akshay_akulaOpen source and no markup is the right default for a gateway. The caching question above is the one I would want answered before swapping models though.
- ceroxylon>The gateway adds under 1 ms for BYOK requestsAmazing! Really brilliant idea, thank you for sharing this project. There is so much ground to cover in the LLM gateway / routing / reporting world, and this is a great start. The Tinker implementation is my favorite part, fine tuning is much better than a sea of context files.
- cheema33I have not tried it yet. Is it similar to LiteLLM? If so, what sets it apart?
- 0xbadcafebeeYou started it a week ago? I look forward to checking back in 3 weeks when you've exited for $1B
- 23davidSuper interesting and congrats on the release. Curious if you initially had this in Python and then rewrote in Rust?
- ashermaniaFinally an open source tool doing this!
- jing09928[dead]