We finally have an $NBIS investor who can tell the difference between two inference services. Nebius needs to have 70%+ cache hit rate and need to come up with specialize post training for Kimi K3 to match @FireworksAI_HQ is at. There’s something missing in their orchestration for KV cache placement, session affinity, expert placement and/or batching policy. Or on the inference optimization side, something non-optimal in their compilation or kernels.
在 X 查看 ↗