$NVDA just showed why Groq 3 LPX could become a major piece of the agentic AI stack. Groq 3 LPX hit ~3,431 tokens/sec at 100K context versus ~870 for the next result showing why Nvidia is splitting heavy context processing onto Rubin and latency-sensitive token generation onto LPX. Nvidia clearly redesigning inference around how agents actually work letting it own more of the stack as agentic AI scales.
View on X ↗