Nvidia Launches Groq 3 LPX Rack for Low-Latency AI
Nvidia announced full production of the Groq 3 LPX rack, a low-latency inference system deploying later this year at neocloud provider Nebius.
Nvidia announced that its Groq 3 LPX rack has entered full production and will be deployed later this year at the neocloud provider Nebius. The product is the result of a $20 billion acquisition of assets from chip startup Groq in December, marking the largest purchase in the company's history.
Designed for low-latency inference to improve the responsiveness of AI agents, particularly in coding applications, the rack packages 256 Samsung-manufactured chips. Nvidia reports the system can deliver 3,400 tokens per second. The technology will operate alongside Vera central processors and Rubin graphics processors.
CEO Jensen Huang previously indicated that 25% of data center space for coding applications would be allocated to Groq chips, while the remaining 75% would utilize Vera Rubin systems. This deployment places Nvidia in direct competition with Advanced Micro Devices and Cerebras, the latter of which powers the Ultrafast mode for OpenAI.