Groq
Ultra-fast LLM inference API running open models at extreme speed.
About Groq
Groq's LPU hardware runs Llama, Gemma and other open models at speeds up to 800 tokens per second, enabling low-latency AI applications requiring real-time response times.