Groq

Other Freemium

Ultra-fast LLM inference API running open models at extreme speed.

About Groq

Groq's LPU hardware runs Llama, Gemma and other open models at speeds up to 800 tokens per second, enabling low-latency AI applications requiring real-time response times.