Groq

Inference provider running open-weight models (Llama, Mixtral, Qwen) on custom silicon called an LPU rather than the usual GPUs. The concrete result: answers streaming at 300 to 1000 tokens per second where a classic provider gives you 30 to 80. The API speaks the same dialect as OpenAI, so wiring Groq in means swapping a URL and a key in existing code. Free tier with no credit card, roughly 30 requests per minute, enough to prototype on.

Strengths

Limitations

Best for

Official site

View on Coeurdar