Most AI chat tools charge a fixed monthly subscription whether you use them or not. OpenFastChat charges close to the underlying Groq inference cost and shows the exact price after every reply, so light users pay very little and heavy users stay in control with an optional monthly compute cap.
Prices are per token and shown live in the app after every reply. The table below is a guide to which model fits which job.
| Model | Best for | Notes |
|---|---|---|
| Llama 3.1 8B Instant | Fast everyday chat and drafting | Lowest per-token cost, great for high volume |
| Llama 3.3 70B Versatile | Balanced general assistant and coding | Strong quality at a moderate per-token price |
| Qwen3 32B | Multilingual chat and coding | Cheap per token, solid all-rounder |
| Kimi K2 | Long documents and agentic tasks | Large context, higher per-token price |
| GPT-OSS 120B | Heavier reasoning and writing | Open-weight large model, mid-range pricing |