Gemma 2 9B
Gemma 2 9B
Section titled “Gemma 2 9B”Provider: Groq (Google) Context Window: 8,000 tokens Quality: Medium Price: $0.20 per 1M input tokens
Overview
Section titled “Overview”Gemma 2 9B is Google’s compact open-source model, optimized for speed on Groq’s hardware. It’s the fastest option available through Fastlane.
Benchmarks
Section titled “Benchmarks”| Benchmark | Score | Percentile |
|---|---|---|
| MMLU | 71.3% | Top 30% |
| HumanEval | 59.1% | Top 40% |
| MATH | 42.3% | Top 45% |
| GPQA | 26.8% | Top 45% |
| HellaSwag | 86.4% | Top 30% |
| ARC-Challenge | 84.2% | Top 30% |
Source: Google AI technical reports
Best For
Section titled “Best For”- Ultra-fast responses
- Simple Q&A
- Classification and tagging
- Maximum throughput workloads
Pricing
Section titled “Pricing”| Tier | Input | Output |
|---|---|---|
| BYOK | $0.20/1M | $0.20/1M |
| Marketplace | $0.206/1M | $0.206/1M |
Routing Recommendation
Section titled “Routing Recommendation”- Accurate mode: Never selected
- Balanced mode: Only for very simple prompts
- Cheap mode: Frequently selected
- Eco mode: Top choice - minimal energy, fastest inference
curl http://127.0.0.1:8787/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"model": "fastlane-auto", "messages": [{"role": "user", "content": "Classify this email as spam or not"}]}'