Complete reference for all GitHub Models supported by md-evals, including capabilities, pricing, and recommendations.
GitHub Models provides free and low-cost access to four premier language models through Azure AI Inference. All models are available at no cost to GitHub users during the public preview phase.
✅ All models are currently free during public preview. Check GitHub Models availability for pricing updates.
| Model | Provider | Context | Temp Range | Speed | Quality | Best For | Cost |
|---|---|---|---|---|---|---|---|
| claude-3.5-sonnet | Anthropic | 200k | 0.0–2.0 | Fast | Excellent | Complex reasoning, instruction following | Free |
| gpt-4o | OpenAI | 128k | 0.0–2.0 | Medium | Excellent | General purpose, balanced performance | Free |
| deepseek-r1 | DeepSeek | 64k | 0.0–1.0 | Fast | Good | Code generation, cost-efficient | Free |
| grok-3 | xAI | 128k | 0.0–2.0 | Medium | Very Good | Real-time reasoning, latest capabilities | Free |
Model ID: claude-3.5-sonnet
Provider: Anthropic
Release: October 2024
| Aspect | Value |
|---|---|
| Context Window | 200,000 tokens (~150k words) |
| Max Output | 4,096 tokens |
| Temperature Range | 0.0–2.0 |
| Accuracy | 99% on standard benchmarks |
| Reasoning | Excellent (multimodal reasoning) |
✅ Ideal for:
- Complex instruction following
- Detailed analysis and summarization
- Long-context document processing (e.g., code reviews)
- Multi-step reasoning tasks
- SKILL.md evaluation (detailed system prompts)
❌ Not ideal for:
- Ultra-fast, latency-critical tasks
- Simple classification (overkill for simple tasks)
defaults:
model: "claude-3.5-sonnet"
provider: "github-models"
temperature: 0.7 # 0 = deterministic, 1 = balanced, 2 = creativeprompt = """
Analyze the following skill definition and provide:
1. Summary of what the skill does
2. Potential edge cases
3. Suggestions for improvement
SKILL.md content:
[... long content ...]
"""- Input Speed: ~1500 tokens/sec
- Output Speed: ~80 tokens/sec
- Avg Response Time: 1.5–2.5 seconds
- Free Tier Rate Limit: 15 requests/min
Model ID: gpt-4o
Provider: OpenAI
Release: May 2024
| Aspect | Value |
|---|---|
| Context Window | 128,000 tokens (~100k words) |
| Max Output | 4,096 tokens |
| Temperature Range | 0.0–2.0 |
| Accuracy | 98% on standard benchmarks |
| Reasoning | Very good (strong on math/logic) |
✅ Ideal for:
- General-purpose evaluations
- Balanced speed and quality
- Mathematical reasoning
- Code analysis
- When model diversity is needed (A/B testing providers)
❌ Not ideal for:
- Ultra-long context (>128k tokens)
- Open-ended creative writing
defaults:
model: "gpt-4o"
provider: "github-models"
temperature: 1.0 # Balancedprompt = """
Evaluate this code snippet for:
1. Correctness
2. Performance
3. Security issues
Code:
[... code ...]
"""- Input Speed: ~2000 tokens/sec
- Output Speed: ~100 tokens/sec
- Avg Response Time: 1.0–2.0 seconds
- Free Tier Rate Limit: 15 requests/min
Model ID: deepseek-r1
Provider: DeepSeek
Release: January 2025
| Aspect | Value |
|---|---|
| Context Window | 64,000 tokens (~50k words) |
| Max Output | 2,048 tokens |
| Temperature Range | 0.0–1.0 |
| Accuracy | 97% on standard benchmarks |
| Reasoning | Good (optimized for code) |
✅ Ideal for:
- Code generation and analysis
- Cost-optimized evaluations
- Programming task evaluation
- Quick feedback loops
- High-volume evaluation runs
❌ Not ideal for:
- Very large context requirements
- Long, open-ended creative writing
defaults:
model: "deepseek-r1"
provider: "github-models"
temperature: 0.5 # Lower for code, more deterministicprompt = """
Generate a Python function that:
- Accepts a list of numbers
- Returns the sum of even numbers
- Handles edge cases
Include tests.
"""- Input Speed: ~2500 tokens/sec (fastest input)
- Output Speed: ~120 tokens/sec
- Avg Response Time: 0.8–1.5 seconds (fastest overall)
- Free Tier Rate Limit: 15 requests/min
Model ID: grok-3
Provider: xAI
Release: December 2024
| Aspect | Value |
|---|---|
| Context Window | 128,000 tokens (~100k words) |
| Max Output | 4,096 tokens |
| Temperature Range | 0.0–2.0 |
| Accuracy | 98% on standard benchmarks |
| Reasoning | Very good (real-time aware) |
✅ Ideal for:
- Current events/real-time reasoning
- Balanced speed and quality
- Testing provider diversity
- General-purpose tasks
- When you need latest reasoning capabilities
❌ Not ideal for:
- Ultra-long context (>128k)
- Extremely sensitive tasks (newer model)
defaults:
model: "grok-3"
provider: "github-models"
temperature: 0.7prompt = """
Explain the latest AI developments and their implications.
Current date: [system provides current date]
"""- Input Speed: ~1800 tokens/sec
- Output Speed: ~90 tokens/sec
- Avg Response Time: 1.2–2.2 seconds
- Free Tier Rate Limit: 15 requests/min
| Feature | Claude | GPT-4o | DeepSeek | Grok-3 |
|---|---|---|---|---|
| Text Input | ✅ | ✅ | ✅ | ✅ |
| Long Context (200k) | ✅ | ❌ | ❌ | ❌ |
| System Prompts | ✅ | ✅ | ✅ | ✅ |
| Temperature Control | ✅ | ✅ | ✅ | ✅ |
| Max Tokens Limit | ✅ | ✅ | ✅ | ✅ |
| Aspect | Claude | GPT-4o | DeepSeek | Grok-3 |
|---|---|---|---|---|
| Instruction Following | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐ |
| Code Generation | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐ |
| Math Reasoning | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐ | ⭐⭐⭐ | ⭐⭐⭐⭐ |
| Creative Writing | ⭐⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐ | ⭐⭐⭐⭐ |
| Consistency | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐ |
| Limit | Value |
|---|---|
| Requests per minute | 15 |
| Requests per hour | 300 |
| Requests per day | 1,000 |
| Concurrent connections | 3 |
| Cost | $0 (free during public preview) |
If you exceed rate limits, md-evals automatically:
- Detects the 429 HTTP error
- Raises
RateLimitErrorwith retry-after info - Suggests caching or serial execution
Workaround: Run with fewer parallel workers
# Default: 5 parallel workers (uses 5 req/min)
md-evals run eval.yaml -n 1 # Use 1 worker (15 req/min available)START
↓
Need context > 128k tokens?
├─ YES → Use claude-3.5-sonnet (200k context)
├─ NO
↓
Need fastest speed?
├─ YES → Use deepseek-r1
├─ NO
↓
Need best code generation?
├─ YES → Use gpt-4o or deepseek-r1
├─ NO
↓
Need real-time awareness?
├─ YES → Use grok-3
├─ NO → Use claude-3.5-sonnet (safest choice)
A/B Testing Multiple Models
# eval.yaml
defaults:
provider: "github-models"
temperature: 0.7
treatments:
CONTROL_CLAUDE:
model: "claude-3.5-sonnet"
skill_path: null
TREATMENT_CLAUDE:
model: "claude-3.5-sonnet"
skill_path: "./SKILL.md"
CONTROL_GPT:
model: "gpt-4o"
skill_path: null
TREATMENT_GPT:
model: "gpt-4o"
skill_path: "./SKILL.md"
tests:
# ... your testsRun with: md-evals run eval.yaml
Cost-Optimized Evaluation
defaults:
model: "deepseek-r1" # Fastest, lowest latency
provider: "github-models"Long Document Processing
defaults:
model: "claude-3.5-sonnet" # Only 200k context available
provider: "github-models"GitHub Models provides token counts for billing and analysis. Counts are:
- Approximate: ±10% accuracy for typical responses
- Per-model: Different tokenizers for each model
- Streamed: Extracted from response metadata automatically
Example output:
Model: claude-3.5-sonnet
Prompt: "Explain AI in 100 words"
Response: "AI is... [50 words]"
Input tokens: 9
Output tokens: 51
Total tokens: 60
Token counts appear in results:
md-evals run eval.yaml -o markdown
# Shows: "Total tokens used: 1,234" in reportAll models are free during public preview on GitHub Models.
After preview ends, expected pricing:
- Claude 3.5 Sonnet: ~$0.003 input / $0.015 output (per 1k tokens)
- GPT-4o: ~$0.003 input / $0.012 output (per 1k tokens)
- DeepSeek R1: ~$0.0014 input / $0.0042 output (per 1k tokens)
- Grok-3: ~$0.002 input / $0.006 output (per 1k tokens)
Cost = (input_tokens * input_price + output_tokens * output_price) * model_count * test_count
Example:
- 100 input tokens per request
- 50 output tokens per response
- 1 model (claude-3.5-sonnet)
- 100 tests
- 2 treatments (control + with skill)
Cost ≈ (100 × $0.003 + 50 × $0.015) × 1 × 100 × 2 ≈ $2.10
Track token usage in results:
md-evals run eval.yaml -o json | jq '.summary.total_tokens'Q: Which model should I start with?
A: Start with claude-3.5-sonnet. It has the largest context window and best instruction-following. It's the safest choice for evaluating skills.
Q: Can I switch models between runs?
A: Yes! Change the model field in eval.yaml or use md-evals run --model gpt-4o
Q: Are token counts exact?
A: Approximate (±10%). Different models use different tokenizers. Exact counts require an extra API call, which we avoid for performance.
Q: What's the difference between prompt and response tokens?
A: Input tokens = your prompt, Output tokens = model's response. You pay for both.
Q: Can I use GitHub Models in my own application?
A: Yes! Azure AI Inference SDK is open. See Design documentation.
Q: Where can I report issues with models?
A: Report to GitHub Models Issues
- Getting Started Guide — Set up your token
- Example Configurations — Real eval configs
- Troubleshooting — Common problems