AI Provider Scorecard
Evaluate and compare AI vendors against your organization's values and technical requirements. Adjust weights by mission type — the ranking shifts accordingly.
Select your organization type to apply recommended weights, or adjust the sliders manually. The composite score reflects your priorities — not a universal ranking.
Organization type
Provider scores
| Provider | Score | Tier | |||||
|---|---|---|---|---|---|---|---|
Anthropic (Claude) US · Private | 8.7 | 7.7 | 5.7 | 8.0 | 8.6 | 8.1 | Recommended |
Mistral France · GDPR-governed | 7.3 | 8.0 | 7.7 | 7.3 | 5.5 | 6.6 | With caution |
Google (Gemini) US · Public | 5.0 | 2.7 | 3.7 | 5.3 | 8.7 | 6.5 | With caution |
OpenAI (ChatGPT) US · Private | 3.7 | 3.3 | 3.7 | 4.7 | 8.7 | 6.2 | With caution |
DeepSeek China · State-adjacent | 2.0 | 2.0 | 7.0 | 3.0 | 7.8 | 5.7 | Not recommended |
xAI (Grok) US · Private | 3.0 | 1.7 | 6.7 | 0.3 | 7.3 | 5.1 | Not recommended |
Meta AI (Llama) US · Public | 6.0 | 2.0 | 1.3 | 4.0 | 6.7 | 5.0 | Not recommended |
The icon columns are each provider's fixed scores — hover an icon for details; they don't change with the sliders above. Capability = average of SWE-bench Verified, general reasoning (ARC-AGI-2, GPQA Diamond, AIME, LMArena), and Terminal-Bench 2.0. Score = your weighted average of the four ethics dimensions and capability, blended by the ethics-vs-capability slider above. Scores as of early 2026.
Methodology: read the full article →