Models and providers
Current Model Family Selection Matrix
Turn current Anthropic, OpenAI, and Google model catalogs into a dated candidate list, then select with workload evidence instead of product rank or memory.
- Format
- Decision Path
- Level
- Intermediate
- Audience
- Practitioner, Developer, Operator, Leader
- Owner
- Project42 Editorial
- Review cadence
- Every 30 days
- Prerequisites
- A defined workload with quality, safety, latency, and cost targets; Access to current provider model and lifecycle documentation; A representative evaluation set
Use the dated snapshot as discovery, not a recommendation
As verified on 2026-07-26, the official catalogs organize current general-purpose choices into capability, balanced, and speed or cost-oriented tiers: Anthropic documents Claude Opus, Sonnet, and Haiku families; OpenAI lists GPT-5.6 Sol, Terra, and Luna alongside earlier and specialized families; Google lists stable and preview Gemini families including Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. These names can change faster than this guide.
Open the linked catalog on every selection or release decision. Record exact model ID, lifecycle state, API or cloud surface, region, modalities, tool and structured-output support, context and output limits, data terms, quota, and price from the deployment actually being considered. Do not infer API availability from a consumer chat product.
Compare identical work under identical controls
Shortlist by required capability and deployment constraints, then run the same versioned cases, prompt, tools, retrieval state, budgets, graders, and stop conditions. Measure critical-case success, unsupported claims, safety behavior, tool trajectory, latency distribution, usage, estimated cost, and operational failure modes.
Prefer the least expensive candidate that clears every blocking requirement with acceptable margin. A provider's general description and public benchmark are evidence about a broad model, not proof for your data, prompts, tools, or users.
Task: [WORKLOAD AND DECISION]
Scope: [PROVIDERS, API/CLOUD SURFACES, REGION, AND DATA CLASS]
Permissions: [WHO MAY TEST, DEPLOY, SPEND, AND ACCESS EVALUATION DATA]
Verified date: [YYYY-MM-DD]
Candidate: [PROVIDER, EXACT MODEL ID, LIFECYCLE, AND DEPLOYMENT]
Required capabilities: [MODALITY, TOOLS, STRUCTURE, CONTEXT, SAFETY, AND RESIDENCY]
Evaluation: [CASES, CRITICAL SLICES, RUBRIC, BASELINE, AND THRESHOLDS]
Observed tradeoffs: [QUALITY, SAFETY, LATENCY, USAGE, COST, AND FAILURES]
Stop conditions: [STALE SOURCE, PREVIEW PROHIBITED, CRITICAL FAILURE, OR BUDGET]
Decision and expiry: [SELECT/HOLD/REJECT -> OWNER -> REVIEW DATE]
Verification: [PINNED CONFIG, REPEATABLE RUN, CANARY, AND POSTCONDITIONS]
Recovery: [ROUTE TO VERIFIED FALLBACK, RECONCILE STATE, AND RERUN GATE]Expected evidence and verification
Expected evidence includes dated first-party catalog records, exact deployable IDs, lifecycle and region checks, a representative versioned evaluation, slice-level results, cost assumptions, a human decision, expiry date, and fallback. Verify the selected model from the real application boundary rather than only in a playground.
Stop promotion when a required feature or region is unavailable, a critical case fails, comparison conditions differ, data terms are unacceptable, or sources are stale. Recover by routing to the last verified model and configuration, reconciling tool or conversation state, and rerunning the unchanged gate before another promotion.