OpenAI, Anthropic's Claude and Google's Gemini all offer powerful models, and new versions arrive every few months. Choosing between them based on headlines or leaderboards is a mistake. The right choice depends on your use case, your data and your constraints — and it will change over time.
Evaluate on your own tasks
Public benchmarks rarely reflect your product. Build a small evaluation set from real examples — fifty to two hundred representative inputs with expected outputs — and score each model on accuracy, format compliance and tone. This single step prevents most costly mistakes.
The criteria that matter
- Quality on your tasks — measured with your evaluation set, not marketing claims
- Latency — especially for chat and interactive features
- Cost at your volume — including input and output tokens and caching options
- Context length — how much information each request must include
- Structured outputs and tool use — reliability when calling functions or returning JSON
- Data handling — retention, training policies, regional hosting and compliance commitments
- Reliability and support — rate limits, uptime history and enterprise agreements
Match models to jobs
Many production systems use more than one model. A larger model handles complex reasoning or long documents, while a smaller, faster model classifies intents, extracts fields or routes requests. Matching model size to the task often cuts costs dramatically without affecting quality.
Design for change
- Put a thin abstraction layer between your application and model APIs
- Version prompts, tools and evaluations alongside your code
- Log inputs, outputs and costs so providers can be compared with real data
- Add automatic fallbacks for outages and rate limits
A practical selection process
- Define success metrics and build an evaluation set
- Shortlist two or three models that meet your data and compliance requirements
- Run the evaluation and review failures by hand
- Estimate cost and latency at expected volume
- Launch with monitoring and re-evaluate every quarter
The best model is the one that performs reliably on your tasks, within your budget and your data rules — and you should expect that answer to change.
- #AI
- #LLM
- #Technology
Tomasz NowakInsights, not noise.
One practical email a month on AI, product and growth. Unsubscribe anytime.




