When OpenAI or Anthropic ship a new model, the launch page will tell you it's stronger at everything. The only way to know what's true for your work is to test it. Here are the three I'd run on any new model, and on two at once when you're choosing between them.
Give it something big
Hand it a whole project, not a one-liner. “Here are a thousand customer reviews. Find the five biggest problems, then build the offer, sales page, and launch plan around them.” Give it your files, your budget, and what you want finished. You're testing whether it can hold the whole job in its head and carry what it learns into the next step.
Give it something you do every week
Pick a recurring task, for me it's content research: give it the accounts to follow, examples you like, and the tools to open those posts, and ask for the strongest ideas and why your audience would care. Run it once, fix what it misses, then schedule it. You're testing how much it takes off your plate without you restarting every time.
Make them compete on something you're stuck on
Say you're getting traffic but no sales. Give both models the same page, the same numbers, and ask what they'd change and why. Check their answers against the real data: which found something useful, and which needed fewer corrections?
Score them on the things that actually matter to you:
Test | Model A | Model B
------------------------------------------------
Useful finding? | |
Errors / hallucinations | |
Corrections needed | |
Time to a usable result | |
Cost of the run | |
Winner for MY work: [which one, and why]
This tells you which model is worth using for your work, instead of trusting whichever company says theirs is best. Vendor strengths are claims to test, not a verdict.