Back to all articles

Tools & Models

A smarter AI model just dropped. Now what?

Five things to try before you hand a smarter AI model the same old work.

By Mollie Amkraut Mueller · · 8 min read

This month, Claude Fable 5 came back, then OpenAI launched GPT-5.6 Sol, Terra and Luna.

At this point, a new top-tier AI model drops roughly every time I finish getting to know the last one. I've spent a slightly unreasonable amount of my maternity leave trying them all.

For most people, though, the problem isn't access to smart models. It's figuring out what to do with each new one before the next one arrives.

The default is to open the new model and give it the same work you gave the old model. Write this email. Summarize this document. Help me brainstorm.

Maybe it does the work a little better. But that doesn't really tell you what changed.

After listening to this episode of The AI Daily Brief and finding Daniel Miessler's list of prompts to run whenever a major new model drops, I realized I have a bit of a ritual for this.

Here's what I recommend doing every time a smarter model arrives.

1. Ask what just became possible for you

Don't start by handing the model a task. Start by asking it where it can raise the ceiling on your work.

This works best if it can access some real context about you: ChatGPT's memory, a connected Google Drive, your desktop files, a project folder, a GitHub repository, or whatever contains the actual mess of your work.

If it doesn't have that, give it a five-minute voice dump. A ramble with lots of context is more useful than a beautifully written but underspecified prompt.

Then try:

"Look across everything you know or can access about my work. Based on the capabilities of this new model, rank the five tasks or projects where it could create a meaningfully better result, or make something possible that wasn't before. Use evidence from my actual work, explain why each is newly viable, and give me a ready-to-run first test."

The difference between asking "What can you do?" and "What can you now do for me?" is enormous.

2. Give it a question big enough to expose its intelligence

Daniel Miessler's prompt list includes questions so large that I would previously have felt faintly ridiculous asking an AI to answer them.

That's rather the point.

A few adapted versions worth trying:

I've already run the first two with Sol. The self-model audit caught that my AI was still treating old business plans and an old maternity-leave timeline as current. That stale context was quietly shaping every recommendation it gave me.

A smarter model is a good excuse to clean up what your AI thinks it knows about you, then use that improved context to ask a much bigger question.

3. Run your personal litmus test

Every serious AI tester seems to have one oddly specific task they give every new model.

Ethan Mollick has his otter on an airplane using WiFi. He has run versions of that prompt for years, so you can watch AI image generation progress from melted piles of fur to photorealistic otters operating laptops.

Claire Vo has built a more formal four-part benchmark that tests every new model on writing a PRD, building prototypes, debugging code and sounding human.

Mine is an app I've wanted for years that would make it incredibly easy to capture the little memories of my kids and eventually turn them into something beautiful. Models have been able to build functional apps for a while. None has yet solved the trust, continuity and behavior-change parts well enough that I would actually keep using it.

Your test doesn't need to be scientific. It just needs to be:

Run it every time. Once all the models can do it easily, make the test harder.

4. Reopen something that never quite worked

One of the best ways to see what has changed is to return to an old project that a previous model almost got right.

Not a total disaster. An idea that was promising, but kept falling short because the execution, product judgment or technical lift wasn't quite there.

This week I rebuilt an old app of mine called Board of Advisors with Sol. In conversation, it figured out that this shouldn't really be a conventional standalone app at all. It should be an MCP server, essentially a connector that lets people use it directly through ChatGPT.

It then built most of the connector system and a much better-looking interface around it. I haven't fully tested the finished result yet, so the jury is still out. But it understood the product I had in my head and made an architectural leap that previous versions hadn't.

Try:

"Here is an old project or attempt that never quite worked. Here is what disappointed me about it. Reconsider the underlying problem from first principles. Is the old approach still the right one? What would you build now that wasn't practical or possible with the models available then?"

Don't only ask the new model to improve the old answer. Give it permission to question the entire approach.

5. Give it a bar, boundaries and permission to keep going

Smarter models can work longer and take more initiative. That is wonderful right up until they confidently change the wrong thing or take an action you wanted to approve.

OpenAI's current prompting guidance recommends giving the model the context it needs, the few constraints that genuinely matter and a clear point where it must stop for approval.

Nathaniel's episode also covers a newer pattern: giving the model a measurable bar for success, then letting it work in a loop. It creates something, checks it against the bar, finds the biggest gap, fixes it and checks again.

A reusable version is:

"Here is the outcome I want: [outcome].

We are done when: [specific, observable success criteria].

Keep these things unchanged: [important constraints].

Do not send, publish, purchase or take any external action without my approval.

After your first pass, test the result against the success criteria. Identify the biggest remaining gap, fix it and repeat until the bar is met or you are genuinely blocked."

"Make it high quality" isn't a bar.

"A new user can complete the core action without any explanation" is a bar.

"Research this thoroughly" isn't a bar.

"Every material claim is linked to a source, and conflicting evidence is surfaced" is a bar.

The clearer the finish line, the more useful the model's newfound tenacity becomes.

The most important part: raise your ambition

The easiest thing to do with a smarter model is give it your existing to-do list.

That can still save time. But it misses the bigger opportunity.

The most interesting work may not currently be on your list at all, because you've assumed it would be too hard, too technical, too expensive or too time-consuming for you to attempt.

So finish your new-model ritual with this:

"Look across my responsibilities, ambitions and ideas. What work am I not attempting because I assume it is beyond my skills, time or resources? Which of those assumptions might no longer be true because of this new model? Pick the most valuable possibility and help me design the smallest real test I could run this week."

The real test of a smarter model isn't whether it writes a better email.

It's whether it changes what you believe you're capable of taking on.