I’ve been using AI writing models in my work for several years now, so I’ve seen my fair share of changes in what they can (and cannot) do. There’s always something new being added, and it’s often noticeable when a feature has changed. Yet, whenever a new model rolls out and I give it a test run, I rarely come away surprised by how different it actually feels.
I’ve had pretty much the same experience with newer ChatGPT releases and the different Claude models I’ve tried. Each one has its own quirks, and I sometimes prefer one over another for a particular task. Even then, I still haven’t had a new release make me feel like the actual writing has taken a major step forward.
The improvements are bigger in technical work
On the OpenAI side of things, GPT-5.6 Sol scored better than earlier models on coding and knowledge-work benchmarks. Then came GPT-6 Astra on September 3, although access was still rolling out when I wrote this. I haven’t spent enough time with Astra yet to make a definitive judgment about it.
Some of Astra’s biggest reported gains are in computer use and software engineering, and that’s where I can see newer models making a much bigger difference. With technical work, the result is often more straightforward to assess because there is usually a clearer outcome, a bit like a maths problem where you can check whether the answer is right. For instance, if a developer asks AI to build part of a web application, they can run the code and see whether it behaves as intended or still needs more work.
Anthropic has been busy with new releases too, with Claude Sonnet 5 arriving in June and Opus 5 following close behind in July. Claude Fable 5.1 then came out on September 1 alongside Mythos 5.1, with the latter currently available only under restricted research access. Anthropic says Fable 5.1 is its strongest widely available model for coding and knowledge work.
When I look at that progress as a writer, it’s less obvious because semantics are only one part of what makes a piece of writing work. A paragraph can convey the intended meaning and follow the brief, yet still sound nothing like something I would write. As a result, I may end up spending quite a while rewriting the copy even when the model has technically done what I asked.
The companies are trying to improve writing too
I don’t want to give the impression that OpenAI and Anthropic are completely ignoring writing quality, because that wouldn’t be fair. OpenAI describes GPT-6 Astra as its best model for following existing templates and says it can match a user’s writing style more closely. The release notes also suggest Astra is better at avoiding unnecessary information being repeated in the final output.
Anthropic has shared similar feedback about Fable 5.1, and an early tester quoted in the launch material found that the model followed writing guidance better than Fable 5. That feedback, however, came from a customer featured by Anthropic rather than an independent test, so I’d treat it as an early indication instead of assuming every writer will notice the same difference.
Following writing instructions better is precisely the kind of progress I care about, because I mostly just want the models to do what I’ve asked without me having to keep correcting the same things. Newer releases are getting better at that, just not enough for the difference to feel especially big yet.
Familiar writing habits still come back
Most people are familiar with the basic process by now, but large language models, or LLMs, are trained on enormous amounts of text. During pretraining, an autoregressive model learns statistical patterns in that text and predicts what comes next by working with tokens, which are small pieces of text such as words, parts of words or punctuation. Post-training then shapes how well the model follows instructions and the kinds of responses it tends to give.
Learning those statistical patterns doesn’t mean ChatGPT has one article template hidden somewhere that it pulls out every time I ask for 1,500 words. The problem is subtler, and after you’ve edited enough AI-generated copy, some of those recurring patterns start to jump out at you.
Once you start noticing them, familiar paragraph structures and transitions can show up across completely different subjects. Research into AI writing has found that different model families can leave recognisable patterns in their prose, so the idea that each one has its own writing habits isn’t just something I’ve picked up on myself.
ChatGPT, in particular, can’t seem to help itself when it comes to short sentences. I can specifically tell it that I don’t want little punchy sentences scattered through an article, then find another one several paragraphs later. Although I wouldn’t say every ChatGPT response does this, it happens often enough in my own work that I actively check for it.
Claude has its own proclivities too, and I find that some of them keep coming back. In my experience, the responses tend to run longer and qualify things more carefully. The bot can also take the scenic route to make a point that could have been said much faster.
My AI-ism list has helped
Over time, I’ve built an ever-growing list of what I like to call AI-isms, mainly so I can keep track of the patterns that keep showing up. These include certain words and sentence structures, along with other writing habits I repeatedly find in machine-generated copy. I’ve written before about some of the less obvious signs of AI writing, and quite a few have found their way onto that list.
Again, modern models are better at following those instructions than earlier ones, and I have noticed some improvement. Thankfully, the dreaded em dash is one habit I see far less often now. I can give ChatGPT or Claude a fairly detailed set of writing rules and get noticeably closer to what I want within the first few attempts than I could before.
I’ve still had to build that setup brick by brick, spending hours refining prompts and giving the model examples of previous writing so it has a better sense of what I’m after. Technically, what I’m doing is a process called prompt engineering, which is different from fine-tuning. Fine-tuning involves further training a model on a dataset, but I may come back to that another time.
The ever-so-familiar AI-ism problem hasn’t disappeared completely, and new habits still find their way in too. I can ban a particular sentence construction and see it reappear later in the same article. Sometimes the wording changes enough to hide it, but the same underlying habit is still there, just wearing a different disguise. Once I spend too long correcting these things, some of the time AI was supposed to save starts disappearing.
AI works better as a second pair of eyes
I’ll give credit where it’s due because AI has in fact made parts of my process easier in small yet impactful ways. I use ChatGPT and Claude to organise messy notes, and I find both particularly good at telling me when an explanation has become harder to follow than it needs to be.
That second use has been particularly helpful because I can paste in something I’ve written and ask where the logic starts becoming difficult to follow. Both models can often pinpoint the exact part that needs more work.
The strange (and annoying) part is what happens when I then ask either model to write the replacement itself. ChatGPT may diagnose the weakness correctly and then fall straight back into its old habits when writing the new paragraph. I find Claude can do much the same, sometimes fixing the logic but leaving me with wording that feels too polished or overexplained. In the end, I get more out of these models when they help me inspect my writing than when I ask them to take over completely.
Research tools make a bigger difference to me
Another benefit I’ve found with AI is research, where the tools have made a much clearer difference to my workflow. ChatGPT’s Deep Research has probably transformed that process more than switching between general-purpose models ever has. The tool can search through a large number of sources and put everything neatly into a documented report with citations or links I can check afterwards. I still cross-check the original material because I always want the final call on what makes it into the article.
Perplexity is useful for a similar reason, and I tend to use it when I want another way into a research-heavy topic. What I like about Perplexity is having models like GPT, Claude, Grok and Kimi all in one place, so I can switch between them without bouncing across different platforms.
I personally don’t rely on Claude as much for research, but I find it particularly good at helping me organise my thoughts and reason through an idea. Sometimes I’ll move something over from ChatGPT just to see what Claude has to say, especially when I feel like I’ve been going in circles with the same problem.
ChatGPT’s image tools have found a place in my workflow, too. I can generate a visual from a written prompt or edit an image I already have, which can be handy when I need something to accompany an article.
Another model, another rewrite
Switching between models can improve individual sentences or help me handle a difficult paragraph better, but as I’ve found time and again, the overall writing process hasn’t changed an awful lot. Sometimes Claude can give me a fresher take on a section than ChatGPT, and other times the reverse is true. Perplexity earns its keep in a different way, particularly for research, because I can access several models in one place.
Ultimately, I care more about the features that genuinely save me time than the name of the latest model, especially when I still find myself making several rounds of corrections before the copy is where I want it.
