WEDNESDAY INSIGHT
One development worth understanding, and what it means for you.

Good morning.
Two years ago ChatGPT made things up about four times out of ten when it did not know an answer. Today it does it closer to nine. Accuracy went up. Honesty about the gaps went down. Here is what the numbers actually say, and the habit that protects you either way.
First time reading? Get your own free subscription here.
AI INSIGHT
The Free Version Got Worse and Nobody Announced It
Two years ago, when ChatGPT hit a question it could not answer, it made something up about four times out of ten. Today's ChatGPT does it closer to nine times out of ten.
That is not a typo. It comes from a test called AA-Omniscience, run by the research firm Artificial Analysis, which measures something most people never think to check. When a model does not know the answer, does it admit it, or does it bluff? A lower score is better. Engadget pulled the current numbers together this week, comparing ChatGPT and Claude head to head.
Here is what makes it strange. On raw knowledge, the two are close. On bluffing, they are not in the same building.

The number that matters is not the one you would check
Accuracy gets the headlines. It moves a couple of points per release and gets a press announcement. The bluffing rate is the one that decides whether you can trust an answer on a Tuesday afternoon when you have no time to verify it.
The current scores, lower being better:
Top paid tiers: Claude bluffs 55 percent of the time it does not know. ChatGPT, 89 percent.
Free tiers, which is what most people are using: Claude bluffs 37 percent of the time. ChatGPT, 85 percent.
ChatGPT against its own past: the version most people were using two years ago bluffed 38 percent of the time. Its replacement does it 89 percent of the time.
Read that last line again 🙃. On this measure, ChatGPT has gotten dramatically worse at admitting it does not know something, while getting better at almost everything else.
That combination is worse than either problem alone. A model that is often wrong teaches you to check. A model that is usually right and never hesitates teaches you to stop checking, which is exactly when it costs you. I wrote about why this happens in I Asked AI to Help Me Write About AI Hallucinations. It Hallucinated. The short version is that these tools predict, they do not know, and confidence is not evidence of anything.
The honest complication
This is not a clean win for Claude, and I should say plainly that I use Claude every day to help produce this newsletter. So here is the number that cuts the other way.
On the free tier, ChatGPT is measurably more knowledgeable. It scores 46 percent on raw accuracy against Claude's 38 percent. At the top paid tier they are basically tied, 61 percent for Claude and 59 percent for ChatGPT.
So free ChatGPT knows more. It also makes things up more than twice as often when it hits the edge of what it knows. Which of those two things matters more depends entirely on whether you are checking its work.
There is one more piece of context. This year ChatGPT started running ads on its free and cheapest paid tiers. Anthropic went the other way with Claude, adding features to the free tier and leaving ads out. Those are two different bets about what a free user is for.
What this actually means for you
A benchmark is not your kitchen table. These tests hammer the models with obscure questions designed to find the edge of what they know. Your recipe substitution or your packing list is nowhere near that edge, and both assistants handle ordinary requests fine. The scores will also move with the next release, and probably in a few months.
So the takeaway is not a scoreboard. It is this: the assumption that a newer version is a more reliable version has now been measured, and it did not hold. If you moved to the latest model and quietly relaxed your standards because it felt sharper, that instinct was working against you.
Newer is a release date. It is not a promise.
Try This With AI
What it does: Forces the assistant to separate what it actually knows from what it is filling in, before you act on any of it.
Before you begin: Have a real question ready, ideally one with a number, a date, a rule, or a price in it.
"Answer my question below. Then, before anything else, list every part of your answer you are not fully confident about, and tell me how I could check each one myself. If you do not know something, say so instead of guessing. My question: [your question here]"
Customize it: Swap the bracketed line for whatever you were about to ask anyway. This works best on questions where being wrong costs you money or time.
Then ask this: "Now rank those uncertain parts from most likely to be wrong to least likely, and tell me which one would cause me the most trouble if I acted on it."
READER POLL
When was the last time you fact-checked something your AI told you?
WHERE TO GO NEXT
More on this topic, from sources worth your time:
AI Doesn't Know What Day It Is. It Won't Say So. -- The same confidence problem, applied to dates and timeframes, plus the one habit that catches it.
Galaxy.ai -- Runs thousands of AI models under one login, which makes it practical to ask the same question twice and see whether the answers agree.
Incogni -- Now that ads are arriving in free AI tiers, this pulls your personal information back from the data brokers who feed that machine.
FROM OUR PARTNER
Stop overpaying! Get Tello's same reliable 5G network for under $25
Tired of paying for perks you don't actually use? With Tello, you can build your own plan and get the same reliable 5G coverage for way less.
Advertising Disclosure: We evaluate all recommendations of products and services independently. Clicking on links provided on this page may result in AI for Daily Living earning compensation, which supports independent publishers like us.



