Your model shipped days ago. Its knowledge stopped months ago. Both clocks, ticking live, for every major lab.

85 points•joozio•15 days ago•50 comments•

50 comments

gjskngnf14 days ago
I remember when the US captured Venezuelan president Maduro, and when I posed a prompt related to this, the model said that’s pure fiction. I told it to double check. Still didn’t want to entertain the idea. It only acquiesced when I specifically directed it to check Reuters. I haven’t noticed this problem in months. Model cutoff seems to be less of a problem these days.
super25614 days ago
It's a "problem" of compute, I think. If you query without an account on ChatGPT you will see the model look up less stuff and research less, than when you have a paid account and choose "medium" or "high" in the effort slider.

Which makes sense, because of you have looked into search and crawlers you notice that search is actual quite expensive (which is why e.g. Kagi charges a few bucks for search every month).

Catloafdev14 days ago
It's not strictly compute, because this has noticeably improved in open-weight models too, such as Gemma and Qwen. I suspect they noticed this issue and adjusted their training to be better about it over time.
NegativeLatency14 days ago
Gets me with AWS stuff on claude all the time, fortunately there's a official amazon MCP for their docs which helps a lot, but I still have to occasionally tell it to check the docs/mcp.
InsideOutSanta14 days ago
Came here to say the same thing. Models used to rely heavily on world knowledge from their training data. They are now much better at tool use and deciding when to research a topic, rather than just answering from memory.

I wonder how much that extends to using LLMs for programming. I assume most knowledge of programming language syntax still comes from training data.

NegativeLatency14 days ago
I find they generally do ok, but a few lines in an AGENTS.md or manual prompting to verify stuff against current docs/source, and check for current version of software helps a lot.
tannertech14 days ago
ChatGPT once told me I was the target of a sophisticated nation state misinformation campaign when I linked it a Reuters article
dpoloncsak13 days ago
There are a lot of people who would say that is exactly what Reuters is. Any American news outlet, really.
Isamu14 days ago
>the model said that’s pure fiction.

Were you expecting your model to be updated on current events? Why?

Also the specific event you are referring to is a statistically very improbable event, prior to its actually happening.

>It only acquiesced when I specifically directed it to check Reuters.

Do all models do this? They check in with Reuters? Why would a model think that you asking about an extremely improbable event warranted reaching out to Reuters?

SirMaster14 days ago
If OpenAI is going to call Astra AGI, then I would expect it to be able to update it's weights to new knowledge, because a generally intelligent being can indeed do this.

I can teach myself to play an instrument, and I'm not just building this huge lookup table that I have to access every time I play the instrument. I am updating the weights in my neurons.

Until AI can do this it's not AGI in my book.

gjskngnf14 days ago
I was not expecting model weights to be updated on current events.

It’s clearly warranted because a model that trusts its weights on current events will give an outdated answer. Extremely improbable events happen all the time.

mywittyname14 days ago
He asked it to double check. It's reasonable to expect the LLM to handle that trivial task.
jasonjmcghee14 days ago
It still matters, but in the age of good reasoning, tool use, and web search, this is much less of a problem than it used to be.
lukewarm70714 days ago
in my chat with gemini it could not differentiate between current events and fiction.

if you point it to the web it got the point, but started treating everything like fiction. so it simply started making up possible scenarios and playing them off as real answers when asked for factual information.

i could not tell what the issue was or how to fix it because the reasoning is encrypted. the obfuscation model spat out something like: 'the user is asking for details about a fictional scenario in which the usa has assassinated the leader of iran'

i really don't like the way big ai companies are going. encrypted thinking, guardrails, adversarial personality, moralizing. it is creating something anti-human.

spindump893014 days ago
It can be quite hard to determine what needs a tool call or not. LLMs are not well calibrated to what they know and don't know, and tool calls can add latency and extra costs. There are lots of things that are "obvious" right until they aren't - especially political events and disasters.
dominotw14 days ago
all the reasoning still comes from pretraining data
bigmadshoe14 days ago
I think they mean reasoning their way to the need for a web search.
binlog14 days ago
Says who? Models can also use results from tool calls in their reasoning loops.
BraamV9 days ago
This is why I am not worried about AI taking over the world. Every thread is frozen in time. Interesting that Astra has 2 months earlier cutoff than Fable 5.1.. Does that indicate much bigger (and therefore longer) training run?
ctkhn14 days ago
For general purpose use this is interesting, but if I'm just using an LLM for coding, does this matter at all? I would hope something like a new java version after a model's publish date can be handled and understood by the model through tool calls and context even if it's not explicitly in the training data, the same way the LLM doesn't have my existing code or the plan to change it baked in from training.
prng202114 days ago
The inference time results are quite different. These models haven’t just been trained with tons of docs, blog entries and videos about Java 8. They’ve also been trained on tons of code examples from simple to complex real world ones.

You’re not going to get all that just by pointing the LLM to the recently released Java 27 documentation. That information is also potentially adding tons of content to your context, which is already filled with tons of other data (your code, other recently released libraries it has to get documentation for, etc).

ctkhn14 days ago
Gotcha, that makes a lot of sense. Thank you
smt8814 days ago
My experience with Claude is that it doesn’t handle its own training cutoff properly. It responds as though its cutoff is today. Gemini is even worse about this.

I’m going to add this site to my agents files so that they’re explicitly aware of their own limitations.

ryanschaefer14 days ago
Do people prefer the new flat style LLMs are producing? I don’t mind it as much as the gradient theme they were pumping out previously.

Read the full thread on Hacker News →

Related stories