Have you considered stratifying based on different verticals or domains? For example, perhaps "healthcare" or "coding" is fresher than "poetry" because there's more economic value there?
I did a sub experiment for Opus 5 across several sub-domains (esp coding) specifically and surprisingly didn't notice any major differences in recall. It's possible that this is true for other models and Opus 5 is just the exception thou.
This is such a cool and clever result!! I love how much insight you extract from the model just by doing some inference!! :)
This is interesting analysis and cool. Is there a possibility that models are trained not to reveal and probably obfuscate the cut-off knowledge?
It seems very unlikely bc doing so would make the model directly worse at recalling information.
Have you considered stratifying based on different verticals or domains? For example, perhaps "healthcare" or "coding" is fresher than "poetry" because there's more economic value there?
I did a sub experiment for Opus 5 across several sub-domains (esp coding) specifically and surprisingly didn't notice any major differences in recall. It's possible that this is true for other models and Opus 5 is just the exception thou.
fascinating!!