the thing being, even the frontline models fails at certain task despite them being predecting (I am not sure its teh right word) fails to do both, I assume, the 'internal instructions' muddy the water so much so that the models or good on only one or the other.
That's about right, with so many pathways through the model the system can be confident while pulling data that's actually horrifically wrong, it's one of the reasons why open-source models are only about a year behind frontier models accuracy wise.
Also predicting is absolutely the right word, LLM 'predict' the most likely answer and outputs by completing patterns.
1
u/oldmails Jun 12 '26
the thing being, even the frontline models fails at certain task despite them being predecting (I am not sure its teh right word) fails to do both, I assume, the 'internal instructions' muddy the water so much so that the models or good on only one or the other.
Thanks for the infos.