Ask HN: What is one simple thing LLMs are insanely bad at?
I am looking for ideas on what to train a specialized model for!What is one simple thing you repeatedly ask ChatGPT, Claude, or another model to do that it still somehow messes up?
21 points by davidest - 29 comments
nhl toronto scores nhl hockey toronto scores "nhl hockey" toronto score today nhl "hockey score toronto" "hockey" who won toronto
etc.
Somehow being good at semantic search makes them bad at keyword search, for whatever reason.
I've had some luck on the web app side if I use playwright or similar for the model to interact with but still far from efficient.
More of an image model than a LLM model tho
I found this can work with AI. You get it to generate a lot more at first, and then do several passes over it to compress and squeeze out the noise while keeping the core information. With AI, at least with my prompts, it takes some effort (on my end) to get it to really really cut down the noise and not cut everything out.
Tasks it writes are typically too easy but also it utterly fails to see how a different model might misunderstand a vague part of the prompt.
Anno 1800 was a recent one I had trouble with, using Claude Opus. Completely made up game mechanics. Rainbow Six Siege, too.
You wouldn't tolerate this kind of duplicity from a human coworker, but AI is so fast and efficient at lying, so it's OK.
Oh, you said simple. Speaking like a human