Chatbots and large language models can execute a seemingly countless number of tasks, from writing emails and reports to generating code and analyzing data. However, they still primarily act only in response to user prompts, and rely on their own predictive models to generate text. That’s why they can be so good at some tasks, such as simulating human writing, and surprisingly bad at others, such as mathematical or logical reasoning.