Can you always trust an AI system’s answers, or are there important limitations you should know?
As impressive as they are, transformer models don’t know how to talk — they know only how to look for patterns in sequential data (such as sentences, statements, or functions). When you train a generative model on enough data, it becomes very good at finding patterns and making predictions, but it does have limitations, and you should never trust the output of a chatbot (or any AI system) entirely.
Even the most advanced AI chatbots can sound convincing, but their confidence isn’t a guarantee of correctness. Always double-check critical information.
AI chatbots are language models tuned for conversation. If you ask a language model for the answer to a basic math problem, it will usually respond confidently with some answer. However, upon checking that answer using a calculator, you may be surprised that it’s often just plain wrong. Currently, generative models are able to make predictions based only on content they’ve previously seen. If you give them a math problem they’ve never seen before (no matter how trivial), they’ll respond based on the answers to similar math questions in their training data rather than by doing the math the way a calculator would.
Have you ever received a confidently wrong answer from an AI? What was your first reaction?
Language models like ChatGPT don’t “calculate” in the traditional sense. Instead, they predict what comes next in a sequence based on patterns in their training data. This means they can mimic math answers they’ve seen before but can’t reliably compute or reason through new problems unless they’ve seen similar ones. For reliable calculations, always use dedicated math tools.
AI chatbots often respond to prompts with paragraphs when a one-word answer will do. The standard ChatGPT response to even a simple question reads like a high school book report, containing an introduction, an analysis of an issue from multiple viewpoints, and a summary.
Why might AI models prefer longer, detailed answers? How could this be useful or problematic?
Although it’s now possible for ChatGPT to access data on the internet, the model behind ChatGPT is finite. As a result, ChatGPT doesn’t know everything. When prompted with questions about obscure or recent topics, the answer the model returns may be wrong.
Practitioners often double-check AI-generated answers, especially when dealing with emerging topics or technical details outside common knowledge.
No machine learning model has had the experience of being human. As a result, the responses generated by the model lack common sense. For example, if you ask ChatGPT how to swim to the moon, it may provide an answer without questioning the absurdity of the question. Any human would first question the value of such a strange question.
How might the lack of common sense in AI affect real-world decision-making?
AI chatbots can answer any question with perfect accuracy and logic.
AI chatbots often respond based on patterns in data, not true understanding or reasoning, which can lead to errors or nonsensical answers.
The accuracy of responses generated by a model depends on many factors, including the training data, context, user input, complexity of the prompt and language, and bias. As the user of a model, you have control over only some of these factors. Where possible, however, you can help the model provide better responses by knowing how best to prompt it and by challenging the model’s output in follow-up prompts.
In customer service, inaccurate AI-generated responses can lead to confusion, loss of trust, and even legal consequences if unchecked.
Because they’re trained largely on text written by people, machine learning models will pick up the biases and preferences that exist in the training data. Creators of models put a lot of effort into eliminating bias — which is a worthy goal but an impossible task.
Unintended consequences and even dangerous situations may result from bias in models. The classic example is when Microsoft released its Tay chatbot to the internet in 2016. Within one day of talking to people, the chatbot went from saying things like “Humans are cool” to making racist and sexist comments.
What safeguards can be put in place to limit AI bias and its consequences?
Bias refers to systematic favoritism or prejudice in AI outputs, often inherited from the training data, which can impact fairness and accuracy.
Common sense is the practical human ability to make sound judgments based on everyday experience, something current AI systems lack.
Put an AI chatbot to the test and see how it handles these challenges:
AI systems, especially language models, have important limitations in math, knowledge, common sense, accuracy, and bias — it’s critical to understand these before relying on their answers.
No machine learning model has had the experience of being human, so AI responses often lack common sense and may not question illogical prompts.
You should never trust the output of a chatbot (or any AI system) entirely.
Why do language models struggle with basic math problems?
Tap to revealThey predict answers based on patterns in training data, not by calculating like a calculator.
What is a limitation of AI chatbots regarding common sense?
Tap to revealAI chatbots lack human experience and often do not question absurd or illogical prompts.
What is bias in AI systems?
Tap to revealBias is systematic prejudice or favoritism that AI inherits from its training data, which can affect fairness and accuracy.
How could understanding the limitations of AI change the way you use or rely on AI tools?
Describe a time when you received an incorrect or biased answer from an AI system. How did you recognize the problem, and what did you do next?
Why might an AI chatbot provide an incorrect answer to a simple math problem?
How confident are you that you can explain why AI systems should not always be fully trusted?