Data drives nearly every decision in AI, but what happens when that data is not as reliable as it seems? From “little white lies” to overlooked details, mistruths can slip into datasets in ways both obvious and subtle. Understanding these mistruths is essential for anyone hoping to build, use, or manage AI systems.
Even data collected by sensors can be flawed—technology reflects the limits and biases of its human creators.
Imagine a car accident: multiple parties, multiple stories, and every detail matters. The driver claims to have been blinded by the sun; the pedestrian saw the driver hesitate; a bystander remembers a screech of tires. Even when everyone “tells the truth” as they know it, the resulting data is a tangled web for both humans and AI to unravel.
Humans are used to seeing data for what it is, in many cases: an opinion. In fact, in some cases, people skew data to the point where it becomes useless, a mistruth. A computer can’t tell the difference between truthful and untruthful data — all it sees is data. One issue that makes it difficult, if not impossible, to create an AI that actually thinks like a human is that humans can work with mistruths, and computers can’t. The best you can hope to achieve is to see the errant data as outliers and then filter it out, but that technique doesn’t necessarily solve the problem because a human would still use the data and attempt to determine a truth based on the mistruths that are there.
A common thought about creating less contaminated datasets is that, instead of allowing humans to enter the data, collecting the data via sensors or other means should be possible. Unfortunately, sensors and other mechanical input methodologies reflect the goals of their human inventors and the limits of what the particular technology is able to detect. Consequently, even machine- or sensor-derived data is also subject to generating mistruths that are quite difficult for an AI to detect and overcome.
The following sections use a car accident as the main example to illustrate five types of mistruths that can appear in data. The concepts that the accident is trying to portray may not always appear in data, and they may appear in different ways than discussed. The fact remains that you normally need to deal with these sorts of issues when viewing data.
Commission
Mistruths of commission are those that reflect an outright attempt to substitute truthful information for untruthful information. For example, when filling out an accident report, someone could state that the sun momentarily blinded them, making it impossible to see someone they hit. In reality, perhaps the person was distracted by something else or wasn’t actually thinking about driving (possibly considering a nice dinner). If no one can disprove this theory, the person might get by with a lesser charge. However, the point is that the data would also be contaminated. The effect is that now an insurance company would base premiums on errant data.
Although it would seem that mistruths of commission are completely avoidable, often they aren’t. Humans tell “little white lies” to save others from embarrassment or to deal with an issue with the least amount of personal effort. Sometimes a mistruth of commission is based on errant input or hearsay. In fact, the sources of errors of commission are so many that it is truly difficult to come up with a scenario where someone could avoid them entirely. Regardless, mistruths of commission are one type of mistruth that someone can avoid more often than not.
What are some reasons people might intentionally or unintentionally introduce mistruths of commission into data?
Omission
Mistruths of omission are those in which a person tells the truth in every stated fact but leaves out an important fact that would change the perception of an incident as a whole. Thinking again about the accident report, say that your car strikes a deer, causing significant damage to your car. You truthfully say that the road was wet; it was near twilight, so the light wasn’t as good as it could be; you were a little late in pressing on the brake; and the deer simply darted out from a thicket at the side of the road. The conclusion would be that the incident is simply an accident.
However, you left out an important fact: You were texting at the time. If law enforcement knew about the texting, it would change the reason for the accident to inattentive driving. You might be fined, and the insurance adjuster would use a different reason when entering the incident into the database. As with the mistruth of commission, the resulting errant data would change how the insurance company adjusts premiums.
Avoiding mistruths of omission is nearly impossible. Yes, people can purposely leave facts out of a report, but it’s just as likely that they’ll simply fail to include all the facts. After all, most people are quite rattled after an accident, so they can easily lose focus and report only those truths that leave the most significant impression. Even if a person later remembers additional details and reports them, the database is unlikely to ever contain a full set of truths.
How could mistruths of omission impact the reliability of datasets used by AI?
Perspective
Mistruths of perspective occur when multiple parties view an incident from multiple vantage points. For example, in considering an accident involving a struck pedestrian, the person driving the car, the person getting hit by the car, and a bystander who witnessed the event would all have different perspectives. An officer taking reports from each person would understandably glean different facts from each one, even assuming that each person tells the truth as each knows it. In fact, experience shows that this is almost always the case, and the info that the officer submits as a report is the middle ground of what each of those involved states, augmented by personal experience. In other words, the report will be close to the truth, but not close enough for an AI.
When dealing with perspective, consider vantage point. The driver of the car can see the dashboard and knows the car’s condition at the time of the accident. This is information that the other two parties lack. Likewise, the person getting hit by the car has the best vantage point for seeing the driver’s facial expression (intent). The bystander might be in the best position to see whether the driver made an attempt to stop, and assess issues such as whether the driver tried to swerve. Each party will have to make a report based on seen data without the benefit of hidden data.
Perspective is perhaps the most dangerous of the mistruths because anyone who tries to derive the truth in this scenario ends up, at best, with an average of the various stories, which will never be fully correct. A human viewing the information can rely on intuition and instinct to potentially obtain a better approximation of the truth, but an AI will always use just the average, which means that the AI is always at a significant disadvantage. Unfortunately, avoiding mistruths of perspective is impossible because no matter how many witnesses you have to the event, the best you can hope to achieve is an approximation of the truth, not the actual truth.
Think about this other scenario that involves perception: You’re a deaf person in 1927. Each week, you go to the theater to view a silent film, and for an hour or more, you feel like everyone else. You can experience the movie in the same way everyone else does; there are no differences. In October of that year, you see a sign saying that the theater is upgrading to support a sound system so that it can display talkies — films with a soundtrack.
The sign says that talkies are the best thing ever, and almost everyone seems to agree, except for you, the deaf person, who is now made to feel like a second-class citizen — different from everyone else and even pretty much excluded from the theater. In the deaf person’s eyes (from the perspective of their lived experience), the change is not “the best thing ever,” but a step backward. This illustrates how perspective shapes data, and why AI struggles to interpret reality without understanding these layers.
Why is perspective such a challenging mistruth for AI to handle compared to commission or omission?
- Learned how mistruths can enter datasets through commission, omission, and perspective.
- Recognized that even sensor data is limited by human design and interpretation.
Flashcard Deck: Five Mistruths in Data
What is a mistruth of commission?
Tap to revealAn intentional or unintentional substitution of truthful information with untruthful information in data.
Define mistruth of omission.
Tap to revealWhen important facts are left out from data, changing the overall perception even though all stated facts are true.
What does mistruth of perspective mean?
Tap to revealDifferent parties view an incident from unique vantage points, resulting in varied and subjective data.
Read the following scenario and identify the types of mistruths present:
- A driver reports being distracted by their phone during an accident, but omits mentioning that their windshield wipers were broken.
- Three witnesses give slightly different accounts of the accident.
- One witness claims the driver was speeding, though the speedometer says otherwise.
Reflect on a time when you encountered conflicting information in data—whether in news, work, or daily life. What mistruths might have been present, and how did you attempt to resolve them?
Insurance companies rely on accident reports to set premiums. If those reports contain mistruths—whether through commission, omission, or perspective—their algorithms can miscalculate risk, costing both the company and customers.
Practitioners in AI and data science spend significant effort cleaning data. Spotting outliers is just the start; understanding mistruths requires human insight and constant vigilance.
An intentional or accidental replacement of the truth with false information.
Leaving out important facts that significantly affect the interpretation of an event or dataset.
Data collected by machines or sensors is always reliable and free from human bias.
Sensor data reflects the motives and limitations of its human creators, and can contain mistruths that are difficult for AI to detect and correct.
Researchers in fields like statistics and psychology have long studied the effects of bias, perspective, and omission on data quality. Methods like triangulation, cross-validation, and anomaly detection help mitigate these issues, but no technique can guarantee perfect truth. Understanding the origin and nature of mistruths is key to designing AI systems that are resilient, adaptable, and transparent.
Which type of mistruth occurs when important facts are intentionally or unintentionally left out from data?
AI systems cannot distinguish truth from mistruth in data; understanding and addressing these five types of mistruths is essential for reliable results.
Even with technological advances, data remains subject to human bias, omission, and perspective—challenges that both AI practitioners and users must recognize.
“A computer can’t tell the difference between truthful and untruthful data — all it sees is data.”
The Shift
- Mistruths enter data in multiple ways, challenging AI’s ability to reason like humans.
- Human intuition often helps navigate mistruths, but AI relies strictly on the data provided.
- Building reliable AI systems requires awareness of commission, omission, perspective, and other data pitfalls.