Why isn’t collecting more data always better for AI—and what questions must we answer before gathering it?
If you get the feeling that everyone is acquiring your data without thought or reason, you’re right. Organizations collect, categorize, and store everyone’s data—seemingly without goal or intent.
By 2025, it’s estimated that 463 exabytes of data will be created each day globally—yet most organizations use only a fraction of it effectively.
Data acquisition has become a narcotic for organizations worldwide, and some leaders seem to think that whoever collects the most data wins. However, collecting data for its own sake accomplishes nothing. Douglas Adams, in The Hitchhiker’s Guide to the Galaxy, illustrates this with a race of supercreatures who build an immense computer to calculate the meaning of “life, the universe, and everything.” The computer’s answer—42—doesn’t help because the real question remains unknown. This parable reminds us: gathering data in unlimited amounts is futile unless you first know the questions you need to answer.
Think about a time when you or your organization collected data without a clear goal. What challenges did you encounter?
The main problem any organization needs to address with data acquisition is which questions to ask and why those questions matter. Tailoring data acquisition to answer the right questions is essential.
The process of collecting and measuring information on variables of interest in a systematic way to answer specific questions or solve problems.
Imagine you run a small retail shop. What information would you need to improve sales? Write down 3–5 specific questions you’d want your data to answer.
- List five questions whose answers could help your shop grow.
- For each question, consider what kind of data you would need to collect.
- How many people walk in front of the store each day?
- How many of those people stop to look in the window?
- How long do they look?
- What time of day are they looking?
- Do certain displays tend to produce better results?
- Which of these displays cause people to enter the store and shop?
Why is it important to validate if each question you ask truly addresses a business need before collecting data?
The use of technology to automatically collect, process, and store data, often at scale, with minimal human intervention.
Collecting all this data by hand would be impossible, which is where automation comes in. However, automation can introduce new problems and unreliable data if not implemented thoughtfully.
- Sensors can collect only the data they’re designed to collect, so you might miss data when sensors aren’t suited for your purpose.
- People create errant data in various ways, which means the data you receive might be false.
- Data can become skewed when the conditions for collecting it are incorrectly defined.
- Interpreting data incorrectly means that AI outputs will also be incorrect.
- Converting a real-world question into an algorithm that a computer can understand is an error-prone process.
Retail analytics platforms use sensors and cameras to track customer movements in stores. If not carefully calibrated, these systems might miss key events or misinterpret staff movements as customer behavior, leading to wrong business decisions.
Want to go deeper? The science behind data quality issues
Data quality is not just about accuracy; it also includes completeness, consistency, and relevance. Automated systems can introduce errors through misconfigured sensors, software bugs, or misinterpretation of context. Data scientists routinely use techniques like validation, cleansing, and cross-checking with multiple sources to mitigate these risks and ensure the data used for AI is as reliable as possible.
Collecting more data always leads to better AI results.
Without clear questions and high-quality, relevant data, collecting more information can actually worsen AI outcomes and lead to false conclusions.
The answer is indeed correct, but they need to know the question in order for the answer to make sense.
Why is it critical to define your questions before acquiring data?
Tap to revealBecause data collection is only useful when it is guided by clear, important questions that address real business needs.
What is a key risk of automating data acquisition?
Tap to revealAutomation can produce unreliable or irrelevant data if not carefully designed and monitored.
What does “skewed data” mean in the context of data acquisition?
Tap to revealData that does not accurately represent reality due to errors in how it was collected or defined.
- Collecting more data is not always better—focus on purposeful questions.
- Automation can help, but also introduces new risks and challenges.
How could unreliable or ill-formed data mislead an organization’s AI-driven decisions?
Reflect on a situation where data collection or analysis led to an unexpected or incorrect conclusion. What could have been done differently to ensure the right data was acquired and used?
Effective data acquisition starts with asking the right questions and understanding the limitations and potential pitfalls of both manual and automated data collection.
Poorly collected or irrelevant data combined with inappropriate algorithms can mislead decision-making and undermine the value of AI.
What strategies can organizations use to ensure that their data acquisition truly serves their business goals instead of just accumulating information?
The process of is only valuable when guided by well-defined questions and clear business needs.
If you automate data collection, you don’t need to worry about the data’s relevance or quality.
Practitioners recommend starting every AI project by first defining the decisions you need to make, then identifying the data that will best support those decisions—never the other way around.
How confident are you that you can explain why collecting more data isn’t always better for AI?
The Shift
- Purposeful data acquisition begins by asking the right questions, not by hoarding information.
- Automation is powerful but can introduce new risks—careful design and validation are essential.
- The quality and relevance of data are as important as the quantity for successful AI outcomes.