Today, data is everywhere—fueling discoveries in science, powering the apps on your phone, and transforming how artificial intelligence (AI) systems learn and interact with the world. But what makes this era’s data so different, and why does it matter for AI?
Every minute, people upload over 500 hours of video to YouTube and send more than 40 million messages via WhatsApp. Most of this information is unstructured, providing vast learning opportunities for AI.
You may have heard big data mentioned in many specialized scientific and business publications, and you may have even wondered what the term really means. From a technical perspective, big data refers to large and complex amounts of computer data, so large and intricate that applications can’t deal with the data by simply using additional storage or increasing computer power.
Extremely large and complex datasets that cannot be managed or processed by traditional data-handling applications.
Big data implies a revolution in data storage and manipulation. It affects what you can achieve with data in more qualitative terms (meaning that in addition to doing more, you can perform tasks better). From a human perspective, computers store big data in different data formats (such as database files and .csv files), but regardless of storage type, the computer still sees data as a stream of ones and zeros (the core language of computers). You can view data as being one of two types, structured and unstructured, depending on how you produce and consume it. Some data has a clear structure (you know exactly what it contains and where to find every piece of data), whereas other data is unstructured (you have an idea of what it contains, but you don’t know exactly how it is arranged).
- Structured data include typical examples such as database tables, in which information is arranged into columns and each column contains a specific type of information. Data is often structured by design. You gather it selectively and record it in its correct place. For example, you may want to place a count of the number of people buying a certain product in a specific column, in a specific table, or in a specific database. As with a library, if you know what data you need, you can find it immediately.
-
TIPUnstructured data consists of images, videos, and sound recordings. You may use an unstructured form for text so that you can tag it with characteristics, such as size, date, or content type. Usually, you don’t know exactly where data appears in an unstructured dataset, because the data appears as sequences of ones and zeros that an application must interpret or visualize.
REMEMBERTransforming unstructured data into a structured form can cost lots of time and effort and can involve the work of many people. Most of the data of the big data revolution is unstructured and stored as is, unless someone renders it structured.
Think about the types of data you interact with every day—how many are structured versus unstructured?
This copious and sophisticated data store didn’t appear suddenly overnight. It took time to develop the technology to store this amount of data. In addition, it took time to spread the technology that generates and delivers data — namely, computers, sensors, smart mobile phones, and the internet and its web services. The following sections help you understand what makes data a universal resource today.
Want to go deeper? The science behind structured vs. unstructured data
Structured data is highly organized and easily searchable using simple, straightforward algorithms. Think of spreadsheets and relational databases—every value is in a predictable place. Unstructured data, however, requires advanced AI techniques like natural language processing (NLP) or image recognition to extract meaning, because it doesn’t follow a pre-defined format. The explosion of digital media (photos, videos, sensor readings) has made unstructured data the dominant type in today’s digital world.
Using data everywhere
Scientists need more powerful computers than the average person because of their scientific experiments. They began dealing with impressive amounts of data years before anyone coined the term big data. At that point, the internet wasn’t producing the vast sums of data that it does today.
Big data isn’t a fad created by software and hardware vendors but has a basis in many scientific fields, such as astronomy (space missions), satellite (surveillance and monitoring), meteorology (storm predictions), physics (particle accelerators), and genomics (DNA sequences).
Although an AI application can specialize in a scientific field — such as IBM’s Watson, which boasts an impressive medical-diagnosis capability because it can learn information from millions of scientific papers on diseases and medicine — the actual AI application driver often has more mundane facets. Actual AI applications are mostly prized for being able to recognize objects, move along paths, or understand what people say and speak to them. Data contribution to the actual AI renaissance that molded it in such a fashion didn’t derive from the classical sources of scientific data.
The internet now generates and distributes new data in large amounts. Our current daily data production is estimated to amount to about 2.5 quintillion (a number with 18 zeros) bytes, with the lion’s share going to unstructured data like video and audio.
Wearable devices like smartwatches are used by millions to track health data. This information, though largely unstructured, is helping AI spot early signs of diseases such as COVID-19—sometimes before symptoms appear.
All this data is related to common human activities, feelings, experiences, and relations. Roaming through this data, an AI can easily learn how reasoning and acting more human-like works. Here are some examples of the more interesting data you can find:
-
Large repositories of faces and expressions from photos and videos posted on social media websites like Facebook, YouTube, and Google: They provide information about gender, age, feelings, and possibly sexual orientation, political orientation, or IQ (see “Face-reading AI will be able to detect your politics and IQ, professor says” at The
Guardian.com). - Privately held medical information and biometric data from smartwatches, which measure body data such as temperature and heart rate during both illness and good health: Interestingly enough, data from smartwatches is seen as a method to detect serious diseases, such as COVID-19, early.
- Datasets of how people relate to each other and what drives their interest from sources such as social media and search engines: For instance, a study from Cambridge University’s Psychometrics Centre claims that Facebook interactions contain a lot of data about intimate relationships.
-
Information on how we speak, which is recorded by mobile phones. For example, OK Google, a function found on Android mobile phones, routinely records questions and sometimes even more, as explained in “Google’s been quietly recording your voice; here’s how to listen to — and delete — the archive” at
https://qz.com/526545/googles-been-quietly-recording-your-voice-heres-how-to-listen-to-and-delete-the-archive.
How do you feel about the idea that your spoken commands or social media posts might be part of massive datasets that help train AI systems?
Every day, users connect even more devices to the internet that start storing new personal data. There are now personal assistants that sit in houses, such as Amazon Echo and other integrated smart home devices that offer ways to regulate and facilitate the domestic environment. These are just the tip of the iceberg because many other common tools of everyday life are becoming interconnected (from the refrigerator to the toothbrush) and able to process, record, and transmit data. The internet of things (IoT) is becoming a reality.
Practitioners note that AI’s progress is now closely tied to the “datafication” of everyday life—where even mundane actions become data points for smart technologies to learn from.
Information that does not follow a specific model or format, such as text, images, or audio, requiring interpretation or processing to extract meaning.
- Big data is too large and complex for traditional storage and processing methods.
- Structured and unstructured data both play crucial roles in AI development.
- The majority of new data generated today is unstructured, such as videos, audio, and social media posts.
This copious and sophisticated data store didn’t appear suddenly overnight. It took time to develop the technology to store this amount of data.
Consider a technology you use daily. What kinds of data might it be generating, and how could that data be used for AI?
Big data—especially unstructured data—has become a universal resource, powering the rapid progress of AI across scientific and everyday domains.
Which type of data makes up the majority of the information generated daily on the internet?
What is structured data?
Tap to revealData organized into a defined format—like tables with specific columns—making it easy to search and process.
Give an example of unstructured data.
Tap to revealPhotos, videos, audio recordings, or free-form text that don’t fit a defined format.
Why is big data important for AI?
Tap to revealIt provides vast, diverse information for AI to learn patterns, improve predictions, and mimic human-like reasoning.
Big data is only valuable in scientific research and has little impact on everyday life.
Big data now comes largely from everyday activities—like social media, smart devices, and daily interactions—making it essential to both science and daily AI applications.
Take a moment to observe the digital services you use in a typical day. Can you spot examples of structured and unstructured data?
- List three apps or devices you use frequently (e.g., messaging app, fitness tracker, social media).
- For each, identify one type of structured data and one type of unstructured data it generates or uses.
- Reflect: Which type do you think is most prevalent, and why?
In what ways do you think the data you generate every day could be valuable to AI systems? Consider both positive uses (like health or safety) and potential risks.
Understanding the difference between structured and unstructured data is crucial for grasping how AI systems learn and make decisions in our data-rich world.