Skip to contentSkip to search

Big Data

VIVA Subject Guide
YouTube video

1 Introduction

Big data are extremely large collections of data that may be analysed to reveal patterns and trends, especially relating to human behaviour.

For example, Amazon - one of the world’s leading e-retailers - collects an enormous amount of information about each customer. Not only, obviously, what a customer has bought, but what other products they have have looked at on their website but not bought, how they made payment for purchases, where they live, what items they have have returned, what products they have ordered more than once, what ratings and/or comments that customers may have posted about specific products. All of this enables them to make recommendations to customers of other items that might appeal to them.

2 Characteristics of big data

  • Volume

The amount of data being stored by the likes of Amazon is enormous - many times the amount of data that can be stored on a normal PC.

  • Variety

The data will come from a large variety of sources - again, Amazon collects information about past purchases, which other products the customer has looked at on their website, etc. etc.. All of this information needs special software and algorithms that collate the different data into useful information.

  • Velocity

For the data to be useful, it needs to be analysed and information provided quickly enough. In the case of Amazon, suggestions need to be presented to customers immediately they access the website.

  • Value

Gathering and analysing data is expensive in terms of particularly the software needed. There is little point in incurring this expense if the company does not benefit by getting value from it. For example, Amazon gains value by being able to make additional sales.

  • Veracity

Veracity means accuracy and truthfulness. If a company is to make use of big data they need to be confident that the data is reliable and that the analysis is accurate.

3 Types of big data

  • Structured data

Structured data is data that is already held in a defined way. An example is the data held by a bank recording receipts and payments from your bank account. This is stored on the bank’s database in a pre-defined way and it easily accessible.

  • Unstructured data

Unstructured data is data that is not held in a pre-defined way and would include, for example, word processor documents and spreadsheets. This is typically not stored in a way that is easily accessible.

  • Semi-structured data

This is data that is not able to be stored in a conventional database, but by the use of markers is stored in a way that can be accessed more easily than unstructured data.

4 General data terms

There are several terms relating to data that you should be aware of for the exam, whether or not the data being held is part of big data or not.

  • Numerical data

This is data that is held as numbers, for example a database containing the marks that all students achieved in an exam. You can be expected to perform some arithmetic on numerical data in the exam, and these calculations are explained in the next chapter.

  • Categorical data

Categorical data can be either nominal or ordinal.

Nominal data (or naming data) is data held that records, for example, students’ names and addresses. Ordinal data is data that records, for example, the ratings given to a product or service by customers when asked to give a rating from 1 (poor) to 5 (excellent).

To be useful, data needs to be analysed and summarised, and you should be aware of the terms descriptive analysis and inferential analysis.

  • Descriptive analysis

Descriptive analysis is taking the recorded data and describing it in various ways. For example, we might have a record of the marks obtained by all students who took one particular exam. One analysis we might perform to is calculate the average mark obtained on that particular exam.

  • Inferential analysis

Suppose we wished to know the average salary earned by the entire population of a country. It would obviously not be practical to ask every person in the country what their salary was, and so what we might do is ask a random sample of (say) 2,000 people and calculate the average for those 2,000. We would then infer that this was a reasonable estimate of the average salary for the entire population.

Practice questions

Big Data

5 questions

Answer the questions one at a time. Your progress is saved so you can leave and come back.

Open chapter practice