Big Data
1 Introduction
Big data are extremely large collections of data that may be analysed to reveal patterns and trends, especially relating to human behaviour.
For example, Amazon - one of the world’s leading e-retailers - collects an enormous amount of information about each customer. Not only, obviously, what a customer has bought, but what other products they have have looked at on their website but not bought, how they made payment for purchases, where they live, what items they have have returned, what products they have ordered more than once, what ratings and/or comments that customers may have posted about specific products. All of this enables them to make recommendations to customers of other items that might appeal to them.
2 Characteristics of big data
Volume
The amount of data being stored by the likes of Amazon is enormous - many times the amount of data that can be stored on a normal PC.
Variety
The data will come from a large variety of sources - again, Amazon collects information about past purchases, which other products the customer has looked at on their website, etc. etc.. All of this information needs special software and algorithms that collate the different data into useful information.
Velocity
For the data to be useful, it needs to be analysed and information provided quickly enough. In the case of Amazon, suggestions need to be presented to customers immediately they access the website.
Value
Gathering and analysing data is expensive in terms of particularly the software needed. There is little point in incurring this expense if the company does not benefit by getting value from it. For example, Amazon gains value by being able to make additional sales.
Veracity
Veracity means accuracy and truthfulness. If a company is to make use of big data they need to be confident that the data is reliable and that the analysis is accurate.
3 Types of big data
Structured data
Structured data is data that is already held in a defined way. An example is the data held by a bank recording receipts and payments from your bank account. This is stored on the bank’s database in a pre-defined way and it easily accessible.
Unstructured data
Unstructured data is data that is not held in a pre-defined way and would include, for example, word processor documents and spreadsheets. This is typically not stored in a way that is easily accessible.
Semi-structured data
This is data that is not able to be stored in a conventional database, but by the use of markers is stored in a way that can be accessed more easily than unstructured data.
4 General data terms
There are several terms relating to data that you should be aware of for the exam, whether or not the data being held is part of big data or not.
Numerical data
This is data that is held as numbers, for example a database containing the marks that all students achieved in an exam. You can be expected to perform some arithmetic on numerical data in the exam, and these calculations are explained in the next chapter.
Categorical data
Categorical data can be either nominal or ordinal.
Nominal data (or naming data) is data held that records, for example, students’ names and addresses. Ordinal data is data that records, for example, the ratings given to a product or service by customers when asked to give a rating from 1 (poor) to 5 (excellent).
To be useful, data needs to be analysed and summarised, and you should be aware of the terms descriptive analysis and inferential analysis.
Descriptive analysis
Descriptive analysis is taking the recorded data and describing it in various ways. For example, we might have a record of the marks obtained by all students who took one particular exam. One analysis we might perform to is calculate the average mark obtained on that particular exam.
Inferential analysis
Suppose we wished to know the average salary earned by the entire population of a country. It would obviously not be practical to ask every person in the country what their salary was, and so what we might do is ask a random sample of (say) 2,000 people and calculate the average for those 2,000. We would then infer that this was a reasonable estimate of the average salary for the entire population.
Big Data
5 questionsAnswer the questions one at a time. Your progress is saved so you can leave and come back.
Open chapter practice

