Chapter 4
Data and Information in a Digital World
1 The nature of information
1.1 Introduction
This lecture was recorded under the previous syllabus and originally formed part 2 of a combined Technology and Information chapter. Its opening infrastructure discussion now belongs to Chapter 3; its data material remains a useful foundation. The five ways of using data and the four data competencies are structured explicitly in these notes, and the legal reference is updated to the UK GDPR and Data Protection Act 2018 as amended by the Data (Use and Access) Act 2025.
Good decision-making implies looking at the available data and information and trying to respond to it rationally. Decisions made from a position of ignorance are little more than gambling.
This chapter uses two technical terms: data and information. These can be described as:
Data = raw fact.
For example, a list of the heights of a city’s population or a list of all sales invoices still outstanding could be perfectly accurate, but the unprocessed lists are difficult to use.
Information = data with meaning.
The data has been processed in some way so that it becomes useful and informative. For example:
People’s heights: averages for relevant groups, measures of variation, and charts showing the distribution. The data is now more useful and informative.
Invoices outstanding: sort and total by customer so that you know who owes what. Also sort by date so that you know which debts are older and may need to be chased.
The real challenge of information technology is not collecting data: it is displaying it in a useful format at the right time to the right people.
Knowledge = information ‘in someone’s head’.
Only then can the information be used.
Explicit knowledge: captured (written down or otherwise recorded) and searchable. ‘We know what we know’. For example, product specifications, receivables ledger, inventory quantities and turnover.
Tacit knowledge: not formally captured and recorded; not safeguarded; not shareable nor searchable. For example, a sales person’s skills with customers or a designer’s approach to successful design. When the employee goes, so does their tacit knowledge.
The challenge is to transform tacit knowledge into explicit knowledge so that it can be widely used in the organisation and is not dependent on one person’s presence.
2 The qualities of good information
Good information is often said to have the following properties (mnemonic = ACCURATE)
Accurate: sufficiently accurate not to be misleading. Not all information has to be accurate to the last dollar or cent, and increasing accuracy usually introduces delays and increases costs.
Complete: incomplete information is likely to be misleading.
Note: “Cleaning data” or “data cleansing” are the terms used to describe the process of detecting, correcting or removing corrupt and inaccurate records. Incomplete, inaccurate, irrelevant and out-of-date data must be detected, corrected or removed. Not only is this process important for the purposes of analysing data and making decisions but there are ethical and legal reasons why inaccurate or irrelevant data should not be held.
Cost-beneficial: the benefit arising from using the information should be greater than the cost of providing it.
User-targeted: ensure that the information helps the user to make decisions and to perform his or her job well. Avoid information overload.
Relevant: relevant to the person and to the organisation.
Authoritative: the provenance of the information should be stated and should be reliable. Note that much of the information on the Internet is not authoritative. You might be unsure about its source, whether it is still relevant and has been updated or whether it is deliberately distorted.
Timely: the information should be provided soon enough to be useful in decision-making. Sometimes it is needed almost instantly (for example when controlling machines); in other cases it might not be needed for days or weeks. Faster provision will usually increase its cost.
Easy-to-use: well laid out, well labelled, convenient menu systems and short-cut keys.
It should be realised that people at different levels in an organisation require different types of information
Information for top management is often forward-looking (for planning purposes) as well as historical, makes use of many estimates, and needs to be sourced externally as well as internally (eg what are our competitors doing?). It is often highly summarised and presented in graphical formats. It is often ‘ad hoc’ meaning that very different reports and information can be required at short notice.
Information for operational staff is almost always internal, historical and very detailed and very accurate. It is almost always routine. It is primarily used for recording transactions and making simple decisions.
Middle managers have a mix of informational needs.
A good IT system has to be capable of supplying suitable information to every level.
3 Big data
There are many definitions of the term ‘big data’ but most suggest something like the following:
“Extremely large collections of data (data sets) that may be analysed to reveal patterns, trends, and associations, especially relating to human behaviour and interactions.”
The scale, variety or speed of these data sets may exceed the practical capability of conventional storage and processing methods.
In 2001 analyst Doug Laney described three dimensions that became known as the 3Vs of big data:
Volume
Variety
Velocity
These characteristics, and sometimes additional ones, have been generally adopted as essential qualities of big data.
The commonest fourth ‘V’ that is sometimes added is Veracity: is the data true? Can its accuracy be relied upon?
3.1 Volume
The volume of big data held by large companies such as Walmart (supermarkets), Apple and eBay is measured in multiple petabytes. What’s a petabyte? It’s 1015 bytes (characters) of information. A typical disc on a personal computer (PC) holds 109 bytes (a gigabyte), so the big data depositories of these companies hold at least the data that could typically be held on 1 million PCs, perhaps even 10 to 20 million PCs.
These numbers probably mean little even when converted into equivalent PCs. It is more instructive to list the types of data that large organisations typically hold:
Who collects it | What they typically hold about you |
|---|---|
Retailers | Via loyalty cards: every purchase – what, when, where, how you paid, coupons used. Via websites and apps: every product viewed, every page visited, everything ever bought. |
Social media | Friends and contacts, postings, your location when posting, photographs (which can be scanned for identification), and whatever else you choose to reveal. |
Mobile phone companies | Numbers called, texts sent, every location your phone has been while switched on (to within a few metres), browsing habits, voicemails. |
Internet and browser providers | Every site and page visited, searches entered, downloads, emails (routinely scanned to profile your interests). |
Banks and payment systems | Every receipt and payment; card payment details (amount, date, retailer, location); ATMs used. |
3.2 Variety
Some of the variety of information can be seen from the examples listed above. In particular, the following types of information are held:
Browsing activities: sites, pages visited, membership of sites, downloads, searches
Financial transactions
Interests
Buying habits
Reaction to ads on the internet or to advertising emails
Geographical information
Information about social and business contacts
Text
Numerical information
Graphical information (such as photographs)
Oral information (such as voice mails)
Technical information, such as jet engine vibration and temperature analysis
This data can be both structured and unstructured:
Structured data
This data is stored within defined fields (numerical, text, date etc) often with defined lengths, within a defined record, in a file of similar records. Structured data requires a model of the types and format of business data that will be recorded and how the data will be stored, processed and accessed. This is called a data model. Designing the model defines and limits the data that can be collected and stored, and the processing that can be performed on it.
An example of structured data is found in banking systems, which record the receipts and payments from your current account: date, amount, receipt/payment, short explanations such as payee or source of the money.
Structured data is easily accessible by well-established database structured query languages.
Unstructured data
Unstructured data refers to information that does not have a pre-defined data-model. It comes in all shapes and sizes and this variety and irregularities make it difficult to store it in a way that will allow it to be analysed, searched or otherwise used. An often quoted statistic is that 80% of business data is unstructured, residing in word processor documents, spreadsheets, PowerPoint files, audio, video, social media interactions and map data.
3.3 Velocity
Information must be provided quickly enough to be of use in decision making. There would be little point in a store analysing a customer's in-store behaviour and sending a tailored offer once the customer had already left. If facial recognition is going to be used by shops and hotels, it has to be more or less instant so that guests can be welcomed by name.
You will understand that the volume and variety conspire against the third, velocity. Methods have to be found to process huge quantities of non-uniform, awkward data in real-time.
4 Big data analytics
Conventional databases cannot store or process data at this scale, so specialised technology is used: distributed processing frameworks (such as the open-source Apache Hadoop and its successors) spread the data and the processing across clusters of many machines, and cloud providers offer vast 'data warehouses' and 'data lakes' as a service, so that even modest organisations can rent big-data capability by the hour.
The processing of big data is generally known as big data analytics and includes:
Data mining: analysing data to identify patterns and establish relationships such as associations (where several events are connected), sequences (where one event leads to another) and correlations.
Predictive analytics: a type of data mining which aims to predict future events. For example, the chance of someone being persuaded to upgrade a flight.
Text analytics: scanning text such as emails and word processing documents to extract useful information. It could simply be looking for key-words that indicate an interest in a product or place.
Voice analytics: as above with audio.
Statistical analytics: used to identify trends, correlations and changes in behaviour.
Machine learning (see earlier in this chapter) is the engine behind much of this: predictive analytics is applied machine learning, and text and voice analytics rely on ML models trained to understand language.
5 Extraction, transformation and loading (ETL)
5.1 Introduction
Data is often collected in a number of different database systems, perhaps managed by different organisations, but could be particularly useful if pieces of it were brought together into a single database.
For example, a supermarket will record purchases of products every hour of every day for every branch. A meteorological company will, similarly, record weather conditions and forecasts. It might be interesting to look for relationships between these two repositories of data. Perhaps it would show something like:
Sunny weather or sunny weather forecasted = a strong demand for barbecue supplies
Cold weather or cold weather forecasted = strong demand for ‘comfort food’.
Bringing together elements from separate databases is the heart of ETL.
5.2 Extraction
Data from different sources will inevitably come in many different forms and patterns so extraction has to be carefully tailored for each source. For example, one database might hold dates as ddmmyy, another as mmddyy, another as ddmmyyyy and so on. All of these formats will have to be catered for - and dates are relatively simple (almost standardised) data compared to most of the data that is held in databases.
The data also needs to be validated ie is it correct? There is no point extracting data for future use if it is obviously incorrect.
5.3 Transformation
This process converts the data that has been extracted from its original, probably varied, format into the strict format needed by the new database. So, for example, all dates might be changed into the format dd/mm/yyyy. Similarly, address data might have both ZIP codes (USA) and postal codes (UK) and these will both need to be catered for in the new records.
5.4 Loading
This is the process of writing the data to the new database. It might be a relatively simple process to write the data if it a once-off exercise, but often the ETL process will be carried on every day or every month so that the information is up-to-date. Also, historical data must be preserved so that trends and patterns can be investigated. How much historical information is to be retained? For example, two-year monthly comparatives imply 24 sets of data and when a new month is loaded the oldest is lost and every month’s data moves back one month.
6 Business intelligence and data visualisation
Business intelligence is the technology driven process of:
collecting data
analysing data
presenting information
to help directors, managers and all employees to make informed business decisions.
Generally, the information presented will have historical, current and predictive elements. So, for example, to make an informed business decision about whether or not a production facility should be closed down, managers would need to know:
Historical sales figures, revenues and costs
Current sales figures, revenues and costs
Future, predicted sales figures, revenues and costs.
To complete the picture, information would also be needed about competitor activity, new products that might supplant current production, customers’ preferences, economic outlook
6.1 Data visualisation
The final step of business intelligence – presenting information – has become a discipline in its own right. Data visualisation means presenting data graphically – charts, maps, interactive graphics – so that patterns, trends and outliers communicate themselves far faster than tables of figures can.
Dashboards bring an organisation's KPIs together on one screen, updated in real time, often with the ability to 'drill down' from a headline figure to the detail behind it.
Self-service BI tools (for example, Microsoft Power BI or Tableau) let managers and accountants build their own reports and visualisations directly, without waiting for the IT department to program them.
Interactive visualisations let the user change the question – filter by region, slide a date range, switch products – and see the picture update instantly.
Good visualisation follows the qualities of good information: the right chart for the message, honestly scaled axes, and no more detail than the audience needs. A badly designed chart – a truncated axis exaggerating growth, a 3-D pie chart obscuring proportions – misleads exactly as effectively as it informs, which makes visualisation an ethical issue as well as a technical one. For finance, visualisation is central to the 'communicate insight to influence' step of the information to impact framework (Chapter 2).
7 Dangers of big data
Despite the examples of the use of big data in commerce, particularly for marketing and customer relationship management, there are some potential dangers and drawbacks.
Cost: It is expensive to establish the hardware and analytical software needed, though these costs are continually falling.
Regulation: Some countries and cultures worry about the amount of information that is being collected and have passed laws governing its collection, storage and use. Breaking a law can have serious reputational and punitive consequences.
Loss and theft of data: Apart from the consequences arising from regulatory breaches as mentioned above, companies might find themselves open to civil legal action if data were stolen and individuals suffered as a consequence.
Incorrect data (veracity): If data is incorrect or out of date incorrect conclusions are likely. Even if the data is correct, some correlations might be spurious leading to false positive results.
Employee monitoring: Data collection can allow employees to be monitored in detail every second of the day. For example, sensors in name badges can record employee movements, whom each employee talks to and even in what tone of voice; warehouse systems can measure each worker's rate of picking minute by minute. This information could be used to reduce stress and improve how people work together – but it can just as easily be used to put employees under constant, oppressive surveillance.
7.1 Data protection law
Around the 1990s many governments became worried about the amount of information being held about their citizens and the potential for damage or misuse. Many countries therefore enacted data protection legislation to govern the collection and use of data.
For example, in the UK the UK GDPR and the Data Protection Act 2018, as amended by the Data (Use and Access) Act 2025 (and in the EU the General Data Protection Regulation, GDPR, on which they are based), impose obligations not to collect more personal data than is needed, not to hold it for longer than needed, and to keep it accurate and secure. Individuals have rights including access, correction and, in certain circumstances, deletion. Breaches can attract very large fines.
Obviously the huge amounts of data now being collected and held (see the big data section above) mean that this problem is potentially more serious.
7.2 Ethical, social and security dangers
In addition to the obvious privacy issues, the following are ethical, social and security dangers:
Incorrect data leading to incorrect decisions. For example, incorrect information on financial affairs can lead to individuals being refused credit.
Theft of information (such as the theft of credit card information, the hacking of emails and industrial espionage).
Unauthorised alteration of information.
Fraudulent websites (eg ‘phishing’ sites can look like legitimate bank sites and can induce people to enter their account and PIN numbers).
Time-wasting by employees as they browse the internet.
Downloading or sending offensive material.
Violations of copyright laws.
Denial of service attacks (DoS) where there are attempts to make a machine or network unavailable to its intended users. For example, a site is bombarded with automated requests for access and the site fails.
Computer viruses which might simply be a nuisance or which are designed to cripple machines and systems.
Physical dangers, such as floods, fire or terrorist attacks can mean that organisations cannot continue to function.
Innocent harm being done. For example, there have been several recent cases of banks updating their software and errors in the updates caused on-line banking and cash machines to fail for several days.
Ransomware attacks, where criminals encrypt an organisation's data and demand payment to restore it.
It is therefore essential that organisations:
Ensure that data is collected legally.
Ensure that the data is input completely and accurately.
Ensure that data is held safely and securely – including physical security.
Ensure that data can be accessed only for authorised purposes by authorised individuals
Ensure that software is properly tested.
Arrange regular backups of data.
Use virus checkers.
Use firewalls to prevent unauthorised access from outside (by hackers).
Have suitable standby arrangement (disaster recovery plan) should the computer system be severely damaged.
7.3 Ethics of data usage
Complying with data protection law is the floor, not the ceiling. The syllabus expects finance professionals to think about the ethics of data usage – questions the law may not yet answer:
Consent and expectation: was the data collected for the purpose it is now being used for? Customers who shared a delivery address did not thereby consent to psychological profiling.
Fairness and bias: does analysis of the data treat groups of people fairly, or does it systematically disadvantage some – in pricing, credit or employment?
Transparency: would the organisation be comfortable explaining publicly what it does with the data it holds?
Proportionality: is the benefit to the organisation proportionate to the intrusion on the individual – particularly with employee monitoring?
Finance often owns, analyses or reports finance-controlled datasets and uses data owned by other business functions, so these are finance's ethical questions, not just the IT department's. They connect directly to the ethical principles in Chapter 12.
8 The five ways of using data
8.1 Introduction
Everything so far in this chapter points one way: machines now collect and process data better and more cheaply than people do. The role of the finance professional is therefore shifting from producing information to using it – using data to create and preserve value for the organisation. IT is not an end in itself: in a profit-seeking business it should enhance long-term profitability, and in a not-for-profit organisation it should cut costs or improve the services provided.
Analysis of data can be done at increasing levels of sophistication:
Descriptive analytics: what happened? (routine reporting)
Diagnostic analytics: why did it happen? (drilling into causes, e.g. variance analysis)
Predictive analytics: what is likely to happen? (forecasting)
Prescriptive analytics: what should we do about it? (recommending or automating the response)
The syllabus identifies five ways in which the finance function uses data to create and preserve value.
8.2 1. Decision-making
The most fundamental use. Good decisions need good information: whether to close a production facility requires historical, current and predicted revenues and costs, competitor activity and the economic outlook – exactly what business intelligence assembles. Data-driven decision-making replaces gut feel with evidence at every level, from daily operational choices (which supplier, which price) to strategic ones (which market to enter).
8.3 2. Understanding the customer
The data trail customers leave – purchases, browsing, locations, reactions to offers – lets an organisation understand who its customers actually are and how they behave. Customers can be segmented into groups (by age, location, life stage, buying habits) and each segment's profitability measured. Finance's contribution is to quantify it: which customers and segments actually make the organisation money – something revenue figures alone do not reveal once the costs of serving each segment are counted.
8.4 3. Developing the customer value proposition
Understanding customers is only valuable if the organisation acts on it: designing products, services, prices and channels that give each segment what it values. Analysing consumer behaviour shows which features are valued and which are not; rapid design and production technology turns that insight into products quickly. A clothing retailer such as Zara designs thousands of products a year and moves each from design to shops in a few weeks – impossible without data on what is selling flowing straight back into design decisions. Data can also improve the proposition itself: personalised recommendations, tailored offers, products that learn from use.
8.5 4. Enhancing operational efficiency
Automating processes removes labour cost and error from routine transactions – in the factory, the warehouse and the finance function itself.
Better forecasting aligns purchasing and production with real demand, cutting waste and stock-outs.
Coordination with suppliers and logistics: shared, real-time data is what makes just-in-time inventory possible – orders, parts and deliveries synchronised end to end (see also extranets, earlier in this chapter).
KPI monitoring: technology delivers key performance indicators daily or in real time – yesterday's sales, margins, output, bed occupancy – so problems are corrected while they still matter. KPIs are developed further in the interface chapters (Chapters 4 to 12).
8.6 5. Monetising data
Finally, data itself can be turned directly into revenue:
New products and services added to existing offerings, identified from analysis of customer behaviour.
Entirely new business models: streaming subscriptions instead of broadcast schedules, online marketplaces instead of high-street stores – businesses whose model rests on innovative use of data.
Sharing data with alliance partners so that both parties analyse a larger, more reliable data set.
Selling data: data that is not personal can be sold outright – satellite-derived geodata on topography, soil and geology, for example, commands high prices from agriculture, construction and infrastructure planners. Selling personal data is tightly restricted by data protection law.
The five ways of using data: decision-making; understanding the customer; developing the customer value proposition; enhancing operational efficiency; and monetising data – all subject to the ethics of data usage. Each maps onto finance's primary activities: data feeds the analysis that produces insight, insight is communicated to influence decisions, and decisions are implemented to create value.
9 The four data competencies
9.1 Introduction
To use data in these five ways, finance professionals need a defined set of data competencies. The syllabus names four:
Competency | What it involves |
|---|---|
1. Data strategy and planning | Deciding what data the organisation needs and how it will get, govern and use it |
2. Data engineering, extraction and mining | Getting data out of source systems and into usable form (ETL, data platforms, data mining) |
3. Data modelling, manipulation and analysis | Structuring and analysing the data to answer business questions (analytics, ML, BI) |
4. Data and insight communication | Turning analysis into influence: visualisation, dashboards, narrative reporting |
9.2 1. Data strategy and planning
This starts with an assessment of data needs: what decisions does the organisation have to make, and therefore what data does it need, at what quality, how often, from where? It also covers planning how data will be governed – who owns each data set, how quality is maintained, how long data is kept, who may access it. Finance is well placed to lead here, because finance knows what the organisation is trying to achieve and what information the decisions require. This is where finance's data competency has its clearest competitive advantage.
9.3 2. Data engineering, extraction and mining
The plumbing: extracting data from source systems, transforming it into consistent form and loading it into a common store – the ETL process described earlier – plus building and running the big-data platforms and mining the data for patterns. This is the most technical competency, and the one where finance most needs to work alongside IT specialists and data scientists. Finance's role is to specify what data is needed and to check that what arrives is complete and accurate, not to build the pipelines itself.
9.4 3. Data modelling, manipulation and analysis
Structuring data so it can be analysed, and applying the analytical techniques – from spreadsheet models and variance analysis through business intelligence to big data analytics and machine learning. This competency is shared: data scientists bring sophisticated statistical and ML techniques, but they generally know little about finance or the business. Their skills have to be steered by finance professionals so that the right questions are asked, sensible hypotheses are tested and results are interpreted commercially – a spurious correlation is only obvious to someone who understands the business.
9.5 4. Data and insight communication
Analysis creates no value until someone acts on it. The fourth competency is communicating data and insight so that it influences decisions: choosing the message, the audience and the format; using data visualisation and dashboards; and telling the story behind the numbers in plain language. This is the competency employers most often say is missing in technical specialists – and it is where accountants, as trained communicators of financial information, add distinctive value.
9.6 Finance and the data scientists
Pulling the four competencies together: finance's comparative advantage lies at the two ends of the chain – knowing what data the business needs (competency 1) and communicating insight to influence decisions (competency 4). Data scientists' advantage lies in the middle – engineering and advanced analysis (competencies 2 and 3). The finance professional acts as the link between the business, the IT function and the data scientists: determining the data needed, asking the relevant questions, validating and interpreting the analysis, and then making sure the insight is actually used – so that, in the language of Chapter 2's information to impact framework, data ends up improving the organisation's performance.
Learn the two lists as lists – the five ways of using data and the four data competencies are exactly the kind of named frameworks that objective test questions ask you to identify, complete or match to examples.
10 Test your knowledge
Two quick checks before you move on: work through the flashcards to fix this chapter’s key terms and definitions, then sit the objective questions for exam-style practice. Both mark themselves and explain the answers as you go.
Data and Information in a Digital World
18 questionsAnswer the questions one at a time. Your progress is saved so you can leave and come back.
Open chapter practice

