All guides

Oblique Strategies

Reading the Data

Lenses for making sense of data before drawing conclusions: who, what, where and when; the six Vs; data to knowledge; analytics techniques; and the friction in the system.

Updated Sep 11, 202613 min read
data
analytics
evidence
thinking

Most bad decisions start with data that was never properly looked at. These lenses slow that step down. Each one is a checklist of questions to ask of a dataset, a report or a dashboard before anyone acts on it.

Start with the first two, which are about provenance. Move to the six Vs when the volume itself is the problem, and to the techniques when you need to choose a method.

  • Who: When using an evidence-based or critical approach to understanding the problem space, it is important to consider the individuals or groups who will be involved in the process. This may include data scientists, analysts, and domain experts who have the necessary skills and expertise to understand and analyze the data. It is also important to consider the stakeholders who will be impacted by the results of the analysis and to involve them in the process.
  • What: The "what" refers to the data and information that will be analyzed in order to understand the problem space. This may include data from a variety of sources, such as sensor data, text data, and image data. The data must be relevant to the problem or question being addressed, and should be of high quality and accuracy.
  • Where: The "where" refers to the location or context in which the data is collected and analyzed. It is important to consider the specific environment or context in which the data was collected, as this can have a significant impact on the results of the analysis.
  • When: The "when" refers to the time at which the data was collected and analyzed. It is important to consider the timeliness of the data, as well as any temporal patterns or trends that may be present in the data. For example, in some cases, data collected at different times may be analyzed together, while in other cases, it may be necessary to analyze data from different time periods separately.

By considering these factors, an evidence-based or critical approach can be used to gain a more comprehensive and accurate understanding of the problem space.1

  • Volumes: The "volumes" refer to the amount of data that will be analyzed in order to understand the problem space. It is important to consider the size and scale of the data, as well as the computational resources required to process and analyze the data. It is important to have the right computational resources and data storage to handle the volume of data being analyzed and to be aware of the limitations and challenges that may arise from working with large amounts of data.
  • Frequencies: The "frequencies" refer to the rate at which the data is generated or collected. It is important to consider the temporal aspect of the data, as well as any patterns or trends in the data that may be related to the frequency. For example, data collected at high frequencies may reveal patterns and trends that are not visible in data collected at lower frequencies.
  • Effects: The "effects" refer to the impact or outcome of the analysis on the problem space. It is important to consider the potential implications of the analysis and to evaluate the effectiveness of the analysis in achieving the desired outcome.
  • Trends: The "trends" refer to patterns or changes in the data over time. It is important to consider the temporal aspect of the data and to identify and analyze any trends that may be present in the data. This can help to understand the underlying patterns and relationships in the data and can provide insights into the problem space.

By considering these factors, an evidence-based or critical approach can be used to gain a more comprehensive and accurate understanding of the problem space. The ability to analyze trends and frequencies, as well as the effects of the analysis on the problem space, can lead to more accurate conclusions and more effective decision-making.

  • Data: Data refers to raw facts and figures that are collected and recorded. Data can come from a variety of sources, such as sensor data, text data, and image data. Data is unorganized and in its raw form, and it needs to be processed and analyzed to extract meaningful information.
  • Information: Information is data that has been processed and organized in a way that makes it meaningful and useful. Information is often presented in the form of tables, charts, and graphs and it can be used to make decisions and take action.
  • Knowledge: Knowledge is information that has been analyzed and interpreted, and it has been contextualized in a way that allows it to be applied to a specific problem or situation. It is the understanding of a subject or a field and it is gained through learning and experience.
  • Learning: Learning is the process of acquiring knowledge, skills, values, and attitudes through education and experience. It is the process of gaining understanding and insight into information and knowledge, and it enables individuals to improve their performance and make better decisions.

This process emphasizes that data alone is not enough to make informed decisions, but by progressing through each stage, the data becomes more valuable, and it can be used to make more informed decisions and drive learning. The process of turning data into information, knowledge, and learning is a continuous cycle, with new data constantly being collected and analyzed to improve understanding and drive decision-making.2

  • Volume: Volume refers to the amount of data that is generated and stored. It can be measured in terms of storage space, number of records, or number of transactions. With the growing amount of data generated by various sources, organizations must have the ability to handle large volumes of data.
  • Velocity: Velocity refers to the speed at which data is generated, captured and processed. It can be measured in terms of the rate at which data is created, the time it takes to process data, or the frequency of data updates. As the data is generated and captured at a faster rate, organizations must have the ability to process the data in real-time.
  • Variety: Variety refers to the different types and formats of data that are generated. It can be structured data, semi-structured data, or unstructured data. Variety can include data from various sources such as social media, IoT devices, and transactional systems. Organizations must have the ability to handle and process different types of data.
  • Variability: Variability refers to the degree to which data can change over time. It can be measured in terms of the rate of change, the amount of change, or the nature of the change. Organizations must have the ability to handle and process data that changes over time.
  • Veracity: Veracity refers to the trustworthiness and accuracy of data. It can be measured in terms of the completeness, consistency, and correctness of data. Organizations must have the ability to ensure that the data is accurate and reliable.
  • Value: Value refers to the benefit or usefulness of data to an organization. It can be measured in terms of the insights that can be gained from data, the decisions that can be made, and the actions that can be taken. Organizations must have the ability to extract value from data by leveraging advanced analytics and visualization tools.

These concepts are often referred to as the "6 V's of Big Data"3, and they represent the challenges and opportunities that organizations face when dealing with big data. By understanding and managing the volume, velocity, variety, variability, veracity, and value of data, organizations can derive insights and make better decisions, which can lead to better performance and increased competitiveness.

In data analytics, several key techniques and methodologies are employed to derive insights from data.4 These include:

  • Aggregation: Summarizing and combining data from various sources to provide an overall view.
  • Experimentation: Running controlled experiments to test hypotheses and understand cause-and-effect relationships.
  • Prediction: Using statistical models and machine learning algorithms to forecast future trends or outcomes based on historical data.
  • Clustering: Grouping data points based on similarities to identify patterns or segments within the data.
  • Decision Trees: A decision support tool that uses a tree-like model of decisions and their possible consequences, including chance event outcomes.
  • Accumulation: Collecting and storing data over time to analyze trends and patterns.
  • Derivative: Calculating the rate of change in a variable relative to another, often used in time series analysis.
  • Funnel: Analyzing the customer journey to identify stages where drop-offs occur and optimize conversion rates.

To implement these analytics techniques effectively, a range of tools and platforms can be utilized:

  • Apache Spark: A powerful open-source engine for large-scale data processing, enabling quick processing of large data sets.
  • BigQuery: Google Cloud's serverless, highly scalable, and cost-effective multi-cloud data warehouse designed for business agility.
  • Snowflake: A cloud-based data warehousing platform that provides a scalable and flexible architecture for storing and analyzing large volumes of data.
  • Apache Kafka: A distributed streaming platform that is used for building real-time data pipelines and streaming applications.
  • Apache Flink: A stream processing framework that processes data streams in a real-time and stateful manner.
  • Spark Streaming: An extension of Apache Spark that enables scalable, high-throughput, fault-tolerant stream processing of live data streams.
  • Amazon DQ: A data quality service that provides tools to ensure that your data is accurate, complete, and reliable, often used in conjunction with other analytics tools.

The standard techniques, in roughly the order you meet them.5

  • Forecast: Forecasting is a machine learning technique that uses historical data to make predictions about future events or outcomes. This can be used in a variety of applications, such as financial forecasting, weather forecasting, and demand forecasting. Common techniques include time series analysis and regression analysis.
  • Linear: Linear machine learning models are used for tasks such as regression and classification. These models are based on linear equations and are used to find the best-fitting line or plane that describes the relationship between the input and output variables. Linear models are simple and easy to interpret but may not be able to capture complex relationships in the data.
  • Classification: Classification is a machine learning technique used to assign data points to specific categories or labels. This can be used for tasks such as image classification, email filtering, and fraud detection. Common techniques include decision trees, random forests, and support vector machines.
  • K-Cluster: K-clustering is a technique used to group similar data points together. This can be used for tasks such as customer segmentation, anomaly detection, and image compression. The "k" refers to the number of clusters that the data is divided into. K-means and hierarchical clustering are common techniques.
  • Pivot: Pivot is a technique used to reorganize data in a way that makes it easier to analyze and understand. This can be used to turn rows into columns and columns into rows, making it easier to identify patterns and trends in the data. Pivot tables and cross-tabulation are common techniques.
  • Grouping: Grouping is a technique used to group data points based on specific criteria. This can be used for tasks such as data summarization and data visualization. Grouping can be done by aggregating data by mean, median, sum, and other statistical measures.
  • Statistics: Totals and Counts: Totals and counts are statistical measures used to summarize data. Totals are used to find the sum of a set of numbers, while counts are used to find the number of observations in a dataset. Totals and counts are often used in conjunction with other statistical measures, such as mean and standard deviation, to gain a more comprehensive understanding of the data.
  • Viscosity: In data and information systems, viscosity refers to the resistance to change in the data flow and processes. High viscosity can occur when there are multiple layers of data silos and a lack of standardization in data management, making it difficult to make changes or updates to the system. To overcome viscosity, it may be necessary to implement a data governance framework and standardize data management processes.
  • Friction: In data and information systems, friction refers to the resistance to change caused by different stakeholders and groups within the system. It can occur when there are conflicting priorities and a lack of communication between different departments or groups. To overcome friction, it may be necessary to establish clear communication channels and align the priorities of different stakeholders.
  • Automation: Automation in data and information systems refers to the use of technology to automate repetitive tasks and processes, such as data entry and data analysis. Automation can improve efficiency, reduce errors, and increase scalability in the system.
  • Immediacy: Immediacy in data and information systems refers to the ability to access and process data in real-time. This can be achieved through the use of technologies such as in-memory databases and real-time analytics. Immediacy can improve decision-making and reduce the time it takes to respond to changes in the system.
  • Niche: In data and information systems, a niche refers to a specific area or function that the system is designed to support. By focusing on a specific niche, the system can be optimized to meet the specific needs of that area, resulting in improved performance and increased efficiency.
  • Governance: Governance in data and information systems refers to the policies, procedures, and standards that are in place to ensure the security, integrity, and compliance of the data. Governance can help to ensure that the system is operating effectively and efficiently and that any risks are identified and managed.
  • Learning: Learning in data and information systems refers to the ability of the system to adapt and improve over time. This can be achieved through the use of machine learning algorithms, which can analyze data and improve the performance of the system. Learning can help to improve the accuracy and effectiveness of the system and can reduce the need for manual intervention.

Footnotes🔗

  1. The who, what, where and when are Kipling's "six honest serving-men" (with why and how). Kipling, R. (1902). The Elephant's Child. In Just So Stories. Macmillan. ↩

  2. Ackoff, R. L. (1989). From data to wisdom. Journal of Applied Systems Analysis, 16, 3–9. The hierarchy is reviewed in Rowley, J. (2007). The wisdom hierarchy: Representations of the DIKW hierarchy. Journal of Information Science, 33(2), 163–180. doi:10.1177/0165551506070706 ↩

  3. The original three Vs: Laney, D. (2001). 3D Data Management: Controlling Data Volume, Velocity, and Variety (META Group research note, 6 February 2001). Veracity, variability and value were added by later authors; see Gandomi, A., & Haider, M. (2015). Beyond the hype: Big data concepts, methods, and analytics. International Journal of Information Management, 35(2), 137–144. doi:10.1016/j.ijinfomgt.2014.10.007 ↩

  4. For a business-facing treatment of these tasks see Provost, F., & Fawcett, T. (2013). Data Science for Business. O'Reilly. ↩

  5. The standard reference for these methods is Hastie, T., Tibshirani, R., & Friedman, J. (2009). The Elements of Statistical Learning (2nd ed.). Springer. Free PDF from the authors. ↩

Reading the Data | Push Manifesto · Push Manifesto