Module 1: How data is made
Topic 1: How Actors Shape Data
a. Definition of factors
From a social science perspective, actors are individuals or groups (such as
organisations, institutions, or communities) that act, make decisions, and interact
with others within a social context.
They observe and understand the world. They do so through the lens of
their pre-existing knowledge, presumptions, biases, and rationalities. This lens
is shaped by their cultural, social, educational, disciplinary, and professional
environments.
Actors also view the world with a certain goal-orientation: they observe not just
passively but with the intention of acting, solving problems, or achieving specific
objectives => their observations are selective and influenced by what is important
to their goals => a connection between their thoughts and actions.
b. Factors Influencing the Shaping of Data
- Different characteristics:
o Objective (size, colour, smell) vs Subjective (grace, beauty)
o History
o Purposes
o Price
o Relationship (possession)
=> need to decide which elements or entities to include, which aspects and properties
to include, as well as which relationships.
c. Mapping the Effect of Characteristics and Backgrounds
The different characteristics and backgrounds of actors can influence their perspective
on the world and its elements, and these perspectives are reflected in the data they
make:
Cultural background
Social background
Economic background
Political background
Educational background, and – particularly relevant for scientific research –
disciplinary background
Past experience
* Influenced by rational decisions
These decisions include determining the purpose of the data, identifying the most
important aspects, and considering how the data might be used or modified in the
future. They also involve selecting the most appropriate tools and methods for
collecting and analysing data.
Interest and purpose: determine which elements or entities are selected for data
collection => act as a filter
Relevance and significance: decisions must be made about which aspects are
considered relevant and significant enough to be included. These decisions
shape what the data ultimately represents and can be based on different types of
reasoning:
o Instrumental: meaning that relevance is judged based on whether an
aspect serves a specific purpose or goal.
o Essentialist: where certain aspects are viewed as inherently important or
unimportant, regardless of a specific goal. These decisions are often influenced
by cultural norms, traditions, or disciplinary perspectives. For example, in early
medical research, certain health indicators, such as women's pain levels, were
often considered less significant, not necessarily for instrumental reasons but
due to longstanding biases in medical data collection.
=> Both of these premises influence what is considered "valid data", shaping
not only how information is collected but also what kind of knowledge can be
produced from it.
Variability: In data creation, some aspects are seen as changeable and worth
measuring over time, while others are treated as fixed and unchanging. These
assumptions shape how data is structured and what patterns can be observed.
These choices influence how trends are analyzed, what changes are considered
significant, and what remains invisible in the data. Eg: Health research: when
tracking patients with chronic diseases, researchers often repeatedly measure
variables such as blood pressure and weight, on the assumption that these factors
are likely to change and are important indicators of health trends. By contrast, they
may treat a patient’s genetic background or date of birth as fixed data points that
do not require repeated measurement. => trends and changes in blood pressure are
visible and analysed, while any potential long-term shifts in how genetic
information is understood or classified remain invisible in the dataset.
Alternatives: Since interests and purposes do not strictly dictate which elements or
aspects must be represented in data, data-makers must actively choose what to
include and how to represent it. These choices shape the final dataset, but they are
also influenced by whether and to what extent alternative options were
considered. Eg: company decides to visualise its sales data. For example,
suppose a business wants to present its product sales performance. The data team
could use a simple bar chart to show total sales per product, highlighting the most
popular products. However, if they consider other options, they might use a
grouped bar chart to compare sales across different regions or a bubble chart to
show the relationship between sales volume, revenue and customer
demographics. Each visualisation brings different aspects of the data into
focus, revealing patterns or disparities that a single approach might miss. Thus, the
final dataset and its interpretation are shaped by the alternatives considered and
chosen, demonstrating how active decision-making about representation
influences the possible insights.
Futher processing: The way data is structured is heavily influenced by how it is
expected to be processed later. Choices about which elements and aspects are
included often depend on the tools, techniques, and analytical methods that will
be used.
Topic 2: Data & Information
a. Model
Definiton: A model is a simplified version of something and it helps us to
understand or explain it better.
Key features:
o The Mapping Feature, which means that models represent something real or
imaginary
o The Reduction Property, which describes how models simplify reality by focusing
on what’s important and leaving out unnecessary details.
o The Pragmatic Property, which highlights that models are created for a specific
purpose.
*Note: models are simplifications and may not always capture every aspect of reality.
Benefits:
o Make complex ideas easier to understand by simplifying them and focusing on
key details.
o Allow researchers to test theories, predict outcomes, and collaborate across
different fields.
Challenges:
o Rely on assumptions, and if these are wrong, the model’s accuracy is affected.
o Over-rely on models, forgetting that they are only simplified versions of reality
Risks in using
o Modeling assumptions.
o Modeling authority: the people who create the models have control over what
is included. The specific perspective could therefore introduce bias in the detail
highlighted in the model.
o The decision premises: we often trust model results a little too much and
sometimes forget that they simplify a specific aspect of reality.
b. Data & Information
Data refers to simplified symbols or numbers that represent parts of reality.
However, these raw symbols, like numbers or letters, do not carry meaning by
themselves. For example, the number "3" can represent three apples. Information
adds context and meaning. It is a little more complex. For example, "There are
three red apples on the table."\
=> Data comprises raw, unprocessed facts and observations about the world, it
becomes information when we organise it and add meaning.
Both are created as simplified models of reality, focusing on specific details while
leaving out others.
Four dimensions of Information - The semiotic concept of information:
Semiotics is a field of study that deals with the use of signs for meaning-making.
This concept explains what information is. It says that something is information if
it has the four dimensions you see in the picture:
1. On the left hand side, there is the syntactic (ngữ pháp) dimension. The
syntax, refers to the structuring. For information, this means the structure of the
data.
2. The pragmatic dimension on the right hand side: the effect of the data, which
can depend on the context in which the data is perceived as well as who it is
used by.
3. The semantic dimension at the top: What is the message we get from the data?
4. The sigmatic dimension at the bottom: information relates to something in the
real world, such as an object or concept. That is also why computers can
process data but lack the ability to grasp the full complexity of information as
humans do, such as deciding between something real and something imagined.
Turning Data into Information (Datafication) :
o Semiosis: starts with observations, which basically means that people observe
the world, pick out specific elements, and adapt them to fit their needs =>
turning mental ideas into symbols, like expressing thoughts through words.
o Formalisation: this process follows this rules. It creates data that computers
can process — but only as data. Remember the four dimensions of information
that we just discussed. Computers work with symbols based on their structure
but don’t understand their meaning or the context. They can’t interpret
meaning or consider the impact of information. To help with this, metadata —
essentially "data about data" — is used to provide context or meaning.
However, for computers, metadata is still just a set of symbols, allowing them
to compare it syntactically without truly understanding it.
o Ultimately, computers process data, not information. Humans must interpret
processed data, assign meaning, and embed it in social contexts to transform it
into information that can be understood or acted upon.
Risks of Datafication: data is never neutral — it reflects human decisions,
priorities, and biases
Topic 3: How Structure Influences Meaning
a. How to Structure Data
Catergorisation: ontological distinctions.
o At the most basic level, categories are conceptual tools we use to make
distinctions between things, ideas, or entities in the world. These distinctions
help us structure our understanding of reality.
o In essence, they are ontological distinctions, meaning they define what "kinds
of things" exist in the world from a particular perspective.
o However, these categories are not always neutral. They often reflect cultural,
social, or disciplinary assumptions about the world. The choice of categories
shapes what is considered important or relevant, and influences how we
perceive reality.
Classification: organising the world
o Classification is when you put things into groups based on how they fit in
with other things. It is a way of arranging things into categories.
Classification makes the world simpler by helping us to understand, analyse
and use it more easily.
o For example, in healthcare data, patients may be classified based on age
groups, medical conditions, or treatment plans. These classifications allow
healthcare professionals to recognize patterns and make informed
decisions.
o Classifications also come with assumptions. These are ideas or beliefs that
might affect how we classify things. The decision of how to classify depends
on what we think is most important. For example, if we were to base our
classification on a person's salary, we might be more aware of differences in
wealth, but we might also miss other important differences, like cultural or
geographical background.
Typification: recognising patterns
o Typification is when you create types or models that show the usual
characteristics of a category. When we typify something, we identify the
most common or expected attributes that define members of a category. This
helps us to make generalisations (sự tổng quát hóa) and form expectations
about data.
o For example, we might think of a "typical" student as someone between 18 and
25, studying full-time for a degree. This helps institutions to make policies or
design systems that cater to most people, but it can mean that those who don't
fit this mould - like older students or part-time learners - are overlooked.
o When we turn data into types, we often create averages, norms, or models that
guide our analysis. But typifications can also exclude unusual cases or
groups, slightly changing data-driven decisions and results so that they might
not show the full diversity of a population.
Systematisation: creating systems of organisation
o Systematisation is the process of creating systems that organise categories in
a logical way. It involves arranging categories and types into a clear
framework, often with hierarchies or relationships between them.
o Think of systematisation as building a map of categories, with each category
connected to others in important ways. These systems create a
comprehensive picture of how various entities relate to one another.
Identification: tying data to individuals, groups and things
o The process of identification links the abstract categories and classifications to
real people or things.
o For example, if a patient's data says they are diabetic, their treatment will be
based on that. But identifying someone based on categories can also be
limiting. It can overlook the fact that every person has a different background
and different circumstances.
o In digital systems, identification is important for tracking, personalising and
categorising individuals based on the data they produce. This can lead to
highly specific and tailored services, but there are also ethical concerns about
privacy and consent.
Topic 4: Qualitative vs. Quantitative Data
a. Main Objective of Qualitative Research
- Qualitative research aims at gaining a deep understanding of a topic by exploring the
characteristics or categories of a specific phenomenon.
- It is particularly valuable when little is known about a subject, as it helps to form the
basis of further research. From the results, hypotheses can be generated that may later
be tested using quantitative methods.
b. Definition
Qualitative data refers to non-numerical information such as interviews, newspaper
articles, photographs or detailed observations that are systematically recorded. For
example, if you’re studying how people feel about a new policy, you might interview
them to learn their thoughts and emotions. This type of data can be used to uncover
how a phenomenon is constituted or how people experience and interpret it.
Quantitative data, on the other hand, is numerical and can be counted or measured.
Think of “quantity” as a description for an amount of something. Therefore, it is very
useful when you want to test specific hypotheses to see if one thing causes another,
get statistical proof for something, or find patterns and connections, like figuring out
why something happens or how big a trend is. Most of the time, this kind of research
looks at lots of data, which makes it easier to say something about a larger group of
people. Typically, quantitative data is collected through standardised methods like
coded questionnaires, censuses, or physical measurements. For example, you could
use a survey to ask people to rate their satisfaction on a scale from one to ten.
Statistics like population sizes or percentages are other examples of quantitative data.
*Note abt quantitive data:
- The way questions are asked or how the data is analysed can significantly influence
the results.
- When people look at the data to analyse it, their own ideas and ways of thinking can
influence how they understand it
- Statistical effects or special circumstances of the sample can lead to
misinterpretations of data, revealing, for example, correlations where there are none.
b. Mixed-Methods – Combining Qualitative and Quantitative Data
- It is used to get a more complete picture of a certain phenomenon, using the
strengths of both qualitative and quantitative data.
- So, with a mixed-methods approach you can for example identify trends and
understand people’s motifs at the same time. For instance, you might use a survey
with a set of defined answer choices to get quantitative data, and also include open-
ended questions where people can explain their own answers in their own words for
qualitative insights.
c. Categories in Qualitative and Quantitative Data
- Categories help organise qualitative data by grouping similar characteristics, like
colors, cities, or patterns, making it easier to observe and analyse phenomena.
- Express numerical or measurable groups or values for quantitative data
=>These groupings provide structure to descriptive information and serve as a
foundation for both qualitative and quantitative analysis.