Introduction
The digital revolution divided social scientists and researchers into two categories:
- Enthusiasts for the availability of data
- Sceptics about the reliability of data and the issues about collecting and data ownership
The availability of big amounts of data, with low prices and times for the collecting, brings some issues: in the first place, the planning of the strategy for the data collection and analysis seems to be less accurate; at the same time, data becomes old in a shorter period.
Digital data (= digital traces of human behavior and opinions recorded by a wide set of digital services operating in different domains of society) are also complex:
- They reshaped the distinction between online and offline research
- Their use between people is evolving: people know that their digital presence is observable, and so they modify their behavior to adapt to this visibility
- Digital data is becoming more complex (e.g., questionnaire: with the outcome of questions, we can now collect metadata - data about data - and paradata - data about process)
+ Digital data from survey and interviews are cross-sectional without a longitudinal temporal dimension; + Most social science datasets are coarse aggregations of variables because of the limitations in what can be asked from self-reported instruments. If we do not include our digital social life in our research practices, our capacity to understand human societies is greatly diminished.
Digital research using digital data and methods
Self-reported and behavioral data
Surveys are an example of self-reported data = when someone reports on something they have done or on what they think or believe vs. observational or behavioral data. One of the problems of self-reported data is the social desirability: there are social norms governing some attitudes and behaviors and people sometimes misrepresent themselves to appear to comply with these norms. But they have been the main way of conducting social research.
Since 1990, psychologists have distinguished between two systems of thought: System 1 (S1) - made up of intuitive thoughts of great capacity, based on associations acquired through experience and quickly and automatically calculates information / System 2 (S2) - involves low-capacity reflective thinking, based on rules acquired through culture or formal learning and calculates information in a relatively slow, effortful, and controlled manner. To these systems of thoughts are associated two processes: Type 1 (fast, automatic, unconscious) and Type 2 (slow, conscious, controlled) > Daniel Kahneman, Thinking, Fast and Slow (2011) > the "dual model" is now the most used to understand human behavior. Humans are rational with limits, their decisions are influenced by systematic error and biases originated by how our cognitions and emotions work.
Self-reported methods are not the best option to study the influences of environments and unconscious thought on human behavior > the use of behavioral data is complementary to them because they have a great deal of information on social relations and people’s embeddedness BUT there are still biases generated by the design and the aims of digital platforms (affordances). The large majority of data collected by social researchers have been static (collected in a given time) because the collection of longitudinal data was very expensive > it has been difficult to study social events with a longitudinal perspective. Digital data introduces this opportunity.
Big Data
A common description of big data refers to the three Vs > Volume (quantity of data produced), Velocity (fast-moving nature of digital data, that are often produced “on the fly”), Variety (multiplicity of formats that data can have in the digital world). Big data implies some complications for researchers: they are generated from a vast range of invisible processes with incomparable dimensions. They are not the design of a researcher who already has in mind an idea of theoretical framework and an analytical strategy > they are organic, created by different actors in the context not of research, but of producing or delivering goods or services vs. designed data (e.g., questionnaires).
The secondary datasets analysis (traditionally used, is the reuse of existing datasets collected by official institutions or other researchers) has in common with big data this idea of repurposing of data = data collected originally for other aims are repurposed for specific research goals BUT digital data add a level of complexity: the lack of transparency on how the data were collected is a problem for researchers.
Three positive features of big data:
- Increased size and resolution: not only we can include in our research a big number of cases and participants (moving from sample to population), but we can also have a larger set of data points per individual (which is very interesting for the measurement of stratified variables).
- Big data is long - that is, longitudinal: refers to their lengthy time span.
- Non-reactive heterogeneity, including behavioral: they are not, in most cases, collected by means of direct elicitation of people > The invisibility of digital data collection raises ethical concerns but is also an opportunity to study and capture behavioral information (better than traditional social scientific instruments, as we already discussed before).
These features should not divert attention from methodological problems such as random errors (caused by unknown and unpredictable changes in the measurement, tend to distribute according to a normal or Gaussian distribution so that if you increase the size of data you reduce them) and systematic errors (that result from the way data are created and very large datasets might therefore blind researchers to this kind of error).
The construct validity problem
When we work with data designed by someone else for some other purpose, we have to reverse the traditional steps of social research: we already have variables that we need to interpret as measures of a property, that are relevant to an indicator linked to the concept we want to study. We also have to consider that a single indicator is not able to express an abstract concept, that needs several indicators to be fully covered + We have challenges in the construct validity of our indicators = the validity with which we can make generalizations from research operations (because we are operating with measures that we didn’t design and that were never intended to represent the construct of interest).
To demonstrate construct validity:
- Traditionally, researchers utilized multitrait-multimethod matrix (MTMM) = construct validity is established when measures of the same construct measured using different methods correlate more strongly with one another compared with measures of different constructs measured with the same method and different methods.
- Currently, the dominant approach is the confirmatory factor analysis (CFA) = if the measure is good and fit established, and fits the data better than other plausible alternatives, is thought to be construct-valid.
Both are problematic in big data dimension: for example, considering behavioral traces, if fewer than three behavioral indicators exist for a construct, the CFA model is not identified. This is a challenge for researchers: the most convincing option is to assess construct validity by testing whether variables constructed from digital data are correlated with other accepted measures of the construct.
Representativeness and access
In the social sciences, we study large populations, so we need to draw only samples of observation. The best thing is to have a random sample: that permits us to have a good idea of how accurate our estimates are. If we have a non-random sample, we need to convince ourselves and others that it closely resembles what we would obtain with a truly random sample.
In case of a non-random sample, we need to consider also the transportability issue of pattern found in one subset of the population to other parts of it or to other populations > big data from digital platforms do not have a base of users that can be considered a representative sample of the entire population, simply because not everyone is a user of every platform + deep web = is the part of the WWW that is not indexed by traditional search engines (like Google) (for example, everything that is password protected) and dark web = websites that are not accessible using normal browsers because they use the Tor encryption: we need to consider traditional search engines as a sample.
> Problem of access to data produced by third parties other than researchers (governmental institutions, private companies, NGO), which is constituted by many factors: willingness to share data; legal restrictions (not only the legal framework that regulates the use and sharing of data, like the European GDPR – NB. European countries tend to be stricter than the US, and that generates another problem of “academic inequality” – but also the terms of use present in many datasets); business motives (for researchers is important to share data to obtain validation by the scientific community BUT companies want to keep data private in order not to give any advantage to competitors (data are important economic assets).
> Issue of opacity of digital data = many digital infrastructures hide the role played by the algorithms, but we have to remember that what we see is already the outcome of analytical operations.
Native or complex digital methods
Digitalized methods = digital transposition of existing methods
Native digital methods = new methods built around a specificity of digital technologies in order to take advantage of their features (e.g., Rogers, “search of research”). The challenge will be to develop complex mixed digital methods, the combination of different methods, both at the collection and at the analytical level.
Digital structured, unstructured and semi-structured data
Structured data = made up of clearly defined data types whose pattern makes them easily searchable. They usually reside in relational database management systems (RDMBMS). They can be human- or machine-generated. This format is searchable both with human-generated queries and via algorithms and mature analytics tools exist for it.
Unstructured data = data that is usually not easy to search, including formats like audio, video and SM posting. Unstructured data analysis is a nascent industry, still not mature. They have internal structure, but are not structured via predefined data models or schema. They may be textual or not, human-generated (e.g., social media) or machine-generated (e.g., satellite imagery, digital surveillance). Unstructured data make up most of social scientific interesting data, and analytics tools for mining them are developing, most of them based on machine learning (= analytics that not only work at computer speeds, but also automatically learn from their activity and user decisions) > the analysis of unstructured data with machine learning allows organizations to analyze digital communications for compliance, track high-volume customer conversations in social media and gain new marketing intelligence.
Semi-structured data: they are becoming common (e.g., Email: has some internal structure, thanks to its metadata, but its message field is unstructured; XML; JSON, NoSQL databases = they differ from relational databases because they do not separate the organization schema from the data).
Unobtrusive vs obtrusive methods
The distinction between data that are results of an obtrusive method and of an unobtrusive method is important because people tend to react to researchers' measurements > Hawthorne effect = individuals modify their behavior when they are aware that they are observed + social desirability = people modify their behavior in order to comply with social norms (that is particularly evident in SM, where people present themselves as their positive image).
The increased opportunity of online unobtrusive methods (like SM network analysis) has generated concerns about covert research = often methods prevent subjects from knowing that their behaviors and communications are being observed and recorded (remember: we need informed consent!) > the use of digital data of an unobtrusive nature has been more problematic, especially when we analyze platforms designed for social networking, in which people don’t have the expectation that what they share will become public + another element of complication is the mutating legal context and the differences between countries (e.g., GDPR, that requires informed consent every time that data have been collected for one use and they are repurposed for another).
Web scraping and news sources
The objective of web scraping is to build a corpus for further analysis and for text processing and mining (starting from WWW pages) > it has some standard steps:
- Crawling: extraction of the web content and data reorganization, obtaining a list of URL address links
- Parsing web pages to extract metadata and content
- Scraping and preparing data for late analysis
The focus is on web pages that use HTTP (hypertext transfer protocol - the rule with which web pages store content), as news sites, and social media.
HTML is a special form of XML, a data storage format (for storing and transporting data in a structured way), and has its own type of tags and attributes, recognized by browsers > HTML files can be transformed to XML files for portability purpose, because XML languages store the information as plain text and use user-defined and comprehensible > BUT when we deal with large datasets or data, with a strong hierarchical structure, it may take up a lot of memory to try to import and manipulate the data. There are different formats of XML, like .docx, .odt, .epub.
Another standard for data storage and interchanges found frequently on the web is JavaScript Object Notation (JSON), which has some preferable features. The crucial property is that plain text is unstructured data, at least for computer programs that simply read a text file line by line.
Depending on the technique that has been used to collect data, there are specific tools that are preferable:
- R: the advantage is that we can use all the technologies from within R.
- XPath query language: used to select pieces of information from documents in HTML, XML, through a series of filtering and extraction steps.
- JSON documents are more lightweight.
- In order to extract information from AJAX-enriched web pages it is preferable to use Selenium.
Besides web scraping, we have text mining, which provides solutions for the automatic categorization of text and is particularly useful when analyzing web data, often unlabelled and unstructured text. Once we have obtained web texts, researchers have different options for their analysis, for example, they can choose between quantitative methods and qualitative methods.
Social media data
Social scientists that want to use social media data have to answer three questions:
- Which SM platform is the most appropriate in relation to my research questions? They have to collect evidence that a platform is normally preferred for such and such activity, and then they can select it.
- What are the criteria and their rationale for selecting and collecting data from this platform? For example, if we are studying political behavior on Twitter, how do we translate the available possible actions that Twitter users can perform (like likes) into politically meaningful behavior?
- What level of resolution (= amount of data) and analysis is required?
When thinking about SM data, it is useful to think about several general sources of biases:
- Access biases: no SM platform gives unlimited access to its data to a third party like researchers and the limitations are often not clear and can be sources of systematic errors (e.g., Twitter has a limit on the number of hourly tweets that can be downloaded about a topic)
- SM populations: we know a little about SM demographic characteristics, but we can assume that the SM populations are not representative of the general population
- Sampling biases that are common in every sample-based research (e.g., if we collect tweets based on their geolocalization, we are excluding everyone that has preferred to not use a geolocalization)
- Nonhuman agents, like bots (computer-generated users programmed to post and distribute specific content)
- Distortions of human behavior: for example, Google only stores the final researches submitted by users, after auto-completion is done, excluding what people have typed > we cannot use these data to study human framing of searches.
Collecting data from APIs
When we talk about APIs, we are referring to web services or web APIs, and they are important because they allow the retrieval and processing of data, providing data in various formats (JSON has become the most popular). Standardization of APIs helps programmers familiarize with the mechanics of an API quickly: the more popular API standards are REST (= representational state transfer > resources are referenced and representation of these resources, that are documents like HTML, XML or JSON file, are exchanged) and SOAP.
Many web services are open to anybody, but sometimes APIs require the user to register and provide an individual key when making a request to the web service, and authentication is used to trace data usage and to restrict access + related to authentication is authorization, that means granting an application access to authentication details. There are a large number of tutorials available online.
Understanding social media data
We need to think about which different entities are contained in social media data and how their relations can carry information that can be explicit or implicit > from a conceptual point of view, the information that we can collect on SM concerns:
- Users (ID for example)
- Content/Resources (text, image, video, weblink)
- Metadata (geolocation)
- Groups, a combination of the users, content, and metadata.
We can now identify the explicit relationships between such entities:
- Users-Users: most SM platforms contain information about relationships between users (e.g., the information about who follows whom).
Scarica il documento per vederlo tutto.
Scarica il documento per vederlo tutto.
Scarica il documento per vederlo tutto.
Scarica il documento per vederlo tutto.
Scarica il documento per vederlo tutto.
-
Riassunto esame Digital media , Prof. Ceravolo Flavio, libro consigliato Appunti Digital Epistemology, Ceravolo
-
Riassunto esame Social media e Digital marketing, prof Rea, libro consigliato Social media marketing, Tuten, Solomo…
-
Riassunto Esame Digital Marketing, prof Rea, libro consigliato Social media ROI, Cosenza
-
Riassunto esame Digital media, Prof. Locatelli Elisabetta, libro consigliato Media digitali. Digital media , Balbi,…