<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>data collection on Methods Bites</title>
    <link>https://socialsciencedatalab.mzes.uni-mannheim.de/tags/data-collection/</link>
    <description>Recent content in data collection on Methods Bites</description>
    <generator>Hugo -- gohugo.io</generator>
    <lastBuildDate>Wed, 25 Mar 2026 00:00:00 +0000</lastBuildDate>
    
        <atom:link href="https://socialsciencedatalab.mzes.uni-mannheim.de/tags/data-collection/index.xml" rel="self" type="application/rss+xml" />
    
    
    <item>
      <title>Survey Recruitment on LinkedIn: A Step-by-Step Guide to Targeted Outreach</title>
      <link>https://socialsciencedatalab.mzes.uni-mannheim.de/article/linked-in/</link>
      <pubDate>Wed, 25 Mar 2026 00:00:00 +0000</pubDate>
      
      <guid>https://socialsciencedatalab.mzes.uni-mannheim.de/article/linked-in/</guid>
      <description><![CDATA[
        </p>
<!-- Optional: One paragraph on learning objectives -->
<p>LinkedIn’s professional user base and precise targeting capabilities present social science researchers with a valuable opportunity to recruit survey participants from highly specific populations, particularly those relevant to research on employment and industry dynamics. While platforms such as Facebook and Instagram are commonly used for recruitment of survey participants, LinkedIn’s potential in survey participant recruitment remains underutilized. Its professional focus, career-centric context, and only minimal off-topic content make it ideal for reaching the potential labor force population. This is particularly relevant for studies requiring insights from specific industries, occupations, or education levels.
In this Methods Bites Tutorial <a href="https://www.linkedin.com/in/zaza-zindel">Dr. Zaza Zindel</a> (German Centre for Integration and Migration Research (DeZIM)) and <a href="https://www.linkedin.com/in/lisa-de-vries-38465423a">Dr. Lisa de Vries</a> (FernUniversität in Hagen) provide a step-by-step guide on how to use LinkedIn ads to recruit participants for survey research.</p>
<p>After reading this blog post, readers will be able to:</p>
<ul>
<li>understand how LinkedIn advertisements can be used for survey participant recruitment,</li>
<li>set up targeted LinkedIn advertising campaigns, and</li>
<li>evaluate campaign performance and adapt recruitment strategies accordingly.</li>
</ul>
<!-- If applicable: Note with references on related materials -->
<!-- 
Generate an overview of the article 
Note: Section anchors are generated automatically from section headings as 
      hyphenated lower-case labels without special characters; e.g. a section
      title "Why R?" will result in the anchor "#why-r".
-->
<div id="overview" class="section level3">
<h3>Overview</h3>
<ol style="list-style-type: decimal">
<li><a href="#first-section"><strong>Introduction to LinkedIn ads as survey recruitment method</strong></a></li>
<li><a href="#second-section"><strong>Recruiting via LinkedIn – a step-by-step guide</strong></a>
<ol style="list-style-type: decimal">
<li><a href="#second-first-subsection">LinkedIn Campaign Manager</a>
<ol style="list-style-type: decimal">
<li><a href="#second-first-first-subsubsection">Personal LinkedIn account</a></li>
<li><a href="#second-first-second-subsubsection">LinkedIn ad account</a></li>
<li><a href="#second-first-third-subsubsection">LinkedIn page</a></li>
</ol></li>
<li><a href="#second-second-subsection">Setting up an ad campaign</a>
<ol style="list-style-type: decimal">
<li><a href="#second-second-first-subsubsection">Campaign level</a></li>
<li><a href="#second-second-second-subsubsection">Ad set level</a></li>
<li><a href="#second-second-third-subsubsection">Ad level</a></li>
</ol></li>
</ol></li>
<li><a href="#third-section"><strong>Ethical considerations</strong></a></li>
<li><a href="#fourth-section"><strong>Application Example</strong></a>
<ol style="list-style-type: decimal">
<li><a href="#fourth-first-subsection">Ad campaign details</a>
<ol style="list-style-type: decimal">
<li><a href="#fourth-first-first-subsubsection">Ad creatives</a></li>
<li><a href="#fourth-first-second-subsubsection">Campaign setup</a></li>
</ol></li>
<li><a href="#fourth-second-subsection">Campaign results</a></li>
<li><a href="#fourth-third-subsection">Sample composition</a></li>
</ol></li>
<li><a href="#fifth-section"><strong>Summary and key takeaways</strong></a></li>
<li><a href="#further-reading"><strong>Further reading</strong></a></li>
</ol>
</div>
<div id="first-section" class="section level3">
<h3>Introduction to LinkedIn ads as survey recruitment method</h3>
<p>Social media platforms have become an increasingly common tool for survey participant recruitment, allowing researchers to reach large populations quickly, cost-effectively, and with fine-grained targeting options. Most existing methodological work in this area, however, has focused on general-purpose platforms such as Facebook and Instagram. Systematic reviews of social media recruitment consistently show that these platforms dominate the field, whereas LinkedIn appears only sporadically and is rarely the central focus of empirical or methodological studies (e.g., <a href="https://doi.org/10.12758/mda.2022.15">Zindel 2023</a>; <a href="https://doi.org/10.7748/nr.2023.e1859">Jones et al. 2023</a>). This imbalance is striking, because LinkedIn offers a professionalized user base, rich profile information, and targeting parameters that are directly relevant to research on labor market and workplace related topics.</p>
<p>LinkedIn is a professional networking platform where people connect, share content, and participate in industry-specific communities. With more than one billion registered members worldwide (<a href="https://about.linkedin.com/">LinkedIn 2025</a>), it has become a central hub for working professionals, job seekers, and industry leaders. Because LinkedIn is explicitly work-oriented – and because user profiles contain rich educational and occupational information – it offers considerable potential for social science research. The platform mirrors real-world professional networks and career trajectories, making it particularly interesting for studies on labor market behavior, employment trends, and professional attitudes. A central strength of LinkedIn is the ability to reach narrowly defined populations based on industry, occupation, seniority, education, or geographic region. These segments are often difficult or impossible to access through traditional recruitment channels or general social media platforms. The granularity of profile information and targeting options enables researchers to sample very specific occupational groups or sectors, which is valuable for both descriptive and analytical research on employment and organizations. But next to the numerous possibilities of survey recruitment on LinkedIn, researchers must evaluate potential challenges regarding data protection, representativeness, and further potential bias.</p>
<p>To date, researchers use LinkedIn mostly either as an object of study – for example, to examine recruitment practices in labor markets, employer branding, or self-presentation strategies (e.g., <a href="https://doi.org/10.5958/2321-5763.2020.00010.4">Hosain &amp; Liu 2021</a>; <a href="https://doi.org/10.1177/0894439310386567">Caers &amp; Castelyns 2010</a>) – or as a data source for analyzing profile characteristics and network structures (e.g., <a href="https://doi.org/10.1016/j.chb.2018.08.033">Banerji &amp; Reimer 2019</a>; <a href="https://doi.org/10.1108/ER-07-2013-0086">Zide et al. 2014</a>). Some more recent work has begun to use LinkedIn to complement survey data, for instance by linking respondents’ answers to information from their LinkedIn profiles or from organizational pages (e.g., <a href="https://doi.org/10.1093/jssam/smae029">Al Baghal et al. 2024</a>). These studies illustrate that LinkedIn can provide rich contextual information on careers and organizations, but they typically focus on the use of existing profile data rather than on recruitment processes.</p>
<p>Researchers have also started experimenting with LinkedIn as a direct recruitment tool for survey-based studies. Documented approaches include promoting surveys through public (e.g., <a href="https://doi.org/10.1177/0193945917740706">Stokes et al. 2017</a>; <a href="https://doi.org/10.1093/swr/svaf002">Keemink et al. 2025</a>) or direct messages and personalized outreach (<a href="https://doi.org/10.1108/QRJ-04-2024-0085">Griffiths et al. 2025</a>; <a href="https://doi.org/10.1016/j.mex.2021.101393">Kozłowski et al. 2021</a>). However, methodological guidance on how to employ LinkedIn in a systematic and reproducible way remains limited. In particular, paid LinkedIn advertisements are still rarely documented in the academic literature. Existing accounts tend to be reflective case studies rather than detailed tutorials, and they often provide only high-level descriptions of campaign configurations and outcomes (e.g., <a href="https://doi.org/10.1186/s12889-023-16852-9">Kohl et al., 2023</a>).</p>
</div>
<div id="second-section" class="section level3">
<h3>Recruiting via LinkedIn – a step-by-step guide</h3>
<p>This section outlines a practical workflow for designing and implementing LinkedIn advertising campaigns for survey recruitment.</p>
<div id="second-first-subsection" class="section level5">
<h5>LinkedIn campaign manager</h5>
<p>Before launching any advertisements, several prerequisites must be met. Researchers require:</p>
<ol style="list-style-type: decimal">
<li>a personal LinkedIn account,</li>
<li>a LinkedIn ad account, and</li>
<li>a LinkedIn page.</li>
</ol>
<p>Each component serves a distinct purpose in the LinkedIn advertising ecosystem. The creation of all three components is free of charge.</p>
<div id="second-first-first-subsubsection" class="section level6">
<h6>Personal LinkedIn account</h6>
<p>Advertising on LinkedIn requires an active personal account associated with a real individual (i.e., a verifiable professional identity). During registration, the user must provide standard information such as name, location, job title, and current employer. For authenticity and compliance with LinkedIn’s advertising policies, the account should not be newly created: it must be at least 24 hours old, have at least one verified connection, and represent a verifiable professional identity. These criteria ensure that advertisements are traceable to legitimate users rather than anonymous entities or automated accounts. Importantly, these personal accounts are used for access and accountability within LinkedIn’s advertising system, but they are not displayed as the public-facing identity of the ads. Instead, the advertisements are shown as being placed by the LinkedIn Page (described below), meaning it is not externally visible which individual account initiated or manages a given campaign.</p>
</div>
<div id="second-first-second-subsubsection" class="section level6">
<h6>LinkedIn ad account</h6>
<p>Once a personal profile exists, an advertising account can be established via the <a href="https://business.linkedin.com/advertise">LinkedIn Advertise platform</a>. Selecting “Create ad” or “Get Started” initiates the process (see Figure <a href="#fig1">1</a>).</p>
<a name="fig1"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig1"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInMarketingSolution.png" alt="Starting page of LinkedIn Advertise website." width="70%" />
<p class="caption">
Figure 1: Starting page of LinkedIn Advertise website.
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://business.linkedin.com/advertise">LinkedIn Advertise</a>
</sup></sub>
</p>
</div>
<p>The platform then prompts the user to assign an account name and link the account to a LinkedIn Page (see Figure <a href="#fig2">2</a>). If no appropriate institutional page is available, a new one must be created.</p>
<a name="fig2"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig2"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInAdsAccount.png" alt="LinkedIn ads account set up." width="70%" />
<p class="caption">
Figure 2: LinkedIn ads account set up.
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager/new-advertiser">LinkedIn Campaign Manager</a>
</sup></sub>
</p>
</div>
<p>After this step, the ad account provides access to campaign creation, audience definition, and budgeting functions.</p>
</div>
<div id="second-first-third-subsubsection" class="section level6">
<h6>LinkedIn page</h6>
<p>A LinkedIn page functions as the institutional or organizational identity under which advertisements are displayed. Depending on the project’s context, the page may represent:</p>
<ul>
<li>a company or organization (e.g., a research institute),</li>
<li>a showcase page (a subpage linked to a main institutional account), or</li>
<li>an educational institution (e.g., an University).</li>
</ul>
<p>Unlike platforms such as Facebook (Meta), LinkedIn does not currently offer a dedicated project or educational page type, meaning researchers must register under one of the existing categories. Each page requires basic information such as a name, a public URL, the industry, and the size of the organization (see Figure <a href="#fig3">3</a>). In addition, it is possible to indicate further information (i.e., logo and a short organizational slogan). For research recruitment, using an institutional page (for example, a university department or research group) can increase credibility and transparency in the eyes of potential participants.</p>
<a name="fig3"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig3"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInPage.png" alt="Set up LinkedIn page – example: educational institute." width="60%" />
<p class="caption">
Figure 3: Set up LinkedIn page – example: educational institute.
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager/new-advertiser">LinkedIn Campaign Manager</a>
</sup></sub>
</p>
</div>
</div>
</div>
<div id="second-second-subsection" class="section level5">
<h5>Setting up an ad campaign</h5>
<p>The LinkedIn advertising system follows a three-tier hierarchical structure:</p>
<a name="fig4"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig4"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInAdvertisingSystem.png" alt="LinkedIn advertising system" width="80%" />
<p class="caption">
Figure 4: LinkedIn advertising system
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: Own illustration.
</sup></sub>
</p>
</div>
<p>Each tier governs a different aspect of the set-up, from defining overall objectives to specifying detailed audience characteristics and ad creatives. Understanding this structure is essential for organizing an effective and transparent recruitment effort.</p>
<div id="second-second-first-subsubsection" class="section level6">
<h6>Campaign level</h6>
<p>At the highest level, the campaign establishes the overarching goal and organizational framework of the recruitment project. LinkedIn offers three types of objectives (see Figure <a href="#fig5">5</a>):</p>
<ul>
<li>Awareness (increasing visibility),</li>
<li>Consideration (encouraging engagement), and</li>
<li>Conversion (prompting specific actions).</li>
</ul>
<p>For survey recruitment, the “Website visits” option under the Consideration category is generally most suitable, as it directs users to an external landing page, typically the survey entry point.</p>
<a name="fig5"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig5"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInObjective.png" alt="Choosing a campaign objective" width="60%" />
<p class="caption">
Figure 5: Choosing a campaign objective
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager">LinkedIn Campaign Manager</a>
</sup></sub>
</p>
</div>
<p>Once the campaign objective has been chosen, LinkedIn prompts the user to select a campaign type. Two options are available: a manual setup (“Classic”) and an AI-assisted setup (“Accelerate”) (see Figure <a href="#fig6">6</a>). While the latter automatically adjusts parameters to optimize engagement, we strongly recommend avoiding it. Algorithmic optimization introduces unknown and uncontrolled sources of bias into audience selection, thereby comprising the transparency and replicability required in scientific research. Manual configuration ensures that all targeting and budgeting decisions remain under the advertiser’s, that is, the researcher’s control and can be documented in detail.</p>
<a name="fig6"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig6"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInCampaignType.png" alt="Selecting a campaign type" width="70%" />
<p class="caption">
Figure 6: Selecting a campaign type
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager">LinkedIn Campaign Manager</a>
</sup></sub>
</p>
</div>
<p>After selecting the campaign type, researchers must determine the ad creation process, that is, how individual advertisement will be assembled and optimized within the campaign. Again, LinkedIn offers two routes: Flexible ad creation and standard ad creation (see Figure <a href="#fig7">7</a>).
In the flexible ad creation mode, the platform automatically generates and tests different combinations of ad elements (e.g., images, headlines, and descriptions) and allocates the budget to those combinations that achieve the best predicted performance. This process resembles A/B testing but is fully automated. Although potentially efficient, this approach lacks transparency. The platform does not disclose the criteria used to define “best performance” nor how these predictions are generated. Consequently, it introduces algorithmic bias and variation that are not under research control.
In research contexts, the standard ad creation mode is clearly preferable. This mode requires researchers to manually configure each ad creatively. All text and visual elements, as well as audience and delivery settings, are explicitly defined by the researchers. This manual process ensures that recruitment materials remain consistent and that all campaign parameters are fully reproducible and can be reported in a methods section.</p>
<a name="fig7"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig7"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInHowCreate.png" alt="Selecting an ad creation process" width="70%" />
<p class="caption">
Figure 7: Selecting an ad creation process
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager">LinkedIn Campaign Manager</a>
</sup></sub>
</p>
</div>
<p>After determining the creation route, the campaign configuration also includes the definition of schedule and budget parameters (see Figure <a href="#fig8">8</a>). Researchers may specify:</p>
<ul>
<li>a start and end date for the campaign,</li>
<li>a total (lifetime) budget, and</li>
<li>whether budget optimization across campaigns should be enabled.</li>
</ul>
<p>Again, we recommend disabling all forms of automated budget optimization. Although these features can improve marketing efficiency by reallocating funds to higher-performing ads, they rely on algorithmic assessments that are not accessible to researchers and could bias the recruited sample in unknown and unobservable ways.</p>
<a name="fig8"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig8"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInGroupBudget.png" alt="Campaign details" width="70%" />
<p class="caption">
Figure 8: Campaign details
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager">LinkedIn Campaign Manager</a>
</sup></sub>
</p>
</div>
</div>
<div id="second-second-second-subsubsection" class="section level6">
<h6>Ad set level</h6>
<p>At ad set level, the researcher defines the target population, that is, the group of LinkedIn users who will be exposed to the survey advertisement. LinkedIn provides numerous targeting attributes derived from the members’ professional profiles, company pages, and inferred behavioral or interest signals (see Figure <a href="#fig9">9</a>). Table <a href="#table1">1</a> provides a comprehensive overview of the main attribute categories and typical options within each. Where known, information on how LinkedIn obtains or infers that information is indicated.</p>
<center>
<a name="table1"></a>
<caption>
Table 1: LinkedIn Attributes and Data Sources
</caption>
<iframe src="/article/linked-in/Table1.htm" width="100%" height="840px" data-external="1">
</iframe>
<div style="text-align: centre">
<p>
<sub><sup>
Note: Data availability and functionality may vary by region due to privacy regulations, particularly with the Europeam Economic Area and Switzerland.
</sup></sub>
</p>
</div>
</center>
<a name="fig9"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig9"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInAudience.png" alt="Definition of target audience" width="70%" />
<p class="caption">
Figure 9: Definition of target audience
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager">LinkedIn Campaign Manager</a>
</sup></sub>
</p>
</div>
<p>As seen in Table <a href="#table1">1</a>, researchers have various options for target parameters. Therefore, is important that targeting parameters must never be used to discriminate against individuals based on protected characteristics such as gender, age, or actual or perceived race or ethnicity. Researchers should also comply with platform-specific restrictions and regional regulations that limit the use of sensitive targeting criteria.
The option “enable audience expansion” should be deactivated (see Figure <a href="#fig9">9</a>). This algorithmic extension may deliver ads to users outside the intended target population and thus undermine the transparency and controllability of the sample frame.
LinkedIn’s Campaign Manager also displays an estimated performance forecast, showing projected target audience size, impressions, clicks, and costs (see Figure <a href="#fig10">10</a>). These values can assist in budget planning but are based on proprietary models and should be interpreted as approximate indicators only, not as precise predictions.</p>
<a name="fig10"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig10"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInForecast.png" alt="Forecasted results" width="25%" />
<p class="caption">
Figure 10: Forecasted results
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager">LinkedIn Campaign Manager</a>
</sup></sub>
</p>
</div>
<p>Once the campaign objectives and audience have been defined, the next step is the selection of an appropriate advertisement format (see Figure <a href="#fig11">11</a>). LinkedIn currently offers several formats that differ in visual layout, degree of interactivity, and expected engagement levels. The choice of format should align with the overall research design, recruitment objectives, and desired intensity of participation involvement. Available formats include, among other:</p>
<ul>
<li>Single image ads: static image with headline and text,</li>
<li>Carousel ads: multiple scrollable image cards that allow the presentation of different aspects or topics within one advertisement, and</li>
<li>Video ads: short clip (15-30 seconds) combining visuals and text to communicate information dynamically.</li>
</ul>
<p>For most survey recruitment, these formats are generally the most suitable due to their high visibility and clear message presentation.</p>
<a name="fig11"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig11"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInAdFormat.png" alt="Ad format" width="65%" />
<p class="caption">
Figure 11: Ad format
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager">LinkedIn Campaign Manager</a>
</sup></sub>
</p>
</div>
<p>The integration of URL tracking parameters (e.g., UTM codes) into the destination link is recommended to enable transparent performance monitoring (see Figure <a href="#fig12">12</a>). A typical example would be:</p>
<pre class="text"><code>utm_source=[SOCIAL MEDIA PLATFORM]&amp;utm_medium=[PAYMENT EVENT]
&amp;utm_campaign=[CAMPAIGN NAME]&amp;utm_content=[AD ID].</code></pre>
<p>This method complies with data protection requirements and does not collect personal information. It simply allows researchers to distinguish traffic from different platforms or creatives in their web analytics.</p>
<a name="fig12"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig12"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInURL.png" alt="URL tracking parameters" width="65%" />
<p class="caption">
Figure 12: URL tracking parameters
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager">LinkedIn Campaign Manager</a>
</sup></sub>
</p>
</div>
<p>Ad placement is then specified (see Figure <a href="#fig13">13</a>). The LinkedIn feed should be selected as the primary placement to ensure visibility in a professional context. The LinkedIn Audience Network, which extends ad delivery to partner websites and apps, should be deactivated in research campaigns to maintain control and transparency regarding where exactly recruitment occurs.</p>
<a name="fig13"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig13"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInPlacements.png" alt="Placements option" width="65%" />
<p class="caption">
Figure 13: Placements option
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager">LinkedIn Campaign Manager</a>
</sup></sub>
</p>
</div>
<p>The budgeting model determines how funds are distributed over time (Figure <a href="#fig14">14</a>). LinkedIn provides three options:</p>
<ul>
<li>Daily budget: specifies maximum daily expenditure, ensuring consistent ad delivery,</li>
<li>Lifetime budget: sets a total expenditure limit for the campaign period, and</li>
<li>Combined budget: applies both daily and lifetime limits for maximum control.</li>
</ul>
<p>Short-term campaigns benefit from daily budgets to maintain stable exposure, while longer recruitment periods may require lifetime or combined budget settings. Key performance metrics such as impressions, click-through rate (CTR), and cost per click (CPC) should be monitored throughout the campaign and reported alongside sample characteristics.</p>
<a name="fig14"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig14"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInBudgetSchedule.png" alt="Budget and schedule options" width="60%" />
<p class="caption">
Figure 14: Budget and schedule options
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager">LinkedIn Campaign Manager</a>
</sup></sub>
</p>
</div>
<p>The bidding strategy determines how LinkedIn allocates ads in its auctions system (Figure <a href="#fig15">15</a>). LinkedIn currently offers three bid types:</p>
<ul>
<li>Maximum delivery: an automated bidding option where LinkedIn’s system sets the bid with the goal of spending the full budget and maximizing the selected key result,</li>
<li>Cost cap: an automated option where LinkedIn sets bids while trying to keep the average cost per key result close to a user-specified target (e.g., CPC or CPM), and</li>
<li>Manual bidding: a more hands-on option where advertisers set an explicit bid value that is used in the ad auction.</li>
</ul>
<p>From a research perspective, there is a trade-off between cost-efficiency and methodological control. Fully automated options such as maximum delivery can be convenient and may help use the budget efficiently, but they also shift additional decision-making to opaque algorithms. Where fine-grained cost control and transparency are priorities, cost cap or even manual bidding may be preferable, provided that campaigns are monitored closely, and bid levels are adjusted if necessary. Regardless of the chosen strategy, researchers should document the bid type and any changes during fieldwork so that recruitment conditions can be accurately reported.</p>
<a name="fig15"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig15"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInBidding.png" alt="Bidding options" width="70%" />
<p class="caption">
Figure 15: Bidding options
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager">LinkedIn Campaign Manager</a>
</sup></sub>
</p>
</div>
<p>Although LinkedIn supports conversion tracking via tracking pixels (see Figure <a href="#fig16">16</a>), this method is ethically and legally problematic in research contexts. Tracking pixels typically collect behavioral data across websites (e.g., page views, time spent, subsequent actions) and can be used to link on-platform interactions with off-platform behavior. Under data protection frameworks such as the General Data Protection Regulation (GDPR), such tracking requires a valid legal basis and, in most cases, informed consent: participants must be clearly informed before any tracking takes place about what data are collected, for what purpose, on which legal basis, how long they are stored, and with whom they may be shared. Only if users have actively agreed (e.g., via a consent banner or consent form on the landing page) may tracking pixels be activated.
However, for scientific campaigns, this form of tracking is usually difficult to justify. Social media platforms and advertising networks have repeatedly been affected by data breaches and other security incidents (e.g., <a href="https://www.theguardian.com/uk-news/2023/jul/15/revealed-metropolitan-police-shared-sensitive-data-about-victims-with-facebook">Das 2023</a>; <a href="https://themarkup.org/pixel-hunt/2022/06/16/facebook-is-receiving-sensitive-medical-information-from-hospital-websites">Feathers et al. 2022</a>; <a href="https://themarkup.org/pixel-hunt/2022/11/22/tax-filing-websites-have-been-sending-users-financial-information-to-facebook">Fondrie-Teitler et al. 2022</a>). When third-party tracking is enabled, additional actors gain access to behavioral data and identifiers, increasing the potential attack surface. Even if the probability of a breach is difficult to quantify, the possible consequences for participants – for example, the unintended disclosure of sensitive behavioral patterns or re-identification risks when log data are leaked – can be substantial. From a risk-minimization perspective, avoiding non-essential tracking is therefore a central protective measure.
We therefore recommend that researchers do not use LinkedIn’s conversion tracking or similar third-party tracking pixels for recruitment. Instead, evaluation should rely on aggregate campaign statistics (impressions, clicks, CPC) provided by the platform, combined with survey-based indicators (e.g., completion rates and a self-reported recruitment channel question on the first page of the questionnaire, asking participants how they became aware of the study).</p>
<a name="fig16"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig16"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInConversion.png" alt="Conversion tracking" width="80%" />
<p class="caption">
Figure 16: Conversion tracking
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager">LinkedIn Campaign Manager</a>
</sup></sub>
</p>
</div>
</div>
<div id="second-second-third-subsubsection" class="section level6">
<h6>Ad level</h6>
<p>At the ad level, the creative elements that LinkedIn users are exposed to are defined.
The design of these elements can draw on established findings from research on invitation letters in survey methodology. In line with this literature, ad content should be convincing and informative, and should convey a clear perceived leverage for participation, for example by highlighting the social value of the study, the relevance of the topic for the target population, or the opportunity to contribute one’s perspective (e.g., Dillman et al. 2014 ; <a href="http://www.jstor.org/stable/3078721">Groves et al. 2000</a>). On LinkedIn, as on other social media platforms, an ad competes with a large volume of other digital content. Consequently, the ad must first attract attention within a very short time frame and then provide sufficient relevance and clarity to motivate users to click on the ad and proceed to the survey landing page.
The introductory text (Figure <a href="#fig17">17</a>) – typically the most prominent text element in the ad – should succinctly communicate the purpose of the study and its relevance for the targeted audience. A concise structure is advisable: one sentence specifying what the study is about and who it addresses, followed by a sentence that conveys why participation is meaningful (e.g., expected benefits, contribution to knowledge, or policy relevance).</p>
<a name="fig17"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig17"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInAdDesign.png" alt="Ad design  - Text and media input" width="80%" />
<p class="caption">
Figure 17: Ad design - Text and media input
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager">LinkedIn Campaign Manager</a>
</sup></sub>
</p>
</div>
<p>The visual content serves as the primary attention grabber and can function as an additional targeting device (Figure <a href="#fig17">17</a>). On LinkedIn, visuals are often embedded in a professional and organizational context, for example by depicting typical work situations (e.g., team meetings, hybrid collaboration, industry-specific workplaces), abstract graphics related to the topic (e.g., charts, icons for digitalization or skills), or neutral employer-branding style imagery. When these visuals are clearly aligned with a specific survey topic – for instance, leadership and management practices, remote work arrangements, skills development, or employee well-being – they can increase click-through rates and the number of completed questionnaires, because they immediately signal topical relevance to certain occupational groups or sectors. At the same time, these visuals may amplify self-selection on work-related attitudes or experiences (e.g., job satisfaction, perceived stress, views on organizational change), as they are particularly salient for individuals with above-average involvement or strong opinions (<a href="https://doi.org/10.12758/mda.2025.05">Donzowa et al. 2025</a>; <a href="https://doi.org/10.1080/13645579.2025.2597303">Zindel et al., 2025</a>). This implies a trade-off: if recruitment within a narrowly defined topic area is prioritized, topic-specific professional visuals are advantageous; if concerns about attitudinal bias are central, more neutral or general workplace imagery may be preferable.
In addition to the visual itself, LinkedIn offers the option to set an image alt text (Figure <a href="#fig18">18</a>). This alt text enhances accessibility for users who rely on screen readers and is displayed if the image cannot be rendered. From a survey recruitment perspective, the alt text should provide a concise, neutral description of what is shown in the image and, where appropriate, can briefly restate the topic of the study (e.g., “Illustration of professionals discussing climate policies in an office setting”). Overly promotional language or keyword stuffing should be avoided. The goal is to ensure that the meaning of the visual remains understandable even in the absence of the actual image and that the alt text is consistent with the overall framing of the study (cf. [World Wide Web Consortium [W3C], 2025[(<a href="https://www.w3.org/TR/WCAG21/" class="uri">https://www.w3.org/TR/WCAG21/</a>)]).</p>
<a name="fig18"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig18"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInAdDesignOption.png" alt="Ad design - Further options" width="85%" />
<p class="caption">
Figure 18: Ad design - Further options
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager">LinkedIn Campaign Manager</a>
</sup></sub>
</p>
</div>
<p>The destination URL is the link to the survey website and leads directly to the survey entry page, without intermediary pages or redirects. This entry page must present the informed consent statement and essential information on data protection, study sponsors, and contact details before data collection begins. From a user-experience perspective, the number of steps between the ad (i.e., the survey invitation) and the first survey question should be minimized, as each additional step potentially increases drop-off probabilities (cf. <a href="https://doi.org/10.1093/acprof:oso/9780199747047.001.0001">Tourangeau et al. 2013</a>).
Besides the introductory text and the image, LinkedIn ads typically include a headline and an optional ad description. The headline is usually displayed in a highly prominent position and should therefore be short, specific, and benefit oriented. For survey recruitment, headlines that clearly state the topic and target group (e.g., “Short survey on remote work for HR professionals”) are generally preferable to vague or purely promotional formulations.
The ad description provides additional information to people who see the ad. However, this field is not visible in most scenarios and will only appear for a small portion of LinkedIn members, for example in certain placements on the LinkedIn Audience Network (if enabled). Accordingly, essential information about the study should not be placed exclusively in the description but should already be contained in the introductory text and headline. The description can be used as a secondary field to elaborate on the headline, clarify what happens after the click, indicate approximate survey length, or highlight a specific leverage (e.g., “10-minute academic survey, anonymous, results will inform future workplace policies”). Given that descriptions may be truncated or omitted depending on device and placement, key information should appear at the beginning of the text and should not be critical for understanding the study.
The call-to-action (CTA) button should be consistent with the overall framing of the ad and align with user expectations on a professional platform. LinkedIn offers several standard CTA options, including “Learn more”, “Sign up”, “Register”, and others. For survey recruitment, a CTA such as “Learn more” is typically most appropriate, as it signals low-threshold, information-oriented action compatible with participation in a research study. In settings where the survey is embedded in a broader program or panel, “Sign up” or “Register” may also be suitable.</p>
<a name="fig19"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig19"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInAdPreview.png" alt="Ad preview" width="50%" />
<p class="caption">
Figure 19: Ad preview
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager">LinkedIn Campaign Manager.</a> Image by Africa Studio via <a href=“https.//stock.adobe.com/115404199“>stock.adobe.com.</a>
</sup></sub>
</p>
</div>
<p>Prior to launch, the preview function in Campaign Manager should be used to verify appearance across devices and placements (see Figure <a href="#fig19">19</a>). Across all ad components, internal consistency is crucial. The introductory text, headline, description, visual elements (including alt text), destination URL, and CTA should convey a coherent message and set realistic expectations about the survey experience. A consistent, methodologically informed ad design that is at the same time adapted to platform-specific norms on LinkedIn potentially increases the likelihood of attracting suitable respondents rather than merely generating clicks.</p>
</div>
</div>
</div>
<div id="third-section" class="section level3">
<h3>Ethical considerations</h3>
<p>When recruiting survey participants via social media, particular attention should be paid to ethical and data protection issues. Most important, all procedures must comply with applicable data protection frameworks (e.g., GDPR), including a clear specification of the legal basis for processing, information on storage periods, and contact details of the responsible data controller and, where applicable, the data protection officer. Moreover, recruitment messages and landing pages should be transparent about the study’s purpose, the institutions involved, and the intended use of the data. Participants should be informed about their data protection rights (e.g., rights of access, rectification, erasure, restriction of processing, and complaint) in a concise and comprehensible way.
Where surveys are anonymous, this should be explicitly communicated and technically ensured, for instance by avoiding the collection of directly identifying information and by separating contact data from survey responses if recontact is planned. Where anonymity cannot be fully guaranteed – for example because organization-level or very narrow professional targeting makes the population of potential invitees very small – this limitation should be stated clearly. In such cases, it should be emphasized that confidentiality will nonetheless be preserved, for example through secure data handling and the exclusive reporting of results in aggregated form.
Given that recruitment on LinkedIn and similar platforms can take place at very specific levels (e.g., by employer, job title, industry niche, or narrowly defined professional segments), the risk of de facto identifiability of respondents may be higher than in general-population web surveys. This applies in particular when small organizations or rare occupations are involved. Ethical scrutiny is therefore essential. Targeting strategies should avoid extremely small and potentially identifiable groups where this is not strictly necessary for the research aim, and analysis and reporting should follow minimum cell-size rules to prevent the indirect identification of individuals or specific organizations.
It is strongly recommended to seek approval or at least formal advice from the relevant ethics committee at the home institution or from an external ethics body before fieldwork begins. An ethics review helps to ensure that the sampling and recruitment strategy is appropriate, that privacy and confidentiality risks are minimized, that the informed consent procedure is adequate, and that the study aligns with existing ethical guidelines and professional codes of conduct.</p>
</div>
<div id="fourth-section" class="section level3">
<h3>Application example</h3>
<p>This section illustrates how the LinkedIn recruitment workflow described above can be applied in practice. It summarizes the design, implementation, and outcomes of an exemplary advertising campaign conducted as part of a master’s seminar at Bielefeld University in summer term 2024 titled “Sociological methods - quantitative: Survey-based measurement of inclusion and diversity in companies: Development and implementation of employee surveys”, taught by Zaza Zindel and Lisa de Vries.<br />
Within this seminar, students developed a questionnaire to measure inclusion and diversity in workplace settings. The instrument comprised 37 items, including the following topics:</p>
<ul>
<li>Working life (7 items)</li>
<li>Career orientation (3 items)</li>
<li>Diversity (3 items)</li>
<li>Discrimination (4 items)</li>
<li>Health (6 items)</li>
<li>Sociodemographic (13 items)</li>
<li>Closing question (1 item).</li>
</ul>
<p>The target population consisted of employees in Germany aged 18 years or older. The survey was programmed using the software LimeSurvey. To reach the target population without cooperation from specific companies, survey recruitment was carried out via LinkedIn ads. A total budget of 500 Euro was allocated for the campaign.</p>
<div id="fourth-first-subsection" class="section level4">
<h4>Ad campaign details</h4>
<p>Following the three-tier campaign structure described above, the campaign was configured at the campaign, ad set, and ad levels. A dedicated LinkedIn page titled “Empirische Sozialforschung – Uni Bielefeld” (English translation: “Empirical Social Research – Bielefeld University” was created to host the ads and serve as the institutional identity for the recruitment process (see Figure <a href="#fig20">20</a>).</p>
<a name="fig20"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig20"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInEmpiricalPage.png" alt="LinkedIn page: Empirical social research – Bielefeld University" width="80%" />
<p class="caption">
Figure 20: LinkedIn page: Empirical social research – Bielefeld University
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager">LinkedIn Campaign Manager.</a> Images by <a href=“https.//stock.adobe.com/115404199“>Monster Ztudio</a> and <a href=“https.//stock.adobe.com/390374250“>Siberian Art</a> via stock.adobe.com.
</sup></sub>
</p>
</div>
<div id="fourth-first-first-subsubsection" class="section level6">
<h6>Ad creatives</h6>
<p>Two separate advertisements in the single image format were developed. They used identical textual content but different ad images to examine how visuals influence engagement (Figures <a href="#fig21">21</a> and <a href="#fig22">22</a>).</p>
<p><strong>Headline:</strong> We want your opinion: Take part in our survey on the German labor market!
<strong>Introductory text:</strong> Help us better understand the German labor market! We are looking for participants for a short survey that is part of a scientific study conducted by Bielefeld University. The survey is aimed at all employed people aged 18 and over in Germany.</p>
<a name="fig21"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig21"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInAd1.png" alt="Targeting ad 1 used to recruit participants on LinkedIn" width="60%" />
<p class="caption">
Figure 21: Targeting ad 1 used to recruit participants on LinkedIn
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager">LinkedIn Campaign Manager.</a> Image by mitay20 via <a href=“https.//stock.adobe.com/307640218“>stock.adobe.com</a>.
</sup></sub>
</p>
</div>
<a name="fig22"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig22"></span>
<img src="../../../../../../../../article/linked-in/images/LinkedInAd2.png" alt="Targeting ad 2 used to recruit participants on LinkedIn" width="60%" />
<p class="caption">
Figure 22: Targeting ad 2 used to recruit participants on LinkedIn
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.linkedin.com/campaignmanager">LinkedIn Campaign Manager.</a> Image by Africa Studio via <a href=“https.//stock.adobe.com/115404199“>stock.adobe.com.</a>
</sup></sub>
</p>
</div>
</div>
<div id="fourth-first-second-subsubsection" class="section level6">
<h6>Campaign setup</h6>
<p>Both advertisements employed identical targeting criteria to ensure comparability across creatives. To recruit broadly from employed individuals in Germany, targeting was restricted to:</p>
<ul>
<li>Location: Germany</li>
<li>Profile Language: German</li>
<li>Age: 18+ years</li>
</ul>
<p>No additional targeting criteria were applied, and audience expansion was disabled. The placement was set to LinkedIn Feed only and audience network placements was disabled. A lifetime budget of 500 EUR was defined, optimization goal was landing page clicks, and the bidding strategy was maximum delivery until all the budget is spent.</p>
</div>
</div>
<div id="fourth-second-subsection" class="section level4">
<h4>Campaign results</h4>
<p>The LinkedIn recruitment campaign ran for five consecutive days, from 6 to 10 July 2024, and was managed through LinkedIn Campaign Manager using the institutional page created for this project.
Across both advertisements, the campaign reached a substantial professional audience and generated a sizeable number of completed interviews within a modest budget. Table <a href="#table2">2</a> summarizes the performance metrics for both advertisements. In total, the two ads produced 56,990 impressions, resulting in 301 link clicks. Of those who clicked on the advertisement, 143 started the interview and 98 completed the questionnaire. With a total spend of 500 EUR, this corresponds to an overall cost per click (CPC) of 1.66 EUR and a cost per completed interview (CI) of 5.10 EUR. These results suggest that LinkedIn Ads can serve as an efficient recruitment channel for workplace-related surveys, producing a relatively high number of completed questionnaires at moderate cost.</p>
<center>
<a name="table2"></a>
<caption>
Table 2: Results of ad and survey performance
</caption>
<br>
<iframe src="/article/linked-in/Table2.htm" width="70%" height="270px" data-external="1"></iframe>
<div style="text-align: centre">
<p>
<sub><sup>
Note: CTR = Click-through-Rate, CPC = Cost per click, CI = completed interview.
</sup></sub>
</p>
</div>
</center>
<p>Both ads achieved a click-through rate (CTR) of roughly 0.5% (0.54% for Ad 1 and 0.52% for Ad 2). The overall CPC was 1.66 EUR, with a slightly lower CPC for Ad 1 (1.58 EUR) than for Ad 2 (1.73 EUR). At the same time, Ad 2 produced more completed interviews (57 compared to 41) and therefore a slightly lower cost per completed interview (5.03 EUR for Ad 2 versus 5.20 EUR for Ad 1). These differences are numerically small, but they show that the two creatives did not behave identically in terms of cost structure, even though they were identical in text and targeting.
The delivery patterns of the two ads are also differed. Ad 1 was shown to a larger number of unique users (reach 7,269) and with a lower average frequency (3.45 displays per user). Ad 2, by contrast, reached fewer unique users (5,189) but was displayed more often to those who were reached (average frequency 6.15). Thus, the total number of impressions was higher for Ad 2, even though its reach was lower. Within the constraints of the shared targeting, this indicates that LinkedIn’s delivery system concentrated Ad 2 more strongly on a subset of the potential audience, while Ad 1 was distributed more broadly.</p>
</div>
<div id="fourth-third-subsection" class="section level4">
<h4>Sample composition</h4>
<p>The composition of the achieved sample reflects the characteristics of those LinkedIn users who clicked on the advertisements, proceeded to the survey, and completed the questionnaire. Table <a href="#table3">3</a> shows the distribution of the sociodemographic variables. Respondents are distributed across all age groups included in the questionnaire, with the largest shares in the ranges 45–54 years (31.6%) and 35–44 years (29.6%) (age mean: 45.84; Range: 23-69). Regarding gender, 62 respondents identify as female and 32 as male, while four respondents did not answer this question. No respondent selected the response option for another gender identity. Education attainment in the sample is strongly concentrated at the higher end of the scale: Most respondents have a higher secondary school diploma (85.7 %), further 10.2 percent hold a secondary school diploma.
If we compare the composition of sociodemographic variables with the whole population in Germany, we see that the sample contain relatively few young people (<a href="https://ergebnisse.zensus2022.de/datenbank/online/statistic/1000A/table/1000A-1017">Statistisches Bundesamt 2025a</a>). Even if 21.4 percent of the sample are 55 years and older, only 10.4 percent are 60 years and older, and no respondents are older than 70. In 2023, 29,8 percent of the population were 60 years and older based on Zensus (<a href="https://ergebnisse.zensus2022.de/datenbank/online/statistic/1000A/table/1000A-1017">Statistisches Bundesamt 2025a</a>). These differences may be due to the fact that LinkedIn is used particularly by people of working age who are neither in education nor retired. Moreover, the percentage of women and the educational attainment in the sample is higher than in the whole population in Germany (<a href="https://ergebnisse.zensus2022.de/datenbank/online/statistic/1000A/table/1000A-1017">Statistische Ämter des Bundes und der Länder 2025a</a>; <a href="https://www.destatis.de/DE/Themen/Arbeit/Arbeitsmarkt/Erwerbstaetigkeit/Tabellen/eckwerttabelle.html">Statistische Ämter des Bundes und der Länder 2025b</a>).</p>
<center>
<a name="table3"></a>
<caption>
Table 3: Sociodemographic variables
</caption>
<br>
<iframe src="/article/linked-in/Table3.htm" width="90%" height="430px" data-external="1"></iframe>
<div style="text-align: centre">
<p>
<sub><sup>
Source: Dataset “DEI Companies” recruited on LinkedIn; own calculations.
</sup></sub>
</p>
</div>
</center>
<p>Table <a href="#table4">4</a> presents the occupational variables. Most respondents report full-time employment (75.5%), while 19.4 percent work part-time. Only four respondents indicated marginal or no employment. In 2024 in Germany, the unemployment rate was 3.1 percent, and the part-time rate was 29,1 percent (<a href="https://www.destatis.de/DE/Themen/Arbeit/Arbeitsmarkt/Erwerbstaetigkeit/Tabellen/eckwerttabelle.html">Statistisches Bundesamt 2025b</a>; <a href="https://www.destatis.de/DE/Themen/Querschnitt/Gleichstellungsindikatoren/teilzeitquote-f25.html">Statistisches Bundesamt 2025c</a>). Firm size is also skewed toward larger organizations: 58.2 percent of the respondents work in companies with more than 200 employees, and 24.5 percent work in firms with between 20 and 200 employees. Only 15.3 percent of respondents report working in small firms with fewer than 20 employees. Based on Destatis, in 2023, nearly the half of the German population work in large companies (<a href="https://www.destatis.de/DE/Themen/Branchen-Unternehmen/Unternehmen/Kleine-Unternehmen-Mittlere-Unternehmen/aktuell-beschaeftigte.html">Statistisches Bundesamt 2025d</a>). With regard to leadership roles, 24.5 percent indicate that they hold a leadership position, while 59.2 percent do not. Notably, 16 respondents did not answer this question, resulting in a higher share of missing values than the other occupational variables. The distribution across industry sectors covers a broad range of fields. The largest groups of respondents work in manufacturing (15.3%), information and communication (15.3%), and human health and social work activities (15.3%). Other sectors represented in the sample include scientific and technical activities, education, financial and insurance activities, public administration, and several smaller categories. Even if the share of people working in the manufacturing sector is smaller in the sample than in the general population, that a relatively high percentage work in the area of services (e.g., education or human health and social work activities) is also true in the general population (<a href="https://www.destatis.de/DE/Themen/Wirtschaft/Konjunkturindikatoren/Lange-Reihen/Arbeitsmarkt/lrerw13a.html">Statistisches Bundesamt 2025e</a>; <a href="https://www.destatis.de/DE/Themen/Arbeit/Arbeitsmarkt/Erwerbstaetigkeit/Tabellen/arbeitnehmer-wirtschaftsbereiche.html#fussnote-2-122230">Statistisches Bundesamt 2025f</a>).</p>
<center>
<a name="table4"></a>
<caption>
Table 4: Occupational variables
</caption>
<br>
<iframe src="/article/linked-in/Table4.htm" width="65%" height="625px" data-external="1"></iframe>
<div style="text-align: centre">
<p>
<sub><sup>
Source: Dataset “DEI Companies” recruited on LinkedIn; own calculations.
</sup></sub>
</p>
</div>
</center>
<p>The discussed patterns are consistent with previous evidence that LinkedIn use is concentrated among higher-income professionals and knowledge-intensive occupations (<a href="https://doi.org/10.1177/0163443712468605">Van Dijck, 2013</a>; <a href="https://doi.org/10.1177/0002764217717559">Blank &amp; Lutz, 2017</a>) and underlines that LinkedIn-based samples capture a selective segment of the working-age population.</p>
</div>
</div>
<div id="fifth-section" class="section level3">
<h3>Summary and key takeaways</h3>
<p>LinkedIn provides a distinctive infrastructure for survey recruitment in professional and occupational populations. Its career-focused user base, comparatively low noise-to-signal ratio, and rich profile data on education, industry, and employment make it particularly suitable for studies on labor markets, workplace dynamics, and professional attitudes. In contrast to general-purpose platforms such as Instagram or TikTok, LinkedIn is designed around work, networking, and expertise. As a result, it tends to reach individuals of working age, often with medium to high levels of education and strong labor market attachment, which is reflected in the sample composition of the application example presented in this tutorial. This step-by-step guide provided an overview of how to use LinkedIn ads to recruit participants for survey research.
The application example from Bielefeld University illustrates that with a moderate budget, LinkedIn can yield a substantial number of completed interviews at reasonable cost, while also highlighting typical patterns such as the overrepresentation of highly educated respondents and specific age groups.</p>
<p>Taken together, the key takeaways from this tutorial can be summarized as follows:</p>
<ol style="list-style-type: decimal">
<li>Alternative recruitment channel: LinkedIn offers a viable alternative to conventional sampling approaches and general social media platforms for reaching employed and occupationally defined populations.</li>
<li>Distinct user structure: Compared to platforms such as Instagram or TikTok, LinkedIn tends to reach individuals in mid-career and higher educational strata, which is advantageous for many labor market-related studies.</li>
<li>Fine-grained targeting: Detailed targeting by job title, industry, company size, skills, and other professional attributes enables access to niche populations that are otherwise difficult to sample.</li>
<li>Selective coverage and bias: Only certain groups are reachable via LinkedIn, and topic-specific visuals or messaging can reinforce self-selection, requiring careful interpretation of substantive findings.</li>
<li>Heightened ethical responsibilities: Because targeting can be highly specific, ethical issues around privacy, identifiability, and fair treatment are particularly salient and call for robust consent procedures, conservative reporting practices, and prior ethics review.</li>
<li>Potential stability over time: As a professional networking platform, LinkedIn may be less prone to rapid shifts in usage patterns than entertainment-oriented social networks, potentially providing a relatively stable environment for recurring or longitudinal recruitment; however, this requires continued empirical monitoring.</li>
<li>Need for further evidence: Future studies should systematically evaluate recruitment efficiency, sample quality, and bias across different designs and topics and explore, where ethically permissible, whether and how organizational or contextual information can be incorporated without compromising participant protection.</li>
</ol>
</div>
<div id="further-reading" class="section level3">
<h3>Further reading</h3>
<p>Researchers interested in using social media platforms for survey recruitment may benefit from the following literature, which covers methodological, ethical, and practical aspects of recruiting survey respondents via social media advertisements. Together, these publications provide guidance on campaign design and performance, ethical and legal considerations, data quality, and the assessment of non-probability samples.</p>
<ul>
<li>Höhne, J. K., Claassen, J., Kühne, S., &amp; Zindel, Z. (2025). Social media ads for survey recruitment: Performance, costs, user engagement. <em>International Journal of Market Research</em>. <a href="https://doi.org/10.1177/14707853251367805" class="uri">https://doi.org/10.1177/14707853251367805</a></li>
<li>Pötzschke, S., Weiß, B., Daikeler, J., Silber, H., &amp; Beuthner, C. (2023). <em>A guideline on how to recruit respondents for online surveys using Facebook and Instagram: Using hard-to-reach health workers as an example</em>. Mannheim: GESIS – Leibniz Institute for the Social Sciences (GESIS Survey Guidelines). <a href="https://doi.org/10.15465/gesis-sg_en_045" class="uri">https://doi.org/10.15465/gesis-sg_en_045</a></li>
<li>Rohr, B., Felderer, B., Silber, H., Daikeler, J., Roßmann, J., &amp; Schröder, J. (2024). <em>When are non-probability surveys fit for my purpose?</em> Mannheim: GESIS – Leibniz Institute for the Social Sciences (GESIS Survey Guidelines). <a href="https://doi.org/10.15465/gesis-sg_en_050" class="uri">https://doi.org/10.15465/gesis-sg_en_050</a></li>
<li>Zimmermann, B. M., Willem, T., Bredthauer, C. J., &amp; Buyx, A. (2022). Ethical issues in social media recruitment for clinical studies: Ethical analysis and framework. <em>Journal of Medical Internet Research, 24</em>(5), e31231. <a href="https://doi.org/10.2196/31231" class="uri">https://doi.org/10.2196/31231</a></li>
<li>Zindel, Z., Kühne, S., Perrotta, D., &amp; Zagheni, E. (2025). Ad images in social media survey recruitment: What they see is what we get. <em>International Journal of Social Research Methodology</em>, 1–20. <a href="https://doi.org/10.1080/13645579.2025.2597303" class="uri">https://doi.org/10.1080/13645579.2025.2597303</a></li>
<li>Zindel, Z. (2026). Should we worry about problematic response behaviour in social media surveys? Understanding the impact of social group cues in recruitment. <em>Survey Methods: Insights from the Field</em>. <a href="https://doi.org/10.13094/SMIF-2026-00015" class="uri">https://doi.org/10.13094/SMIF-2026-00015</a></li>
<li>Zindel, Z. (2023). Social media recruitment in online survey research: A systematic literature review. <em>Methods, Data, Analyses, 17</em>(2), 207–248. <a href="https://doi.org/10.12758/mda.2022.15" class="uri">https://doi.org/10.12758/mda.2022.15</a></li>
</ul>
<!-- Add something about the instructor -->
</div>
<div id="about-the-authors" class="section level3">
<h3>About the authors</h3>
<p>Dr. Zaza Zindel <a href="mailto:zindel@dezim-institut.de"><svg aria-hidden="true" role="img" viewBox="0 0 512 512" style="height:1em;width:1em;vertical-align:-0.125em;margin-left:auto;margin-right:auto;font-size:inherit;fill:currentColor;overflow:visible;position:relative;"><path d="M64 112c-8.8 0-16 7.2-16 16v22.1L220.5 291.7c20.7 17 50.4 17 71.1 0L464 150.1V128c0-8.8-7.2-16-16-16H64zM48 212.2V384c0 8.8 7.2 16 16 16H448c8.8 0 16-7.2 16-16V212.2L322 328.8c-38.4 31.5-93.7 31.5-132 0L48 212.2zM0 128C0 92.7 28.7 64 64 64H448c35.3 0 64 28.7 64 64V384c0 35.3-28.7 64-64 64H64c-35.3 0-64-28.7-64-64V128z"/></svg></a> <a href="https://linkedin.com/in/zaza-zindel"><svg aria-hidden="true" role="img" viewBox="0 0 448 512" style="height:1em;width:0.88em;vertical-align:-0.125em;margin-left:auto;margin-right:auto;font-size:inherit;fill:currentColor;overflow:visible;position:relative;"><path d="M416 32H31.9C14.3 32 0 46.5 0 64.3v383.4C0 465.5 14.3 480 31.9 480H416c17.6 0 32-14.5 32-32.3V64.3c0-17.8-14.4-32.3-32-32.3zM135.4 416H69V202.2h66.5V416zm-33.2-243c-21.3 0-38.5-17.3-38.5-38.5S80.9 96 102.2 96c21.2 0 38.5 17.3 38.5 38.5 0 21.3-17.2 38.5-38.5 38.5zm282.1 243h-66.4V312c0-24.8-.5-56.7-34.5-56.7-34.6 0-39.9 27-39.9 54.9V416h-66.4V202.2h63.7v29.2h.9c8.9-16.8 30.6-34.5 62.9-34.5 67.2 0 79.7 44.3 79.7 101.9V416z"/></svg></a> is a research associate at the German Centre for Integration and Migration Research (DeZIM). She is a social scientist working on survey methodology, data science, and social media research to study vulnerable, rare, and underrepresented populations and their experiences of discrimination. Her work combines methodological innovation with substantive applications using survey data, experiments, and new data sources, drawing on computational and data science approaches.</p>
<p>Dr. Lisa de Vries <a href="mailto:lisa.devries@fernuni-hagen.de"><svg aria-hidden="true" role="img" viewBox="0 0 512 512" style="height:1em;width:1em;vertical-align:-0.125em;margin-left:auto;margin-right:auto;font-size:inherit;fill:currentColor;overflow:visible;position:relative;"><path d="M64 112c-8.8 0-16 7.2-16 16v22.1L220.5 291.7c20.7 17 50.4 17 71.1 0L464 150.1V128c0-8.8-7.2-16-16-16H64zM48 212.2V384c0 8.8 7.2 16 16 16H448c8.8 0 16-7.2 16-16V212.2L322 328.8c-38.4 31.5-93.7 31.5-132 0L48 212.2zM0 128C0 92.7 28.7 64 64 64H448c35.3 0 64 28.7 64 64V384c0 35.3-28.7 64-64 64H64c-35.3 0-64-28.7-64-64V128z"/></svg></a> <a href="https://linkedin.com/in/lisa-de-vries-38465423a"><svg aria-hidden="true" role="img" viewBox="0 0 448 512" style="height:1em;width:0.88em;vertical-align:-0.125em;margin-left:auto;margin-right:auto;font-size:inherit;fill:currentColor;overflow:visible;position:relative;"><path d="M416 32H31.9C14.3 32 0 46.5 0 64.3v383.4C0 465.5 14.3 480 31.9 480H416c17.6 0 32-14.5 32-32.3V64.3c0-17.8-14.4-32.3-32-32.3zM135.4 416H69V202.2h66.5V416zm-33.2-243c-21.3 0-38.5-17.3-38.5-38.5S80.9 96 102.2 96c21.2 0 38.5 17.3 38.5 38.5 0 21.3-17.2 38.5-38.5 38.5zm282.1 243h-66.4V312c0-24.8-.5-56.7-34.5-56.7-34.6 0-39.9 27-39.9 54.9V416h-66.4V202.2h63.7v29.2h.9c8.9-16.8 30.6-34.5 62.9-34.5 67.2 0 79.7 44.3 79.7 101.9V416z"/></svg></a> is acting professor at FernUniversität in Hagen. Her research focuses on discrimination and inequality based on sexual orientation and gender from a quantitative perspective. Her research interests are sexual and gender minorities, labor market and educational inequalities, global developments of social acceptance and equality, and the inclusion of minority populations in research and surveys.</p>
</div>
]]>
      </description>
    </item>
    
    <item>
      <title>Using TikTok ads for survey recruitment: a step-by-step approach</title>
      <link>https://socialsciencedatalab.mzes.uni-mannheim.de/article/tiktok/</link>
      <pubDate>Fri, 15 Sep 2023 01:00:00 +0100</pubDate>
      
      <guid>https://socialsciencedatalab.mzes.uni-mannheim.de/article/tiktok/</guid>
      <description><![CDATA[
        </p>
<p>TikTok’s rapid growth and diverse user base present social science researchers with a unique opportunity to study a large and varied population, gaining valuable insights into their attitudes and behaviors. Unlike platforms such as Facebook and Instagram, TikTok’s potential in survey recruitment has been relatively underexplored. The platform’s cost-effective reach and detailed targeting parameters make it particularly appealing for reaching traditionally hard-to-reach or rare populations. Furthermore, with its video-centric format and predominantly young user base, TikTok provides a means to engage and attract respondents from younger generations to participate in online surveys. In this <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/categories/tutorials/">Methods Bites Tutorial</a>, <a href="https://twitter.com/ZazaZindel">Zaza Zindel</a> and <a href="https://twitter.com/Simon_LuetkeW">Simon Lütkewitte</a> (Bielefeld University) provide a step-by-step guide on how to use TikTok ads for survey recruitment.</p>
<!-- Optional: One paragraph on learning objectives -->
<p>After reading this blog post, readers should be able to:</p>
<ul>
<li>understand how to use TikTok ads for survey participant recruitment</li>
<li>set up targeted TikTok ad campaigns</li>
<li>analyze and evaluate the performance of ad campaigns and make necessary adjustments</li>
</ul>
<!-- If applicable: Note with references on related materials -->
<p><em>Note:</em> This blog post supplements Zaza Zindel’s workshop in the <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/page/events/">MZES Social Science Data Lab</a>. The original workshop materials, including slides and scripts, are available from our <a href="https://github.com/SocialScienceDataLab/social-media-ads-survey">GitHub</a>. A live recording of the workshop is available on our <a href="https://www.youtube.com/watch?v=9_z5QgMVoMU">YouTube Channel</a>.</p>
<!-- 
Generate an overview of the article 
Note: Section anchors are generated automatically from section headings as 
      hyphenated lower-case labels without special characters; e.g. a section
      title "Why R?" will result in the anchor "#why-r".
-->
<div id="overview" class="section level3">
<h3>Overview</h3>
<ol style="list-style-type: decimal">
<li><a href="#first-section"><strong>Introduction to social media survey recruitment</strong></a></li>
<li><a href="#second-section"><strong>Recruiting via TikTok – a step-by-step guide</strong></a>
<ol style="list-style-type: decimal">
<li><a href="#first-subsection">The TikTok Ads Manager</a></li>
<li><a href="#second-subsection">Setting up an ad campaign</a>
<ol style="list-style-type: decimal">
<li><a href="#first-subsubsection">Campaign level</a></li>
<li><a href="#second-subsubsection">Ad group level</a></li>
<li><a href="#third-subsubsection">Ad level</a></li>
</ol></li>
</ol></li>
<li><a href="#third-section"><strong>Application example</strong></a>
<ol style="list-style-type: decimal">
<li><a href="#third-first-subsection">Ad campaign details</a></li>
<li><a href="#third-second-subsection">Campaign results</a></li>
<li><a href="#third-third-subsection">Sample composition</a></li>
</ol></li>
<li><a href="#fourth-section"><strong>Ethical considerations</strong></a></li>
<li><a href="#fifth-section"><strong>Summary and key takeaways</strong></a></li>
<li><a href="#further-reading"><strong>Further reading</strong></a></li>
</ol>
</div>
<div id="first-section" class="section level3">
<h3>Introduction to social media survey recruitment</h3>
<p>Recruiting research participants is a crucial component of any survey study. However, traditional recruitment methods, such as random walk procedures or random digit dialing, can be time-consuming and expensive. With the growing popularity of social media, platforms like TikTok provide a unique opportunity to quickly and effectively recruit a diverse range of survey participants. TikTok is a popular social media app that allows users to create and share short-form videos with a global audience. With a self-reported number of over 1 billion active users, the platform has gained immense popularity, particularly among younger generations (<a href="https://www.tiktok.com/business/en">TikTok, n.d.</a>). TikTok is relevant and interesting for social science research due to its unique format and engaging nature. The app provides a window into the cultural and social trends of different communities, making it an attractive platform for studying human behavior and attitudes. Additionally, TikTok’s algorithm allows for a personalized feed, meaning that users are served content based on their interests and behaviors. This feature makes it an interesting platform for reaching and studying specific populations.</p>
<p>As a potential tool for survey recruitment, TikTok ads offer several advantages. The platform’s vast and diverse user base provides researchers with access to a broad range of potential participants, while its targeting options allow them to reach very specific populations that may be difficult to reach through more traditional recruitment channels. Unlike social media platforms like Facebook and Instagram, TikTok’s structure revolves around short-form videos set to music, offering a unique and visually engaging format that researchers can utilize to capture users’ attention and encourage their participation. By leveraging visually appealing videos, researchers can create compelling survey invitations, enhancing the overall survey experience. Furthermore, TikTok’s predominantly young user base sets it apart from other platforms, making it particularly valuable for studies targeting younger demographics. This demographic distinction allows researchers to gain insights from a generation with unique attitudes, behaviors, and perspectives, ultimately contributing to a more comprehensive understanding of societal trends and dynamics.</p>
</div>
<div id="second-section" class="section level3">
<h3>Recruiting via TikTok – a step-by-step guide</h3>
<p>In the following section, we will walk you through the step-by-step process of setting up advertisements on TikTok. We will highlight the important factors you need to consider in order to generate a diverse and comprehensive sample for a web survey.</p>
<div id="first-subsection" class="section level5">
<h5>The TikTok Ads Manager</h5>
<p>Having an active TikTok account is the first step towards advertising and recruiting survey participants on TikTok. This is because only individuals and organizations with registered business accounts are allowed to run ads on the platform. To register for a business account, go to the <a href="https://www.tiktok.com/business/en">TikTok for Business website</a> and proceed by clicking the “Create Now” button (see Figure <a href="#fig1">1</a>).</p>
<a name="fig1"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig1"></span>
<img src="/../../../../../article/tiktok_files/images/TikTokforBusiness.png" alt="Starting page of TikTok for Business website." width="80%" />
<p class="caption">
Figure 1: Starting page of TikTok for Business website.
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.tiktok.com/business/en">Tik Tok for Business</a>
</sup></sub>
</p>
</div>
<p>From there, follow the prompts to sign up for a new account. You will need to provide basic information such as your email address and a password to create an account. Once you have an account, you can proceed to set up your business profile (see Figure <a href="#fig2">2</a>). This includes providing information such as your company name, industry and country. Make sure to accurately and clearly represent your business, as this will help to build trust with potential survey participants.</p>
<a name="fig2"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig2"></span>
<img src="/../../../../../article/tiktok_files/images/BusinessAccount2.PNG" alt="TikTok for Business profile." width="40%" />
<p class="caption">
Figure 2: TikTok for Business profile.
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://www.tiktok.com/business/en">Tik Tok for Business (after login)</a>
</sup></sub>
</p>
</div>
</div>
<div id="second-subsection" class="section level5">
<h5>Setting up an ad campaign</h5>
<p>When creating an advertising campaign on TikTok, it is important to familiarize yourself with the campaign structure. The campaign settings are structured into three distinctive hierarchical levels: campaign level, ad group level, and ad level (see Figure <a href="#fig3">3</a>). Within the ad campaign, you create ad groups, which serve as subdivisions allowing for further targeting refinement. Each ad group can contain multiple individual ads that serve as the creative content displayed to TikTok users. This means that within a single ad campaign, you can create multiple ad groups, and within each ad group, there can be multiple ads. Each level requires making informed choices and providing necessary information to ensure the effectiveness of your ad campaign.</p>
<a name="fig3"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig3"></span>
<img src="/../../../../../article/tiktok_files/images/hierarchy.png" alt="TikTok ad hierarchy." width="80%" />
<p class="caption">
Figure 3: TikTok ad hierarchy.
</p>
</div>
<div id="first-subsubsection" class="section level6">
<h6>Campaign level</h6>
<p>At the campaign level, the first step is to choose the campaign objective that best aligns with your advertising goal. This decision will determine how TikTok will optimize the ad delivery. TikTok offers three main categories of campaign objectives: awareness, consideration, and conversion (see Figure <a href="#fig4">4</a>). Each category serves a distinct purpose in terms of user interaction with your ad and the desired actions you want users to take.</p>
<a name="fig4"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig4"></span>
<img src="/../../../../../article/tiktok_files/images/setup3_1.png" alt="Selecting an advertising objective." width="80%" />
<p class="caption">
Figure 4: Selecting an advertising objective.
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://ads.tiktok.com/">Tik Tok Ads Manager (after login)</a>
</sup></sub>
</p>
</div>
<p>Firstly, awareness campaigns are aimed at sparking interest and enhancing the visibility of a brand, product, or service. These campaigns are designed to reach as many people as possible and create a lasting impression of a brand. Secondly, consideration campaigns are focused on driving engagement and interest. This includes encouraging users to visit a website, interact with content, or follow an account. These campaigns are designed to provide users with more information about a brand, product, or service, and to motivate them to take a specific desired action. Lastly, conversion campaigns are designed to drive users towards a specific action, such as making a purchase or downloading an app.</p>
<p>When creating a TikTok ad campaign to recruit survey participants, we highly recommend setting the campaign objective to “Traffic”, which is categorized under consideration. This objective focuses on driving traffic to a designated landing page, such as your survey website. By selecting this objective, your ad campaign will be optimized to effectively direct users to your survey website, increasing the likelihood of attracting potential participants to take part in your survey. TikTok’s machine learning algorithms will be leveraged to identify users within the target audience who are most likely to take the desired action of clicking on the ad and visiting your landing page.</p>
<p>After selecting the campaign objective, you can proceed to define some essential settings of your campaign. These settings include providing a campaign name and defining a maximum budget for the entire campaign (see Figure <a href="#fig5">5</a>). While the maximum budget is not mandatory, we recommend setting it as a precautionary measure to prevent your campaign from exceeding the spending limit of your research project.</p>
<a name="fig5"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig5"></span>
<img src="/../../../../../article/tiktok_files/images/Setup3.2.png" alt="Defining campaign name and campaign budget." width="70%" />
<p class="caption">
Figure 5: Defining campaign name and campaign budget.
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://ads.tiktok.com/">Tik Tok Ads Manager (after login)</a>
</sup></sub>
</p>
</div>
</div>
</div>
<div id="second-subsubsection" class="section level5">
<h5>Ad group level</h5>
<p>At this level, you can create multiple ad groups within a campaign, each with its own settings tailored to the target audience, thereby enhancing ad effectiveness. You can specify various settings unique to each ad group. These include the optimization location, placements, targeting, budget and schedule, bidding, and optimization. First, set the optimization location, i.e. where to direct traffic. For survey recruitment, the optimization location should be set to the survey website where participants can complete the survey. Under the traffic objective, ad placement in Germany is limited to TikTok and Pangle (see Figure <a href="#fig6">6</a>). While TikTok is the primary platform where ads will appear, Pangle is an ad network that provides access to third-party apps. We recommend to choose TikTok instead of Pangle, as TikTok has a larger user base and provides more advanced targeting options, which can help you reach a more specific audience. For other countries, other ad placements are available (for an overview see <a href="https://ads.tiktok.com/help/article/placements-available-locations">TikTok (n.d.)</a>). Additionally, within the placements section, you have the option to customize your ad settings to control the level of user engagement. Consider deactivating comments to avoid harmful or violent reactions, disabling video downloads to protect your content, and restricting video sharing to ensure that your ads are exclusively viewed within the TikTok platform.</p>
<a name="fig6"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig6"></span>
<img src="/../../../../../../article/tiktok_files/images/Setup5.PNG" alt="Defining placement options." width="70%" />
<p class="caption">
Figure 6: Defining placement options.
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://ads.tiktok.com/">Tik Tok Ads Manager (after login)</a>
</sup></sub>
</p>
</div>
<p>Targeting is another crucial step at the ad group level where you define the audience that your ad will be shown to. TikTok provides a range of targeting options, including demographic information, interests, behaviors, and custom audiences. You should define the target audience based on your research objectives and select the appropriate targeting options to effectively reach your desired audience. Take advantage of the wide range of targeting options available when running TikTok ads. Here are some examples:</p>
<ul>
<li>Demographic targeting: Gender (binary), age groups, and language preferences (see Figure <a href="#fig7">7</a>).</li>
<li>Geographic targeting: Countries/regions you can target using “Location” depending on the country/region associated with your TikTok For Business account. For a comprehensive list of all available locations per country, see <a href="https://ads.tiktok.com/help/article/placements-available-locations">TikTok (n.d.)</a>.</li>
<li>Interest-based targeting: Reach users based on their inferred interests, derived from their activity on TikTok (see Figure <a href="#fig8">8</a>).</li>
<li>Behavioral targeting: Reach users based on their behavior on TikTok, such as the videos they have liked, shared, or commented on (see Figure <a href="#fig8">8</a>).</li>
</ul>
<a name="fig7"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig7"></span>
<img src="/../../../../../article/tiktok_files/images/Setup6_1.PNG" alt="Defining demographic targeting parameters." width="70%" />
<p class="caption">
Figure 7: Defining demographic targeting parameters.
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://ads.tiktok.com/">Tik Tok Ads Manager (after login)</a>
</sup></sub>
</p>
</div>
<a name="fig8"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig8"></span>
<img src="/../../../../../article/tiktok_files/images/Setup7_1.PNG" alt="Defining interest- and behavior-based targeting." width="70%" />
<p class="caption">
Figure 8: Defining interest- and behavior-based targeting.
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://ads.tiktok.com/">Tik Tok Ads Manager (after login)</a>
</sup></sub>
</p>
</div>
<p>It is crucial to ensure compliance with TikTok’s advertising policies and avoid any form of discriminatory targeting towards protected groups. TikTok maintains strict advertising policies to uphold community standards and legal requirements. Advertisers are expected to follow these policies, which are regularly updated, in order to maintain the quality and integrity of the TikTok platform. For a detailed description of all prohibited and restricted content, please refer to <a href="https://ads.tiktok.com/help/article/tiktok-advertising-policies-ad-creatives-landing-page">TikTok (n.d.)</a>.</p>
<p>The next step at the ad group level is to set a budget and schedule. You should define a daily budget for each ad group and specify the display schedule for each group. TikTok offers an estimation of an ad group’s potential reach based on the selected targeting options. This estimation can assist in determining an appropriate budget for your campaign (see Figure <a href="#fig9">9</a>).</p>
<a name="fig9"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig9"></span>
<img src="/../../../../../article/tiktok_files/images/Setup11.PNG" alt="Estimated audience reach." width="25%" />
<p class="caption">
Figure 9: Estimated audience reach.
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://ads.tiktok.com/">Tik Tok Ads Manager (after login)</a>
</sup></sub>
</p>
</div>
<p>Once you have set the budget and schedule for your ad group, the next step is to define the options in Bidding &amp; Optimization (see Figure <a href="#fig10">10</a>). This step determines how the ad will be delivered and optimized for maximum effectiveness. On TikTok, you have the choice between two optimization goals: Click and landing page view. The default and recommended optimization goal is Click, which focuses on driving users to click on the ad and visit the website. Alternatively, the landing page view optimization goal targets users who not only click on the ad but also waits for the landing page to load completely. This is particularly useful for campaigns that require users engagement with the content on the landing page.</p>
<a name="fig10"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig10"></span>
<img src="/../../../../../article/tiktok_files/images/Setup10_1.PNG" alt="Defining bidding and optimization options." width="70%" />
<p class="caption">
Figure 10: Defining bidding and optimization options.
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: <a href="https://ads.tiktok.com/">Tik Tok Ads Manager (after login)</a>
</sup></sub>
</p>
</div>
<p>Furthermore, you also have the option to select between two bid strategies: lowest cost and bid cap. The lowest cost bidding strategy is a spend-based approach, where TikTok’s algorithm automatically adjusts the bid to generate as many results as possible at the lowest cost per result. This strategy is ideal for most campaigns. However, if you want to have more control over the budget allocation, you can choose the bid cap option. With bid cap, you can set a maximum bid amount for each link click, ensuring that the cost per result remains below your predetermined limit. This strategy can be particularly beneficial for campaigns with limited budgets or when you want to ensure that you are not overpaying for each link click.</p>
</div>
<div id="third-subsubsection" class="section level5">
<h5>Ad level</h5>
<p>On the ad level, you can create a captivating advertisement that will be displayed to the target audience. The goal is to make the ad highly appealing and drive engagement among your target population. To achieve this, your ad should be eye-catching, and effectively convey the purpose of the survey. Given that TikTok is a platform centered around videos, we highly recommend incorporating a short and captivating video into your ad. The video should grab viewers’ attention and motivate them to take action. While the ad format can also be a single image or image carousel, it is advisable to include at least some dynamic elements to align with the platform’s style. Moreover, you must keep in mind that TikTok ads are primarily optimized for mobile devices. Hence, the ad should be visually appealing and easily viewable on smartphones, ensuring a seamless user experience.</p>
<a name="fig11"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig11"></span>
<img src="/../../../../../article/tiktok_files/images/Ad_example.png" alt="Example ad design." width="80%" />
<p class="caption">
Figure 11: Example ad design.
</p>
</div>
<div style="text-align: right">
<p>
<sub><sup>
Source: Image by Macrovector via stock.adobe.com</a>
</sup></sub>
</p>
</div>
<p>In addition, you have the option to choose from various call-to-action buttons, such as “Learn More” or “Sign Up”, which are designed to encourage users to take action. When users click on the button, they will be directed to the survey (see Figure <a href="#fig11">11</a>). Furthermore, it is important to ensure that the ad text and description convey a clear and concise message that motivates users to participate in the survey.</p>
<p>Once the ad campaign is created, it can be launched. However, it is important to note that TikTok will review the ads to ensure compliance with their guidelines and policies. It is essential to avoid displaying prohibited products or services, such as live animals, adult sexual products, casinos, military gear, or political candidates or parties, in the ad creatives or on the landing pages. A comprehensive list of prohibited content can be found at <a href="https://ads.tiktok.com/help/article/tiktok-advertising-policies-industry-entry">TikTok (n.d.)</a>.</p>
<p>Once the ads are live, you can utilize TikTok’s ad manager tool to track their performance. Monitoring the number of views, clicks, and conversions is essential to assess the reach, engagement, and overall effectiveness of the ads. This valuable information can be used to make necessary adjustments to the ad settings and enhance the performance of future ads.</p>
</div>
</div>
<div id="third-section" class="section level3">
<h3>Application example</h3>
<p>The following section presents the results of an exemplary TikTok advertising campaign conducted as part of a master’s seminar at Bielefeld University. In this seminar, students were tasked with designing a questionnaire aligned with their individual research interests. The main topics were attitudes toward refugees and youth crime in the UK and Germany. To efficiently and cost-effectively generate a substantial sample size, the study was promoted through TikTok ads.</p>
<div id="third-first-subsection" class="section level4">
<h4>Ad campaign details</h4>
<p>In accordance with the previously discussed hierarchical structure, various settings were configured at the three different levels of the advertising campaign (see Table <a href="#table1">1</a>).</p>
<a name="table1"></a>
<caption>
Table 1: Setup of example campaign
</caption>
<iframe src="hierarchy/index.html" frameborder="no" width="100%" height="840px" frameborder="no">
</iframe>
<p>At the campaign level, we opted for a consideration advertising objective, specifically “Traffic”, with the aim of driving traffic to our survey website. We set the maximum budget of 400.00 EUR for the entire study. Moving on to the ad group level, we created three distinct ad groups: one targeting English-speaking individuals in the UK (ad group: UK), another targeting German-speaking individuals in Germany (ad group: DE1), and a third focusing on English speaking-individuals in Germany (ad group: DE2). While these ad groups differed in terms of location and language targeting, certain settings were kept consistent across all three: age range, gender, maximum budget per ad group, schedule, delivery, bid strategy, optimization goal, and billing event. For each ad group, we allocated a maximum budget of 130.00 EUR, aiming to evenly distribute our budget across the ad groups. The recruitment phase was scheduled to run from August 6th, 2022, at 16:00 to August 9th, 2022, at 16:00. We configured all ad groups to deliver ads throughout the day, and the optimization goal for each group was to maximize clicks on the advertisements. To minimize costs while maximizing the number of clicks, we selected the “lowest cost” bid strategy. Based on our settings, the billing event was set to cost per click (CPC).</p>
<p>We created a total of 15 ads, with five ads designed under each of the three ad groups. These ads included either a video or an image with animations. As an example, Video <a href="#video1">1</a> illustrates an ad tailored to address the topic of juvenile delinquency, and was specifically targeted to the English-speaking target population. In addition to the visuals, each ad featured our custom identity (<span class="citation">@SozPolStudyUniBielefeld</span>), license-free background music provided through TikTok, and concise yet meaningful ad texts, such as “Take our short survey!” in the corresponding language of the target group. To encourage user engagement, we utilized the “Learn more” call-to-action button. When users clicked on this button, they were redirected to our survey page, and we employed URL parameters to track from which respective ad participation in the survey took place.</p>
<p><a name="video1"></a></p>
<center>
<video width="500" height="500" controls>
  <source src="/../../../../../article/tiktok_files/Ads/YD1UK-YD1DE2.mp4" type="video/mp4">
</video>
</center>
<br>
<caption>Video 1: Example of an ad video.</caption>
</br>
<div style="text-align: right">
<p>
<sub><sup>
Source: Video by cottonbro via Canva.com</a>
</sup></sub>
</p>
</div>
</div>
<div id="third-second-subsection" class="section level4">
<h4>Campaign results</h4>
<p>Table <a href="#table2">2</a> provides an overview of the ad campaign’s performance metrics. Over the course of the 4-day campaign, it reached a total of 117,626 TikTok users. The videos or animated images were played 189,100 times, with the ads being displayed 205,071 times, as reported by the TikTok Ad Manager. The campaign incurred a total cost of 387.48 EUR and generated 3,603 clicks on the call-to-action button, representing a click-through rate of 1.76 percent. The average cost per click was 0.11 EUR, which is comparable to recruitment via other social media platforms like Facebook or Instagram (cf. e.g., <a href="https://doi.org/10.12758/mda.2022.15">Zindel 2022</a>).</p>
<p>Analyzing the ad performance reveals variation depending on the target groups. Ads placed in Germany (ad groups: D1 &amp; D2) reached a higher overall number of impressions compared to the UK. This discrepancy can be attributed to the less developed TikTok ads market in Germany compared to the UK, resulting in lower costs for link clicks due to reduced ad competition. Consequently, the same ad budget resulted in varying levels of impressions. Figure <a href="#fig12">12</a> illustrates the cost per click (CPC) by ad group. Initially, during the learning phase of the advertising algorithm, ads in the UK were slightly more expensive than those in Germany. Moreover, the ad group targeting English speakers in Germany (ad group: DE2) exhibited the highest cost per click over the short duration of the campaign. This can be attributed to the relatively lower proportion of TikTok users in Germany who use the platform in English compared to those who use it in German. Targeting a smaller user base incurs higher costs to reach a significant number of users, as reflected in the number of impressions in Table <a href="#table2">2</a>.</p>
<a name="fig12"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig12"></span>
<img src="/../../../../../article/tiktok_files/images/CPC.png" alt="Average cost per click (CPC) for each ad group." width="70%" />
<p class="caption">
Figure 12: Average cost per click (CPC) for each ad group.
</p>
</div>
<p>Examining the metrics distribution in Table <a href="#table2">2</a> highlights significant variations in ad performance. Some ads were displayed and viewed more frequently than others within the same ad group. For instance, in ad group DE2, the ad YD2DE2, one of two ads that focused on the topic of youth delinquency (“YD” for short), had the highest allocation rate determined by the advertising algorithm, resulting in the highest number of clicks on the call-to-action button within this group. When considering the click-through-rate (CTR) for all ads, which represents the conversion rate from impressions to clicks, ads served more frequently tended to have comparatively higher CTRs. However, it is important to note that these ads may not necessarily have the highest CTR within their respective ad groups. For example, in the UK ad group, the ad RF2UK had a CTR of 2.66 percent, surpassing the most frequently displayed ad YD2UK with a CTR of 2.22 percent. This discrepancy arises from the fact that ads that initially garner significant attention in the initial learning phase of the advertising algorithm may not necessarily sustain the highest level of interest from the target audience throughout the entire ad runtime <a href="https://ads.tiktok.com/help/article/learning-phase?lang=en">TikTok (n.d.)</a>.</p>
<a name="table2"></a>
<caption>
Table 2: Results of ad performance.
</caption>
<iframe src="table2_new/index.html" frameborder="no" width="100%" height="400px" frameborder="no">
</iframe>
<p>Table <a href="#table3">3</a> presents the conversion of link clicks per ad into started and completed interviews, along with the corresponding total cost and cost per completed interview. Out of the 3,603 link clicks, 1,846 (51.29 percent) resulted in started interviews, and a total of 500 interviews (13.88 percent) were completed. With a total campaign cost of 387.48 EUR, the average cost per completed interview was 0.77 EUR. Notably, variations can be observed across the ad groups. The ad group targeting German-speaking individuals in Germany (DE1) received the highest number of link clicks (n=1,676), resulting in the highest number of started interviews (n=983; 58.65 percent), as well as completed interviews (n=265; 15.81 percent). This specific target group was also the most cost-effective to reach, with an average cost of 0.49 EUR per completed interview. In contrast, the average cost per completed interview was 0.99 EUR for the DE2 ad group and 1.23 EUR for the UK group. This disparity highlights the differences in costs discussed earlier, which are influenced by the bidding system and the ad algorithm.</p>
<a name="table3"></a>
<caption>
Table 3: Results of survey performance.
</caption>
<iframe src="table3/index.html" frameborder="no" width="100%" height="400px" frameborder="no">
</iframe>
</div>
<div id="third-third-subsection" class="section level4">
<h4>Sample composition</h4>
<p>The final sample comprises three subsamples derived from the three different ad groups: English-speaking individuals in the UK (ad group: UK), German-speaking individuals in Germany (ad group: DE1), and English-speaking individuals in Germany (ad group: DE2). Group DE1 accounts for 265 completed interviews, representing 53.00 percent of all interviews. Group DE2 includes 130 interviews (26.00 percent) and group UK has 105 interviews (21.00 percent). The average age of the respondents is 37.22 years (see Figure <a href="#fig13">13</a>), which is higher than expected considering TikTok’s popularity among younger generations. However, it is important to note that the average age in our sample still falls below the overall average age of 44.7 in the German population (<a href="https://de.statista.com/statistik/daten/studie/1084430/umfrage/durchschnittsalter-der-bevoelkerung-in-deutschland/#:~:text=Zum%20Ende%20des%20Jahres%202021,der%20m%C3%A4nnlichen%20Bev%C3%B6lkerung%20in%20Deutschland">Statistische Ämter des Bundes und der Länder, 2022</a>) and the median age for the UK (<a href="https://de.statista.com/statistik/daten/studie/200671/umfrage/durchschnittsalter-der-bevoelkerung-in-grossbritannien/">UN DESA, 2022</a>).</p>
<a name="fig13"></a>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:fig13"></span>
<img src="../../../../../../../../article/tiktok_files/images/Age.png" alt="Age distribution in overall sample." width="50%" />
<p class="caption">
Figure 13: Age distribution in overall sample.
</p>
</div>
<p>The average age within the DE1 ad group is 47.74 years, whereas it is 23.99 years for the DE2 ad group and 25.53 years for individuals in the UK. This indicates a strong selection bias towards older individuals in the German-speaking group, while the other ad groups exhibit a strong bias towards younger participants. Table <a href="#table4">4</a> compared the age distribution of each ad group with the ad allocation information provided by TikTok. These selection effects are primarily influenced by the ad algorithm. Therefore, for studies aiming to achieve a demographically balanced age distribution, it is recommended to further stratify ad groups and run specific ads targeted at different age groups to mitigate similar age biases.</p>
<a name="table4"></a>
<caption>
Table 4: Age group distribution by ad groups.
</caption>
<iframe src="table4/index.html" frameborder="no" width="100%" height="480px" frameborder="no">
</iframe>
<p>Regarding gender distribution, a larger proportion (61.52 percent) of male participants was identified in the sample, while only 27.05 percent identified as female (see Table <a href="#table5">5</a>). Additionally, 60 individuals (11.43 percent) in the sample reported identifying with a different gender. A comparison of the sample data with the information from the TikTok Ads-Manager reveals that this effect can also be partially attributed to the allocation of the ads by the advertising algorithm. The DE1 ad group was predominantly delivered to men, while the DE2 and UK ad groups did not exhibit such a strong gender effect in their ad display. It appears that a large proportion of respondents, assumed to be female, chose the third response option, “another gender”. However, it is important to note that we lack individual gender information within the TikTok platform.</p>
<a name="table5"></a>
<caption>
Table 5: Gender distribution by ad groups.
</caption>
<iframe src="table5/index.html" frameborder="no" width="100%" height="320px" frameborder="no">
</iframe>
</div>
</div>
<div id="fourth-section" class="section level3">
<h3>Ethical considerations</h3>
<p>When considering the use of TikTok for survey recruitment and advertisements in general, several ethical implications merit careful consideration. TikTok, a Chinese company owned by ByteDance with established ties to the Chinese state, raises pertinent concerns. Purchasing advertising space on this platform contributes to TikTok’s financial revenue, strengthening its market position as one of the fastest-growing social media platforms. Due to its indirect connection to the Chinese state, numerous government agencies, such as the European Commission and the European Council, have banned TikTok from their official devices. Moreover, some countries have opted for a complete prohibition of TikTok due to concerns encompassing cyber security, privacy, and misinformation.</p>
<p>With the approach outlined in this blog post, TikTok does not have access to the survey data at any time. The social media platform remains incapable of identifying which users have participated in our survey. To ensure that the survey data remains inaccessible to third parties at all times, we strongly advise against the utilization of tracking tools provided by TikTok, as well as other platforms.</p>
<p>Nonetheless, it is worth noting that TikTok creates interest profiles of users based on their interaction with advertised content – a practice also common among other social media platforms. However, considering TikTok’s indirect link to the Chinese state, the possibility that these interest profiles could be accessed by Chinese authorities cannot be disregarded. This potential vulnerability is especially concerning for marginalized population groups, such as LGBTQI* individuals. While TikTok may lack specific information regarding the gender or sexual identity of its users, it can infer from users’ interactions with content related to LGBTQI* topics that these users at least hold an interest in such matters.</p>
<p>For all researchers intending to establish financial collaborations with TikTok –- whether to recruit survey participants or promote other content –- we strongly recommend a preliminary consultation with their respective institutions, employers, and ethics committees. This step will help determine the feasibility of such projects within the ethical framework.</p>
</div>
<div id="fifth-section" class="section level3">
<h3>Summary and key takeaways</h3>
<p>Recruiting survey participants through TikTok ads offers several advantages for researchers. TikTok’s vast and diverse user base allows access to a broad range of potential participants, and its targeting options enable reaching specific populations that may be challenging to reach through traditional methods. By following our presented step-by-step guide for setting up advertisements on TikTok, you can effectively generate a diverse and inexpensive sample for your web surveys. The example of a TikTok advertising campaign demonstrates the platform’s potential in efficiently and cost-effectively recruiting survey participants. We delved into its unique features, demographics, and engagement levels, considering its suitability for reaching a diverse audience. Additionally, we examined the advantages and disadvantages of utilizing TikTok, highlighting its ability to generate rapid and widespread exposure while also acknowledging the challenges it presents, such as limited insights in the algorithmic ad allocation and potential for high selectivity.</p>
<p>Overall, the key takeaways from this tutorial are as follows:</p>
<ul>
<li>TikTok offers a vast user base with diverse demographics, making it a promising platform for survey recruitment efforts aimed at reaching a wide audience.</li>
<li>Ad campaigns on TikTok are structured into campaign levels, ad group levels, and ad levels, each requiring informed choices and necessary information for effectiveness.</li>
<li>TikTok provides various targeting options, including demographics, interests, behaviors, and custom audiences, to reach the desired audience.</li>
<li>Ad creatives on TikTok should be eye-catching, incorporate videos or dynamic elements, and have clear ad text and call-to-action buttons.</li>
<li>It is essential to carefully monitor and address potential ethical concerns, such as data privacy and user consent, when utilizing TikTok for survey recruitment.</li>
<li>Monitoring the performance of TikTok ads through the ad manager tool is crucial for assessing reach, engagement, and effectiveness, especially when it becomes apparent that certain ad groups outperform others.</li>
<li>The distribution of ads by the algorithm may lead to overrepresentation of different subgroups. For demographically more comparable population distributions, it is recommended to stratify the ads using the targeting parameters.</li>
</ul>
</div>
<div id="further-reading" class="section level3">
<h3>Further reading</h3>
<p>Here are several publications that offer a valuable introduction to the topic of social media ads for survey respondent recruitment.These publications provide valuable insights into the use of social media ads for recruiting survey respondents and serve as a great starting point for exploring this subject further:</p>
<ul>
<li>Iannelli, L., Giglietto, F., Rossi, L., &amp; Zurovac, E. (2020). Facebook Digital Traces for Survey Research: Assessing the Efficiency and Effectiveness of a Facebook Ad–Based Procedure for Recruiting Online Survey Respondents in Niche and Difficult-to-Reach Populations. Social Science Computer Review, 38(4), 462–476. <a href="doi:https://doi.org/10.1177/0894439318816638" class="uri">doi:https://doi.org/10.1177/0894439318816638</a></li>
<li>Kühne, S., &amp; Zindel, Z. (2020). Using Facebook and Instagram to Recruit Web Survey Participants: A Step-by-Step Guide and Application. Survey Methods: Insights from the Field. <a href="doi:https://doi.org/10.13094/SMIF-2020-00017" class="uri">doi:https://doi.org/10.13094/SMIF-2020-00017</a></li>
<li>Neundorf, A., &amp; Öztürk, A. (2021, July 26). How to Improve Representativeness and Cost-effectiveness in Samples Recruited through Meta: A Comparison of Advertisement Tools. url:<a href="https://doi.org/10.31219/osf.io/3g74n" class="uri">https://doi.org/10.31219/osf.io/3g74n</a></li>
<li>Neundorf, A., &amp; Öztürk, A. (2021, December 9). Recruiting Research Participants through Facebook Advertisements: A Handbook. url:<a href="https://doi.org/10.31219/osf.io/87rg3" class="uri">https://doi.org/10.31219/osf.io/87rg3</a></li>
<li>Zindel, Z. (2022). Social Media Recruitment in Online Survey Research: A Systematic Literature Review. methods, data, analyses, 0, 42. <a href="doi:https://doi.org/10.12758/mda.2022.15" class="uri">doi:https://doi.org/10.12758/mda.2022.15</a></li>
</ul>
<p>We would also like to highlight the excellent webinar series of the <a href="https://www.gla.ac.uk/research/az/democracyresearch/dataandmethods/socialmediaasaresearchtool/webinarseries/">DEMED project</a>. This series offers a range of insightful presentations that delve into the advantages of social media for survey research.</p>
<!-- Add something about the instructor -->
</div>
<div id="about-the-authors" class="section level3">
<h3>About the authors</h3>
<p>Zaza Zindel <a href="mailto:zaza.zindel@uni-bielefeld.de"><svg aria-hidden="true" role="img" viewBox="0 0 512 512" style="height:1em;width:1em;vertical-align:-0.125em;margin-left:auto;margin-right:auto;font-size:inherit;fill:currentColor;overflow:visible;position:relative;"><path d="M64 112c-8.8 0-16 7.2-16 16v22.1L220.5 291.7c20.7 17 50.4 17 71.1 0L464 150.1V128c0-8.8-7.2-16-16-16H64zM48 212.2V384c0 8.8 7.2 16 16 16H448c8.8 0 16-7.2 16-16V212.2L322 328.8c-38.4 31.5-93.7 31.5-132 0L48 212.2zM0 128C0 92.7 28.7 64 64 64H448c35.3 0 64 28.7 64 64V384c0 35.3-28.7 64-64 64H64c-35.3 0-64-28.7-64-64V128z"/></svg></a> <a href="https://twitter.com/ZazaZindel"><svg aria-hidden="true" role="img" viewBox="0 0 512 512" style="height:1em;width:1em;vertical-align:-0.125em;margin-left:auto;margin-right:auto;font-size:inherit;fill:currentColor;overflow:visible;position:relative;"><path d="M459.37 151.716c.325 4.548.325 9.097.325 13.645 0 138.72-105.583 298.558-298.558 298.558-59.452 0-114.68-17.219-161.137-47.106 8.447.974 16.568 1.299 25.34 1.299 49.055 0 94.213-16.568 130.274-44.832-46.132-.975-84.792-31.188-98.112-72.772 6.498.974 12.995 1.624 19.818 1.624 9.421 0 18.843-1.3 27.614-3.573-48.081-9.747-84.143-51.98-84.143-102.985v-1.299c13.969 7.797 30.214 12.67 47.431 13.319-28.264-18.843-46.781-51.005-46.781-87.391 0-19.492 5.197-37.36 14.294-52.954 51.655 63.675 129.3 105.258 216.365 109.807-1.624-7.797-2.599-15.918-2.599-24.04 0-57.828 46.782-104.934 104.934-104.934 30.213 0 57.502 12.67 76.67 33.137 23.715-4.548 46.456-13.32 66.599-25.34-7.798 24.366-24.366 44.833-46.132 57.827 21.117-2.273 41.584-8.122 60.426-16.243-14.292 20.791-32.161 39.308-52.628 54.253z"/></svg></a> is a research assistant and a doctoral researcher in sociology at Bielefeld University (Germany) with a specialization in survey research. Her dissertation centers around the utilization of social media as a means to recruit rare populations for web surveys. Her research interests encompass survey methodology, the potential of social media for empirical social research, and the exploration of new technologies to improve statistical representation of marginalized, vulnerable, or rare population groups.</p>
<p>Simon Lütkewitte <a href="mailto:simon.luetkewitte@uni-bielefeld.de"><svg aria-hidden="true" role="img" viewBox="0 0 512 512" style="height:1em;width:1em;vertical-align:-0.125em;margin-left:auto;margin-right:auto;font-size:inherit;fill:currentColor;overflow:visible;position:relative;"><path d="M64 112c-8.8 0-16 7.2-16 16v22.1L220.5 291.7c20.7 17 50.4 17 71.1 0L464 150.1V128c0-8.8-7.2-16-16-16H64zM48 212.2V384c0 8.8 7.2 16 16 16H448c8.8 0 16-7.2 16-16V212.2L322 328.8c-38.4 31.5-93.7 31.5-132 0L48 212.2zM0 128C0 92.7 28.7 64 64 64H448c35.3 0 64 28.7 64 64V384c0 35.3-28.7 64-64 64H64c-35.3 0-64-28.7-64-64V128z"/></svg></a> <a href="https://twitter.com/Simon_LuetkeW"><svg aria-hidden="true" role="img" viewBox="0 0 512 512" style="height:1em;width:1em;vertical-align:-0.125em;margin-left:auto;margin-right:auto;font-size:inherit;fill:currentColor;overflow:visible;position:relative;"><path d="M459.37 151.716c.325 4.548.325 9.097.325 13.645 0 138.72-105.583 298.558-298.558 298.558-59.452 0-114.68-17.219-161.137-47.106 8.447.974 16.568 1.299 25.34 1.299 49.055 0 94.213-16.568 130.274-44.832-46.132-.975-84.792-31.188-98.112-72.772 6.498.974 12.995 1.624 19.818 1.624 9.421 0 18.843-1.3 27.614-3.573-48.081-9.747-84.143-51.98-84.143-102.985v-1.299c13.969 7.797 30.214 12.67 47.431 13.319-28.264-18.843-46.781-51.005-46.781-87.391 0-19.492 5.197-37.36 14.294-52.954 51.655 63.675 129.3 105.258 216.365 109.807-1.624-7.797-2.599-15.918-2.599-24.04 0-57.828 46.782-104.934 104.934-104.934 30.213 0 57.502 12.67 76.67 33.137 23.715-4.548 46.456-13.32 66.599-25.34-7.798 24.366-24.366 44.833-46.132 57.827 21.117-2.273 41.584-8.122 60.426-16.243-14.292 20.791-32.161 39.308-52.628 54.253z"/></svg></a> is a research assistant at Bielefeld University (Germany) specializing in empirical social science research, particularly quantitative methods . He is also a PhD student at the Bielefeld Graduate School for History and Sociology (BGHS). His research focuses on values and attitudes within the sphere of sports. Methodologically, he is interested in applying statistical methods such as multilevel modeling, matching techniques, and social network analysis. In his dissertation project, he utilizes quantitative methods to examine the relationship between sports participation and gender-related attitudes and beliefs.</p>
</div>
]]>
      </description>
    </item>
    
    <item>
      <title>Collection, Management, and Analysis of Twitter Data</title>
      <link>https://socialsciencedatalab.mzes.uni-mannheim.de/article/twitter-research-track/</link>
      <pubDate>Thu, 02 Jun 2022 01:00:00 +0100</pubDate>
      
      <guid>https://socialsciencedatalab.mzes.uni-mannheim.de/article/twitter-research-track/</guid>
      <description><![CDATA[
        </p>
<p>As a highly relevant platform for political and social online interactions, researchers increasingly analyze Twitter data. As of 01/2021, Twitter renewed its API, which now includes access to the full history of tweets for academic usage. In this <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/categories/tutorials/">Methods Bites Tutorial</a>, <a href="https://twitter.com/ankuepfer">Andreas Küpfer</a> (Technical University of Darmstadt &amp; MZES) presents a walkthrough of the collection, management, and analysis of Twitter data.</p>
<p>After reading this blog post and engaging with the applied exercises, readers will be able to:</p>
<ul>
<li>complete the academic research track application process for the Twitter API.</li>
<li>crawl tweets using customized queries based on the R package <code>academictwitteR</code> <span class="citation">(Barrie and Ho <a href="#ref-BarrieHo2021" role="doc-biblioref">2021</a>)</span>.</li>
<li>apply a selection of pre-processing steps to these tweets.</li>
<li>take decisions in order to minimize reprodubcibility issues with Twitter data and to comply with the policies.</li>
</ul>
<p><em>Note:</em> This blog post provides a summary of Andreas’ workshop in the <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/page/events/">MZES Social Science Data Lab</a>. The original workshop materials, including slides and scripts, are available from our <a href="https://github.com/SocialScienceDataLab/twitter-api-bert-method">GitHub</a>.
A live recording of the workshop is available on our <a href="https://www.youtube.com/watch?v=Gzl0lpQ7S7w">YouTube channel</a>.</p>
<div id="overview" class="section level3">
<h3>Overview</h3>
<ol style="list-style-type: decimal">
<li><a href="#introduction-to-social-media-twitter-api-v2"><strong>Introduction to social media &amp; Twitter API v2</strong></a></li>
<li><a href="#academic-research-track-application-process"><strong>Academic research track application process</strong></a>
<ol style="list-style-type: decimal">
<li><a href="#prerequisites">Prerequisites</a></li>
<li><a href="#application">Application</a></li>
<li><a href="#after-the-application">After the application</a></li>
</ol></li>
<li><a href="#using-the-api"><strong>Using the API</strong></a>
<ol style="list-style-type: decimal">
<li><a href="#postman-as-a-playground">Postman-as-a-playground</a></li>
<li><a href="#which-package-should-i-choose">Which package should I choose?</a></li>
<li><a href="#academictwitter-a-code-walkthrough-using-r">academictwitteR: a code walkthrough using R</a></li>
</ol></li>
<li><a href="#wpreparing-for-methods-working-with-textual-data"><strong>Preparing for methods: working with textual data</strong></a></li>
<li><a href="#reproducibility-of-research-basbed-on-twitter-data"><strong>Reproducibility of research based on Twitter data</strong></a></li>
<li><a href="#further-readings"><strong>Further readings</strong></a></li>
</ol>
</div>
<div id="introduction-to-social-media-and-twitter-api-v2" class="section level3">
<h3>Introduction to social media and Twitter API v2</h3>
<p>Social media posts are full of potential for data mining and analysis. Despite problems tackling fake accounts and bots on the platform, it can be a very fruitful source to tackle research questions in a bandwidth of disciplines, including social sciences <span class="citation">(e.g., Barberá <a href="#ref-barberá_2015" role="doc-biblioref">2015</a>; Nguyen et al. <a href="#ref-NGUYEN2021100922" role="doc-biblioref">2021</a>; Valle-Cruz et al. <a href="#ref-cruz2022" role="doc-biblioref">2022</a>; Sältzer <a href="#ref-doi:10.1177/1354068820957960" role="doc-biblioref">2022</a>)</span>. Recognizing this potential also for commercial usage, platform providers increasingly restrict free access to such data.</p>
<p>Especially Twitter is an important data source with its richness of social and political interactions. As well Twitter did not offer a free-of-charge option to implement a full archive search of all tweets and users. Back then, the free version of API v1.1 was very limited with a maximum of 3,200 tweets or the past seven days of tweets. In addition, the range of available meta data<a href="#fn1" class="footnote-ref" id="fnref1"><sup>1</sup></a> as well as implemented query<a href="#fn2" class="footnote-ref" id="fnref2"><sup>2</sup></a> options were rather small. These limitations were lifted by the introduction of the redeveloped and rearranged Twitter API v2 in January 2021.<a href="#fn3" class="footnote-ref" id="fnref3"><sup>3</sup></a> For academic purposes, they opened up access to all available tweets and other objects posted on Twitter without any monetary costs for the researcher.</p>
<p>While this blog post focuses on the retrieval of textual data, Twitter content certainly offers more. Looking at social network interactions (e.g., followers, likes, …) is just one of the opportunities beyond text to reveal valuable information. This can be, for example, the usage of follower networks to estimate ideological positions <span class="citation">(e.g., Barberá <a href="#ref-barberá_2015" role="doc-biblioref">2015</a>)</span> or measuring the importance of a user in a social network based on social interaction data.</p>
</div>
<div id="academic-research-track-application-process" class="section level3">
<h3>Academic research track application process</h3>
<p>As Application Programming Interfaces (APIs) are powerful tools which allow access to vast databases full of information, companies offering them are increasingly careful about who is allowed to use them. While the previous version of the Twitter API provided access without a dedicated application (for a detailed description, see this Methods Bites <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/article/collecting-and-analyzing-twitter-using-r/">tutorial</a>), the novel version requires you to go through an application process where you have to provide several details about you and your project with Twitter. This information includes data regarding yourself as well as the research project where you intend to work with Twitter data.</p>
<div id="prerequisites" class="section level5">
<h5>Prerequisites</h5>
<p>Before getting access, you have to fulfill several formal prerequisites to be eligible for application:</p>
<ul>
<li>You are either a master’s student, a doctoral candidate, a post-doc, a faculty member, or a research-focused employee at an academic institution or university.</li>
<li>You have a clearly defined research objective, and you have specific plans for how you intend to use, analyze, and share Twitter data from your research.</li>
<li>You will use this access for non-commercial purposes.<a href="#fn4" class="footnote-ref" id="fnref4"><sup>4</sup></a></li>
</ul>
<p>Furthermore, you need a Twitter account which is also used to log in to the Twitter Developer Platform after a successful application. This portal lets you configure your API projects, keep an eye on your monthly tweet cap<a href="#fn5" class="footnote-ref" id="fnref5"><sup>5</sup></a>, and more. A more detailed explanation of prerequisites can be found on the <a href="https://developer.twitter.com/en/products/twitter-api/academic-research/application-info">Twitter API academic research track Track</a> website.</p>
</div>
<div id="application" class="section level5">
<h5>Application</h5>
<p>The whole process can be initiated by clicking <a href="https://developer.twitter.com/en/portal/petition/academic/is-it-right-for-you">Apply</a> on the official <a href="https://developer.twitter.com/en/products/twitter-api/academic-research">Twitter API academic research track</a> Website. You’ll be asked to log in with your personal Twitter account.</p>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:unnamed-chunk-1"></span>
<img src="../../../../../article/twitter-research-track/twitter_application.png" alt="Twitter Application Steps for academic research track API access." width="100%" />
<p class="caption">
Figure 1: Twitter Application Steps for academic research track API access.
</p>
</div>
<div style="text-align: right">
<p><sub><sup>
Source: <a href="https://developer.twitter.com/en/portal/petition/academic/is-it-right-for-you">Twitter API Application Process</a>
</sub></sup></p>
</div>
<p>The figure above visualizes the steps you have to complete before your application can finally be submitted for Twitter’s internal review:</p>
<ol style="list-style-type: decimal">
<li><strong>Basic Info</strong>: such as phone number verification and country selection</li>
<li><strong>Academic Profile</strong>: such as link to an official profile (department website or similar) and academic role</li>
<li><strong>Project Details</strong>: such as information about findings, description of the project itself, and how the API should be used there (e.g. methodologies and how the outcomes will be shared)</li>
<li><strong>Review</strong>: provides an overview of the previous steps</li>
<li><strong>Terms</strong>: developer agreement and policy</li>
</ol>
<p>Before starting, it is recommended to carefully read which kind of career levels, projects and data behaviors are not allowed to use the API and thus have a high chance of receiving a refusal for their application. To give an example, if you plan to share the content of tweets publicly, you most probably won’t get access to the API as this would violate the Twitter rules. Again, more detailed information about this can be found on the <a href="https://developer.twitter.com/en/products/twitter-api/academic-research/application-info">Twitter API academic research track</a> and <a href="https://developer.twitter.com/en/developer-terms/more-on-restricted-use-cases">Developer Terms</a> information guides.</p>
<p>Step one requests generic information about your Twitter account while in step two you have to provide information about your academic profile. This includes a link to a publicly available record on an official department website and information regarding the academic institution you are working in. The third step is the most sophisticated one: your research project. It asks for short paragraphs about the project in general, what and how Twitter data is used there, and how the outcome of your work is shared with the public. The last two steps, review and terms, do not require any user-specific input but provide an overview of all filled-in information as well as the chance to read the developer agreement and policy.</p>
</div>
<div id="after-submitting-your-application" class="section level5">
<h5>After submitting your application</h5>
<p>After submitting your application, you receive a decision via the e-mail address connected with your Twitter account (usually) within a few days. However, according to Twitter, this process can take up to two weeks.</p>
<p>You application may be rejected for two common reasons: First, you do violate the policy at one point according to the information given, or second, you do not meet the <a href="#before-the-application">requirements (as described above)</a>. Further explanations what can be the next steps after a rejection can be found in the <a href="https://developer.twitter.com/en/support/twitter-api/developer-account">Developer Account Support FAQ</a>.</p>
<p>As of writing this blog post (May 2022), submitting a reapplication for access using the same account is not possible.</p>
</div>
</div>
<div id="using-the-api" class="section level3">
<h3>Using the API</h3>
<p>After your successful application, the <a href="https://developer.twitter.com/en/portal/dashboard">Twitter Developer Portal</a> is there to manage projects and environments (which belong to a project), generate API keys (“credentials” for API access), get an overview of real-time monthly tweet cap usage, check available API endpoints and their specifics and more.</p>
<p>After the creation of a project, an environment can be added and API keys generated.</p>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:unnamed-chunk-2"></span>
<img src="../../../../../article/twitter-research-track/twitter_apikeys.png" alt="API keys of an environment" width="50%" />
<p class="caption">
Figure 2: API keys of an environment
</p>
</div>
<div style="text-align: right">
<p><sub><sup>
Source: <a href="https://developer.twitter.com/en/docs/tutorials/step-by-step-guide-to-making-your-first-request-to-the-twitter-api-v2">Twitter API Application Guide</a>
</sub></sup></p>
</div>
<p>The following keys are generated automatically and used depending on the API interface (e.g. the R package) at hand:</p>
<ul>
<li>API key <span class="math inline">\(\approx\)</span> username (also called consumer key)</li>
<li>API key secret <span class="math inline">\(\approx\)</span> password (also called consumer secret)</li>
<li>Bearer token <span class="math inline">\(\approx\)</span> special access token (also called an authentication token)</li>
</ul>
<p>It is crucial to <strong>keep them private and not push them to GitHub or similar!</strong> Otherwise someone else could gain access to your API account. Instead, store them somewhere locally or directly within an environment variable. The package we’re going into detail later on this blog post is guiding you safely through this process.</p>
<p>However, in case you’re plan to use them in other applications, you can store your keys in different ways. The most common way in R is to add them to the <code>.Renviron</code> file. To do this with comfort, install the R package <code>usethis</code> and call its method <code>usethis::edit_r_environ()</code> which lets you edit the <code>.Renviron</code> in the home directory of your computer. In the following you can add tokens (or anything else you want to keep stored locally) using this format:</p>
<pre class="bash"><code>Key1=value1
Key2=value2
# ...</code></pre>
<p>After saving the file you can access values by calling <code>Sys.getenv("Key1")</code> within your R application. More best practices on managing your secrets can be found on the website <a href="https://bookdown.org/paul/apis_for_social_scientists/best-practices.html">APIs for Social Scientists</a>.</p>
<div id="postman-as-a-playground" class="section level5">
<h5>Postman-as-a-playground</h5>
<p><a href="https://www.postman.com/">Postman</a> is an easily accessible application to try out different queries, tokens, and more. Without any programming knowledge, you get the API results immediately. <a href="https://developer.twitter.com/en/docs/tutorials/postman-getting-started">Here</a> you can find an official tutorial to use Postman with the Twitter API.</p>
<p>However, there are several reasons why Postman cannot replace a package and programming code.</p>
<ul>
<li>Building flexible <a href="https://developer.twitter.com/en/docs/twitter-api/tweets/search/integrate/build-a-query">queries</a> (e.g., a list of users to retrieve tweets from)</li>
<li>Handling large responses which come split up during <a href="https://developer.twitter.com/en/docs/twitter-api/pagination">pagination</a></li>
<li>Handle <a href="https://developer.twitter.com/en/docs/twitter-api/rate-limits#v2-limits">rate limit restrictions</a></li>
<li>Transforming responses into manageable data structures (e.g., dataframe and comma-separated values)</li>
</ul>
<p>All of these tasks can be handled by a suitable package in your favorite programming language.</p>
</div>
<div id="which-package-should-i-choose" class="section level5">
<h5>Which package should I choose?</h5>
<p>It has to be noted that there are dozens of packages out there but only some of them already integrated the academic research track of the Twitter API. A selection of packages is listed below:</p>
<ul>
<li><strong><a href="https://github.com/cjbarrie/academictwitteR">academictwitteR</a> (R)</strong>:
The package offers customizable fucntions for all common v2 API endpoints. Additionally, it smoothly guides the developer through all critical steps (e.g. authentication or data processing) of the API interaction.</li>
<li><strong><a href="https://github.com/MaelKubli/RTwitterV2">RTwitterV2</a> (R)</strong>:
Although <code>RTwitterV2</code> as of now has less API endpoints included than <code>academictwitteR</code> it still is a valuable alternative which covers all basic functionalities.</li>
<li><strong><a href="https://github.com/ropensci/rtweet">rtweet</a> (R)</strong>:
<code>rtweet</code> does not support the academic research track yet, however it offers much basic functionality by using the previous API version. A dedicated Methods Bites blog post introducing <code>rtweet</code> in detail can be found <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/article/collecting-and-analyzing-twitter-using-r/">here</a>.</li>
<li><strong><a href="https://github.com/twitterdev/search-tweets-python/tree/v2">searchtweets-v2</a> (Python)</strong>:
This is the official package developed and maintained by Twitter available for Python. The library offers flexible functions which even handle very specialized requests but one has to dive deeper into the technical aspects of the API.</li>
<li><strong><a href="https://github.com/tweepy/tweepy">tweepy</a> (Python)</strong>:
<code>tweepy</code> is the most common package for Python and backed up by a large developer community. As a bonus, it includes many examples of how to use the various features offered by the package.</li>
</ul>
<p>Which package you pick should depend on your preferred programming language as well as whether the feature list of a package fits your research purpose.</p>
</div>
<div id="academictwitter-a-code-walkthrough-using-r" class="section level5">
<h5>academictwitteR: a code walkthrough using R</h5>
<p>In this blog post <a href="https://github.com/cjbarrie/academictwitteR">academictwitteR</a> <span class="citation">(Barrie and Ho <a href="#ref-BarrieHo2021" role="doc-biblioref">2021</a>)</span> (available for R) is used to demonstrate a simple scenario of retrieving tweets from German members of the parliament. The name <code>academictwitteR</code> is derived by the Twitter API academic research track for which it is developed for.</p>
<p>We will start by first loading all the needed R packages for the walkthrough:</p>
<details>
<p><summary>Code: R packages used in this tutorial</summary></p>
<pre class="r"><code>## Save package names as a vector of strings
pkgs &lt;- c(&quot;dplyr&quot;, &quot;academictwitteR&quot;, &quot;quanteda&quot;, &quot;purrr&quot;)

## Install uninstalled packages
lapply(pkgs[!(pkgs %in% installed.packages())], install.packages)

## Load all packages to library and adjust options
lapply(pkgs, library, character.only = TRUE)</code></pre>
</details>
<p><br>
After loading the packages, we need to share our API Bearer Token with <code>academictwitteR</code>. The following code will guide you through the process to store the key in an R-specific environment file (<code>.Renviron</code>) which we introduced earlier in this blog post:</p>
<pre class="r"><code>academictwitteR::set_bearer()</code></pre>
<pre class="bash"><code>## Instructions:
## ℹ 1. Add line: TWITTER_BEARER=YOURTOKENHERE to .Renviron 
##      on new line, replacing YOURTOKENHERE with  actual bearer token
## ℹ 2. Restart R</code></pre>
<p>After restarting R, everything is initialized and we can load a table of Twitter user IDs from German MPs into R:</p>
<pre class="r"><code>german_mps &lt;- read.csv(&quot;data/MP_de_twitter_uid.csv&quot;,
                       colClasses=c(&quot;user_id&quot;=&quot;character&quot;))
head(german_mps)</code></pre>
<pre class="bash"><code>##              user_id                   name party
## 1           44608858       Marc Henrichmann   CDU
## 2 819914159915667456      Stephan Pilsinger   CSU
## 3         1391875208   Markus Alexander Uhl   CDU
## 4          569832889 Sigmar Hartmut Gabriel   SPD
## ...</code></pre>
<p>To prevent replication issues with your work, it is recommended to use the Twitter user ID (e.g. <em>819914159915667456</em>) instead of the user handle (e.g. <em>@StephPilsinger</em>) as the user handle can be changed by the user over time. This would result in not being able anymore to recrawl tweets of these users. In case you only have access to the handle, there is a v2 API endpoint to receive a user object from a handle: <em>/2/users/by/username/:username</em></p>
<p>Databases and lists of Twitter users can be retrieved from the following sources:</p>
<ul>
<li>The Twitter Parliamentarian Database <span class="citation">(Vliet, Törnberg, and Uitermark <a href="#ref-vanVliet2020TheTP" role="doc-biblioref">2020</a>)</span></li>
<li>Public Twitter lists (e.g. <a href="https://twitter.com/i/lists/912241909002833921" class="uri">https://twitter.com/i/lists/912241909002833921</a>)<a href="#fn6" class="footnote-ref" id="fnref6"><sup>6</sup></a></li>
<li>legislatoR R Package <span class="citation">(Göbel and Munzert <a href="#ref-göbel_munzert_2021" role="doc-biblioref">2021</a>)</span></li>
</ul>
<p>Afterward, we are ready to crawl our first tweets using a simple wrapper function (<code>get_tweets_from_user()</code>) asking for a single <code>user_id</code>. <code>get_all_tweets()</code>, which is called inside this function is the heart of our code. It manages the generation of queries for the API, working with rate limits as well as storing the data in <code>JSON</code>-files (which can be transformed later).</p>
<p>In case you look for specific content, tweet types, or even topics, you can add another parameter to the package function: <code>query</code>. It allows you to narrow down your search by using specific strings. To give an example, one could look for English retweets containing the keywords <em>putin</em> or <em>selenskij</em> having a geo-location attached. This can be achieved by simply assigning the following string to the <code>query</code> parameter:</p>
<pre class="r"><code>(putin OR selenskyj) -is:retweet lang:en has:geo</code></pre>
<p>Beyond that, there exist many more parameters to individualize the crawling method. All of them are documented in the <a href="https://cran.r-project.org/web/packages/academictwitteR/academictwitteR.pdf">official <code>academictwitteR</code> CRAN documentation</a> of the package.
However, in this tutorial I only restrict my search to a Twitter user ID as well as a start and end date for the tweets we are interested in:</p>
<pre class="r"><code># function to retrieve tweets in a specific time period of a single user
# (list of user IDs would be possible but one should keep
# the max. query string of 1024 characters in mind)

get_tweets_from_user &lt;- function(user_id) {
  # Another option is to add &quot;query&quot; parameter
  academictwitteR::get_all_tweets(
    users = user_id,
    start_tweets = &quot;2021-01-01T00:00:00Z&quot;,
    end_tweets = &quot;2021-09-30T00:00:00Z&quot;,
    data_path = &quot;data/raw/&quot;,
    n = 100)
}</code></pre>
<p>The function is then called for each <code>user_id</code> in the dataframe by using <code>walk()</code> from the <a href="https://purrr.tidyverse.org"><code>purrr</code></a> package (the <code>purrr</code> package allows you to work with functions and vectors):</p>
<pre class="r"><code>purrr::walk(german_mps[[&quot;user_id&quot;]], get_tweets_from_user)</code></pre>
<p>To import the tweets into a workable format, call <code>bind_tweets()</code> from <code>academictwitteR</code>. It consolidates all available files in the given <code>data_path</code> and organizes them into the requested format (in our case tidy). In addition, only a relevant fraction of columns is selected in the code below by using <code>select()</code> from the <a href="https://dplyr.tidyverse.org"><code>dplyr</code></a>-package.</p>
<pre class="r"><code># concatenate all retrieved tweets into one dataframe and select which columns
# should be kept
# Another option: set parameter &quot;user&quot; to TRUE to retrieve user information
tweets_df &lt;- academictwitteR::bind_tweets(data_path = &quot;data/raw/&quot;,
                                          output_format = &quot;tidy&quot;) %&gt;%
  dplyr::select(
    tweet_id,
    text,
    author_id,
    user_username,
    created_at,
    sourcetweet_type,
    sourcetweet_text,
    lang
  )</code></pre>
<p>Finally, I store the tweets in a single .csv-file:</p>
<pre class="r"><code>write.csv(tweets_df, &quot;data/raw/tweets_german_mp.csv&quot;, row.names = FALSE)</code></pre>
<p>Congratulations! You successfully applied to the academic research track, got admitted, and crawled a selection of tweets using the R package <code>academictwitteR</code>.</p>
</div>
</div>
<div id="preparing-for-methods-working-with-textual-data" class="section level3">
<h3>Preparing for methods: working with textual data</h3>
<p>You are now ready to move on! The usual steps applied to textual data (lowercasing, stopwords removal, stemming, …) depending on the method at hand can be used for pre-processing your tweets. Additional fine-tuning of these steps could involve the removal of, e.g. party IDs, URLs, user mentions or similar. Such steps can be easily done by using regular expressions (<code>regex</code>). Regular expressions are used to extract patterns in texts which then can be used for further analysis, removal or replacement. Many tutorials available on the web (e.g. <a href="https://regexone.com/">RegexOne interactive tutorial</a>) make it straightforward to learn how to bring such expressions into action within your domain.</p>
<p>The following code provides a first starting point for applying pre-processing steps. The code relies on the package <a href="https://quanteda.io"><code>quanteda</code></a> which is an R package that is often used when working with text data in R. If you want to dive deeper into <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/article/advancing-text-mining/">text mining</a> and <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/article/quantitative-analysis-of-political-text/">text analysis</a>, Methods Bites has more blog posts on these topics.</p>
<pre class="r"><code>tweet_corpus &lt;- quanteda::corpus(tweets_df[[&quot;text&quot;]],
                                 docnames = tweets_df[[&quot;tweet_id&quot;]])</code></pre>
<p>The code first transforms the dataframe of tweets into another data format, <code>corpus</code>, keeping the <code>tweet_id</code> as an identifier attached to each tweet <code>text</code>.
Having the tweets in the corpus format makes it easy to apply pre-processing steps after tokenizing. The following list shows a selection of common methods. However, it is important that the decision, on which methods are applied, heavily relies on the following text processing approach:</p>
<ul>
<li><code>remove_punct</code>: removes all punctuation</li>
<li><code>remove_numbers</code>: removes all numbers</li>
<li><code>dfm_tolower()</code>: applies lowercasing</li>
<li><code>dfm_remove(stopwords("german"))</code>: removes German stopwords which occur very frequently</li>
<li><code>dfm_wordstem(language = "german")</code>: applies German stemming (e.g., <em>wurden</em> <span class="math inline">\(\rightarrow\)</span> <em>wurd</em>)</li>
</ul>
<pre class="r"><code># &quot;2020 wurden in Berlin ca. 18.800 Miet-
# in Eigentumswohnungen umgewandelt. #Umwandlungsverbot&quot;
dfm &lt;-
  quanteda::dfm(tweet_corpus %&gt;%
                  quanteda::tokens(
                    remove_punct = TRUE,
                    remove_numbers = TRUE)) %&gt;%
  quanteda::dfm_tolower() %&gt;% # removes capitialization
  quanteda::dfm_remove(
    stopwords(&quot;german&quot;)) %&gt;% # removes German stopwords
  quanteda::dfm_wordstem(
    language = &quot;german&quot;) # transforms words to their German wordstems
# &quot;wurd berlin ca miet- eigentumswohn umgewandelt #umwandlungsverbot&quot;</code></pre>
<p>The function <code>dfm</code> (called above) returns a <a href="https://quanteda.io/reference/dfm.html">sparse document-feature matrix</a> which could be a fruitful starting point for first-word frequency analysis:</p>
<pre class="bash"><code>head(dfm)

## Document-feature matrix of: 6 documents, 87 features (79.77% sparse)
## and 0 docvars.
##                       features
## docs                  leb plotzlich mehr schablon gut bos pass 😉 #esk #miet
##   44608858            1   1         1    1        1   1   1    1  1    1
##   819914159915667456  0   0         0    0        0   0   0    0  0    0
##   1391875208          0   0         0    0        0   0   0    0  0    0
##   569832889           0   0         1    0        0   0   0    0  0    1
## ...</code></pre>
<p>You are finally at the step of applying further methods to tackle your research question and getting deeper insights into your crawled tweets. There is much more to explore: You can find further <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/tags/text-as-data/">text-as-data tutorials</a> on our blog.</p>
</div>
<div id="reproducibility-of-research-based-on-twitter-data" class="section level3">
<h3>Reproducibility of research based on Twitter data</h3>
<p>As reproducible results are one of the major requirements of research projects, it has to be discussed how this could affect your work with Twitter data. The Twitter development agreement includes a clear statement of what researchers are allowed to publish along with their work:</p>
<blockquote>
<font size="-1">
“<em>Academic researchers are permitted to distribute an unlimited number of Tweet IDs and/or User IDs if they are doing so on behalf of an academic institution and for the sole purpose of non-commercial research. For example, you are permitted to share an unlimited number of Tweet IDs for the purpose of enabling peer review or validation of your research.</em>”<a href="#fn7" class="footnote-ref" id="fnref7"><sup>7</sup></a>
</font>
</blockquote>
<p>This means that the content of tweets must not be shared publicly. As tweets can be deleted or accounts can be suspended this certainly states an issue for subsequent researchers attempting to replicate the findings as they won’t be able to recrawl such tweets via the API. However, there are also platforms like <a href="https://polititweet.org/">polititweet.org</a>, which track <em>public figures</em> and based on that justify the publication even of deleted tweets:</p>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:unnamed-chunk-3"></span>
<img src="../../../../../article/twitter-research-track/politweet.png" alt="&lt;a href=&quot;https://polititweet.org&quot;&gt;polititweet.org&lt;/a&gt; section of the landing page" width="100%" />
<p class="caption">
Figure 3: <a href="https://polititweet.org">polititweet.org</a> section of the landing page
</p>
</div>
<div style="text-align: right">
<p><sub><sup>
Source: <a href="https://polititweet.org/">polititweet.org landing page</a>
</sub></sup></p>
</div>
<p>To conclude, this makes the decision of how to share what kind of data not easier. But one still has to ensure to choose the best available option to share his or her data without violating the Twitter rules which are as of today to at least share tweet IDs amongst the community.</p>
</div>
<div id="conclusion" class="section level3">
<h3>Conclusion</h3>
<p>This blog post provides a first glimpse into the academic research track Twitter API and the information richness of Twitter data. As there certainly will be further updates and changes to the API in the future, there are plenty of easy-to-use packages that build on active user communities. The community is there to keep the packages updated accordingly to the current Twitter API version. While there exist a lot of powerful packages to tackle the data gathering step, researchers still need to think carefully about how to further process the crawled information depending on their research question and method as well as how to make their research accessible to the community in an <em>open science</em> approach.</p>
</div>
<div id="further-readings" class="section level3">
<h3>Further readings <a name="furtherreadings"></a></h3>
<ul>
<li><a href="https://developer.twitter.com/en/docs">Official Twitter API Documentation</a></li>
<li><a href="https://doi.org/10.1007/978-1-4614-9372-3">Shamanth Kumar, Fred Morstatter, and Huan Liu. 2013. Twitter Data Analytics. Springer Publishing Company, Incorporated.</a></li>
<li><a href="https://socialsciencedatalab.mzes.uni-mannheim.de/article/collecting-and-analyzing-twitter-using-r/">Denis Cohen and Simon Kühne. 2019. Collecting and Analyzing Twitter Data Using R. Methods Bites.</a></li>
</ul>
</div>
<div id="about-the-author" class="section level3">
<h3>About the author</h3>
<p>Andreas Küpfer <a href="mailto:andreas.kuepfer@tu-darmstadt.de"><svg aria-hidden="true" role="img" viewBox="0 0 512 512" style="height:1em;width:1em;vertical-align:-0.125em;margin-left:auto;margin-right:auto;font-size:inherit;fill:currentColor;overflow:visible;position:relative;"><path d="M502.3 190.8c3.9-3.1 9.7-.2 9.7 4.7V400c0 26.5-21.5 48-48 48H48c-26.5 0-48-21.5-48-48V195.6c0-5 5.7-7.8 9.7-4.7 22.4 17.4 52.1 39.5 154.1 113.6 21.1 15.4 56.7 47.8 92.2 47.6 35.7.3 72-32.8 92.3-47.6 102-74.1 131.6-96.3 154-113.7zM256 320c23.2.4 56.6-29.2 73.4-41.4 132.7-96.3 142.8-104.7 173.4-128.7 5.8-4.5 9.2-11.5 9.2-18.9v-19c0-26.5-21.5-48-48-48H48C21.5 64 0 85.5 0 112v19c0 7.4 3.4 14.3 9.2 18.9 30.6 23.9 40.7 32.4 173.4 128.7 16.8 12.2 50.2 41.8 73.4 41.4z"/></svg></a> <a href="https://twitter.com/ankuepfer"><svg aria-hidden="true" role="img" viewBox="0 0 512 512" style="height:1em;width:1em;vertical-align:-0.125em;margin-left:auto;margin-right:auto;font-size:inherit;fill:currentColor;overflow:visible;position:relative;"><path d="M459.37 151.716c.325 4.548.325 9.097.325 13.645 0 138.72-105.583 298.558-298.558 298.558-59.452 0-114.68-17.219-161.137-47.106 8.447.974 16.568 1.299 25.34 1.299 49.055 0 94.213-16.568 130.274-44.832-46.132-.975-84.792-31.188-98.112-72.772 6.498.974 12.995 1.624 19.818 1.624 9.421 0 18.843-1.3 27.614-3.573-48.081-9.747-84.143-51.98-84.143-102.985v-1.299c13.969 7.797 30.214 12.67 47.431 13.319-28.264-18.843-46.781-51.005-46.781-87.391 0-19.492 5.197-37.36 14.294-52.954 51.655 63.675 129.3 105.258 216.365 109.807-1.624-7.797-2.599-15.918-2.599-24.04 0-57.828 46.782-104.934 104.934-104.934 30.213 0 57.502 12.67 76.67 33.137 23.715-4.548 46.456-13.32 66.599-25.34-7.798 24.366-24.366 44.833-46.132 57.827 21.117-2.273 41.584-8.122 60.426-16.243-14.292 20.791-32.161 39.308-52.628 54.253z"/></svg></a> is a graduate of the Mannheim Master in Data Science and a doctoral researcher at the Technical University of Darmstadt. His interdisciplinary research interests include text-as-data, applying machine learning technologies, and substantial inference in the fields of political communication and political competition.</p>
</div>
<div id="references" class="section level3 unnumbered">
<h3>References</h3>
<div id="refs" class="references">
<div id="ref-barberá_2015">
<p>Barberá, Pablo. 2015. “Birds of the Same Feather Tweet Together: Bayesian Ideal Point Estimation Using Twitter Data.” <em>Political Analysis</em> 23 (1): 76–91. <a href="https://doi.org/10.1093/pan/mpu011">https://doi.org/10.1093/pan/mpu011</a>.</p>
</div>
<div id="ref-BarrieHo2021">
<p>Barrie, Christopher, and Justin Chun-ting Ho. 2021. “AcademictwitteR: An R Package to Access the Twitter Academic Research Product Track V2 Api Endpoint.” <em>Journal of Open Source Software</em> 6 (62): 3272. <a href="https://doi.org/10.21105/joss.03272">https://doi.org/10.21105/joss.03272</a>.</p>
</div>
<div id="ref-göbel_munzert_2021">
<p>Göbel, Sascha, and Simon Munzert. 2021. “The Comparative Legislators Database.” <em>British Journal of Political Science</em>, 1–11. <a href="https://doi.org/10.1017/S0007123420000897">https://doi.org/10.1017/S0007123420000897</a>.</p>
</div>
<div id="ref-NGUYEN2021100922">
<p>Nguyen, Thu T., Shaniece Criss, Eli K. Michaels, Rebekah I. Cross, Jackson S. Michaels, Pallavi Dwivedi, Dina Huang, et al. 2021. “Progress and Push-Back: How the Killings of Ahmaud Arbery, Breonna Taylor, and George Floyd Impacted Public Discourse on Race and Racism on Twitter.” <em>SSM - Population Health</em> 15: 100922. <a href="https://doi.org/https://doi.org/10.1016/j.ssmph.2021.100922">https://doi.org/https://doi.org/10.1016/j.ssmph.2021.100922</a>.</p>
</div>
<div id="ref-doi:10.1177/1354068820957960">
<p>Sältzer, Marius. 2022. “Finding the Bird’s Wings: Dimensions of Factional Conflict on Twitter.” <em>Party Politics</em> 28 (1): 61–70. <a href="https://doi.org/10.1177/1354068820957960">https://doi.org/10.1177/1354068820957960</a>.</p>
</div>
<div id="ref-cruz2022">
<p>Valle-Cruz, David, Vanessa Fernandez, Asdrubal Lopez-Chau, and Rodrigo Sandoval Almazan. 2022. “Does Twitter Affect Stock Market Decisions? Financial Sentiment Analysis During Pandemics: A Comparative Study of the H1n1 and the Covid‐19 Periods.” <em>Cognitive Computation</em> 14 (January). <a href="https://doi.org/10.1007/s12559-021-09819-8">https://doi.org/10.1007/s12559-021-09819-8</a>.</p>
</div>
<div id="ref-vanVliet2020TheTP">
<p>Vliet, Livia van, Petter Törnberg, and Justus Uitermark. 2020. “The Twitter Parliamentarian Database: Analyzing Twitter Politics Across 26 Countries.” <em>PLoS ONE</em> 15.</p>
</div>
</div>
</div>
<div class="footnotes">
<hr />
<ol>
<li id="fn1"><p>Meta data serves as a explanatory information such as topical indicators or the language of the tweet which should explain and enrich the actual tweet, image or main object retrieved from the API<a href="#fnref1" class="footnote-back">↩</a></p></li>
<li id="fn2"><p>Queries are filter operators to narrow down the amount of tweets which should be retrieved<a href="#fnref2" class="footnote-back">↩</a></p></li>
<li id="fn3"><p>API stands for Application Programming Interface and allows, simply speaking, the communication between software.<a href="#fnref3" class="footnote-back">↩</a></p></li>
<li id="fn4"><p><a href="https://developer.twitter.com/en/products/twitter-api/academic-research/application-info">Twitter Developer Platform product page</a><a href="#fnref4" class="footnote-back">↩</a></p></li>
<li id="fn5"><p>There is a maximum of tweets which can be retrieved via the API which gets resetted once in a month.<a href="#fnref5" class="footnote-back">↩</a></p></li>
<li id="fn6"><p>Use such lists with caution as they may do not come from verified sources.<a href="#fnref6" class="footnote-back">↩</a></p></li>
<li id="fn7"><p>You can find a detailed description of the content redistribution of Twitter data in the <a href="https://developer.twitter.com/en/developer-terms/policy#:~:text=Academic%20researchers%20are%20permitted%20to,purpose%20of%20non%2Dcommercial%20research.">official developer policies</a>.<a href="#fnref7" class="footnote-back">↩</a></p></li>
</ol>
</div>
]]>
      </description>
    </item>
    
    <item>
      <title>Survey data collection from start to finish</title>
      <link>https://socialsciencedatalab.mzes.uni-mannheim.de/article/survey-data-collection/</link>
      <pubDate>Mon, 11 Apr 2022 01:00:00 +0100</pubDate>
      
      <guid>https://socialsciencedatalab.mzes.uni-mannheim.de/article/survey-data-collection/</guid>
      <description><![CDATA[
        </p>
<!-- Article text -->
<p>Surveys have long been a staple of social science research on individuals’ attitudes and behaviors. In recent years, however, we have witnessed a strong shift from secondary analyses of large general social surveys toward smaller, more targeted primary data collections. This development has been accompanied by the increasing availability of affordable and easy-to-implement surveys using online access panels. While the entry barriers to original survey-based research are now likely lower than ever before, it still comes with notable methodological, administrative, and logistic challenges. To help aspiring survey researchers navigate this process, this <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/categories/instructionals/">Methods Bites Instructional</a> by <a href="https://twitter.com/hellyer_josh">Joshua Hellyer</a> (MZES, University of Mannheim) provides a comprehensive guide to survey data collection with online access panels.</p>
<p>Reading this blog post, you will learn:</p>
<ul>
<li>The steps involved in planning and conducting a survey experiment</li>
<li>Best practices in reproducibility</li>
<li>Tips on completing your application for ethical approval</li>
<li>Pros and cons of working with an online access panel</li>
</ul>
<p><em>Note:</em> This blog post provides a summary of Joshua’s and <a href="https://twitter.com/JohannaGereke">Johanna Gereke</a>’s input talk “Survey data collection from start to finish: Designing &amp; executing reproducible research with an online access panel” in the <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/page/events/">MZES Social Science Data Lab</a> in Spring 2022. The original workshop materials, including slides and scripts, are available from our <a href="https://github.com/SocialScienceDataLab/survey-data-collection">GitHub</a>.
A live recording of the talk is available on our <a href="https://youtu.be/NZhzWVYeCLI">YouTube Channel</a>.</p>
<div id="overview" class="section level3">
<h3>Overview</h3>
<ol style="list-style-type: decimal">
<li><a href="#introduction"><strong>Introduction</strong></a></li>
<li><a href="#research-design"><strong>Research Design</strong></a>
<ol style="list-style-type: decimal">
<li><a href="#sampling-and-power">Sampling and Power</a></li>
<li><a href="#ethical-approval">Ethical Approval</a></li>
</ol></li>
<li><a href="#pre-registration"><strong>Pre-Registration</strong></a></li>
<li><a href="#data-collection"><strong>Data Collection</strong></a>
<ol style="list-style-type: decimal">
<li><a href="#online-access-panels">Online Access Panels</a></li>
<li><a href="#programming">Programming</a></li>
<li><a href="#pilot-testing-and-data-collection">Pilot Testing and Data Collection</a></li>
</ol></li>
<li><a href="#data-and-code-sharing"><strong>Data and Code Sharing</strong></a></li>
<li><a href="#conclusion"><strong>Conclusion</strong></a></li>
</ol>
</div>
<div id="introduction" class="section level3">
<h3>Introduction</h3>
<p>In this blog post, I will walk you through the process of fielding your own survey experiment using an online access panel. Online access panels are pre-selected groups of internet users who are paid to participate in various surveys, often including market research as well as scientific studies. You may have heard of panel providers like <a href="https://www.bilendi.de/">Bilendi/Respondi</a>, <a href="https://www.dynata.com/">Dynata</a>, <a href="https://www.kantarpublic.com/de">Kantar</a>, and <a href="https://yougov.de/">YouGov</a> that are frequently used in social science research. They have become popular among social scientists because they are a relatively fast, easy, and cheap way to reach a sample of the general population, but there are certainly pros and cons that you should consider, as I will describe in this post.</p>
<p>Throughout this post, I will use a recent data collection as an example, a survey experiment that I conducted with <a href="https://johannagereke.com">Johanna Gereke</a> (whose expertise I rely on extensively for this piece) and several other colleagues about demographic threat and how it affects group boundaries in the German context <span class="citation">(Gereke et al. <a href="#ref-gereke_demographic_2022" role="doc-biblioref">2022</a>)</span>. In some ways, this project is unique: it was completed as part of a replication seminar at the University of Mannheim, in which we worked with 12 Master’s and PhD students who helped plan the survey, analyze data, and write the first draft of our paper. But in many other ways, our experiences should apply to a wide range of potential surveys, and I hope that you will find it useful no matter the size of your research team or the topic you plan to study.</p>
<p>So, what steps are required to plan and execute a survey experiment? Your experience may vary depending on the complexity of your planned survey as well as the population you plan to sample from, but you can get an overview of our recent experience in the timeline shown in the figure below. This project was completed relatively quickly, going from initial conception to journal submission in just over 6 months. This timeline was intentionally fast, as we wanted the students to experience the whole data collection process in a single semester, but this is also an example of just how fast you can collect data using an online access panel. As you can see, the most time-consuming part of our research was designing the survey, but also note that ethical approval can take some time (more on this later). While your workflow may look a bit different, I would generally recommend proceeding in a similar order. In particular, ethical approval must come early in the process as you need this to collect any data. Tackling elements like this (and also procurement) that are out of your control early on can save you from delays later.</p>
<div class="figure" style="text-align: center"><span style="display:block;" id="fig:graphic1"></span>
<img src="/../../../../../article/survey-data-collection_files/figure-html/graphic1-1.png" alt="Project timeline." width="100%" />
<p class="caption">
Figure 1: Project timeline.
</p>
</div>
</div>
<div id="research-design" class="section level3">
<h3>Research Design</h3>
<p>For our seminar, we chose to first replicate and then extend Maria Abascal’s recent work on group boundaries in the U.S. context <span class="citation">(Abascal <a href="#ref-abascal_contraction_2020" role="doc-biblioref">2020</a>)</span>. She finds that when white Americans are confronted with information about demographic threat, i.e. that they will no longer make up a majority of the U.S. population, they are more exclusive about who they see as “White”. We tested whether the same effect could be found in Germany, randomly assigning each of our participants to see one of two demographic projections that either predicted that people with a migration background would make up an equal share of the population in the future (treatment), or that they would remain a minority (control). After seeing this, participants rated pictures of 18 people as having a migration background or not. We hypothesized that following Abascal, the demographic threat treatment would make people more likely to classify people who appear ethnically ambiguous as having a migration background.</p>
<div id="questionnaire" class="section level5">
<h5>Questionnaire</h5>
<p>Because our experiment was based on previous research, designing our survey was perhaps more straightforward than it otherwise would have been. However, adapting the questionnaire to the German context still required a lot of consideration.</p>
<p>When writing your questionnaire, it is always good to start by looking in the literature for examples of how other people have asked similar questions. While we had to write some of the more unusual questions ourselves, we could borrow some language from other German-language surveys (like the <a href="https://www.diw.de/en/diw_01.c.615551.en/research_infrastructure__socio-economic_panel__soep.html">Socio-Economic Panel (SOEP)</a> or the <a href="https://www.gesis.org/en/allbus/allbus-home">German General Social Survey (ALLBUS)</a>) for more common demographic questions. This will also give you an idea of what scales are commonly used for your topic. It is also important to ensure that every question provides a complete range of mutually exclusive answer options, although you may occasionally want to constrain the options available. For example, we considered adding an “I don’t know” option when asking people to consider whether someone had a migration background, but we instead decided to force them to choose between “yes” and “no” because this is more useful for our purposes.</p>
<p>Another important consideration is the order of the questions and answer options. This may seem tedious, but order can have serious implications for the results of your study. In our case, we chose to ask questions about respondents’ immigration attitudes before showing them the treatment because we were concerned that the treatment could actually change someone’s attitudes in the short term. On the other hand, asking a question like this at the beginning of a survey might prime participants to think about immigration, which might not be ideal if you were trying to conceal the true intent of your study. These kind of effects are called “order effects” in survey terminology, and it’s worth carefully considering how they might affect your project <span class="citation">(Strack <a href="#ref-strack_order_1992" role="doc-biblioref">1992</a>)</span>. Overall, questionnaire design is an art and I cannot describe every step in detail here, so you should consult resources like the <a href="https://www.gesis.org/en/gesis-survey-guidelines/instruments/questionnaire-design">GESIS Survey Guidelines</a> for more information as you write your own.</p>
</div>
<div id="sampling-and-power" class="section level5">
<h5>Sampling and Power</h5>
<p>While you plan your research, you will also want to consider your target sample. Here, you not only need to decide on the population you hope to sample from, such as Germans without migration background (in our case), but also the sample size you need in order to make statistically sound claims. Calculating sample size may require you to balance many practical and theoretical considerations, as described in this <a href="https://psyarxiv.com/9d3yf/">recent paper</a> <span class="citation">(Lakens <a href="#ref-lakens_sample_2021" role="doc-biblioref">2021</a>)</span>. Some of these considerations might be how large of a sample you can afford to reach, effect sizes reported in previous research, and what the smallest substantively interesting effect size would be.</p>
<p>Ideally, you would determine your sample size with a power calculation, although these calculations can be complicated because of the many <a href="https://stats.oarc.ucla.edu/other/mult-pkg/seminars/intro-power/">factors</a> that influence power. Some calculations may require parameters like expected effect size that are difficult to estimate a priori, but pilot testing and previous research in your field can help you make an informed guess. I cannot cover all of the techniques you could use in such a short blog post, so you should explore previous work using similar designs and see how they have justified sample sizes for an idea of what methods might be most appropriate. When it comes to software, there are many options, fortunately including several that are free of charge. First, statistical software often includes built-in power commands, like “<a href="https://www.stata.com/manuals/pss-2power.pdf">power</a>” in Stata or “<a href="https://cran.r-project.org/web/packages/pwr/pwr.pdf">pwr</a>” in R. Beyond these options, there are free programs that were created specifically for power calculations, like <a href="https://www.psychologie.hhu.de/arbeitsgruppen/allgemeine-psychologie-und-arbeitspsychologie/gpower.html">G*Power</a> and <a href="https://sites.google.com/site/optimaldesignsoftware/home">Optimal Design</a>.</p>
<p>There are also alternatives to traditional methods of power calculation, which can be especially useful for complex experimental designs. Because we were conducting a replication and had access to the original data source, we used a simulation-based technique. We duplicated Abascal’s data to make a large population from which we then drew many samples, assuming that the effect size she finds is the true effect. For a power of 0.8, treatment effects in 80% of these samples should be significant. (For more information on this approach, see <span class="citation">Arnold et al. (<a href="#ref-arnold_simulation_2011" role="doc-biblioref">2011</a>)</span>.) A drawback of this method is that it required us to make assumptions about the effect size and variance based on U.S. data, which may not accurately predict responses in the German context.</p>
<p>Determining your target population, on the other hand, is mainly based on theoretical considerations. We pursued a sample that was representative of the adult German population without migration background in terms of gender, age, education, employment status, and region (East/West). To ensure balance on age and gender, we implemented quotas by gender and age, meaning that our sample would include a fixed number of men and women in each of three age groups. Note that your survey provider may be able to help you restrict your sample to certain demographic groups, but you may also need to implement screening questions to ensure the right sample composition. In our case, we had to screen for migration background and program quotas ourselves.</p>
<p>It is also important to note that despite our efforts to create a representative sample, samples drawn from online access panels may suffer from selection bias <span class="citation">(Bethlehem <a href="#ref-bethlehem_selection_2010" role="doc-biblioref">2010</a>)</span>. Participants must have Internet access, find the provider’s website, and opt into taking the survey, and people who do this may differ from those who do not on variables both observed and unobserved. While some have proposed methods like propensity scores to correct for this bias <span class="citation">(Schonlau et al. <a href="#ref-schonlau_selection_2009" role="doc-biblioref">2009</a>)</span>, this should be considered a weakness of online access panels, especially when extrapolating results to the general population.</p>
</div>
</div>
<div id="ethical-approval" class="section level3">
<h3>Ethical Approval</h3>
<p>Once you have generally decided on your research design and your survey population, it is a good idea to start your application for ethical approval.</p>
<p>Based on our experience, the application itself is generally not too time-consuming unless your design brings up many ethical issues, but the approval process can take several weeks or even months depending on your institution. It has historically taken us about 4-6 weeks to get approval at the University of Mannheim. You should definitely plan ahead, because it is critical that you have this approval before beginning the data collection. This is not only to ensure that you are acting within the ethical and legal guidelines of your institution and your field of research, but also because many journals require ethical approval for any work they publish.</p>
<p>When do you need to apply for ethical approval? Your institution’s rules may vary, but the <a href="https://www.uni-mannheim.de/en/about/organization/bodies-and-committees/committees-and-councils/ethics-committee/">University of Mannheim</a> requires it for any research on humans that:</p>
<ul>
<li>involves any personal or personally identifiable data</li>
<li>deceives subjects</li>
<li>involves psychological or physical health risks</li>
<li>triggers strong emotions or asks about traumatic experiences</li>
<li>manipulates subjects’ self-image</li>
<li>involves minors, or</li>
<li>presents risks to human dignity, life, health, and peaceful coexistence.</li>
</ul>
<p>In our case, the first two bullets applied. Any demographic information about a respondent counts as personal data, even if it could not be used to identify any specific person, so your research will almost certainly collect some personal data. Additionally, we employed deception by withholding the purpose of the study from respondents until the end of the survey, which we had to address by debriefing all participants after the survey was completed. If you are unsure whether your study requires ethical approval or not, it is a good idea to ask the staff person associated with your institution’s ethics committee. You should also look at their website, which might have additional information to help you plan your application. (Researchers in Mannheim can find an application checklist and other information <a href="https://www.uni-mannheim.de/en/about/organization/bodies-and-committees/committees-and-councils/ethics-committee/">here</a>.)</p>
<p>Your application should explain the ethical issues that will arise as part of conducting your research. You should explain why these features (like deception or data collection) are necessary for your design, and explain how you are minimizing any potential harms associated with them. You will also want to include as much information about your survey as you can, possibly including any planned debriefing text as well as the invitation to the survey, and maybe even the entire questionnaire depending on the expectations of your institution.</p>
<p>Another area that you may need to address in your application is data protection. This is perhaps especially important to researchers working in the European Union due to the <a href="https://gdpr-info.eu">General Data Protection Regulation (GDPR)</a> passed in 2016. I am not a lawyer and cannot provide a legal opinion on what regulations may apply to your research, but luckily the <a href="https://www.berd-nfdi.de">BERD @ NFDI</a> in Mannheim has put together a useful tool called the <a href="https://www.berd-nfdi.de/servicestools/legal-questions-in-data-science/">Interactive Virtual Assistant</a> that can give you more information on this topic (unfortunately only in German for now, but an English version is planned). In general, you should try to limit the amount of personal data you collect as much as possible, and make a plan to store this data securely. You will also want to avoid collecting any information that could be used to uniquely identify any of your respondents (or “personally identifiable information”, also called PII). Note that PII is not limited to names or addresses, but also includes combinations of variables that could identify a single person (such as a combination of ethnicity, age, and postal code). The ethics committee may be especially critical of designs that require the collection of such information.</p>
</div>
<div id="pre-registration" class="section level3">
<h3>Pre-Registration</h3>
<p>Once you have received ethical approval and have firmly decided on your design, you should consider pre-registering your research. A pre-registration is simply a report of your hypotheses, data source, and planned research design that you write and upload to a repository before starting data collection. This is meant to prevent practices like selective reporting and p-hacking, and to disclose confirmatory and exploratory analyses, which makes your decision-making more transparent to those who read and evaluate your work <span class="citation">(Nosek et al. <a href="#ref-nosek_preregistration_2018" role="doc-biblioref">2018</a>)</span>. This is certainly not a required step, but one that can help you plan and motivate your research, and that promotes principles of open science. One important thing to note is that pre-registering your design does not forbid you from performing exploratory analyses or even making design changes later on; it merely requires that you disclose that you have done so <span class="citation">(Simmons, Nelson, and Simonsohn <a href="#ref-simmons_pre-registration_2021" role="doc-biblioref">2021</a>)</span>.</p>
<p>If you would like to pre-register your study, you can find several useful templates on <a href="https://osf.io/zab38/">OSF</a>. A pre-registration will only require you to briefly describe your hypotheses, design, and planned analysis, information that you may have already collected for your ethics application. Once you have filled out your template, you can upload your pre-registration to a repository like <a href="https://osf.io/">OSF</a> or <a href="https://aspredicted.org/">AsPredicted</a> (among others), where it will be given a timestamp to verify when it was posted. Once posted, your pre-registration can be embargoed, making it invisible for several months while you complete your research. Note that you can also <a href="https://help.osf.io/hc/en-us/articles/360019930333-Create-a-View-only-Link-for-a-Project">create an anonymized link</a> to your pre-registration if you would like to include it in a manuscript submitted for peer review. If you are looking for an example, here is <a href="https://osf.io/2ygjd">our pre-registration</a>, although note that ours uses the “Replication Recipe” template designed specifically for replications.</p>
</div>
<div id="data-collection" class="section level3">
<h3>Data Collection</h3>
<div id="online-access-panels" class="section level5">
<h5>Online Access Panels</h5>
<p>You have planned your research, you received ethical approval, and you pre-registered your study – now you can finally begin collecting data. But we are still missing one key element: where is the data coming from? This is where an online access panel comes in. Online access panels have been used to study a variety of topics in the social sciences in recent years, from <a href="https://ejpr.onlinelibrary.wiley.com/doi/full/10.1111/1475-6765.12401">support for democracy</a> to <a href="https://journals.sagepub.com/doi/full/10.1177/0010414017740590">immigration attitudes</a>, among many other examples. They have also been a popular method of studying <a href="https://link.springer.com/article/10.1007/s11150-020-09529-4">reactions</a> to the COVID-19 pandemic, as they are relatively cheap, quick to deploy, and do not require in-person contact. Despite these benefits, there are also drawbacks that you should be aware of before collecting data. I have listed some of the main pros and cons below.</p>
<p>Pros:</p>
<ul>
<li>Fast data collection</li>
<li>Relatively low cost</li>
<li>Possibility of accessing a representative (non-student) sample</li>
<li>International data collection possible</li>
</ul>
<p>Cons:</p>
<ul>
<li>Only include internet users</li>
<li>Self-selection into participation can lead to biased population estimates <span class="citation">(Bethlehem <a href="#ref-bethlehem_selection_2010" role="doc-biblioref">2010</a>)</span></li>
<li>Some users may respond carelessly to get incentives</li>
<li>Potential loss of naivete (but not as much as MTurk, see <span class="citation">Chandler et al. (<a href="#ref-chandler_online_2019" role="doc-biblioref">2019</a>)</span>)</li>
<li>May be difficult to achieve large sample of minority groups</li>
</ul>
<p>If you decide that an online access panel might be a good fit for your research, your first step will be to compare prices and collect different offers.</p>
<p>Your institution will likely have rules about how many cost estimates you should collect (we needed three). You will want to contact several firms and provide them with information about your desired sample (especially size and any sample restrictions you want to impose), the timing you are hoping for, and the characteristics of your survey (programmed yourself or by the firm, mobile and/or web, time to complete). Make sure you request a large enough sample that you can run some pilot tests as well (more on this shortly). The providers you can choose from will depend on your desired population, but some of the larger firms in Germany include <a href="https://www.bilendi.de/">Bilendi/Respondi</a>, <a href="https://www.dynata.com/">Dynata</a>, <a href="https://www.kantarpublic.com/de">Kantar Public</a>, and <a href="https://yougov.de/">YouGov</a>. However, this is not a complete list, so you might consider asking colleagues or looking through the literature to see what providers might be a good fit for your planned sample. Throughout this process, you should also be in contact with your institution’s procurement department to make sure you are following all institutional rules. In our case, two providers sent us quotes and a third declined to bid on the project as they could not meet the specifications we asked for (this however still counted as a quote at our institution).</p>
</div>
<div id="programming" class="section level5">
<h5>Programming</h5>
<p>One decision you will have to make while getting quotes from survey companies is whether you plan to program your survey yourself, or whether you would like the company to program it for you. While it is certainly easier to “leave it to the experts”, programming the survey yourself is cheaper and gives you more control over the process, especially when it comes to implementing features like randomization. If you would like to try programming yourself, you might start by asking your institution whether they have a license for survey software. One of the most commonly used tools is <a href="https://www.qualtrics.com/">Qualtrics</a>, which requires little to no programming knowledge for simple setups, but is generally not free (except for surveys with <a href="https://www.qualtrics.com/free-account/?utm_lp=use-case-survey-software">fewer than 100 respondents</a>).</p>
<p>Other tools you may consider are <a href="https://www.otree.org/">oTree</a>, which is free and open source, and well-suited for interactive experiments and behavioral games (although more complicated to program), or <a href="https://www.unipark.com/en/survey-software/">EFS Survey Unipark</a>, a web-based service which is often available through German universities. No matter which software you choose, remember that the time and attention of your participants is valuable, and you should always strive to create an attractive and easy-to-use interface for your participants.</p>
</div>
<div id="pilot-testing-and-data-collection" class="section level5">
<h5>Pilot Testing and Data Collection</h5>
<p>Depending on the complexity of your design, you may also want to consider running some pilot tests before your main data collection to ensure that your survey works as intended. As one example, you might use a pilot test to identify any potential order effects by testing two versions of the survey where the questions are presented in a different order. In the case of our survey, we wanted to check a couple of things: first, that the treatment had an effect on ratings, and second, that the photos were perceived as expected (i.e. native German, migration background, or ambiguous). In our first pilot test, we found that many of our respondents did not understand the treatment, and that we needed to include more ambiguous photos. We then ran a second pilot test where we added a text description to the treatment graphs and required respondents to stay on the treatment page for one minute, which significantly improved respondents’ comprehension. Based on this experience, I would recommend that you plan for more pilot testing than you think you will need.</p>
<p>After our pilot tests, we started our data collection. It only took us about one week to collect just over 1,100 responses, although your experience may vary depending on the specific population you are sampling. Targeting smaller groups or older populations may take longer. Once we closed data collection, we also carefully cleaned the data. One drawback of online access panels is that some participants may speed through the survey in order to earn incentives. Some telltale signs of this are a very short completion time, “straightlining” responses (such as always choosing the first option), or nonsensical answers to open-ended questions. It is thus essential that you look through your data before analysis and screen out any cases that you think might be invalid. In our survey, we dropped 25 responses for a final sample size of 1,077.</p>
</div>
</div>
<div id="data-and-code-sharing" class="section level3">
<h3>Data and Code Sharing</h3>
<p>Once you have collected your data and finished your analysis, it’s time to share your results with the world! Of course you will write up your results in a manuscript, but you should also consider sharing your data and code with other researchers. Like pre-registration, this is optional, and this may not be feasible if your analysis relies heavily on proprietary data, but data sharing is becoming increasingly common in social science research, and may even be requested by some journals. Sharing your data and code fosters transparency and reproducibility, but there are also more self-interested reasons to share: uploading your work can open new opportunities for collaboration and generate more citations for your work.</p>
<p>When you are preparing to share your data and code, it is important to ask yourself: could another researcher reproduce your results without any additional information? In your dataset, this means ensuring clear coding and labeling, perhaps in a separate codebook file. It also means compiling your questionnaires so that people can see the source of the data (and translating them, if necessary). Perhaps most importantly, make sure that you anonymize or delete any personally identifiable information – refer back to your ethics application and be sure that you are doing what you promised to do. If your analysis includes other data sources that you do not own, you should also make sure to exclude these from your dataset (while informing users about where they can be accessed, and providing code that allows users to merge external data to your own). You can find more information on data preparation from <a href="https://www.icpsr.umich.edu/files/deposit/dataprep.pdf">ICPSR</a> and <a href="https://www.gesis.org/datenservices/daten-teilen/how-to-guide-daten-teilen">GESIS</a>. In your code, make sure you include labels to describe what each section does and that you reference all required packages. You should also have multiple people test your code, ideally someone who was not involved in the project, to ensure that everything is clear and that it works (ideally across different systems) as intended.</p>
<p>Once your code and data are replication-ready, you can upload them to a repository of your choice. You want to choose a repository that can ensure long-term preservation and a persistent identifier (such as a URL or DOI) so that people can find your information. Also make sure that your data will be accessible to other researchers, and if needed that it can be accessed anonymously during the peer review process. Some commonly used (and free!) repositories in the social sciences include <a href="https://osf.io/">OSF</a>, <a href="https://dataverse.harvard.edu/">Harvard Dataverse</a>, <a href="https://www.icpsr.umich.edu/web/pages/index.html">ICPSR</a>, and GESIS’ <a href="https://data.gesis.org/sharing/">SowiDataNet|datorium</a>.
You might also check with your university library, as some institutions may have their own repository (like Mannheim’s <a href="https://madoc.bib.uni-mannheim.de/">MADOC</a>).</p>
</div>
<div id="conclusion" class="section level3">
<h3>Conclusion</h3>
<p>If you have made it through all of these steps, congratulations on finishing your first data collection! As you can see, a lot of planning goes into any survey experiment. Despite this, I think it is worth the trouble. Collecting your own data can be a creative, rewarding, and even fun process that allows you to explore entirely new scientific questions. I hope this post has encouraged you to try it out for yourself, and I wish you the best of luck with your research.</p>
</div>
<div id="further-readings" class="section level3">
<h3>Further readings <a name="furtherreadings"></a></h3>
<ul>
<li>Mutz, D. C. (2011). <em>Population-based survey experiments</em>. Princeton University Press.</li>
<li>Auspurg, K., &amp; Hinz, T. (2014). <em>Factorial survey experiments (Vol. 175)</em>. Sage Publications.</li>
<li>Bansak, K., Hainmueller, J., Hopkins, D., &amp; Yamamoto, T. (2021). Conjoint Survey Experiments. In J. Druckman &amp; D. Green (Eds.), <em>Advances in Experimental Political Science</em> (pp. 19-41). Cambridge: Cambridge University Press.</li>
<li>Salganik, M. J. (2019). <a href="https://www.bitbybitbook.com/en/1st-ed/preface/"><em>Bit by bit: Social research in the digital age</em></a>. Princeton University Press.</li>
<li>Callegaro, M., Baker, R. P., Bethlehem, J., Göritz, A. S., Krosnick, J. A. &amp; Lavrakas, P. J. (2014). <em>Online Panel Research: A Data Quality Perspective</em>. Wiley.</li>
</ul>
<!-- Add something about the instructor -->
</div>
<div id="about-the-presenter" class="section level3">
<h3>About the presenter</h3>
<p>Joshua Hellyer <a href="mailto:joshua.hellyer@mzes.uni-mannheim.de"><i class="fa fa-envelope"></i> </a><a href="https://twitter.com/hellyer_josh"><i class="fa fa-twitter"></i></a> is a doctoral researcher at the Mannheim Centre for European Social Research (MZES). His research focuses on discrimination against ethnic and sexual minorities, particularly in the housing and labor markets.</p>
</div>
<div id="references" class="section level3 unnumbered">
<h3>References</h3>
<div id="refs" class="references">
<div id="ref-abascal_contraction_2020">
<p>Abascal, Maria. 2020. “Contraction as a Response to Group Threat: Demographic Decline and Whites’ Classification of People Who Are Ambiguously White.” <em>American Sociological Review</em> 85 (2): 298–322. <a href="https://doi.org/10.1177/0003122420905127">https://doi.org/10.1177/0003122420905127</a>.</p>
</div>
<div id="ref-arnold_simulation_2011">
<p>Arnold, Benjamin F., Daniel R. Hogan, John M. Colford, and Alan E. Hubbard. 2011. “Simulation Methods to Estimate Design Power: An Overview for Applied Research.” <em>BMC Medical Research Methodology</em> 11 (1): 94. <a href="https://doi.org/10.1186/1471-2288-11-94">https://doi.org/10.1186/1471-2288-11-94</a>.</p>
</div>
<div id="ref-bethlehem_selection_2010">
<p>Bethlehem, Jelke. 2010. “Selection Bias in Web Surveys.” <em>International Statistical Review</em> 78 (2): 161–88. <a href="https://doi.org/10.1111/j.1751-5823.2010.00112.x">https://doi.org/10.1111/j.1751-5823.2010.00112.x</a>.</p>
</div>
<div id="ref-chandler_online_2019">
<p>Chandler, Jesse, Cheskie Rosenzweig, Aaron J. Moss, Jonathan Robinson, and Leib Litman. 2019. “Online Panels in Social Science Research: Expanding Sampling Methods Beyond Mechanical Turk.” <em>Behavior Research Methods</em> 51 (5): 2022–38. <a href="https://doi.org/10.3758/s13428-019-01273-7">https://doi.org/10.3758/s13428-019-01273-7</a>.</p>
</div>
<div id="ref-gereke_demographic_2022">
<p>Gereke, Johanna, Joshua Hellyer, Jan Behnert, Saskia Exner, Alexander Herbel, Felix Jäger, Dean Lajic, et al. 2022. “Demographic Change and Group Boundaries in Germany: The Effect of Projected Demographic Decline on Perceptions of Who Has a Migration Background.” <em>Sociological Science</em> 9.</p>
</div>
<div id="ref-lakens_sample_2021">
<p>Lakens, Daniël. 2021. “Sample Size Justification.” PsyArXiv. <a href="https://doi.org/10.31234/osf.io/9d3yf">https://doi.org/10.31234/osf.io/9d3yf</a>.</p>
</div>
<div id="ref-nosek_preregistration_2018">
<p>Nosek, Brian A., Charles R. Ebersole, Alexander C. DeHaven, and David T. Mellor. 2018. “The Preregistration Revolution.” <em>Proceedings of the National Academy of Sciences</em> 115 (11): 2600–2606. <a href="https://doi.org/10.1073/pnas.1708274114">https://doi.org/10.1073/pnas.1708274114</a>.</p>
</div>
<div id="ref-schonlau_selection_2009">
<p>Schonlau, Matthias, Arthur van Soest, Arie Kapteyn, and Mick Couper. 2009. “Selection Bias in Web Surveys and the Use of Propensity Scores.” <em>Sociological Methods &amp; Research</em> 37 (3): 291–318. <a href="https://doi.org/10.1177/0049124108327128">https://doi.org/10.1177/0049124108327128</a>.</p>
</div>
<div id="ref-simmons_pre-registration_2021">
<p>Simmons, Joseph P., Leif D. Nelson, and Uri Simonsohn. 2021. “Pre-Registration: Why and How.” <em>Journal of Consumer Psychology</em> 31 (1): 151–62. <a href="https://doi.org/10.1002/jcpy.1208">https://doi.org/10.1002/jcpy.1208</a>.</p>
</div>
<div id="ref-strack_order_1992">
<p>Strack, Fritz. 1992. “‘Order Effects’ in Survey Research: Activation and Information Functions of Preceding Questions.” In <em>Context Effects in Social and Psychological Research</em>, edited by Norbert Schwarz and Seymour Sudman, 23–34. New York, NY: Springer New York. <a href="https://doi.org/10.1007/978-1-4612-2848-6_3">https://doi.org/10.1007/978-1-4612-2848-6_3</a>.</p>
</div>
</div>
</div>
]]>
      </description>
    </item>
    
    <item>
      <title>Using Web Logs and Smartphone Records for Social Research</title>
      <link>https://socialsciencedatalab.mzes.uni-mannheim.de/article/using-web-logs/</link>
      <pubDate>Tue, 14 Apr 2020 01:00:00 +0100</pubDate>
      
      <guid>https://socialsciencedatalab.mzes.uni-mannheim.de/article/using-web-logs/</guid>
      <description><![CDATA[
        </p>
<p>How can social scientists collect and analyze web logs – records of individuals’ browsing behavior – for their own research? In this <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/categories/instructionals/">Methods Bites Instructional Blog Post</a>, <a href="https://twitter.com/rub3n_luc">Ruben Bach</a> summarizes some key insights of his talk in the <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/">MZES Social Sciences Data Lab</a> in December 2019. The blog post discusses how to obtain and extract information from web logs and related data, shows how they can be used for social research, and concludes with a short discussion of how to handle big data extracted from web logs.</p>
<div id="contents" class="section level5">
<h5>Contents</h5>
<ol style="list-style-type: decimal">
<li><a href="#how-to-use-web-log-and-related-data-for-social-research">How to use web log and related data for social research</a></li>
<li><a href="#how-to-obtain-web-log-and-related-data">How to obtain web log and related data</a></li>
<li><a href="#how-to-handle-big-data">How to handle “big data”</a></li>
<li><a href="#about-the-presenter">About the presenter</a></li>
<li><a href="#further-reading">Further reading</a></li>
<li><a href="#references">References</a></li>
</ol>
</div>
<div id="how-to-use-web-log-and-related-data-for-social-research" class="section level3">
<h3>How to use web log and related data for social research</h3>
<p>Web logs (“browsing histories”) are records of app use and search queries. They are highly interesting data sources for research in the social sciences as they offer detailed insights into human behavior. However, only a few studies have used such data in the social sciences so far.
Examples of such studies include <span class="citation">Stephens-Davidowitz (2014)</span>, who studied racial animus in the 2012 U.S. presidential election, <span class="citation">Peterson, Goel, and Iyengar (2018)</span> and <span class="citation">Flaxman, Goel, and Rao (2016)</span>, who analyzed filter bubbles, echo chambers and partisan polarization in the U.S. and <span class="citation">Guess, Nyhan, and Reifler (2020)</span> who studied the spread of fake news in the U.S. In the Netherlands, <span class="citation">Möller et al. (2019)</span> analyzed online news engagement based on three different modes of news use. In Israel, <span class="citation">Dvir-Girsman (2017)</span> documented that audience homophily is higher among individuals with more extreme ideology and that it is associated with ideological polarization and intolerance. <span class="citation">Bach et al. (2019)</span> showed for Germany that online and mobile device activities predict voting behavior and political preferences to a limited degree only. <span class="citation">Chancellor and Counts (2018)</span> show that internet search data can be used to estimate employment demand in the U.S.</p>
<p>In other disciplines, similar data have been used, for example, to predict influenza activity <span class="citation">(Ginsberg et al. 2009, but see <span class="citation">@lazer2014parable</span>)</span> and to show that users who search for relatively harmless symptoms easily end up searching for serious diseases <span class="citation">(White and Horvitz 2009)</span>. Another study demonstrates how concerns about pregnancy and childbirth change over the course of pregnancy <span class="citation">(Fourney, White, and Horvitz 2015)</span>. Furthermore, several papers show how a variety of user attributes such as socio-demographics can be inferred from web logs, search queries and app records <span class="citation">(see Hinds and Joinson 2018 for a recent overview)</span>. However, many of those studies have been conducted by researchers from computer science. While they often rely on search query data obtained from search engine providers like Bing and Google, they typically focus more on the technical aspects and on the evaluation of the performance of the underlying algorithms. As social scientists, however, we often focus on the theory-driven development of models and the testing of hypotheses about human and societal behavior. Thus, this blog post will focus on the latter.</p>
<p>One challenge that researchers face when designing studies that rely on web logs, records of app use, and search queries is how to get access to such data. Several of the studies mentioned above use large amounts of search query data from search engines like Google or Bing <span class="citation">(Stephens-Davidowitz 2014; Chancellor and Counts 2018; White and Horvitz 2009; Fourney, White, and Horvitz 2015)</span>. These data are, however, usually only available if one teams up with researchers from the respective companies. Another way to obtain data is through commercial providers who keep opt-in panels of users who occasionally answer survey questions in exchange for money. Researchers can pay those vendors in order to get access to their panels and ask participants survey questions. In addition, several of these providers also offer web log and mobile device use records from users who (in exchange for additional pay) agreed to having their online mobile activities monitored. In this blog post, we will mainly focus on this latter way of obtaining data. Before we talk about this topic in more detail, we will briefly summarize what we mean when we speak of web logs, records of app use, and search queries.</p>
<p>To get a better understanding of such data, the table below shows a collection of a few artificial web logs made up for this blog post. Typically, we observe a person identifier (first column) for the person whose records we observe. Second, we have a URL (Uniform Resource Locator) column which tells us which URL this person visited, when (column “Timestamp”) and for how long (“Duration of use”; here, in seconds). The most interesting information in this table is the URL column. Even without a detailed understanding of the specific form of a URL, we can easily see that, in the first row, Person 1 <a href="https://www.wetter.de/deutschland/wetter-mannheim-18224779.html?q=mannheim">visited a web address that seems to inform her about the weather in Mannheim</a>. In addition, we observe when she visited this address and how much time she spent there. The second row tells us that this person likely sent (or received) a message to (from) Peter Mustermann. We cannot, however, observe the content of the message (which we also should not, given obvious privacy reasons). The third row shows that Person 1 then visited the <a href="https://www.facebook.com/CDU/">Facebook page of the CDU</a>. The fourth row shows that she watched a <a href="https://www.youtube.com/watch?v=SpoXEEdsNfE">video on YouTube</a>. If we accessed the web address (or programmed a scraping tool), we could also learn what the video was about. From the visit in the fifth row, we learn that this person read an article about the state of the German economy on <a href="https://www.zeit.de/wirtschaft/unternehmen/2020-01/ifo-index-geschaeftsklima-deutsche-wirtschaft-konjunktur">DIE ZEIT</a>, a German newspaper.</p>
<p>With respect to Person 2, we can also observe what items they <a href="https://www.amazon.de/s?k=mostly+harmless+econometrics">searched on Amazon</a> (sixth row) or <a href="https://twitter.com/realDonaldTrump/status/1222008772102705152">which tweet they saw</a> (seventh row). We also learn that they searched for the voting advice application <a href="https://www.google.com/search?q=wahl+o+mat+hamburg+2020">“Wahl-O-Mat”</a> (ninth row). From the URL in row 9, we can extract the exact words a person entered into a search engine (the “search queries”). From row 10, we also know that Person 2 <a href="https://www.wahl-o-mat.de/hamburg2020/">then actually used this tool</a>. Thus, we already see that we can learn a lot about users’ behavior, their interests and preferences.</p>
<table>
<colgroup>
<col width="8%" />
<col width="65%" />
<col width="13%" />
<col width="11%" />
</colgroup>
<thead>
<tr class="header">
<th>Person ID</th>
<th>URL</th>
<th>Timestamp</th>
<th>Duration of use</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>1</td>
<td><a href="https://www.wetter.de/deutschland/wetter-mannheim-18224779.html?q=mannheim" class="uri">https://www.wetter.de/deutschland/wetter-mannheim-18224779.html?q=mannheim</a></td>
<td>2020-01-21 10:43:45</td>
<td>24</td>
</tr>
<tr class="even">
<td>1</td>
<td><a href="https://www.facebook.com/messages/t/Peter.Mustermann" class="uri">https://www.facebook.com/messages/t/Peter.Mustermann</a></td>
<td>2020-01-22 11:43:45</td>
<td>32</td>
</tr>
<tr class="odd">
<td>1</td>
<td><a href="https://www.facebook.com/CDU/" class="uri">https://www.facebook.com/CDU/</a></td>
<td>2020-01-23 23:43:01</td>
<td>45</td>
</tr>
<tr class="even">
<td>1</td>
<td><a href="https://www.youtube.com/watch?v=SpoXEEdsNfE" class="uri">https://www.youtube.com/watch?v=SpoXEEdsNfE</a></td>
<td>2020-01-24 08:21:45</td>
<td>3625</td>
</tr>
<tr class="odd">
<td>1</td>
<td><a href="https://www.zeit.de/wirtschaft/unternehmen/2020-01/ifo-index-geschaeftsklima-deutsche-wirtschaft-konjunktur" class="uri">https://www.zeit.de/wirtschaft/unternehmen/2020-01/ifo-index-geschaeftsklima-deutsche-wirtschaft-konjunktur</a></td>
<td>2020-01-25 16:14:07</td>
<td>67</td>
</tr>
<tr class="even">
<td>2</td>
<td><a href="https://www.amazon.de/s?k=mostly+harmless+econometrics" class="uri">https://www.amazon.de/s?k=mostly+harmless+econometrics</a></td>
<td>2020-01-26 23:54:45</td>
<td>23</td>
</tr>
<tr class="odd">
<td>2</td>
<td><a href="https://www.sueddeutsche.de/politik/bundeswehr-wehrbeauftragter-bericht-1.4774621" class="uri">https://www.sueddeutsche.de/politik/bundeswehr-wehrbeauftragter-bericht-1.4774621</a></td>
<td>2020-01-27 07:43:45</td>
<td>245</td>
</tr>
<tr class="even">
<td>2</td>
<td><a href="https://twitter.com/realDonaldTrump/status/1222008772102705152" class="uri">https://twitter.com/realDonaldTrump/status/1222008772102705152</a></td>
<td>2020-01-28 01:01:45</td>
<td>56</td>
</tr>
<tr class="odd">
<td>2</td>
<td><a href="https://www.google.com/search?q=wahl+o+mat+hamburg+2020" class="uri">https://www.google.com/search?q=wahl+o+mat+hamburg+2020</a></td>
<td>2020-01-29 17:14:45</td>
<td>32</td>
</tr>
<tr class="even">
<td>2</td>
<td><a href="https://www.wahl-o-mat.de/hamburg2020/" class="uri">https://www.wahl-o-mat.de/hamburg2020/</a></td>
<td>2020-01-30 09:09:45</td>
<td>578</td>
</tr>
<tr class="odd">
<td>2</td>
<td><a href="https://www.notebookcheck.com/Top-10-Ultrabooks-im-Test-bei-Notebookcheck.125873.0.html" class="uri">https://www.notebookcheck.com/Top-10-Ultrabooks-im-Test-bei-Notebookcheck.125873.0.html</a></td>
<td>2020-01-30 11:49:45</td>
<td>243</td>
</tr>
<tr class="even">
<td>2</td>
<td><a href="https://www.google.com/search?q=flug+frankfurt+new+york" class="uri">https://www.google.com/search?q=flug+frankfurt+new+york</a></td>
<td>2020-01-30 12:09:45</td>
<td>67</td>
</tr>
<tr class="odd">
<td>2</td>
<td><a href="https://www.google.com/flights?lite=0#flt=/m/02z0j./m/02_286.2020-01-30*/m/02_286./m/02z0j.2020-02-12;c:EUR;e:1;sd:1;t:f" class="uri">https://www.google.com/flights?lite=0#flt=/m/02z0j./m/02_286.2020-01-30*/m/02_286./m/02z0j.2020-02-12;c:EUR;e:1;sd:1;t:f</a></td>
<td>2020-01-30 18:56:45</td>
<td>456</td>
</tr>
</tbody>
</table>
<p>We observe similar records regarding app use on mobile devices, as shown in the next table. Here, we observe the name of an app instead of a URL. Unfortunately, we cannot observe what users do inside the apps they use. That is, we do not know which articles they read when they open a newspaper app or which videos they watch on YouTube or Netflix. Yet, if somebody frequently opens newspaper apps, we might infer that this person may be interested in politics. Moreover, users who use an app called “Period Tracker” are likely female and we may even learn when they have their period by observing temporal variations in usage of this app. Similarly, somebody who uses an app that informs them about prayer times is likely religious. These are just a few examples that show how observing users’ web logs or app use can reveal information that users may perceive as sensitive personal information <span class="citation">(see Bach et al. 2019 for a study on this topic)</span>.</p>
<table>
<thead>
<tr class="header">
<th>Person ID</th>
<th>App</th>
<th>Timestamp</th>
<th>Duration of use</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>1</td>
<td>Facebook</td>
<td>2020-01-21 10:43:45</td>
<td>24</td>
</tr>
<tr class="even">
<td>1</td>
<td>Kleiderkreisel</td>
<td>2020-01-22 11:43:45</td>
<td>32</td>
</tr>
<tr class="odd">
<td>1</td>
<td>WhatsApp Messenger</td>
<td>2020-01-23 23:43:01</td>
<td>45</td>
</tr>
<tr class="even">
<td>1</td>
<td>Youtube</td>
<td>2020-01-24 08:21:45</td>
<td>3625</td>
</tr>
<tr class="odd">
<td>1</td>
<td>GMX Mail</td>
<td>2020-01-25 16:14:07</td>
<td>67</td>
</tr>
<tr class="even">
<td>2</td>
<td>Period tracker</td>
<td>2020-01-26 23:54:45</td>
<td>23</td>
</tr>
<tr class="odd">
<td>2</td>
<td>Pedometer</td>
<td>2020-01-27 07:43:45</td>
<td>245</td>
</tr>
<tr class="even">
<td>2</td>
<td>WhatsApp Messenger</td>
<td>2020-01-28 01:01:45</td>
<td>56</td>
</tr>
<tr class="odd">
<td>2</td>
<td>WhatsApp Messenger</td>
<td>2020-01-29 17:14:45</td>
<td>32</td>
</tr>
<tr class="even">
<td>2</td>
<td>Youtube</td>
<td>2020-01-30 09:09:45</td>
<td>578</td>
</tr>
<tr class="odd">
<td>2</td>
<td>Netflix</td>
<td>2020-01-30 18:56:45</td>
<td>456</td>
</tr>
</tbody>
</table>
<p>It is important to note that participants of commercial panels can usually pause data collection on their devices temporarily. This might happen if they do not feel comfortable having their activities recorded, e.g. when they do online banking or watch movies from illegal streams. However, our own data collections show large amounts of potentially sensitive information (such as adult content, gambling, and illegal streaming), which suggests that users do not make use of this possibility very often.</p>
</div>
<div id="how-to-obtain-web-log-and-related-data" class="section level3">
<h3>How to obtain web log and related data</h3>
<p>As mentioned earlier, the easiest way to obtain web log and app use data is through commercial vendors who operate online access panels (“non-probability panels”). So far, there seem to be only a handful of providers, such as <a href="https://www.respondi.com/EN/">respondi AG</a> (Germany, UK and France, for example), <a href="https://yougov.co.uk">YouGov</a> (UK and US, amongst others) and <a href="https://www.netquest.com/en/online-surveys-investigation">netquest</a> (mostly Spain and Latin America). For an overview of providers, see for example, <a href="https://wakoopa.com/get-data/">Wakoopa Hub</a>. Since such data have not been available for long, we do not know much about the quality of such data <span class="citation">(for an exception see Revilla, Ochoa, and Loewe 2017)</span>.. In addition to accessing web logs and app use data, researchers can collect survey data for the same individuals. That is, by asking users questions we can enrich the web log data with important information about users’ socio-demographic characteristics, their voting behavior, or their political preferences. This information is usually not directly observable from the web logs and app use records, but likely makes the data much more valuable for social scientific research purposes. However, we should always keep in mind that due to the opt-in nature of these access panels, we need to be careful when making statements about the representativeness of our findings. That is, because most of our statistical estimators (e.g., for variances) rely on true probability sampling, we need to adjust our models and estimators, e.g. by using weights <span class="citation">(for more information on non-probability sampling and non-probability panels see, e.g., Mercer et al. 2017; Cornesse et al. 2020)</span>.</p>
<p>Another way to obtain data on users’ online activities is through <a href="https://trends.google.com">Google Trends</a>. Briefly speaking, this service allows the estimation of the popularity of search queries in Google Search across various regions and languages. Several R packages in R allow users to automatically extract data from Google Trends – see, for instance, <a href="https://cran.r-project.org/web/packages/gtrendsR/gtrendsR.pdf"><code>gtrendsR</code></a>. Further details on Google Trends can be found <a href="https://www.aeaweb.org/conference/2016/retrieve.php?pdfid=772">here</a>. A major drawback of these data is that they do not come with additional information about users and only provide information at aggregate levels. That is, one can learn only about the popularity of a specific search query compared to the popularity of other search queries, which is arguably much less informative than individual-level web logs.</p>
<p>A third way to obtain web log, app, and search query data is through developing one’s own research app. A few projects using this approach were launched in recent years. One of the most prominent examples is the <a href="https://www.iab.de/751/section.aspx/1470">IAB-SMART project</a> <span class="citation">(for details, see Kreuter et al. 2019)</span>. Researchers of the Institute for Employment Research (IAB) in Nuremberg, Germany, developed an app that has since been downloaded by several hundred respondents of the IAB’s panel survey “Labour Market and Social Security”. The app collects information on respondents’ smartphone activities and can access even more information than those described above (including geolocation and address books). While this approach is likely the most elaborate, it is also the most difficult and most costly one to implement: Researchers need to program their own tools and recruit participants on their own or piggyback on an existing study.</p>
</div>
<div id="how-to-handle-big-data" class="section level3">
<h3>How to handle “big data”</h3>
<p>Finally, a note on useful tools and prerequisites for analyzing web logs and records of smartphone use. First of all, the amount of data can quickly exceed the computational power of a standard desktop computer. Four months of web log data used in <span class="citation">Bach et al. (2019)</span>, for example, contained about 38 million observations. Working with data of this size, researchers may have to consider using remote computing services like <a href="https://www.digitalocean.com">Digital Ocean</a>, <a href="https://aws.amazon.com">Amazon AWS</a> or <a href="https://azure.microsoft.com/en-us/">Microsoft Azure</a>, which offer computational resources for little money through virtual servers. Second, understanding URL contents by observing single URLs is straightforward. Analyzing thousands of URLs, however, requires text mining and natural language processing (NLP) techniques if one wants, for example, to select only those URLs that point to news articles. Moreover, in addition to analyzing the title of a news article (which can often be observed from the URL alone), one might also want to analyze the whole content of the article. In such cases, in addition to being able to automatically extract the topic of an article through NLP techniques, knowing how to scrape website contents will likely also be helpful. Some useful materials are linked <a href="#further-reading">below</a>.</p>
</div>
<div id="about-the-presenter" class="section level3">
<h3>About the presenter</h3>
<p>Ruben Bach <a href="mailto: r.bach@uni-mannheim.de "><i class="fa
              fa-envelope"></i> </a>
<a href="https://ruben-bach.com/"><i class="fa
              fa-globe"></i> </a>
<a href=" https://twitter.com/rub3n_luc"><i class="fa
              fa-twitter"></i></a> is a postdoctoral researcher at the University of Mannheim, focusing on social science quantitative research methods. His interests include topics related to big data in the social sciences, machine learning, causal inference, and survey research.</p>
</div>
<div id="further-reading" class="section level3">
<h3>Further reading</h3>
<ul>
<li><a href="https://socialsciencedatalab.mzes.uni-mannheim.de/article/advancing-text-mining/">Meyer, Cosima and Cornelius Puschmann. 2019. <em>Advancing Text Mining with R and quanteda</em>. Methods Bites: Blog of the MZES Social Science Data Lab. Blog Post Tutorial.</a></li>
<li><a href="https://github.com/SocialScienceDataLab/Intro-to-web-scraping-with-R">Munzert, Simon. 2016. <em>Three easy-to-learn tools to scrape data from the Web with R</em>. MZES Social Science Data Lab. Workshop Materials.</a></li>
</ul>
</div>
<div id="references" class="section level3 unnumbered">
<h3 class="unnumbered">References</h3>
<div id="refs" class="references hanging-indent">
<div id="ref-bach2019">
<p>Bach, R. L., C. Kern, A. Amaya, F. Keusch, F. Kreuter, J. Heinemann, and J. Hecht. 2019. “Predicting Voting Behavior Using Digital Trace Data.” <em>Social Science Computer Review</em>. <a href="https://doi.org/10.1177/0894439319882896">https://doi.org/10.1177/0894439319882896</a>.</p>
</div>
<div id="ref-chancellor2018">
<p>Chancellor, S., and S. Counts. 2018. “Measuring Employment Demand Using Internet Search Data.” In <em>Proceeding of the 2018 Chi Conference on Human Factors in Computing Systems</em>, 1–14. CHI ’18. New York, NY, USA: ACM.</p>
</div>
<div id="ref-cornesse2020">
<p>Cornesse, C., A. G. Blom, D. Dutwin, J. A. Krosnick, E. D. De Leeuw, S. Legleye, J. Pasek, et al. 2020. “A Review of Conceptual Approaches and Empirical Evidence on Probability and Nonprobability Sample Survey Research.” <em>Journal of Survey Statistics and Methodology</em>. <a href="https://doi.org/10.1093/jssam/smz041%20">https://doi.org/10.1093/jssam/smz041</a>.</p>
</div>
<div id="ref-dvir2017">
<p>Dvir-Girsman, S. 2017. “Media Audience Homophily: Partisan Websites, Audience Identity and Polarization Processes.” <em>New Media &amp; Society</em> 19 (7): 1072–91.</p>
</div>
<div id="ref-flaxman2016filter">
<p>Flaxman, Seth, Sharad Goel, and Justin M Rao. 2016. “Filter Bubbles, Echo Chambers, and Online News Consumption.” <em>Public Opinion Quarterly</em> 80 (S1): 298–320.</p>
</div>
<div id="ref-Fourney2015">
<p>Fourney, Adam, Ryen W. White, and Eric Horvitz. 2015. “Exploring Time-Dependent Concerns About Pregnancy and Childbirth from Search Logs.” In <em>Proceedings of the 33rd Annual Acm Conference on Human Factors in Computing Systems</em>, 737–46. CHI ’15. New York, NY, USA: ACM. <a href="https://doi.org/https://doi.org/10.1145/2702123.2702427">https://doi.org/https://doi.org/10.1145/2702123.2702427</a>.</p>
</div>
<div id="ref-ginsberg2009detecting">
<p>Ginsberg, Jeremy, Matthew H Mohebbi, Rajan S Patel, Lynnette Brammer, Mark S Smolinski, and Larry Brilliant. 2009. “Detecting Influenza Epidemics Using Search Engine Query Data.” <em>Nature</em> 457 (7232): 1012–4.</p>
</div>
<div id="ref-guess2020exposure">
<p>Guess, Andrew M, Brendan Nyhan, and Jason Reifler. 2020. “Exposure to Untrustworthy Websites in the 2016 Us Election.” <em>Nature Human Behaviour</em>, 1–9.</p>
</div>
<div id="ref-hinds">
<p>Hinds, J., and A. N. Joinson. 2018. “What Demographic Attributes Do Our Digital Footprints Reveal? A Systematic Review.” <em>PLoS One</em> 13: 1–40.</p>
</div>
<div id="ref-KreuterSMART">
<p>Kreuter, Frauke, Georg-Christoph Haas, Florian Keusch, Sebastian Bähr, and Mark Trappmann. 2019. “Collecting Survey and Smartphone Sensor Data with an App: Opportunities and Challenges Around Privacy and Informed Consent.” <em>Social Science Computer Review</em>. <a href="https://doi.org/10.1177/0894439318816389">https://doi.org/10.1177/0894439318816389</a>.</p>
</div>
<div id="ref-lazer2014parable">
<p>Lazer, David, Ryan Kennedy, Gary King, and Alessandro Vespignani. 2014. “The Parable of Google Flu: Traps in Big Data Analysis.” <em>Science</em> 343 (6176): 1203–5.</p>
</div>
<div id="ref-AMercer">
<p>Mercer, Andrew W., Frauke Kreuter, Scott Keeter, and Elizabeth A. Stuart. 2017. “Theory and Practice in Nonprobability Surveys: Parallels between Causal Inference and Survey Inference.” <em>Public Opinion Quarterly</em> 81 (S1): 250–71. <a href="https://doi.org/10.1093/poq/nfw060">https://doi.org/10.1093/poq/nfw060</a>.</p>
</div>
<div id="ref-moller2019explaining">
<p>Möller, Judith, Robbert Nicolai van de Velde, Lisa Merten, and Cornelius Puschmann. 2019. “Explaining Online News Engagement Based on Browsing Behavior: Creatures of Habit?” <em>Social Science Computer Review</em>, 0894439319828012. <a href="https://doi.org/10.1177/0894439319828012">https://doi.org/10.1177/0894439319828012</a>.</p>
</div>
<div id="ref-peterson2018echo">
<p>Peterson, Erik, Sharad Goel, and Shanto Iyengar. 2018. “Echo Chambers and Partisan Polarization: Evidence from the 2016 Presidential Campaign.”</p>
</div>
<div id="ref-revilla2017using">
<p>Revilla, Melanie, Carlos Ochoa, and Germán Loewe. 2017. “Using Passive Data from a Meter to Complement Survey Data in Order to Study Online Behavior.” <em>Social Science Computer Review</em> 35 (4): 521–36.</p>
</div>
<div id="ref-stephens2014cost">
<p>Stephens-Davidowitz, Seth. 2014. “The Cost of Racial Animus on a Black Candidate: Evidence Using Google Search Data.” <em>Journal of Public Economics</em> 118: 26–40.</p>
</div>
<div id="ref-white2009">
<p>White, Ryen W., and Eric Horvitz. 2009. “Cyberchondria: Studies of the Escalation of Medical Concerns in Web Search.” <em>ACM Trans. Inf. Syst.</em> 27 (4): 23:1–23:37. <a href="https://doi.org/https://doi.org/10.1145/1629096.1629101">https://doi.org/https://doi.org/10.1145/1629096.1629101</a>.</p>
</div>
</div>
</div>
]]>
      </description>
    </item>
    
    <item>
      <title>Studying Politics on and with Wikipedia</title>
      <link>https://socialsciencedatalab.mzes.uni-mannheim.de/article/studying-politics-wikipedia/</link>
      <pubDate>Mon, 26 Aug 2019 01:00:00 +0100</pubDate>
      
      <guid>https://socialsciencedatalab.mzes.uni-mannheim.de/article/studying-politics-wikipedia/</guid>
      <description><![CDATA[
        </p>
<p>The online encyclopedia Wikipedia, together with its sibling, the collaboratively edited knowledge base Wikidata, provides incredibly rich yet largely untapped sources for political research. In this <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/categories/tutorials/">Methods Bites Tutorial</a>, <a href="https://twitter.com/denis_cohen">Denis Cohen</a> and <a href="https://twitter.com/Nick_Baumann97">Nick Baumann</a> offer a hands-on recap of <a href="https://twitter.com/simonsaysnothin">Simon Munzert</a>’s (Hertie School of Governance) workshop materials to show how these platforms can inform research on public attention dynamics, policies, political and other events, political elites, and parties, among other things.</p>
<p>After reading this blog post and engaging with the applied exercises, readers should:</p>
<ul>
<li>be able to collect Wikipedia data and Wikidata items using <strong>R</strong></li>
<li>be able to conduct explorative analyses of Wikipedia data using <strong>R</strong></li>
<li>have a basic intuition of the potentials and limitations of using Wikipedia data in research projects</li>
</ul>
<p><em>Note:</em> This blog post provides a summary of Simon’s workshop in the <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/page/events/index.html#munzert-201905">MZES Social Science Data Lab</a> with some adaptations. Simon’s original workshop materials, including slides and scripts, are available from our <a href="https://github.com/SocialScienceDataLab/political-wikipedia-workshop">GitHub</a>.</p>
<div id="contents" class="section level5">
<h5>Contents</h5>
<ol style="list-style-type: decimal">
<li><a href="#wikipedia-for-political-research">Wikipedia for Political Research</a></li>
<li><a href="#collecting-and-analyzing-wikipedia-data">Collecting and Analyzing Wikipedia Data</a>
<ol style="list-style-type: decimal">
<li><a href="#application-1-using-pageviews-to-measure-public-attention">Application 1: Using Pageviews to Measure Public Attention</a></li>
<li><a href="#application-2-using-article-links-to-create-a-network-graph-of-german-mps">Application 2: Using Article Links to Create a Network Graph of German MPs</a></li>
<li><a href="#application-3-using-clickstream-data-to-analyze-referral-patterns">Application 3: Using Clickstream Data to Analyze Referral Patterns</a></li>
</ol></li>
<li><a href="#collecting-data-via-wikidata-queries">Collecting Data via Wikidata Queries</a></li>
<li><a href="#legislator">legislatoR</a>
<ol style="list-style-type: decimal">
<li><a href="#application-1-social-media-adoption-rates">Application 1: Social Media Adoption Rates</a></li>
<li><a href="#application-2-public-attention-to-members-of-the-german-bundestag">Application 2: Public Attention to Members of the German Bundestag</a></li>
</ol></li>
<li><a href="#conclusion">Conclusion</a></li>
<li><a href="#about-the-presenter">About the Presenter</a></li>
<li><a href="#references">References</a></li>
</ol>
</div>
<div id="wikipedia-for-political-research" class="section level3">
<h3>Wikipedia for Political Research</h3>
<p>According to its website, <a href="https://en.wikipedia.org/wiki/Wikipedia:About"><em>“Wikipedia […] is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content”</em></a>. As of July 2019, it comprises <a href="https://en.wikipedia.org/wiki/Wikipedia:Size_of_Wikipedia">more than 48 million articles</a> and is ranked <a href="https://www.alexa.com/topsites">sixth in the list of the most frequently visited websites</a>.</p>
<p>Wikipedia harbors numerous types of data. These include both article contents as well as meta information such as pageviews, clickstreams, links and backlinks, or edits and revision histories. Additionally, Wikipedia’s sibling, the collaboratively edited document-oriented data base <a href="https://en.wikipedia.org/wiki/Wikidata">Wikidata</a>, provides access to over 58 million data items (as of July 2019). Given the broad collection of articles on politicians and institutions from all over the world, Wikipedia offers tremendous potential for (comparative) political research.</p>
<p>In what follows, we will introduce the functionalities of various <strong>R</strong> packages, including <a href="https://cran.r-project.org/web/packages/WikipediR/index.html"><code>WikipediR</code></a>, <a href="https://cran.r-project.org/web/packages/WikidataR/index.html"><code>WikidataR</code></a>, and <a href="https://cran.rstudio.com/web/packages/pageviews/index.html"><code>pageviews</code></a>. In doing so, we will showcase how to connect to Wikipedia and Wikidata APIs, how to efficiently access and parse content, and how to process the retrieved data in order to address various questions of substantive interest. We will also provide an overview of the <code>legislatoR</code> package, a fully relational individual-level data package that comprises political, sociodemographic, and Wikipedia-related data on elected politicians from various consolidated democracies.</p>
<details>
<p><summary> Code: <strong>R</strong> packages used in this tutorial</summary></p>
<pre class="r"><code>## Packages
pkgs &lt;- c(
  &quot;devtools&quot;,
  &quot;ggnetwork&quot;,
  &quot;igraph&quot;,
  &quot;intergraph&quot;,
  &quot;tidyverse&quot;,
  &quot;rvest&quot;,
  &quot;devtools&quot;,
  &quot;magrittr&quot;,
  &quot;plotly&quot;,
  &quot;RColorBrewer&quot;,
  &quot;colorspace&quot;,
  &quot;lubridate&quot;,
  &quot;networkD3&quot;,
  &quot;pageviews&quot;,
  &quot;readr&quot;,
  &quot;wikipediatrend&quot;,
  &quot;WikipediR&quot;,
  &quot;WikidataR&quot;
)

## Install uninstalled packages
lapply(pkgs[!(pkgs %in% installed.packages())], install.packages)

## Load all packages to library
lapply(pkgs, library, character.only = TRUE)

## legislatoR
devtools::install_github(&quot;saschagobel/legislatoR&quot;)
library(legislatoR)</code></pre>
</details>
<p><br /></p>
</div>
<div id="collecting-and-analyzing-wikipedia-data" class="section level3">
<h3>Collecting and Analyzing Wikipedia Data</h3>
<div id="application-1-using-pageviews-to-measure-public-attention" class="section level5">
<h5>Application 1: Using Pageviews to Measure Public Attention</h5>
<p>Pageviews measure the aggregate number of clicks for a given Wikipedia article. Data on pageviews can be collected from different sources. First, this <a href="https://tools.wmflabs.org/pageviews/?project=en.wikipedia.org&amp;platform=all-access&amp;agent=user&amp;range=latest-20&amp;pages=Cat%7CDog">interactive tool</a> provides summary data which allows users to compare various search items’ popularity in a specified period. Secondly, <a href="https://dumps.wikimedia.org/">Wikimedia Downloads</a>, a collection of archived Wikimedia wikis, offers <a href="https://dumps.wikimedia.org/other/pagecounts-raw/">pageviews data through August 2016</a> as well as <a href="https://dumps.wikimedia.org/other/pageviews/">data using a new pageviews definition from May 2015 onward</a>.</p>
<p>The code chunk below demonstrates how to collect and graphically display pageviews data using the <code>pageviews</code> package. We use the command <code>article_pageviews()</code>, where the argument <code>project = "en.wikipedia"</code> specifies that we want to collect pageviews of <code>article = "Donald Trump"</code> from the English Wikipedia. We can only restrict our query to a given language edition; it is not possible to limit queries to pageviews from a specific country. We also specify the argument <code>user_type = "user"</code>, which ensures that we exclude pageviews generated by bots and spiders. Finally, <code>start</code> and <code>end</code> define the period on which we want to collect pageviews data: July 2015 to May 2017. We proceed analogously for <code>article = "Hillary Clinton"</code>.</p>
<details>
<p><summary> Code: Pageviews Data Collection </summary></p>
<pre class="r"><code># get pageviews
trump_views &lt;-
  article_pageviews(
    project = &quot;en.wikipedia&quot;,
    article = &quot;Donald Trump&quot;,
    user_type = &quot;user&quot;,
    start = &quot;2015070100&quot;,
    end = &quot;2017050100&quot;
  )
head(trump_views)

clinton_views &lt;-
  article_pageviews(
    project = &quot;en.wikipedia&quot;,
    article = &quot;Hillary Clinton&quot;,
    user_type = &quot;user&quot;,
    start = &quot;2015070100&quot;,
    end = &quot;2017050100&quot;
  )</code></pre>
</details>
<p><br />
This query allows us to retrieve the pageviews for both Trump’s and Clinton’s Wikipedia articles by date. We can then plot the frequencies of pageviews over time to identify trends in search behaviour. As we can see, the data indicate that Trump attracted considerably more attention than Clinton throughout the 2016 election campaign.</p>
<details>
<p><summary> Code: Plotting Pageviews </summary></p>
<pre class="r"><code># Plot pageviews
plot(ymd(trump_views$date), trump_views$views, col = &quot;red&quot;, type = &quot;l&quot;, xlab=&quot;Time&quot;, ylab=&quot;Pageviews&quot;)
lines(ymd(clinton_views$date), clinton_views$views, col = &quot;blue&quot;)
legend(&quot;topleft&quot;, legend=c(&quot;Trump&quot;,&quot;Clinton&quot;), cex=.8,col=c(&quot;red&quot;,&quot;blue&quot;), lty=1) </code></pre>
</details>
<p><img src="/../../../../../article/studying-politics-wikipedia_files/figure-html/code%203b-1.png" width="672" style="display: block; margin: auto;" /></p>
</div>
<div id="application-2-using-article-links-to-create-a-network-graph-of-german-mps" class="section level5">
<h5>Application 2: Using Article Links to Create a Network Graph of German MPs</h5>
<p>The <a href="https://cran.r-project.org/web/packages/WikipediR/index.html"><code>WikipediR</code></a> package is a wrapper for the MediaWiki API that can be used to retrieve page contents as well as metadata for articles and categories, e.g. information about users or page edit histories. The functionality of the package includes:</p>
<ul>
<li><code>page_content()</code>: Retrieve current article versions (HTML and wikitext as possible output formats)</li>
<li><code>revision_content()</code>: Retrieve older versions of the article; this also includes metadata about the revision history</li>
<li><code>page_links()</code>: Retrieve outgoing links from the page’s content (which Wikipedia articles does the page link to?)</li>
<li><code>page_backlinks()</code>: Retrieve incoming links (which Wikipedia articles link to the page?)</li>
<li><code>page_external_links()</code>: Retrieve outgoing links to external sites</li>
<li><code>page_info()</code>: Page metadata</li>
<li><code>categories_in_page()</code>: What categories is a given page in?</li>
<li><code>pages_in_category()</code>: What pages are in a given category?</li>
</ul>
<p>For our application, we use the <code>page_links()</code> function to extract mutual referrals between the articles on members of the 2017-2021 German Bundestag. We can then use this information to create a network graph of current German MPs. First, we use the <a href="#legislator"><code>legislatoR</code></a> package to retrieve a list of all German MPs of the 2017-2021 German Bundestag, including information on their page IDs and page titles in the German Wikipedia. Using this information, we then extract all <code>page_links()</code> in every MP’s Wikipedia articles. The third step identifies the subset of links for every MP that link to the Wikipedia article of another current MP.</p>
<p>This allows us to finally plot an interactive network using the <code>forceNetwork()</code> command from the <a href="https://cran.r-project.org/web/packages/networkD3/index.html"><code>networkD3</code></a> package. We can save the interactive network graph as an HTML widget, which is included below.</p>
<details>
<p><summary> Code: Creating an Interactive Network Graph Based on Article Links</summary></p>
<pre class="r"><code>## step 1: get info about legislators
dat &lt;- semi_join(
  x = get_core(legislature = &quot;deu&quot;),
  y = filter(get_political(legislature = &quot;deu&quot;), session == 19),
  by = &quot;pageid&quot;
)

## step 2: get page links (max 500 links)
if (!file.exists(&quot;studying-politics-wikipedia/data/wikipediR/mdb_links_list.RData&quot;)) {
  links_list &lt;- list()
  for (i in 1:nrow(dat)) {
    links &lt;-
      page_links(
        &quot;de&quot;,
        &quot;wikipedia&quot;,
        page = dat$wikititle[i],
        clean_response = TRUE,
        limit = 500,
        namespaces = 0
      )
    links_list[[i]] &lt;- lapply(links[[1]]$links, &quot;[&quot;, 2) %&gt;% unlist
  }
  save(links_list, file = &quot;studying-politics-wikipedia/data/wikipediR/mdb_links_list.RData&quot;)
} else{
  load(&quot;studying-politics-wikipedia/data/wikipediR/mdb_links_list.RData&quot;)
}

## step 3: identify links between MPs
# loop preparation
connections &lt;- data.frame(from = NULL, to = NULL)
# loop
for (i in seq_along(dat$wikititle)) {
  links_in_pslinks &lt;-
    seq_along(dat$wikititle)[str_replace_all(dat$wikititle, &quot;_&quot;, &quot; &quot;) %in%
                               links_list[[i]]]
  links_in_pslinks &lt;- links_in_pslinks[links_in_pslinks != i]
  connections &lt;-
    rbind(connections,
          data.frame(
            from = rep(i - 1, length(links_in_pslinks)), # -1 for zero-indexing
            to = links_in_pslinks - 1 # here too
            )
          )
}

# results
names(connections) &lt;- c(&quot;from&quot;, &quot;to&quot;)

# make symmetrical
connections &lt;- rbind(connections,
                     data.frame(from = connections$to,
                                to = connections$from))
connections &lt;- connections[!duplicated(connections), ]


## step 4: visualize connections
connections$value &lt;- 1
nodesDF &lt;- data.frame(name = dat$name, group = 1)

network_out &lt;-
  forceNetwork(
    Links = connections,
    Nodes = nodesDF,
    Source = &quot;from&quot;,
    Target = &quot;to&quot;,
    Value = &quot;value&quot;,
    NodeID = &quot;name&quot;,
    Group = &quot;group&quot;,
    zoom = TRUE,
    opacityNoHover = 3,
    height = 360,
    width = 636
  )</code></pre>
</details>
<div style="position:relative;padding-top:56.25%;">
<p><iframe src="/socialsciencedatalab/studying-politics-wikipedia/network_out.html" frameborder="0" allowfullscreen
    style="position:absolute;top:0;left:0;width:100%;height:100%;" scrolling="no" onload="resizeIframe(this)"></iframe></p>
</div>
<p>Using the underlying <code>connections</code> data set, we can also identify which members of the German parliament share the most nodes with others. Perhaps unsurprisingly, we see the German chancellor Angela Merkel on top of the list, followed by a list of current and former federal ministers and (deputy) party leaders.</p>
<details>
<p><summary> Code: Top 10 MPs by Connections Counts</summary></p>
<pre class="r"><code>nodesDF$id &lt;- as.numeric(rownames(nodesDF)) - 1
connections_df &lt;-
  merge(connections,
        nodesDF,
        by.x = &quot;to&quot;,
        by.y = &quot;id&quot;,
        all = TRUE)
to_count_df &lt;- count(connections_df, name)
arrange(to_count_df, desc(n))</code></pre>
</details>
<pre><code>## # A tibble: 712 x 2
##    name                     n
##    &lt;fct&gt;                &lt;int&gt;
##  1 Angela Merkel           59
##  2 Andrea Nahles           40
##  3 Heiko Maas              38
##  4 Katarina Barley         38
##  5 Peter Altmaier          38
##  6 Wolfgang Schäuble       38
##  7 Wolfgang Kubicki        37
##  8 Hans-Peter Friedrich    34
##  9 Hermann Gröhe           34
## 10 Ursula von der Leyen    33
## # ... with 702 more rows</code></pre>
</div>
<div id="application-3-using-clickstream-data-to-analyze-referral-patterns" class="section level5">
<h5>Application 3: Using Clickstream Data to Analyze Referral Patterns</h5>
<p>Wikipedia articles usually <a href="https://en.wikipedia.org/wiki/Wikipedia:About"><em>“provide links designed to guide the user to related pages with additional information”</em></a>. This allows us to collect <a href="https://meta.wikimedia.org/wiki/Research:Wikipedia_clickstream">clickstream data</a>. Clickstreams yield information on the incoming and outgoing traffic of articles. They capture the articles that refer users to a given article as well as the links within a given article that users click to navigate to other articles. Clickstream data are inherently dyadic: Observations represent referral patterns for article-pairs (previous site → current site). Thus, our quantity of interest is the cumulated number of times this pattern was observed in a given period of time.</p>
<p>Clickstream data are offered as monthly aggregate counts for the major Wikipedia language editions. To obtain the data, we first have to download the raw clickstream data from <a href="https://dumps.wikimedia.org/other/clickstream/">this page</a>, where they are offered as compressed files. After extracting the files, we can load them into <strong>R</strong>.</p>
<p>In the example below, we focus on two party groups of the 8th (2014-2019) European Parliament: the euroskeptic EFDD (Europe of Freedom and Direct Democracy) and the far right ENF (Europe of Nations and Freedom). In particular, we are interested in clickstreams between the two party groups, between the party groups and their member parties, and between the individual member parties.</p>
<p>Toward this end, we download clickstream data from the English Wikipedia for May 2019, the month of the 2019 European Parliament elections. We identify 19 articles of interest and store them in the object <code>articles</code>. Having retrieved and extracted the clickstream data from May 2019, we import the TSV file into <strong>R</strong> using <code>read.table()</code>. Lastly, we subset the data to observations that involve referrals between all available article-pairs of the 19 articles.</p>
<details>
<p><summary> Code: Collecting and Processing Clickstream Data </summary></p>
<pre class="r"><code># retrieve article titles of interest
enf &lt;- &quot;Europe_of_Nations_and_Freedom&quot;
efdd &lt;- &quot;Europe_of_Freedom_and_Direct_Democracy&quot;

enf_parties &lt;- c(
  &quot;Freedom_Party_of_Austria&quot;,
  &quot;Vlaams_Belang&quot;,
  &quot;National_Rally_(France)&quot;,
  &quot;The_Blue_Party_(Germany)&quot;,
  &quot;Lega_Nord&quot;,
  &quot;Party_for_Freedom&quot;,
  &quot;Congress_of_the_New_Right&quot;
)

efdd_parties &lt;- c(
  &quot;Svobodní&quot;,
  &quot;The_Patriots_(France)&quot;,
  &quot;Debout_la_France&quot;,
  &quot;Alternative_for_Germany&quot;,
  &quot;Five_Star_Movement&quot;,
  &quot;Order_and_Justice&quot;,
  &quot;Liberty_(Poland)&quot;,
  &quot;Brexit_Party&quot;,
  &quot;Social_Democratic_Party_(UK,_1990–present)&quot;,
  &quot;Libertarian_Party_(UK)&quot;
)

articles &lt;- c(enf, efdd, enf_parties, efdd_parties)

# import raw clickstream data
cs &lt;-
  read.table(
    &quot;clickstream-enwiki-2019-05.tsv&quot;,
    header = FALSE,
    col.names = c(&quot;prev&quot;, &quot;curr&quot;, &quot;type&quot;, &quot;n&quot;),
    fill = TRUE,
    stringsAsFactors = FALSE
  )
cs$n &lt;- as.integer(cs$n)

# subset
cs &lt;- subset(cs, prev %in% articles &amp; curr %in% articles)</code></pre>
</details>
<p><br />
Next, we aim to analyze aggregate referral patterns. We first assign both previous (<code>prev</code>) and current (<code>curr</code>) articles to one of four categories: Articles on the EFDD and ENF parliamentary groups (one article each), articles on ENF member parties (7 articles), and articles on EFDD member parties (10 articles). We then summarize the data to obtain aggregate referral counts between all category pairs. Lastly, we display these in an interactive Sankey diagram using the <a href="https://cran.r-project.org/web/packages/plotly/index.html"><code>plotly</code></a> package.</p>
<details>
<p><summary> Code: Analyzing and Plotting Clickstream Data </summary></p>
<pre class="r"><code># assign categories
cs &lt;- cs %&gt;%
  mutate(
    curr_cat = ifelse(
      curr == enf,
      &quot;ENF Group&quot;,
      ifelse(
        curr == efdd,
        &quot;EFDD Group&quot;,
        ifelse(curr %in% enf_parties, &quot;ENF Parties&quot;,
               &quot;EFDD Parties&quot;)
      )
    ),
    prev_cat = ifelse(
      prev == enf,
      &quot;ENF Group&quot;,
      ifelse(
        prev == efdd,
        &quot;EFDD Group&quot;,
        ifelse(prev %in% enf_parties, &quot;ENF Parties&quot;,
               &quot;EFDD Parties&quot;)
      )
    )
  )

# summarize data
cs_sum &lt;-  cs %&gt;%
  group_by(curr_cat, prev_cat) %&gt;%
  summarize(n = sum(n)) %&gt;%
  arrange(prev_cat)

# Sankey diagram using plotly
labels &lt;- c(unique(cs_sum$prev_cat), unique(cs_sum$curr_cat))
colors &lt;- ifelse(grepl(&quot;EFDD&quot;, labels), &quot;#24B9B9&quot;, &quot;#2B3856&quot;)
sankey_plot &lt;- plot_ly(
  type = &quot;sankey&quot;,
  orientation = &quot;h&quot;,
  
  node = list(
    label = labels,
    color = colors,
    pad = 15,
    thickness = 15,
    line = list(color = &quot;black&quot;,
                width = 0.5)
  ),
  
  link = list(
    source = as.numeric(as.factor(cs_sum$prev_cat)) - 1L,
    target = as.numeric(as.factor(cs_sum$curr_cat)) + 3L,
    value =  cs_sum$n
  ),
  
  height = 340,
  width = 600
) %&gt;%
  layout(font = list(size = 10))</code></pre>
</details>
<div style="position:relative;padding-top:56.25%;">
<iframe src="/studying-politics-wikipedia/sankey_plot.html" frameborder="0" allowfullscreen style="position:absolute;top:0;left:0;width:100%;height:100%;" scrolling="no" onload="resizeIframe(this)">
</iframe>
</div>
<p>     </p>
<p>The diagram shows that in our data, clickstream dyads involving the articles on the EFDD and ENF parliamentary groups are much more numerous than dyads involving the member parties. Much of this can be attributed to clickstreams between the two party groups, EFDD ↔︎ ENF. Whereas clickstreams between members of the same parliamentary group are also fairly frequent, clickstreams between the member of one group to a member of the respective other group are rare.</p>
<p>Moving beyond clickstreams between the four categories, we can also visualize the full network structure of <em>all</em> individual articles in our data. The code below starts with some preparatory data management and then uses the <a href="https://cran.r-project.org/web/packages/igraph/index.html"><code>igraph</code></a> package to create the network and to customize its graphical display.</p>
<p>In the final section of the code, we use the <a href="https://cran.r-project.org/web/packages/intergraph/index.html"><code>intergraph</code></a>, <a href="https://cran.r-project.org/web/packages/ggnetwork/index.html"><code>ggnetwork</code></a> and <a href="https://cran.r-project.org/web/packages/plotly/index.html"><code>plotly</code></a> packages to produce an interactive HTML5-compatible figure for this blog post. On your own machine, you may skip this section and simply use <code>plot.igraph()</code> on <code>cs_net</code> without transforming the object to a <code>ggplot</code> friendly format.</p>
<details>
<p><summary> Code: Interactive Network Graph </summary></p>
<pre class="r"><code># construct edges
cs_edge &lt;-
  cs %&gt;%
  group_by(prev, curr, prev_cat, curr_cat) %&gt;%
  dplyr::summarise(weight = sum(n)) %&gt;%
  arrange(curr)

# get list of unique articles to construct as nodes
cs_node &lt;- 
  gather(cs_edge,
         `prev`,
         `curr`,
         key = &quot;where&quot;,
         value = &quot;article&quot;) %&gt;%
  ungroup() %&gt;%
  select(article) %&gt;%
  distinct(article)
names(cs_node) &lt;- c(&quot;node&quot;)
cs_node$category &lt;-
  ifelse(cs_node$node == enf,
         &quot;ENF Group&quot;,
         ifelse(
           cs_node$node == efdd,
           &quot;EFDD Group&quot;,
           ifelse(cs_node$node %in% enf_parties, &quot;ENF Parties&quot;,
                  &quot;EFDD Parties&quot;)
         )
  )

# generate graph
set.seed(3)
cs_net &lt;-
  graph.data.frame(cs_edge,
                   vertices = cs_node,
                   directed = F)
cs_net &lt;-
  igraph::simplify(cs_net, remove.multiple = T, remove.loops = T)

# generate colors based on category
V(cs_net)$color &lt;- 
  ifelse(grepl(&quot;EFDD&quot;, V(cs_net)$category), &quot;#24B9B9&quot;, &quot;#2B3856&quot;)

# compute node degrees (#links) and use that to set node size
deg &lt;- igraph::degree(cs_net, mode = &quot;all&quot;)
V(cs_net)$size &lt;- deg / 10

# set labels
V(cs_net)$label &lt;- NA
V(cs_net)$label.cex = 0.5
V(cs_net)$label = ifelse(igraph::degree(cs_net) &gt; 5, V(cs_net)$label, NA)
cs_hc_labels &lt;- as.vector(cs_node$node)

# set edge width based on weight
E(cs_net)$width &lt;- log(E(cs_net)$weight) / 5
E(cs_net)$edge.color &lt;- &quot;gray80&quot;

# transform the network to a ggplot friendly format
# (required to generate interactive graph embedded in blog post)
gg_cs_net &lt;-
  ggnetwork(
    cs_net,
    layout = &quot;fruchtermanreingold&quot;,
    weights = &quot;weight&quot;,
    niter = 50000,
    arrow.gap = 0
  )

cs_plot &lt;- ggplot(gg_cs_net, aes(x = x, y = y, xend = xend, yend = yend)) +
  geom_edges(aes(color = edge.color), size = 0.4, alpha = 0.25) +
  geom_nodes(aes(color = color, size = size)) +
  geom_nodetext(aes(color = color, label = vertex.names, cex = 0.6)) +
  guides(size=FALSE) +
  theme_blank() +
  theme(legend.position = &quot;none&quot;)</code></pre>
</details>
<div style="position:relative;padding-top:56.25%;">
<p><iframe src="/socialsciencedatalab/studying-politics-wikipedia/cs_plot.html" frameborder="0" allowfullscreen
      style="position:absolute;top:0;left:0;width:100%;height:100%;" scrolling="no" onload="resizeIframe(this)"></iframe></p>
</div>
<p> </p>
<p>In the graph above, node diameters indicate the relative weight (total counts) of each article; node colors indicate whether an articles belongs to the EFDD or ENF. We see that members of the same party group tend to share more clickstreams. The Alternative for Germany (AfD), however, shares many connections with members of the ENF. This makes sense when we consider that the AfD has sought closer cooperation with numerous ENF member parties since 2016 with whom it eventually formed the new far right EP group, Identity and Democracy, in June 2019.</p>
<p>Lastly, a word of caution: One should keep in mind that clickstream counts heaviliy depend on how prominently (if at all) outgoing links are placed in a given Wikipedia article. Furthermore, raw counts from an isolated subset of clickstreams (as in the examples above) give no information on the relative importance of a given referral pattern relative to all outgoing referrals of a given article. Users should thus ensure that they use clickstream data in a way that adequately addresses their substantive inquiries.</p>
</div>
</div>
<div id="collecting-data-via-wikidata-queries" class="section level3">
<h3>Collecting Data via Wikidata Queries</h3>
<p>Wikidata is a collaboratively edited knowledge base with <a href="https://www.wikidata.org/wiki/Wikidata:Statistics">over 58 million entries as of July 2019</a>. It harbors various types of database items, including text, numerical quantities, coordinates, and images. There are no language editions, but individual entries can have values in different languages.</p>
<p>Wikidata allows users to submit queries using SPARQL, a query language for data stored in RDF (Resource Description Framework) format (see <a href="https://query.wikidata.org/">this link</a>). Click <a href="https://towardsdatascience.com/a-brief-introduction-to-wikidata-bb4e66395eb1">here</a> for a brief introduction to SPARQL. While basic queries can be used to answer mundane questions (e.g. <em>“what is the capital city of every member of the European Union, and how many inhabitants live there?”</em>), a targeted combination of related queries can be used for systematic data collection.</p>
<p>Instead of submitting explicit SPARQL queries, the example below uses the <code>WikidataR</code> package to combine various queries in order to collect data on the candidates in the <a href="https://en.wikipedia.org/wiki/2019_Conservative_Party_(UK)_leadership_election">2019 leadership election of the UK Conservative Party</a>. Suppose we want to retrieve the following information on each candidate:</p>
<ul>
<li>name</li>
<li>sex</li>
<li>date of birth</li>
<li>political experience</li>
<li>education</li>
<li>official website URL</li>
<li>Twitter accout</li>
<li>Facebook account</li>
</ul>
<p>In Wikidata, entries are stored as <em>items</em> with a unique item ID that starts with “Q”. For instance, the item <a href="https://www.wikidata.org/wiki/Q30325756">2019 Conservative Party (UK) leadership election</a> is stored as “Q30325756”. Items are characterized by a number of <em>statements</em> or <em>claims</em>. Claims start with “P” and detail an item’s properties. For instance, the claim “candidate” is stored as “P726”. Claims have values, which may once again be items. For example, the values of claim “P726” (candidate) of item “Q30325756” (2019 UK Conservative Party leadership election) are 10 items: one entry for each of the 10 candidates running in the leadership election. Take, for example, winning candidate Boris Johnson, who is listed as a candidate under claim “P726”. In turn, the entry on Boris Johnson is stored as item “Q180589”. This item is characterized by numerous claims, including “P1559” (name in native language), “P21” (sex or gender), “P569” (date of birth), “P39” (positions held), “P69” (educated at), “P856” (official website), “P2002” (Twitter username), and “P2013” (Facebook ID).</p>
<p>In order to collect the data for all 10 candidates in the 2019 Conservative Party leadership election, the code chunk below implements the following steps:</p>
<ol style="list-style-type: decimal">
<li>We retrieve <em>item</em> “Q30325756”, i.e., the entry for 2019 Conservative Party (UK) leadership election</li>
<li>We extract <em>claims</em> “P726” of the above item to retrieve the item IDs of all 10 candidates, which we store in the object <code>candidates</code></li>
<li>We save the IDs of the claims of interest, stored in the object <code>claims</code></li>
<li>We then use some nested <code>sapply</code>commands to do the following:
<ul>
<li>Retrieve the <em>item</em> (entry) for each candidate</li>
<li>Extract the eight <em>claims</em> from each candidate <em>item</em></li>
<li>Process the informational value of each extracted claim, depending on whether the claim value is
<ul>
<li>an atomic object (such as web site URLs)</li>
<li>a textual object with auxiliary information (such as names, which come with language information)</li>
<li>a time/date (such as date of birth)</li>
<li>yet another item (such as previous positions, where each position has an own data base entry)</li>
</ul></li>
</ul></li>
</ol>
<details>
<p><summary> Code: Retrieving Items and Claims from Wikidata</summary></p>
<pre class="r"><code># get item based on item id
uk_item &lt;- get_item(&quot;Q30325756&quot;, language = &quot;en&quot;)

# extract candidates
candidates &lt;- extract_claims(uk_item, claims = &quot;P726&quot;)
candidates &lt;- candidates[[1]][[1]]$mainsnak$datavalue$value$id

# collect the following attributes (&quot;claims&quot;) for each candidate
claims &lt;- c(&quot;P1559&quot;, &quot;P21&quot;, &quot;P569&quot;, &quot;P39&quot;, &quot;P69&quot;, &quot;P856&quot;, &quot;P2002&quot;, &quot;P2013&quot;)
names(claims) &lt;- c(&quot;nam&quot;, &quot;sex&quot;, &quot;dob&quot;, &quot;exp&quot;, &quot;edu&quot;, &quot;web&quot;, &quot;twi&quot;, &quot;fbk&quot;)
claims</code></pre>
<pre><code>##     nam     sex     dob     exp     edu     web     twi     fbk 
## &quot;P1559&quot;   &quot;P21&quot;  &quot;P569&quot;   &quot;P39&quot;   &quot;P69&quot;  &quot;P856&quot; &quot;P2002&quot; &quot;P2013&quot;</code></pre>
<pre class="r"><code># retrieve data
uk_data &lt;-
  sapply(candidates,
         function (item) {
           tmp_item &lt;- get_item(item, language = &quot;en&quot;)
           sapply(claims,
                  function(claim) {
                    tmp_claim &lt;- extract_claims(tmp_item, claim)[[1]][[1]]
                    if (any(is.na(tmp_claim))) {
                      return(NA)
                    } else {
                      tmp_claim &lt;- tmp_claim$mainsnak$datavalue$value
                      if (is.atomic(tmp_claim)) {
                        return(tmp_claim)
                      } else if (&quot;text&quot; %in% names(tmp_claim)) {
                        return(tmp_claim$text)
                      } else if (&quot;time&quot; %in% names(tmp_claim)) {
                        tmp_claim &lt;- as.Date(substr(tmp_claim$time, 2, 11))
                        return(tmp_claim)
                      } else if (&quot;id&quot; %in% names(tmp_claim)) {
                        tmp_claim &lt;- tmp_claim$id
                        tmp_claim &lt;- 
                          sapply(tmp_claim, 
                                 get_item, 
                                 language = &quot;en&quot;,
                                 simplify = FALSE,
                                 USE.NAMES = TRUE)
                        tmp_claim &lt;-
                          sapply(tmp_claim,
                                 function (x) {
                                   x[[1]]$labels$en$value
                                 })
                        return(tmp_claim)
                      }
                    }
                  },
                  simplify = FALSE,
                  USE.NAMES = TRUE)
         },
         simplify = FALSE,
         USE.NAMES = TRUE
  )</code></pre>
</details>
<p><br />
The retrieved data are stored in a nested list. At the upper level of the list, we have the ten candidates, named with their respective item IDs. Nested within each of the ten upper-level elements, we have the values of the eight claims, named with the labels we specified above. Claim values may either be atomic (such as date of birth) or vectors (such as “positions held”, which may have multiple entries). Below, we can see the retrieved data for the first candidate on the list, winning candidate Boris Johnson.</p>
<details>
<p><summary> Output: Retrieved Data for Boris Johnson</summary></p>
<pre><code>## $nam
## [1] &quot;Boris Johnson&quot;
## 
## $sex
## Q6581097 
##   &quot;male&quot; 
## 
## $dob
## [1] &quot;1964-06-19&quot;
## 
## $exp
##                                                    Q38931 
##                                         &quot;Mayor of London&quot; 
##                                                  Q1371091 
## &quot;Secretary of State for Foreign and Commonwealth Affairs&quot; 
##                                                 Q28841847 
##       &quot;Member of the Privy Council of the United Kingdom&quot; 
##                                                 Q30524710 
##     &quot;Member of the 57th Parliament of the United Kingdom&quot; 
##                                                 Q30524718 
##     &quot;Member of the 56th Parliament of the United Kingdom&quot; 
##                                                 Q35647955 
##     &quot;Member of the 54th Parliament of the United Kingdom&quot; 
##                                                 Q35921591 
##     &quot;Member of the 53rd Parliament of the United Kingdom&quot; 
##                                                    Q14211 
##                    &quot;Prime Minister of the United Kingdom&quot; 
##                                                  Q3303456 
##                        &quot;Leader of the Conservative Party&quot; 
##                                                 Q77685926 
##     &quot;Member of the 58th Parliament of the United Kingdom&quot; 
##                                                   Q609884 
##                              &quot;First Lord of the Treasury&quot; 
##                                                  Q3315116 
##                          &quot;Minister for the Civil Service&quot; 
##                                                 Q65988624 
##                                  &quot;Minister for the Union&quot; 
## 
## $edu
##                       Q192088                       Q805285 
##                &quot;Eton College&quot;             &quot;Balliol College&quot; 
##                      Q4804780                      Q5413121 
##        &quot;Ashdown House School&quot; &quot;European School, Brussels I&quot; 
## 
## $web
## [1] &quot;http://www.boris-johnson.com&quot;
## 
## $twi
## [1] &quot;BorisJohnson&quot;
## 
## $fbk
## [1] &quot;borisjohnson&quot;</code></pre>
</details>
<p><br /></p>
</div>
<div id="legislator" class="section level3">
<h3>legislatoR</h3>
<p><a href="https://github.com/saschagobel/legislatoR"><code>legislatoR</code></a> is a joint project of <a href="https://github.com/saschagobel">Sascha Göbel</a> and <a href="https://simonmunzert.github.io/">Simon Munzert</a>. It offers a comprehensive relational individual-level database that provides political, sociodemographic, and other Wikipedia-related data on members of various national parliaments, including the all sessions of the Austrian Nationalrat, the German Bundestag, the Irish Dáil, the French Assemblée, and the United States Congress (House and Senate). It currently comprises data of 42,534 elected representatives and holds information for a wide variety of variables, including:</p>
<ul>
<li>sociodemographics (<em>Core</em>)</li>
<li>basic political variables (<em>Political</em>)</li>
<li>records of individual Wikipedia data, including full revision histories (<em>History</em>)</li>
<li>daily user traffic on individual Wikipedia biographies (<em>Traffic</em>)</li>
<li>social media handles and website URLs (<em>Social</em>)</li>
<li>URLs to individual Wikipedia portraits (<em>Portraits</em>)</li>
<li>information on public offices held by MPs (<em>Offices</em>)</li>
<li>MPs’ occupations (<em>Professions</em>)</li>
<li>IDs that link politicians to other files, databases and websites (<em>IDs</em>)</li>
</ul>
<p>The figure below, taken from <span class="citation">Göbel and Munzert (2019)</span>, illustrates the data structure:</p>
<p><img src="/../../../../../article/studying-politics-wikipedia/img/data-structure.png" width="60%" style="display: block; margin: auto;" /></p>
<p>The package provides a relational database. This means that all data sets can be joined with the core data set via one of two keys: the Wikipedia page ID or the Wikidata ID, which uniquely identify individual politicians.</p>
<p><code>legislatoR</code> services the increasing demand for micro-level data on political elites among political scientists, political analysts, and journalists and offers an accessible and rich collection of data on past and present politicians. The inclusion of Wikipedia and other web data allows for the inclusion of detailed information on politicians’ biographies.</p>
<p>To install the current developmental version from <a href="https://github.com/saschagobel/legislatoR">GitHub</a>, we use the <code>devtools</code> package. After installing and loading <code>legislatoR</code>, we can use the <code>ls()</code> command to explore the full functionality of the package.</p>
<details>
<p><summary> Code: Installing <code>legislatoR</code> </summary></p>
<pre class="r"><code>## Install from GitHub
devtools::install_github(&quot;saschagobel/legislatoR&quot;)
library(legislatoR)

## View functionality
ls(&quot;package:legislatoR&quot;)</code></pre>
</details>
<p><br /></p>
<div id="application-1-social-media-adoption-rates" class="section level5">
<h5>Application 1: Social Media Adoption Rates</h5>
<p>To retrieve <code>legislatoR</code> data, we first load the entire core data set of a given national parliament using the <code>get_core()</code> command. In the example below, we focus on the German Bundestag. We immediately <code>right_join()</code> the core data set with the <em>political</em> component using <code>get_political()</code>. We then <code>filter()</code> the data to retain the legislative <code>session</code> of interest (here, the most recent session of the Bundestag, 2017-2021).</p>
<p>In the next step, we <code>left_join()</code> this data set with the <em>social</em> component using <code>get_social()</code>. This gives us full information on the social media accounts of all MPs of the 2017-2021 Bundestag. Whenever MPs do not have an account, this is stored as missing information (<code>NA</code>). Using this information, we can calculate the social media adoption rates in the German parliament.</p>
<details>
<p><summary> Code: Retrieving <code>legislatoR</code> Data and Calculating Social Media Adoption Rates </summary></p>
<pre class="r"><code>## Get social media adoption rates
# get data: Germany
dat_ger &lt;- right_join(
  x = get_core(legislature = &quot;deu&quot;),
  y = filter(
    get_political(legislature = &quot;deu&quot;),
    as.numeric(session) == max(as.numeric(session))
  ),
  by = &quot;pageid&quot;
)
dat_ger &lt;- left_join(x = dat_ger,
                     y = get_social(legislature = &quot;deu&quot;),
                     by = &quot;wikidataid&quot;)
dat_ger$legislature &lt;- &quot;Germany&quot;
dat_ger_sum &lt;- dat_ger %&gt;%
  dplyr::summarize(
    twitter = mean(not(is.na(twitter)), na.rm = TRUE),
    facebook = mean(not(is.na(facebook)), na.rm = TRUE),
    website = mean(not(is.na(website)), na.rm = TRUE),
    session_start = ymd(first(session_start)),
    session_end = ymd(first(session_end)),
    legislature = first(legislature)
  )</code></pre>
</details>
<pre><code>##     twitter  facebook  website session_start session_end legislature
## 1 0.7591036 0.6918768 0.640056    2017-10-24  2021-10-24     Germany</code></pre>
<p>Given the inclusion of <em>political</em> variables in our data set, we could think of numerous feasible extensions. For instance, we could look at social media adoption rates by <code>party</code>. Alternatively, we could use the <code>constituency</code> identifiers to add external data on the rurality/urbanity of German electoral districts and analyze whether politicians competing in urban districts are more likely to maintain social media profiles.</p>
</div>
<div id="application-2-public-attention-to-members-of-the-german-bundestag" class="section level5">
<h5>Application 2: Public Attention to Members of the German Bundestag</h5>
<p>In the second application, we use pageviews to identify peaks in public attention for MPs over time. This is particularly interesting in the context of politically significant events. For instance, we may want to know about public attention to parliamentarians following scandals, around elections, or during election campaigns. The code below illustrates this logic by averaging daily pageviews across all Wikipedia articles on members of the German Bundestag between July 2015 and December 2017.</p>
<details>
<p><summary> Code: Plotting Average Daily Pageviews for German MPs </summary></p>
<pre class="r"><code>## Visualize average pageviews data of German MPs
# get data
ger_traffic &lt;- right_join(
  x = get_traffic(legislature = &quot;deu&quot;),
  y = filter(
    get_political(legislature = &quot;deu&quot;),
    session_end &gt;= as.Date(&quot;2015-07-01&quot;)
  ),
  by = &quot;pageid&quot;
)
ger_traffic &lt;- left_join(x = ger_traffic,
                         y = get_core(legislature = &quot;deu&quot;),
                         by = &quot;pageid&quot;)
ger_traffic &lt;-
  dplyr::select(ger_traffic, pageid, date, traffic, session, party, name)

# aggregate data
ger_traffic$date &lt;- ymd(ger_traffic$date)
ger_traffic_date &lt;- group_by(ger_traffic, date)
ger_traffic_legislators &lt;- group_by(ger_traffic, pageid)
ger_traffic_sum &lt;-
  summarize(ger_traffic_date, mean = mean(traffic, na.rm = TRUE))
ger_traffic_sum &lt;- mutate(
  ger_traffic_sum,
  mean_l1 = lag(mean, 1),
  mean_f1 = lead(mean, 1),
  peak = (mean &gt;= 1.8 * mean_l1 &amp;
            mean &gt; 180)
)

# identify peaks
ger_traffic_peaks &lt;- filter(ger_traffic_sum, peak == TRUE)
ger_traffic_peaks_df &lt;-
  filter(ger_traffic, date %in% ger_traffic_peaks$date)
ger_traffic_peaks_group &lt;-
  group_by(ger_traffic_peaks_df, date) %&gt;% 
  dplyr::arrange(desc(traffic)) %&gt;% 
  filter(row_number() == 1)
ger_traffic_peaks_group &lt;- arrange(ger_traffic_peaks_group, date)
events_vec &lt;-
  c(
    &quot;deceased&quot;,
    &quot;drug affair&quot;,
    &quot;bullying affair&quot;,
    &quot;candidacy for presidency&quot;,
    &quot;deceased&quot;,
    &quot;???&quot;,
    &quot;chancellorchip announcement&quot;,
    &quot;elected president&quot;,
    &quot;???&quot;,
    &quot;policy success&quot;,
    &quot;TV debate&quot;,
    &quot;general election&quot;,
    &quot;threat to resign&quot;,
    &quot;elected speaker of parliament&quot;
  )

# plot
par(oma = c(0, 0, 0, 0))
par(mar = c(0, 4, 0, .5))
par(yaxs = &quot;i&quot;, xaxs = &quot;i&quot;, bty = &quot;n&quot;)
layout(matrix(c(1, 1, 3, 2, 2, 3), 2, 3, byrow = TRUE),
       heights = c(1, 2, 3),
       widths = c(5, 5, 1.8))
# names labels
plot(
  ymd(ger_traffic_sum$date),
  rep(0, length(ger_traffic_sum$date)),
  xlim = c(ymd(&quot;2015-07-01&quot;), ymd(&quot;2018-01-01&quot;)),
  xaxt = &quot;n&quot;,
  ylim = c(0, 8),
  yaxt = &quot;n&quot;,
  xlab = &quot;&quot;,
  ylab = &quot;&quot;,
  cex = 0
)
text(
  ger_traffic_peaks_group$date,
  0,
  ger_traffic_peaks_group$name,
  cex = .75,
  srt = 90,
  adj = c(0, 0)
)
# pageviews time series
par(mar = c(2, 4, 0, .5))
plot(
  ymd(ger_traffic_sum$date),
  ger_traffic_sum$mean,
  type = &quot;l&quot;,
  ylim = c(0, 1.25 * max(ger_traffic_sum$mean)),
  xlim = c(ymd(&quot;2015-07-01&quot;), ymd(&quot;2018-01-01&quot;)),
  xaxt = &quot;n&quot;,
  yaxt = &quot;n&quot;,
  xlab = &quot;&quot;,
  ylab = &quot;mean(pageviews)&quot;,
  col = &quot;white&quot;
)
abline(h = seq(0, 1.5 * max(ger_traffic_sum$mean), 250), col = &quot;lightgrey&quot;)
lines(ymd(ger_traffic_sum$date), ger_traffic_sum$mean, lwd = .5)
dates &lt;- seq(ymd(&quot;2015-07-01&quot;), ymd(&quot;2018-01-01&quot;), by = 1)
axis(1, dates[day(dates) == 1 &amp;
                month(dates) %in% c(1, 4, 7, 10)], labels = FALSE)
axis(1,
     dates[day(dates) == 1 &amp;
             month(dates) %in% c(1)],
     lwd = 0,
     lwd.ticks = 3,
     labels = FALSE)
axis(1,
     dates[day(dates) == 1 &amp;
             month(dates) %in% c(7)],
     labels = as.character(year(dates[day(dates) == 15 &amp;
                                        month(dates) %in% c(7)])),
     tick = F,
     lwd = 0)
axis(2, seq(0, 1.5 * max(ger_traffic_sum$mean), 250), las = 2)
# events labels in time series
for (i in seq_along(events_vec)) {
  text(ger_traffic_peaks_group$date[i],
       ger_traffic_sum$mean[ger_traffic_sum$peak == TRUE][i] + 80,
       i,
       cex = .8)
  points(
    ger_traffic_peaks_group$date[i],
    ger_traffic_sum$mean[ger_traffic_sum$peak == TRUE][i] + 80,
    pch = 1,
    cex = 2.2
  )
}
# election date
# events labels explained
par(mar = c(0, 0, 0, 0))
plot(
  0,
  0,
  xlim = c(0, 5),
  ylim = c(0, 10),
  xaxt = &quot;n&quot;,
  yaxt = &quot;n&quot;,
  xlab = &quot;&quot;,
  ylab = &quot;&quot;,
  cex = 0
)
positions &lt;-
  data.frame(
    events_xpos = 0.45,
    events_ypos = seq(6.5, (6.5 - .5 * length(events_vec)),-.5),
    text_xpos = .5
  )
text(0, 7, &quot;Events&quot;, pos = 4, cex = .75)
for (i in seq_along(events_vec)) {
  text(positions$events_xpos[i], positions$events_ypos[i], i, cex = .8)
  points(positions$events_xpos[i],
         positions$events_ypos[i],
         pch = 1,
         cex = 2.2)
  text(
    positions$text_xpos[i],
    positions$events_ypos[i],
    events_vec[i],
    pos = 4,
    cex = .75
  )
}</code></pre>
</details>
<img src="/../../../../../article/studying-politics-wikipedia_files/figure-html/Code%2013b-1.png" width="960" style="display: block; margin: auto;" />
</details>
<p><br/>
The plot shows 14 notable spikes in daily pageviews. We identify which politician’s article generated the most traffic during each of these events. Furthermore, we add a legend that lists salient political events that likely caused these spikes in attention to MPs Wikipedia entries. For example, spike 14 marks the day on which Wolfgang Schäuble (CDU) was elected speaker of the parliament.</p>
</div>
</div>
<div id="conclusion" class="section level3">
<h3>Conclusion</h3>
<p>Collecting and analyzing Wikipedia data is relatively easy and entirely free. It enables researchers to use and analyze an enormous body of data that offers valuable information for research in political science and beyond. Tools that facilitate the collection, processing, and analysis of Wikipedia data advance rapidly, broadening the realm of possibilities for scientific research. Political science research is increasingly picking up on these developments, as is evident in recent contributions <span class="citation">(Munzert 2015; Göbel and Munzert 2018; Shi et al. 2019)</span> and softwares (such as <a href="https://github.com/saschagobel/legislatoR"><code>legislatoR</code></a>).</p>
<p>However, using Wikipedia data may also come with limitations and pitfalls. As entries can be read and edited by both humans and machines, the accuracy of contents and the validity of metadata are not guaranteed. With respect to the latter, Wikidata adds provenance information to all the data. These can be used to evaluate the validity of the data in question for applied research.
Researchers should also keep in mind that Wikipedia data highly depends on user-driven creation, editing, and use of contents. This may not only lead to systematic selection bias due to data availability but also induce problems of equivalence of data points (e.g., articles on historical political figures likely receive fewer views and edits than articles on active politicians for reasons unrelated to their legislative activity or real-world importance).</p>
<p>These caveats are however all but exclusive to Wikipedia data. They merely underline that Wikipedia data is no exception when it comes to the general necessity of thoroughly scrutinizing and critically assessing the suitability of any given data for addressing substantive research questions.</p>
</div>
<div id="about-the-presenter" class="section level3">
<h3>About the Presenter</h3>
<p>Simon Munzert <a href="mailto:munzert@hertie-school.org"><i class="fa
              fa-envelope"></i> </a>
<a href="https://simonmunzert.github.io/"><i class="fa
              fa-globe"></i> </a>
<a href="https://twitter.com/simonsaysnothin"><i class="fa
              fa-twitter"></i></a> is a lecturer in Political Data Science at the Hertie School of Governance in Berlin, Germany. A former member of the MZES Data and Methods Unit, Simon founded the Social Science Data Lab in 2016. His research focuses on public opinion, political representation, and the role of new media for political processes.</p>
</div>
<div id="references" class="section level3 unnumbered">
<h3 class="unnumbered">References</h3>
<div id="refs" class="references hanging-indent">
<div id="ref-Gobel2018">
<p>Göbel, Sascha, and Simon Munzert. 2018. “Political Advertising on the Wikipedia Marketplace of Information.” <em>Social Science Computer Review</em> 36 (2): 157–75. <a href="https://doi.org/10.1177/0894439317703579">https://doi.org/10.1177/0894439317703579</a>.</p>
</div>
<div id="ref-Gobel2019">
<p>———. 2019. “legislatoR: Political, sociodemographic, and Wikipedia-related data on political elites.” <a href="https://github.com/saschagobel">https://github.com/saschagobel</a>.</p>
</div>
<div id="ref-Munzert2015a">
<p>Munzert, Simon. 2015. “Using Wikipedia Article Traffic Volume to Measure Public Issue Attention.” <em>Working Paper</em>. <a href="https://github.com/simonmunzert/workingPapers/blob/master/wikipedia-salience-v3.pdf">https://github.com/simonmunzert/workingPapers/blob/master/wikipedia-salience-v3.pdf</a>.</p>
</div>
<div id="ref-Shi2019">
<p>Shi, Feng, Misha Teplitskiy, Eamon Duede, and James A. Evans. 2019. “The wisdom of polarized crowds.” <em>Nature Human Behaviour</em> 3 (4): 329–36. <a href="https://doi.org/10.1038/s41562-019-0541-6">https://doi.org/10.1038/s41562-019-0541-6</a>.</p>
</div>
</div>
</div>
]]>
      </description>
    </item>
    
    <item>
      <title>Quantitative Analysis of Political Text</title>
      <link>https://socialsciencedatalab.mzes.uni-mannheim.de/article/quantitative-analysis-of-political-text/</link>
      <pubDate>Mon, 22 Jul 2019 00:00:00 +0100</pubDate>
      
      <guid>https://socialsciencedatalab.mzes.uni-mannheim.de/article/quantitative-analysis-of-political-text/</guid>
      <description><![CDATA[
        </p>
<p>How can we infer actors’ positions, substantive topics, or sentiments from (political) texts? This <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/categories/tutorials/">Methods Bites Tutorial</a> by <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/page/team/">Julian Bernauer</a> summarizes <a href="https://denisetraber.net/">Denise Traber</a>’s workshop in the <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/page/events/">MZES Social Science Data Lab</a> in Spring 2018. Using exemplary sets of political documents (election manifestos and coalition agreements), it showcases tools of QTA for a variety of analytical objectives and demonstrates how to create, process, and analyse a text corpus through a series of hands-on applications.</p>
<p>After reading this blog post and engaging with the applied exercises, readers should:</p>
<ul>
<li>be able to perform some basic preprocessing of text</li>
<li>be able to estimate the sentiment of texts</li>
<li>be able to find topics in texts</li>
<li>be able to estimate (scale) positions of texts</li>
</ul>
<p>You can use these links to navigate across the main sections of this tutotial:</p>
<ol style="list-style-type: decimal">
<li><a href="#tour"><strong>A tour of Quantitative Text Analysis</strong></a></li>
<li><a href="#preprocessing"><strong>(Pre-)processing text</strong></a></li>
<li><a href="#smallcoalition"><strong>A small coalition corpus</strong></a></li>
<li><a href="#sentimentanalysis"><strong>Sentiment analysis using a dictionary</strong></a></li>
<li><a href="#lda"><strong>LDA topic modeling</strong></a></li>
<li><a href="#wordfish"><strong>Wordfish scaling</strong></a></li>
<li><a href="#intraparty"><strong>Estimating intra-party preferences: Comparing speeches to votes</strong></a></li>
<li><a href="#furtherreadings"><strong>Further readings</strong></a></li>
</ol>
<p><em>Note:</em> This blog post presents Denise’s workshop materials in condensed form. The complete workshop materials, including slides and scripts, are available from our GitHub.</p>
<div id="a-tour-of-quantitative-text-analysis" class="section level3">
<h3>A tour of Quantitative Text Analysis <a name="tour"></a></h3>
<p>The workshop started with a few basics: While QTA can be efficient and cheap, it always fails to rely on a correct model of language. It does not free us from reading texts, and validation is key. We learned about the basic distinction between classification (organizing text into categories) and scaling (estimation positions of actors), and its supervised (where hand-coded or other external data is available) and unsupervised (without such data) variants.</p>
</div>
<div id="pre-processing-text" class="section level3">
<h3>(Pre-)processing text <a name="preprocessing"></a></h3>
<p>We relied on the R package <a href="http://quanteda.io"><strong>quanteda</strong></a> developed by Ken Benoit and collaborators, which takes QTA by storm, at least for those working in R. Together with the <a href="https://cran.r-project.org/web/packages/readtext/index.html"><strong>readtext</strong></a> package, it easily allows to get your text data into R, create a so-called “corpus” of texts with the actual content as well as meta-information, and perform various tasks of corpus and text processing (subsetting a corpus, creating a document-feature matrix (dfm), stopword removal) as well as analysis (scaling, classification). A large and increasing number of extras is also available, such as ways to assess text similarity (function <code>textsta_simil()</code>) and lexical diversity (<code>textstat_lexdiv()</code>). Some of these features are demonstrated in an example below. Also see <a href="http://quanteda.io/reference/index.html">this overview</a> by quanteda for a full list of functions and the <a href="https://cran.r-project.org/web/packages/preText/vignettes/getting_started_with_preText.html"><strong>preText</strong></a> package for advise on evaluating pre-processing specifications.</p>
</div>
<div id="a-small-coalition-corpus" class="section level3">
<h3>A small coalition corpus <a name="smallcoalition"></a></h3>
<p>For a few examples from the workshop, consider a small set of three documents: The coalition agreement between the CDU/CSU and the SPD as well as the respective election manifestos from the 2017 Bundestag election. The corpus is created by:</p>
<pre class="r"><code>library(readtext)
library(quanteda)
text &lt;- readtext(paste0(wd, &quot;coalition/*.txt&quot;),
                 docvarsfrom = &quot;filenames&quot;,
                 docvarnames = &quot;Party&quot;)
text$text &lt;- gsub(&quot;\n&quot;, &quot; &quot;, text$text)
coalitioncorpus &lt;- corpus(text, docid_field = &quot;doc_id&quot;)
coalitioncorpus$metadata$source &lt;- &quot;[directory] on [system] by [user]&quot;
summary(coalitioncorpus)</code></pre>
<pre><code>## Corpus consisting of 3 documents:
## 
##           Text Types Tokens Sentences     Party
##     cducsu.txt  4738  26004      1288    cducsu
##  coalition.txt 11660  93214      3763 coalition
##        spd.txt  7650  50298      2402       spd
## 
## Source: [directory] on [system] by [user]
## Created: Wed Nov 11 15:06:41 2020
## Notes:</code></pre>
<p>The code relies on the two packages, <strong>readtext</strong> and <strong>quanteda</strong>, to create a data frame with the text files, using their names for a document-level variable called “Party”. The <code>gsub()</code> command removes whitespace, and <code>corpus()</code> turns the data frame into a corpus, which is a special case of a data frame containing texts, some meta-information and document-level variables, all optimized to perform a variety of quantitative text analysis operations using quanteda.</p>
<p>Further document-level variables are added via:</p>
<pre class="r"><code>docvars(coalitioncorpus, &quot;Year&quot;) &lt;- 2017
docvars(coalitioncorpus, &quot;Party_regex&quot;) &lt;- 
  sub(&quot;[\\.].*&quot;, &quot;&quot;, names(texts(coalitioncorpus)))
docvars(coalitioncorpus)</code></pre>
<pre><code>##                   Party Year Party_regex
## cducsu.txt       cducsu 2017      cducsu
## coalition.txt coalition 2017   coalition
## spd.txt             spd 2017         spd</code></pre>
<p>Note that this uses a regular expression (regex) to alternatively retrieve the party names from the filenames after creating the corpus. For specific analyses, we want to know the distribution of words across documents and create a document-feature matrix (dfm):</p>
<pre class="r"><code>dfm_coal &lt;- dfm(
  coalitioncorpus,
  remove = c(stopwords(&quot;german&quot;),
             &quot;dass&quot;,
             &quot;sowie&quot;,
             &quot;insbesondere&quot;),
  remove_punct = TRUE,
  stem = FALSE
)
dfm_coal[, 1:8]</code></pre>
<pre><code>## Document-feature matrix of: 3 documents, 8 features (16.7% sparse).
## 3 x 8 sparse Matrix of class &quot;dfm&quot;
##                features
## docs            gutes land zeit deutschland liebens lebenswertes gut
##   cducsu.txt        6   48   11         147       1            1  16
##   coalition.txt     1   39   14         195       0            0  13
##   spd.txt           6   44   33          97       0            0  17
##                features
## docs            wohnen
##   cducsu.txt         1
##   coalition.txt     10
##   spd.txt            6</code></pre>
<p>Creating a dfm induces a bag-of-words assumption. This means that the order in which words appear is ignored. A dfm is a means of information reduction and the most efficient way of storing text as data, but allows only analyses under this assumption. We quickly glance at the similarity (function <code>textstat_simil()</code>) and lexical diversity (function <code>textstat_lexdiv()</code>) of texts:</p>
<pre class="r"><code>simil &lt;- textstat_simil(dfm_coal,
                        margin = &quot;documents&quot;,
                        method = &quot;correlation&quot;)
simil </code></pre>
<pre><code>## textstat_simil object; method = &quot;correlation&quot;
##               cducsu.txt coalition.txt spd.txt
## cducsu.txt         1.000         0.968   0.975
## coalition.txt      0.968         1.000   0.986
## spd.txt            0.975         0.986   1.000</code></pre>
<pre class="r"><code>textstat_lexdiv(dfm_coal)[, 1:2]</code></pre>
<pre><code>##        document       TTR
## 1    cducsu.txt 0.3302084
## 2 coalition.txt 0.2487993
## 3       spd.txt 0.2802713</code></pre>
<p>From this, we learn that the SPD manifesto has more similarity to the coalition agreement than that of the CDU/CSU, a notion which somewhat resembles the assessment of the 2017 German coalition. Also, the lexical diversity of the manifestos, measured in terms of types (different words) per token (total words), appears to be higher than the coalition agreement, especially for the CDU/CSU.</p>
</div>
<div id="sentiment-analysis-using-a-dictionary" class="section level3">
<h3>Sentiment analysis using a dictionary <a name="sentimentanalysis"></a></h3>
<p>For sentiment analyis, existing dictionaries are available. It is important to note that these do not necessarily fit the research question at hand. In this example, the German “LIWC” (linguistic inquiry and word count) dictionary is used, but alternatives exist, such as “Lexicoder” for political text. LIWC features the categories “anger”, “posemo” (positive emotion) and “religion”. After obtaining the dictionary and applying it while creating a dfm from the corpus, the share of the texts in the respective categories is displayed. The results indicate that the coalition agreement features less positive emotions as compared to the manifestos and that the SPD manifesto is the most “angry” text, while the CDU/CSU speaks most about religion.</p>
<details>
<p><summary>Code: Using a Dictionary</summary></p>
<pre class="r"><code># Create dictionary
liwcdict &lt;- dictionary(file = paste0(wd, &quot;German_LIWC2001_Dictionary.dic&quot;),
                       format = &quot;LIWC&quot;)

# Create dfm
liwcdfm &lt;- dfm(
  coalitioncorpus,
  remove = c(stopwords(&quot;german&quot;)),
  remove_punct = TRUE,
  stem = FALSE,
  dictionary = liwcdict
)

# Subset and calculate percentage
liwcsub &lt;-
  dfm_select(liwcdfm,
             pattern = c(&quot;Anger&quot;, &quot;Posemo&quot;, &quot;Relig&quot;),
             selection = &quot;keep&quot;)

liwcsub &lt;- convert(liwcsub, to = &quot;data.frame&quot;)
liwcsub$sum &lt;- apply(dfm_coal, FUN = sum, 1)
liwcparties &lt;- data.frame(
  docs = liwcsub$document,
  ShareAnger = liwcsub$Anger / liwcsub$sum,
  SharePosemo = liwcsub$Posemo / liwcsub$sum,
  ShareRelig = liwcsub$Relig / liwcsub$sum
)

liwcparties </code></pre>
</details>
<pre><code>##            docs  ShareAnger SharePosemo  ShareRelig
## 1    cducsu.txt 0.004314995  0.06414959 0.005609493
## 2 coalition.txt 0.004666188  0.05050462 0.004074997
## 3       spd.txt 0.005550042  0.05697063 0.003454993</code></pre>
</div>
<div id="lda-topic-modelling" class="section level3">
<h3>LDA topic modelling <a name="lda"></a></h3>
<p>LDA stands for Latent Dirichlet allocation. In a nutshell, the method represents texts as a mixture of topics, and simultaneously topics as mixtures of words. Fixing the number of topics to <span class="math inline">\(k = 5\)</span>, and using the <strong>topicmodels</strong> package, the command <code>lda()</code> delivers posterior probabilities of the topics for each document.</p>
<details>
<p><summary>Code: LDA Topic Model</summary></p>
<pre class="r"><code>library(topicmodels)

# Preparation
dfm_coal &lt;-
  dfm(
    coalitioncorpus,
    remove = c(
      stopwords(&quot;german&quot;),
      &quot;dass&quot;,
      &quot;sowie&quot;,
      &quot;insbesondere&quot;,
      &quot;b&quot;,
      &quot;z&quot;,
      &quot;a&quot;,
      &quot;u&quot;
    ),
    remove_punct = TRUE,
    remove_numbers = TRUE,
    stem = FALSE
  )

dfm_coal &lt;- dfm_wordstem(dfm_coal, language = &quot;german&quot;)

# Define parameters
burnin &lt;- 1000
iter &lt;- 500
keep &lt;- 50
seed &lt;- 2010
ntopics &lt;- 5

# Run LDA with 5 topics
ldaOut &lt;- LDA(
  dfm_coal,
  k = ntopics,
  method = &quot;Gibbs&quot;,
  control = list(
    burnin = burnin,
    iter = iter,
    keep = keep,
    seed = seed,
    verbose = FALSE
  )
)
# Posterior probabilities of the topics for each document
k &lt;- posterior(ldaOut)</code></pre>
</details>
<p><br />
Interpretation is the difficult part. Each document can be expressed as a mixture of topics, and notwithstanding the precise meaning of the topics, we learn that all texts share content referring to topic 1, while the CDU/CSU manifesto also features topic 2 and the SPD manifesto topic 3 to some extent.</p>
<pre class="r"><code>k$topics</code></pre>
<pre><code>##                       1          2          3          4          5
## cducsu.txt    0.5815482 0.02729874 0.36122496 0.01144536 0.01848272
## coalition.txt 0.6950774 0.03700023 0.06163624 0.10570834 0.10057777
## spd.txt       0.6755292 0.17646590 0.11404313 0.01837605 0.01558576</code></pre>
</div>
<div id="wordfish-scaling" class="section level3">
<h3>Wordfish Scaling <a name="wordfish"></a></h3>
<p>Wordfish scaling derives latent positions from texts based on a bag-of-words assumption. Here is an example relying on a set of Swiss manifestos, using only the sections on immigration. In preparation, a dfm is created while removing stopwords, stemming the remaining words and removing punctuation.</p>
<details>
<p><summary>Code: Preparing Corpus for Wordfish</summary></p>
<pre class="r"><code>manifestos &lt;- readtext(paste0(wd, &quot;manifestos/*.txt&quot;))
manifestocorpus &lt;- corpus(manifestos)
dfm_manifesto &lt;-
  dfm(
    manifestocorpus,
    remove = c(
      &quot;gruen*&quot;,
      &quot;sp&quot;,
      &quot;sozialdemokrat*&quot;,
      &quot;cvp&quot;,
      &quot;fdp&quot;,
      &quot;svp&quot;,
      &quot;fuer&quot;,
      &quot;dass&quot;,
      &quot;koennen&quot;,
      &quot;koennte&quot;,
      &quot;ueber&quot;,
      &quot;waehrend&quot;,
      &quot;wuerde&quot;,
      &quot;wuerden&quot;,
      &quot;schweiz*&quot;,
      &quot;partei*&quot;,
      stopwords(&quot;german&quot;)
    ),
    valuetype = &quot;glob&quot;,
    stem = FALSE,
    remove_punct = TRUE
  )
dfm_manifesto &lt;- dfm_wordstem(dfm_manifesto, language = &quot;german&quot;)</code></pre>
</details>
<p><br />
The function <code>textmodel_wordfish()</code> computes the Wordfish model, originally decribed in an <a href="https://onlinelibrary.wiley.com/doi/full/10.1111/j.1540-5907.2008.00338.x">AJPS article</a> by Slapin and Proksch in 2009. It assumes that the distribution of words across texts follows a Poisson distribution, and can be modeled by document and word fixed effects as well as word-specific weights and document positions. The model is a variant of unsupervised scaling, only requiring the relative location of two texts on the latent dimension. Here, a text of the Swiss People’s Party (SVP) is assumed to be right to that of the Social Democratic Party of Switzerland (SPS). The results make some sense, with the other manifestos aligning as expected on what could be interpreted as a anti-immigration dimension.</p>
<pre class="r"><code>wf &lt;- textmodel_wordfish(dfm_manifesto,
                         dir = c(13, 19),
                         dispersion = &quot;poisson&quot;)
textplot_scale1d(wf)</code></pre>
<p><img src="/../../../../../article/quantitative-analysis-of-political-text_files/figure-html/fish-1.png" width="672" /></p>
<p>Or, with some improvements to the plot:</p>
<details>
<p><summary>Code: Improved Plot of Party Positions</summary></p>
<pre class="r"><code>library(ggplot2)

# Save document scores and confidence intervals in data frame
wfdata &lt;- as.data.frame(predict(wf, interval = &quot;confidence&quot;))

# Add document variables
wfdata$docs &lt;- rownames(wfdata)
wfdata$electionyear &lt;- substr(wfdata$docs, 5, 8)
wfdata$party &lt;- as.factor(substr(wfdata$docs, 1, 3))
wfdata$party &lt;-
  factor(wfdata$party, levels = c(&quot;gps&quot;, &quot;sps&quot;, &quot;cvp&quot;, &quot;fdp&quot;, &quot;svp&quot;))

ggplot(wfdata) +
  geom_pointrange(
    aes(
      x = electionyear,
      y = fit.fit,
      ymin = fit.lwr,
      ymax = fit.upr,
      group = party,
      color = party
    ),
    size = 0.5
  ) +
  geom_line(aes(
    x = electionyear,
    y = fit.fit,
    group = party,
    color = party
  )) +
  theme_bw() +
  labs(title = &quot;Wordfish analysis&quot;,
       y = &quot;Document position&quot;,
       x = &quot;Electionyear&quot;) +
  scale_color_manual(values = c(&quot;green3&quot;,
                                &quot;red1&quot;,
                                &quot;darkorange1&quot;,
                                &quot;dodgerblue4&quot;,
                                &quot;springgreen4&quot;))</code></pre>
</details>
<p><br />
<img src="/../../../../../article/quantitative-analysis-of-political-text_files/figure-html/fish3-1.png" width="672" /></p>
</div>
<div id="further-readings" class="section level3">
<h3>Further readings <a name="furtherreadings"></a></h3>
<ul>
<li>An introductory article to QTA in R, especially relying on <strong>quanteda</strong>: <a href="https://www.tandfonline.com/doi/abs/10.1080/19312458.2017.1387238">Welbers, Kasper, Wouter Van Atteveldt and Kenneth Benoit (2017): Text Analysis in R, <em>Communication Methods and Measures</em> 11(4): 245-65.</a></li>
<li>An application of the methods by the workshop host and co-authors: <a href="https://www.cambridge.org/core/journals/political-science-research-and-methods/article/estimating-intraparty-preferences-comparing-speeches-to-votes/D5812B196E0945B1341AFCD050F24858">Schwarz, Daniel, Denise Traber and Kenneth Benoit (2017): Estimating Intra-Party Preferences: Comparing Speeches to Votes. <em>Political Science Research and Methods</em> 5(2): 379-396.</a></li>
</ul>
</div>
<div id="about-the-presenter" class="section level3">
<h3>About the presenter</h3>
<p><a href="https://denisetraber.net">Denise Traber</a> is a Senior Research Fellow at the University of Lucerne, Switzerland, where she heads an Ambizione research grant project on “The divided people: polarization of political attitudes in Europe” funded by the Swiss National Science Foundation. She has a strong interest in quantitative text analysis, co-organizes the “Zurich Summer School for Women in Political Methodology” and has published the article “Estimating Intra-Party Preferences: Comparing Speeches to Votes” in Political Science Research and Methods in 2017, jointly with Daniel Schwarz and Ken Benoit.</p>
</div>
]]>
      </description>
    </item>
    
    <item>
      <title>Collecting and Analyzing Twitter Data Using R</title>
      <link>https://socialsciencedatalab.mzes.uni-mannheim.de/article/collecting-and-analyzing-twitter-using-r/</link>
      <pubDate>Mon, 15 Jul 2019 01:01:01 +0100</pubDate>
      
      <guid>https://socialsciencedatalab.mzes.uni-mannheim.de/article/collecting-and-analyzing-twitter-using-r/</guid>
      <description><![CDATA[
        </p>
<p>How do you access Twitter’s API, collect a stream of tweets, and analyze the retrieved data? Which potentials, challenges, and limitations for social scientific research come along with using Twitter data? This <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/categories/tutorials/">Methods Bites Tutorial</a> by <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/page/team/">Denis Cohen</a>, based on a workshop by <a href="https://www.simon-kuehne.de">Simon Kühne</a> (Bielefeld University) in the <a href="https://socialsciencedatalab.mzes.uni-mannheim.de/page/events/">MZES Social Science Data Lab</a> in Spring 2019, aims to tackle these questions.</p>
<p>After reading this blog post and engaging with the applied exercises, readers should:</p>
<ul>
<li>be able to collect Twitter data using R</li>
<li>be able to perform explorative analyses of the data using R</li>
<li>have a better understanding of Twitter data, and thus, the potentials and limitations of using it in research projects</li>
</ul>
<p>You can use these links to navigate across the four main sections of this tutotial:</p>
<ol style="list-style-type: decimal">
<li><a href="#about-twitter">About Twitter</a></li>
<li><a href="#collecting-twitter-data">Collecting Twitter Data</a></li>
<li><a href="#analyzing-twitter-data">Analyzing Twitter Data</a></li>
<li><a href="#potential-issues-and-challenges">Potential Issues and Challenges</a></li>
</ol>
<p><em>Note:</em> This blog post presents Simon’s workshop materials in condensed form. The complete workshop materials, including slides and scripts, are available from our <a href="https://github.com/SocialScienceDataLab/twitter">GitHub</a>.</p>
<div id="about-twitter" class="section level3">
<h3>About Twitter</h3>
<p>Twitter is an online news and social networking service, also used for micro-blogging. In everyday use, it mostly serves as a platform for publicly sharing short texts – often along with media content and/or links – in the form of so-called “tweets”. Twitter has approx. 326 Million monthly active users who send about 500 Million tweets each day (see this <a href="https://s22.q4cdn.com/826641620/files/doc_financials/2018/q3/TWTR-Q3_18_InvestorFactSheet.pdf">fact sheet</a>).</p>
<div id="the-basics" class="section level5">
<h5>The Basics</h5>
<ul>
<li>Each user has a profile (page) and can add a photo and information about themselves</li>
<li>Users can <em>follow</em> each other</li>
<li>Users can <em>tweet</em>, i.e., publicly sharing a text/photo/link</li>
<li>Each Tweet is restricted to a maximum of 280 characters</li>
<li>Users can interact with a Tweet via <em>comments (replies), likes,</em> and <em>shares (retweets)</em></li>
<li>Users can interact with other users via <em>direct messaging</em></li>
<li>Users can create a <em>thread:</em> A series of connected tweets</li>
<li>Users use <em>hashtags</em> (e.g., #mannheim) in order to associate their tweets with certain topics and to make them easier to find</li>
<li>Users can search for keywords/hashtags in order to find relevant tweets and users</li>
</ul>
</div>
<div id="twitter-in-social-science-research" class="section level5">
<h5>Twitter in Social Science Research</h5>
<p>Analyzing tweets and social interaction on Twitter can help to answer social science research questions, especially in communication research and political science. Contrary to Facebook (API depreciation/shut-down in April 2018) and Instagram (API depreciation/shut-down in December 2018), Twitter data is (easily) accessible for researchers.</p>
<p>According to the <a href="https://login.webofknowledge.com/">Web of Science</a> database, there are 2,598 journal articles published since 2009 with the word “Twitter” in their titles.</p>
<div class="figure" style="text-align: center"><span id="fig:img1"></span>
<img src="/../../../../../article/collecting-and-analyzing-twitter-using-r/img/twitter_wos_1.PNG" alt="Number of articles by year"  />
<p class="caption">
Figure 1: Number of articles by year
</p>
</div>
<div class="figure" style="text-align: center"><span id="fig:img2"></span>
<img src="/../../../../../article/collecting-and-analyzing-twitter-using-r/img/twitter_wos_2.PNG" alt="Number of articles by discipline"  />
<p class="caption">
Figure 2: Number of articles by discipline
</p>
</div>
</div>
</div>
<div id="collecting-twitter-data" class="section level3">
<h3>Collecting Twitter Data</h3>
<div id="the-twitter-api-platform" class="section level5">
<h5>The Twitter API Platform</h5>
<p>An API (Application Programming Interface) allows users to access (real-time) Twitter data. Twitter offers <a href="https://developer.twitter.com/en/docs.html">a variety of API services</a> – some for free, others not. Using these services, one can search for tweets published in the past, stream tweets in realtime, and manage Twitter accounts and ads. The following exercises will focus on the free-of-charge API service, which is used in the vast majority of research projects: <a href="https://developer.twitter.com/en/docs/tweets/filter-realtime/overview">The Realtime Streaming API</a>.</p>
</div>
<div id="the-streaming-api---collecting-tweets-in-realtime" class="section level5">
<h5>The Streaming API - Collecting Tweets in Realtime</h5>
<p><em>“Establishing a connection to the streaming APIs means making a very long lived HTTP request, and parsing the response incrementally. Conceptually, you can think of it as downloading an infinitely long file over HTTP”</em>. This way, we can receive up to a maximum of 1% of all tweets worldwide. As a query is usually specified by selected keywords or geographic areas, you will be able to collect (almost) all relevant tweets for your research interest. There are three filter parameters that you can use:</p>
<ul>
<li>‘Follow’: Receive tweets of up to 5,000 users</li>
<li>‘Track’: Receive tweets that contain up to 400 keywords</li>
<li>‘Location’: Receive tweets from within a set of up to 25 geographic bounding boxes</li>
</ul>
</div>
<div id="api-authentification" class="section level5">
<h5>API Authentification</h5>
<p>You need to authenticate yourself when making requests to the Twitter API. Twitter uses the <a href="https://oauth.net/">OAuth protocol</a>, an <a href="https://oauth.net/"><em>“open protocol to allow secure authorization in a simple and standard method from web, mobile and desktop applications”</em></a>. This involves five necessary steps:</p>
<ul>
<li>creating a Twitter account</li>
<li>logging in to your Twitter account via <a href="https://developer.twitter.com/" class="uri">https://developer.twitter.com/</a></li>
<li>creating an app</li>
<li>creating keys, access token &amp; secret</li>
</ul>
</div>
<div id="data-collection" class="section level5">
<h5>Data Collection</h5>
<p>There are a number of ways to collect Twitter data, including writing your own script to make continuous HTTP requests, Python’s <code>tweepy</code> package, and R’s <code>rtweet</code> package. The following demonstrates how to collect Twitter data using different Streaming API endpoints and the <a href="https://rtweet.info/reference/stream_tweets.html"><code>rtweet</code></a> package.</p>
</div>
<div id="collecting-tweets-using-the-rtweet-package" class="section level5">
<h5>Collecting Tweets Using the rtweet Package</h5>
<p>As usual, we start with a little housekeeping: Installing required packages per <code>install.packages()</code> and specifying a working directory using <code>setwd()</code>.</p>
<details>
<p><summary>Code: Setup</summary></p>
<pre class="r"><code># Install Packages
install.packages(&quot;rtweet&quot;)
install.packages(&quot;ggmap&quot;)
install.packages(&quot;igraph&quot;)
install.packages(&quot;ggraph&quot;)
install.packages(&quot;tidytext&quot;)
install.packages(&quot;ggplot2&quot;)
install.packages(&quot;dplyr&quot;)
install.packages(&quot;readr&quot;)

# Set Working Directory
setwd(&quot;/PATH&quot;)</code></pre>
</details>
<p><br />
The code chunk below then illustrates three examples of collecting Twitter data using the <code>rtweet</code> package. After loading the required packages into the <code>library()</code>, we specify our authentication token per <code>create_token()</code> (see <a href="#api-authentification">API Authentification</a> above). After this, we are live: Using the function <code>stream_tweets()</code>, we collect:</p>
<ol style="list-style-type: decimal">
<li>a sample of current tweets for <code>timeout = 10</code> seconds.</li>
<li>a sample of current tweets containing either of the keywords <code>q = "trump, donald trump"</code> for <code>timeout = 30</code> seconds.</li>
<li>a sample of current tweets for <code>timeout = 180</code> seconds from a specific location (in this instance, Berlin), restricting our search to a rectangular area defined by coordinates for longitude and latitude. Use, for instance, <a href="https://boundingbox.klokantech.com/" class="uri">https://boundingbox.klokantech.com/</a> to retrieve the coordinates of your choosing.</li>
</ol>
<details>
<p><summary>Code: Data Collection</summary></p>
<pre class="r"><code># Open Libraries
library(rtweet)

# Speficy Authentification Token&#39;s provided in your Twitter App
create_token(
  app = &quot;APPNAME&quot;,
  consumer_key = &quot;consumer_key&quot;,
  consumer_secret = &quot;consumer_secret&quot;,
  access_token = &quot;access_token&quot;,
  access_secret = &quot;access_secret&quot;
)

# Collect a &#39;random&#39; sample of Tweets for 10 seconds
stream_tweets(
  q = &quot;&quot;,
  timeout = 10,
  file_name = &quot;sample.json&quot;,
  parse = FALSE
)
sample &lt;- parse_stream(&quot;sample.json&quot;)
save(sample, file = &quot;sample_live.Rda&quot;)

# Collect Tweets that contain specific keywords for 30 seconds
stream_tweets(
  q = &quot;trump, donald trump&quot;,
  timeout = 30,
  file_name = &quot;trump.json&quot;,
  parse = FALSE
)
trump &lt;- parse_stream(&quot;trump.json&quot;)
save(trump, file = &quot;trump_live.Rda&quot;)

# Collect Tweets from a specific location
stream_tweets(
  c(13.0883,52.3383,13.7612,52.6755), 
  timeout = 180,
  file_name = &quot;berlin.json&quot;,
  parse = FALSE
)
berlin &lt;- parse_stream(&quot;berlin.json&quot;)
save(berlin, file = &quot;berlin_live.Rda&quot;)</code></pre>
</details>
<p><br />
By running the <code>steam_tweets()</code> function, we receive Tweets and related meta-information from the Twitter API. The data is stored in .json format (Java Script Object Notation), though we can store these files as data frames after parsing them per <code>parse_stream()</code>. Each row in the data frame represents a Tweet or Re-Tweet and contains, amongst other, the following information:</p>
<ul>
<li>The content of a Tweet + Tweet-URL + Tweet-ID</li>
<li>User-name + User-ID</li>
<li>Time-stamp</li>
<li>Place, country, geocodes (rarely)</li>
<li>User self-description, residence, no. of followers, no. of friends</li>
<li>URLs to images, videos</li>
</ul>
<p>For illustration, we take the example of the data frame <code>sample</code>, the result of our first query. As we can see, our 10 second sample without specified key words contains 88 variables (columns) and 359 tweets (rows). You can see the full list of variables below, along with a small anonymized portion of five English language tweets (you can only see the first few characters of each tweet, stored in the variable <code>sample$text</code>).</p>
<details>
<p><summary>Code: Viewing the Data</summary></p>
<pre class="r"><code># Dimensions of the data frame
dim(sample)</code></pre>
<pre><code>## [1] 359  88</code></pre>
<pre class="r"><code># Variables in the data frame
names(sample)</code></pre>
<pre><code>##  [1] &quot;user_id&quot;                 &quot;status_id&quot;              
##  [3] &quot;created_at&quot;              &quot;screen_name&quot;            
##  [5] &quot;text&quot;                    &quot;source&quot;                 
##  [7] &quot;display_text_width&quot;      &quot;reply_to_status_id&quot;     
##  [9] &quot;reply_to_user_id&quot;        &quot;reply_to_screen_name&quot;   
## [11] &quot;is_quote&quot;                &quot;is_retweet&quot;             
## [13] &quot;favorite_count&quot;          &quot;retweet_count&quot;          
## [15] &quot;hashtags&quot;                &quot;symbols&quot;                
## [17] &quot;urls_url&quot;                &quot;urls_t.co&quot;              
## [19] &quot;urls_expanded_url&quot;       &quot;media_url&quot;              
## [21] &quot;media_t.co&quot;              &quot;media_expanded_url&quot;     
## [23] &quot;media_type&quot;              &quot;ext_media_url&quot;          
## [25] &quot;ext_media_t.co&quot;          &quot;ext_media_expanded_url&quot; 
## [27] &quot;ext_media_type&quot;          &quot;mentions_user_id&quot;       
## [29] &quot;mentions_screen_name&quot;    &quot;lang&quot;                   
## [31] &quot;quoted_status_id&quot;        &quot;quoted_text&quot;            
## [33] &quot;quoted_created_at&quot;       &quot;quoted_source&quot;          
## [35] &quot;quoted_favorite_count&quot;   &quot;quoted_retweet_count&quot;   
## [37] &quot;quoted_user_id&quot;          &quot;quoted_screen_name&quot;     
## [39] &quot;quoted_name&quot;             &quot;quoted_followers_count&quot; 
## [41] &quot;quoted_friends_count&quot;    &quot;quoted_statuses_count&quot;  
## [43] &quot;quoted_location&quot;         &quot;quoted_description&quot;     
## [45] &quot;quoted_verified&quot;         &quot;retweet_status_id&quot;      
## [47] &quot;retweet_text&quot;            &quot;retweet_created_at&quot;     
## [49] &quot;retweet_source&quot;          &quot;retweet_favorite_count&quot; 
## [51] &quot;retweet_retweet_count&quot;   &quot;retweet_user_id&quot;        
## [53] &quot;retweet_screen_name&quot;     &quot;retweet_name&quot;           
## [55] &quot;retweet_followers_count&quot; &quot;retweet_friends_count&quot;  
## [57] &quot;retweet_statuses_count&quot;  &quot;retweet_location&quot;       
## [59] &quot;retweet_description&quot;     &quot;retweet_verified&quot;       
## [61] &quot;place_url&quot;               &quot;place_name&quot;             
## [63] &quot;place_full_name&quot;         &quot;place_type&quot;             
## [65] &quot;country&quot;                 &quot;country_code&quot;           
## [67] &quot;geo_coords&quot;              &quot;coords_coords&quot;          
## [69] &quot;bbox_coords&quot;             &quot;status_url&quot;             
## [71] &quot;name&quot;                    &quot;location&quot;               
## [73] &quot;description&quot;             &quot;url&quot;                    
## [75] &quot;protected&quot;               &quot;followers_count&quot;        
## [77] &quot;friends_count&quot;           &quot;listed_count&quot;           
## [79] &quot;statuses_count&quot;          &quot;favourites_count&quot;       
## [81] &quot;account_created_at&quot;      &quot;verified&quot;               
## [83] &quot;profile_url&quot;             &quot;profile_expanded_url&quot;   
## [85] &quot;account_lang&quot;            &quot;profile_banner_url&quot;     
## [87] &quot;profile_background_url&quot;  &quot;profile_image_url&quot;</code></pre>
<pre class="r"><code># An anonymized portion of the data frame, only tweets in English
sample$user_id &lt;- seq(1, nrow(sample), 1)
sample$screen_name &lt;- paste(&quot;name&quot;, seq(1, nrow(sample), 1), sep = &quot;_&quot;)
sample &lt;- subset(sample, lang == &quot;en&quot;)
sample[1:5, c(&quot;user_id&quot;, &quot;created_at&quot;, &quot;screen_name&quot;, &quot;text&quot;, &quot;is_quote&quot;, 
              &quot;is_retweet&quot;)]</code></pre>
<pre><code>## # A tibble: 5 x 6
##   user_id created_at          screen_name text          is_quote is_retweet
##     &lt;dbl&gt; &lt;dttm&gt;              &lt;chr&gt;       &lt;chr&gt;         &lt;lgl&gt;    &lt;lgl&gt;     
## 1       1 2019-04-03 10:22:32 name_1      just see..&lt;U+2764&gt;&lt;U+FE0F&gt;~ FALSE    FALSE     
## 2       4 2019-04-03 10:22:33 name_4      I’m slowy gi~ FALSE    TRUE      
## 3      15 2019-04-03 10:22:33 name_15     &quot;\&quot;I didn&#39;t ~ TRUE     FALSE     
## 4      20 2019-04-03 10:22:33 name_20     Hmm. I’ma de~ FALSE    TRUE      
## 5      28 2019-04-03 10:22:33 name_28     I still hear~ TRUE     TRUE</code></pre>
</details>
<p><br /></p>
</div>
</div>
<div id="analyzing-twitter-data" class="section level3">
<h3>Analyzing Twitter Data</h3>
<p>We have now collected Tweets and meta-information based on selected keywords and/or regional parameters. This begs the question what to do with the raw data. Some common research interests include:</p>
<ul>
<li><em>Content Analysis:</em> What kind of topics are users talking about?</li>
<li><em>Sentiment Analysis:</em> What kind of opinions, attitudes, and emotions towards objects are users communicating?</li>
<li><em>Network Analysis:</em> Who is related to whom? Who are important users?</li>
<li><em>Geospatial Analysis:</em> Where are users/Tweets coming from?</li>
</ul>
<p>In the following, we present two quick examples that showcase possible avenues for the analysis of Twitter data.</p>
<div id="example-1-prepping-tweets-for-text-analysis" class="section level5">
<h5>Example 1: Prepping Tweets for Text Analysis</h5>
<p>Quantitative Text Analysis (QTA) typically relies on pre-processed textual data (see <a href="/../../../../../article/quantitative-analysis-of-political-text/">this blog post</a> for our SSDL workshop on QTA by <a href="https://denisetraber.net/">Denise Traber</a>). The code chunk below illustrates how a collection of Tweets can easily be prepared for QTA techniques such as content or sentiment analysis.</p>
<p>In this example, we use our 30 second sample of Tweets containing the key words “donald trump” and/or “trump”, which we stored as <code>trump_live.RDa</code> (see <a href="#collecting-tweets-using-the-rtweet-package">Collecting Tweets Using the rtweet Package</a> above). Tweet contents are stored in the variable <code>trump$text</code>. Using the <code>tidytext</code>and <code>dplyr</code> packages, we then process the Tweets into a text corpus as follows:</p>
<ol style="list-style-type: decimal">
<li>Remove URLs from all tweets using <code>gsub()</code></li>
<li>Remove punctuation, convert to lowercase, and seperate all words using <code>unnest_tokens()</code></li>
<li>Remove <a href="https://en.wikipedia.org/wiki/Stop_words">stop words</a> by first loading a list of stop words from the <code>tidytext</code> package via <code>data("stop_words")</code> and then removing these words from our tweets via <code>anti_join(stop_words)</code></li>
</ol>
<details>
<p><summary>Code: Preparing Data for Content Analysis</summary></p>
<pre class="r"><code># Open Libraries
library(tidytext)
library(dplyr)

# Data Cleaning
# Delete Links in the Tweets
trump$text &lt;- gsub(&quot;http.*&quot;, &quot;&quot;, trump$text)
trump$text &lt;- gsub(&quot;https.*&quot;, &quot;&quot;, trump$text)
trump$text &lt;- gsub(&quot;&amp;amp;&quot;, &quot;&amp;&quot;, trump$text)

# Remove punctuation, convert to lowercase, seperate all words
trump_clean &lt;- trump %&gt;%
  dplyr::select(text) %&gt;%
  unnest_tokens(word, text)

# Load list of stop words - from the tidytext package
data(&quot;stop_words&quot;)

# Remove stop words from your list of words
cleaned_tweet_words &lt;- trump_clean %&gt;%
  anti_join(stop_words)</code></pre>
</details>
<p><br />
Following this, we can calculate the word counts in our Tweet collection and visualize the 15 most frequent words:</p>
<details>
<p><summary>Code: Plotting the 15 Most Frequent Words</summary></p>
<pre class="r"><code># Plot the top 15 words
cleaned_tweet_words %&gt;%
  count(word, sort = TRUE) %&gt;%
  top_n(15) %&gt;%
  mutate(word = reorder(word, n)) %&gt;%
  ggplot(aes(x = word, y = n)) +
  geom_col() +
  xlab(NULL) +
  coord_flip() +
  labs(y = &quot;Count&quot;,
       x = &quot;Unique words&quot;,
       title = &quot;Count of unique words found in tweets&quot;,
       subtitle = &quot;Stop words removed from the list&quot;)</code></pre>
</details>
<p><br /></p>
<pre><code>## Selecting by n</code></pre>
<div class="figure" style="text-align: center"><span id="fig:unnamed-chunk-3"></span>
<img src="/../../../../../article/collecting-and-analyzing-twitter-using-r_files/figure-html/unnamed-chunk-3-1.png" alt="Top 15 Most Frequent Words" width="672" />
<p class="caption">
Figure 3: Top 15 Most Frequent Words
</p>
</div>
<p><br /></p>
</div>
<div id="example-2-analyzing-locations" class="section level5">
<h5>Example 2: Analyzing Locations</h5>
<p>Provided that users share their exact geo-information, we can locate their tweets on geographical maps. For example, in our 3 minute sample of Tweets from <code>berlin</code> (see <a href="#collecting-tweets-using-the-rtweet-package">Collecting Tweets Using the rtweet Package</a> above), this information was available for only two Tweets.</p>
<p>The code below demonstrates how one can view these Tweets mapped onto their exact geographical locations. We first pre-process the data, extracting numerical geo-coordinates for longitude and latitude where this information is available. We then subset the data frame to observations where these information exist. Lastly, using the same geo-coordinates as in the data collection, we set up an empty rectangular map of Berlin. Note that the last step involves accessing the <a href="https://cloud.google.com/maps-platform/#get-started">Google Maps API</a>, for which you will have to register separately (see, e.g., <a href="https://www.r-bloggers.com/geocoding-with-ggmap-and-the-google-api/">this post</a> at R-bloggers).</p>
<details>
<p><summary>Code: Analyzing Locations - Setup</summary></p>
<pre class="r"><code>library(ggmap)
library(dplyr)

# Seperate Geo-Information (Lat/Long) Into Two Variables
berlin &lt;- tidyr::separate(data = berlin,
                          col = geo_coords,
                          into = c(&quot;Latitude&quot;, &quot;Longitude&quot;),
                          sep = &quot;,&quot;,
                          remove = FALSE)

# Remove Parentheses
berlin$Latitude &lt;- stringr::str_replace_all(berlin$Latitude, &quot;[c(]&quot;, &quot;&quot;)
berlin$Longitude &lt;- stringr::str_replace_all(berlin$Longitude, &quot;[)]&quot;, &quot;&quot;)

# Store as numeric
berlin$Latitude &lt;- as.numeric(berlin$Latitude)
berlin$Longitude &lt;- as.numeric(berlin$Longitude)

# Keep only those tweets where geo information is available
berlin &lt;- subset(berlin, !is.na(Latitude) &amp; !is.na(Longitude))

# Set up empty map
berlin_map &lt;- get_map(location = c(lon = mean(c(13.0883, 13.7612)),
                                   lat = mean(c(52.3383, 52.6755))),
                      zoom = 10,
                      maptype = &quot;terrain&quot;,
                      source = &quot;google&quot;)</code></pre>
</details>
<p><br />
Building up on this, we can then map our collected tweets to see from where they were sent: Berlin-Mitte and Potsdam.</p>
<details>
<p><summary>Code: Analyzing Locations - Map</summary></p>
<pre class="r"><code>tweet_map &lt;- ggmap(berlin_map)
tweet_map + geom_point(data = berlin,
                       aes(x = Longitude,
                           y = Latitude),
                           color = &quot;red&quot;,
                           size = 4,
                           alpha = .5)</code></pre>
</details>
<br />
<div class="figure" style="text-align: center"><span id="fig:img3"></span>
<img src="/../../../../../article/collecting-and-analyzing-twitter-using-r/img/tweets_map.png" alt="GPS Coordinates of 2 Tweets"  />
<p class="caption">
Figure 4: GPS Coordinates of 2 Tweets
</p>
</div>
</div>
</div>
<div id="potential-issues-and-challenges" class="section level3">
<h3>Potential Issues and Challenges</h3>
<div id="bias-and-representativity" class="section level5">
<h5>Bias and Representativity</h5>
<p>Twitter users do not represent a random sample from a given population. This is not only due to the presence of bots and company or institutional accounts, but also to the manifold self-selection processes that using Twitter entails:</p>
<p>Population → Internet Users → Twitter Users → Active Twitter Users → Users sharing geo-information</p>
<p>Below are some references that analyze the magnitude and severity of these (and related) problems:</p>
<ul>
<li><a href="https://journals.sagepub.com/doi/10.1177/0894439314558836">Barberá, P. &amp; G. Rivero, 2015: Understanding the Political Representativeness of Twitter Users. Social Science Computer Review 33(6) 712-729.</a></li>
<li><a href="https://journals.sagepub.com/doi/full/10.1177/2053168017720008">Mellon, J. &amp; C. Prosser, 2017: Twitter and Facebook are not representative of the general population: Political attitudes and demographics of British social media users. Research and Politics 2017: 1-9.</a></li>
<li><a href="https://www.nomos-elibrary.de/10.5771/1615-634X-2018-2-140/eine-meinungsstarke-minderheit-als-stimmungsbarometer-ueber-die-persoenlichkeitseigenschaften-aktiver-twitterer-jahrgang-66-2018-heft-2?page=1">Hölig, S., 2018: Eine meinungsstarke Minderheit als Stimmungsbarometer?! Über die Persönlichkeitseigenschaften aktiver Twitterer. M&amp;K Medien- und Kommunikationswissenschaft 66: 140-169.</a></li>
</ul>
</div>
<div id="replicability-and-black-box-twitter" class="section level5">
<h5>Replicability and Black-Box Twitter</h5>
<p>Real-time Twitter data collection is not reproducible and for a given query. Furthermore, you can only hope that Twitter will provide you with a true random sample of Tweets. If completeness is crucial for your research interest, you will have to pay for complete access to all Tweets ever tweeted: <a href="https://developer.twitter.com/en/docs/tutorials/choosing-historical-api.html"><em>“Both Historical PowerTrack and Full-Archive Search provide access to any publicly available Tweet, starting with the first Tweet from March 2006”</em></a>.</p>
</div>
<div id="data-privacy-and-research-ethics" class="section level5">
<h5>Data Privacy and Research Ethics</h5>
<p>Tweets on public Twitter profiles are generally available. There are no measures in place that prevent the collection and analysis of the data, and users’ consent for the collection and processing of their tweets and profile information is usually not required. This practice is also congruous with certain guidelines for academic and commerical social media research (e.g <a href="http://rat-marktforschung.de/fileadmin/user_upload/pdf/R11_RDMS_D.pdf">DGOF Richtlinien zur Social Media Forschung:</a> <em>“In offenen Sozialen Medien bzw. den entsprechenden Bereichen dürfen die personenbezogenen Daten der Teilnehmer grundsätzlich ohne entsprechende explizite Einwilligung auf der Grundlage der gesetzlichen Erlaubnisnorm auch für Zwecke der Markt- und Sozialforschung verarbeitet und genutzt werden.”</em></p>
<p>However, Twitter’s terms of service are not necessarily congruous with German or EU data protection regulations (e.g. DSGVO).
Ultimately, this leaves the ethical and legal questions of how to ensure data privacy to us as researchers. Should we, for instance, further anonymize data, e.g. by separating user IDs from Tweet content and meta-information? What are the implications for open science? Should the full data, including user IDs and geo locations be publicly and permanently shared (e.g., as part of replication materials)? Should the answer to these questions be the same for regular users as opposed to public figures (e.g., politicians)? These questions highlight the urgent need for ongoing discussion about these topics.</p>
</div>
<div id="uncertainty-of-data-access" class="section level5">
<h5>Uncertainty of Data Access</h5>
<p>One should always have in mind that data access is 100% dependent upon Twitter’s willingness to share the data, and therby also on jurisdiction by which Twitter must abide (think, for instance, about Article 13). Data access for research projects through Facebook’s and Instagram’s APIs has previously been shut-down completely with only few weeks notice. Given that, research projects relying on Twitter data are always risky. This applies particularly to research projects that depends on a constant Twitter data influx over a long period of time (e.g., PhD projects).</p>
</div>
<div id="data-storage" class="section level5">
<h5>Data Storage</h5>
<p>Data storage can be an issue when tweets are collected over a long period of time. In many applications, data collections can easily amount to 100-200 GB per month. The use of powerful servers and storage in a relational database (e.g. SQL) are therefore recommended.</p>
</div>
</div>
<div id="conclusion-twitter-in-the-social-sciences" class="section level3">
<h3>Conclusion: Twitter in the Social Sciences</h3>
<p>Collecting Twitter data is comparatively easy and cheap. However, we usually know close to nothing about the users whom we collect tweets from. Thus, the research potential for social science projects is very limited when we are interested in questions that address ‘outside-social-media’ phenomena.</p>
<p>Arguably, we benefit most from Twitter data when we treat it as auxiliary or proxy information, or when we use it in combination with other data sources. Potential applications along those lines include:</p>
<ul>
<li>Variation in fear of crime across different regions</li>
<li>Health monitoring over time</li>
<li>Monitoring highly discussed topics in realtime</li>
<li>Measuring existing stereotypes towards minorities</li>
</ul>
</div>
<div id="further-reading" class="section level3">
<h3>Further Reading</h3>
<ul>
<li><a href="https://mkearney.github.io/nicar_tworkshop/#1">Workshop on using rtweet by its developer Michael W. Kearney</a></li>
<li><a href="https://rtweet.info/index.html">rtweet documentation</a></li>
<li><a href="https://www.earthdatascience.org/courses/earth-analytics/get-data-using-apis/text-mining-twitter-data-intro-r/">Test mining of tweets</a></li>
<li><a href="https://www.datascience.com/blog/beginners-guide-to-shiny-and-leaflet-for-interactive-mapping">Using Shiny and Leaflet</a></li>
<li><a href="https://www.halem-verlag.de/geospatial-analysis-of-social-media-data%e2%80%84-a-practical-framework-and-applications-using-twitter/">Rieder, Y. &amp; S. Kühne, 2018: Geospatial Analysis of Social Media Data - A Practical Framework and Applications. In: Stuetzer, C.M., Welker, M. &amp; M. Egger (Eds.), Computational Social Science in the Age of Big Data. Concepts, Methodologies, Tools, and Applications. DGOF Schriftenreihe, Köln: Herbert van Halem Verlag. URL: http://www.halem-verlag.de/computational-social-science-in-the-age-of-big-data/.</a></li>
</ul>
</div>
<div id="about-the-presenter" class="section level3">
<h3>About the Presenter</h3>
<p>Simon Kühne <a href="mailto:simon.kuehne@uni-bielefeld.de"><i class="fa fa-envelope"></i> </a> <a href="http://simon-kuehne.de/"><i class="fa fa-globe"></i> </a> <a href="https://twitter.com/SimonKuehne"><i class="fa fa-twitter"></i></a> is a post-doc at Bielefeld University. He holds a BA in Sociology and an MA in Survey Methodology from the University of Duisburg-Essen and a PhD in Sociology from Humboldt University of Berlin. His research focuses on survey methodology, social media and online data, and social inequality.</p>
</div>
]]>
      </description>
    </item>
    
  </channel>
</rss>