Notes de cours

LECTURE SUMMARY - Computation Analysis of Digital Communication + EXAM QUESTIONS

Name: LECTURE SUMMARY - Computation Analysis of Digital Communication + EXAM QUESTIONS
SKU: doc_2123504
Rating: 4.00 (1 reviews)
Author: pikayichu

1 vérifier

198 vues 5 fois vendu

Cours
Computational Analysis Of Digital Communication

Établissement
Vrije Universiteit Amsterdam (VU)

LECTURE SUMMARY - Computation Analysis of Digital Communication lecture notes + EXAM QUESTIONS. Please don't mind the English gramma since I wanted to make it as compact as possible :) (from 51 pages to 19 pages!!)

[Montrer plus]

Dernier document publié: 1 année de cela

Aperçu 3 sur 19 pages

Voir l'exemple

Publié le 21 novembre 2022
Fichier mis à jour le 22 novembre 2022
Nombre de pages 19
Écrit en 2022/2023
Type Notes de cours
Professeur(s) Masur
Contenu Toutes les classes

cadc
computational sciences
machine learning
social science
big data
search queries
dictionary
studio r
text analysis
content analysis
prediction
validity
deductive approaches
lexical analysis
dicti
r

Établissement
Vrije Universiteit Amsterdam (VU)
Cours
Master Communicatie Wetenschap
Cours
Computational Analysis Of Digital Communication

1 vérifier

Par: julietteessink • 1 année de cela

pikayichu

Membre depuis 7 année 32 documents vendus

5,89 €

Ajouté

Ajouter au panier

Ajouter au liste de veux

Garantie de satisfaction à 100%
Disponible immédiatement après paiement
En ligne et en PDF
Tu n'es attaché à rien

CADC 2022 - KAYI MAN

LECTURE 1 - What is computational social sciences? 2
PRELIMINARY SUMMARY 4

LECTURE 2 - Text as Data - Basics of Automatic Text Analysis 4
Text as Data - How can we analyze texts with computers? 4
Automated Text Analysis 5
Automatic text analysis steps - van Atteveldt, Welbers, & Van der Velden, 2019 5
Deductive Approaches: Dictionary-based Text Analysis 6
Example - state of union speech corpus 6

LECTURE 3 - Machine Learning - Supervised Text Classification 8
What is machine learning? 8
Supervised text classification 9
Principles of supervised text classification 9
Validation 11
Example - prediction musing genre from lyrics (homework 2/3A) 11
Conclusion 12

LECTURE 4 - Machine Learning - Unsupervised Topic Modeling 13
What is topic modeling? - based on example: nuclear technology from 1945 - 2013 13
Topic Modeling as Dimensionality Reduction 14
Latent Dirichlet Allocation - LDA topic modeling 14
Conclusion and outlook 18
Conclusion 18

EXAMPLE EXAM QUESTION (MULTIPLE CHOICE) 23

EXAMPLE EXAM QUESTION (OPEN FORMAT) 23

, CADC 2022 - KAYI MAN

LECTURE 1 - What is computational social sciences?
● Field of Social Science that uses algorithmic tools and large/unstructured data to understand
human and social behavior
● Computational methods as “microscope”: Methods are not the goal, but contribute to
theoretical development and/or data generation
● Complements rather than replaces traditional methodologies
● Includes methods such as, e.g.,:
○ Advanced data wrangling/data science
○ Combining of different data sets
○ Automated Text Analysis
○ Machine Learning (supervised and unsupervised)
○ Actor-based modeling
○ Simulations
○ …
● To better understand text, results. “how do we understand large data set”
● How can we work with data this large
● To big to put in excel-sheet
● Unsupervised vs. Supervised

TYPICAL WORKFLOW
1. Identification problem/purpose =
2. Data acquisitions = Different way to get existing data
3. Data wrangling = What do i have to transfor, add, delete the data to make it useful
4. Data analysis & modeling = statistical analysis or creating algorithms
5. Reporting = using data to communicate …. ?

WHY IS THIS IMPORTANT NOW?
- Collecting data used to be expensive (surveys, observations)
- Digital age: behaviors of billions are recorded, stored and therefore analyzable
- Digital record of behavior is created by everytime/thing you click/call/pay
- (meta-)data are byproduct of peeps everyday actions aka digital traces
- Big data = often called large-scale records of peeps/businesses

10 CHARACTERISTICS OF BIG DATA (Salganik, 2017, chap. 2.3)
1. Big = scale / volume of current datasets is often impressive
2. Always-on = big data systems are constantly collecting data (FB = always-on)
3. Non reactive = subjects are non reactive and not aware of the collecting (ethical?) of zijn zo
gewend dat het hun behavior niet veranderd
4. Incomplete = most big data sources are incomplete, don't have info that you want to
research. Because data was created for other purposes than research.
5. Inaccessible = Data held by companies/governments are difficult for researchers to access.
6. Non representative = not representative of certain populations
7. Drifting = systems are changing constantly, difficult for long-term study trends. The way they
do it, changes
8. Algorithmically confounded = behavior in big data is not natural; driving by engineering
goals. Weird algorithm implemented by FB, predetermines how the data is gonna look like.
Record produced by the system that is built by platform.
9. Dirty = Big data includes noise (junk, spam)
10. Sensitive = some info that companies/governments have, are sensitive (ethical?)
Privacy issues

, CADC 2022 - KAYI MAN

TYPICAL COMPUTATIONAL RESEARCH STRATEGIES
1. Counting things (how often do peep use phones per day? What topics do news sites cover most?
2. Forecasting and nowcasting (predictions both present and in future; crime prediction…)
3. Approximating experiments (investigate potential nudges to make user select certain news)

ADVANTAGES AND DISADVANTAGES
Advantages of Computational Methods
- Actual behavior vs. self-report (because biased)
- Social context vs. lab setting
- Small N to large N
Disadvantages of Computational Methods
- Techniques often complicated
- Data often proprietary (=eigendomsrecht)
- Samples often biased
- Insufficient metadata (we have data but don’t know who they are)

DEFINITION Van Atteveldt & Peng, 2018
“Computational Communication Science is the
- label applied to the emerging subfield that investigates
- the use of computational algorithms
- to gather and analyze big and often semi- or unstructured data sets
- to develop and test communication science theories”

PROMISES
Three developments fueled the computational methods of communication sciences
1. Vast amounts of digitally available data
2. Improved tools to analyze big data (auto text analysis methods) changes fast!
3. Powerful and cheap processing power & easy computing infrastructure (Github)

ETHICS OF ‘BIG DATA’ AND COMPUTATIONAL RESEARCH
THE “FACEBOOK MOOD MANIPULATION” STUDY (Kramer et al., 2014)
● Massive online experiment (N ~ 700k)
● Main Research Question: Is emotion contagious?
● Experimental groups: positive / negative / control
● Stimulus: Hide (negative / positive / random) messages from FB timeline
● Measurement / dependent variables: sentiment of posts by user

Question: Do you think these studies are problematic? If yes, why?
● No consent is given
● Shared with third party

COMPUTATIONAL TECHNIQUE: SENTIMENT ANALYSIS
● Count occurrences of words in both categories, subtract negative
from positive

Positive words reduced in feed = more negative words used
Negative words reduced in feed = more positive words used

IS THIS GOOD SCIENCE? WHY NOT?
● What’s cool?
○ Potentially interesting research question
○ actual behavior measured as well as self-report measures
● What’s not so cool? A lot…
○ No informed consent, not replicable, manipulation
○ Low internal validity
■ Is sentiment of posts indicative of mood?
■ Does change in sentiment originate in contagion of mood?
○ Low measurement accuracy
■ Are word counts indicative of sentiment?

Les avantages d'acheter des résumés chez Stuvia:

Qualité garantie par les avis des clients

Les clients de Stuvia ont évalués plus de 700 000 résumés. C'est comme ça que vous savez que vous achetez les meilleurs documents.

L’achat facile et rapide

Vous pouvez payer rapidement avec iDeal, carte de crédit ou Stuvia-crédit pour les résumés. Il n'y a pas d'adhésion nécessaire.

Focus sur l’essentiel

Vos camarades écrivent eux-mêmes les notes d’étude, c’est pourquoi les documents sont toujours fiables et à jour. Cela garantit que vous arrivez rapidement au coeur du matériel.

Foire aux questions

Qu'est-ce que j'obtiens en achetant ce document ?

Vous obtenez un PDF, disponible immédiatement après votre achat. Le document acheté est accessible à tout moment, n'importe où et indéfiniment via votre profil.

Garantie de remboursement : comment ça marche ?

Notre garantie de satisfaction garantit que vous trouverez toujours un document d'étude qui vous convient. Vous remplissez un formulaire et notre équipe du service client s'occupe du reste.

Auprès de qui est-ce que j'achète ce résumé ?

Stuvia est une place de marché. Alors, vous n'achetez donc pas ce document chez nous, mais auprès du vendeur pikayichu. Stuvia facilite les paiements au vendeur.

Est-ce que j'aurai un abonnement?

Non, vous n'achetez ce résumé que pour 5,89 €. Vous n'êtes lié à rien après votre achat.

Peut-on faire confiance à Stuvia ?

4.6 étoiles sur Google & Trustpilot (+1000 avis)

84669 résumés ont été vendus ces 30 derniers jours

Fondée en 2010, la référence pour acheter des résumés depuis déjà 14 ans

Commencez à vendre!

Universités et collèges populaires

Livres populaires

Notes de cours

LECTURE SUMMARY - Computation Analysis of Digital Communication + EXAM QUESTIONS

Infos sur le Document

Sujets

École, étude et sujet

1 vérifier

Vendeur

Avis reçus

Aperçu du contenu

Les avantages d'acheter des résumés chez Stuvia:

Qualité garantie par les avis des clients

L’achat facile et rapide

Focus sur l’essentiel

Foire aux questions

Qu'est-ce que j'obtiens en achetant ce document ?

Garantie de remboursement : comment ça marche ?

Auprès de qui est-ce que j'achète ce résumé ?

Est-ce que j'aurai un abonnement?

Peut-on faire confiance à Stuvia ?