AI-Based Emotion Analysis for Text, Image, and Video - A Short Course
A 3-Day Livestream Seminar Taught by Hudson Golino, Ph.D. and Aleksandar Tomašević, Ph.D.
This hands-on seminar presents a practical, end-to-end workflow to study emotions and sentiment in text, images, and video using transformer AI models entirely within R via the transforEmotion package. You will learn three applied workflows: zero-shot emotion classification for text, CLIP-based scoring of images and video, and short structured LLM rationales for interpretation and reporting. No deep learning background or GPU access is required.
By the end, you will be able to select appropriate models for your research question, score multimodal data with transforEmotion, export analysis-ready tables, and evaluate results with transparent, publication-ready methods. You’ll walk away with the tools to:
-
- Analyze text with BERT-style encoders using zero-shot labels you define.
- Detect emotion cues in facial images and short video clips using CLIP-based pipelines.
- Generate concise, structured explanations with small LLMs to improve interpretability.
Starting November 4, this seminar will be presented as a 3-day synchronous, livestream workshop via Zoom. Each day will feature lecture sessions with hands-on exercises. Live attendance is recommended for the best experience. If you can’t join in real time, recordings will be available within 24 hours and accessible for four weeks after the seminar.
Closed captioning is available for all live and recorded sessions. Captions can be translated to a variety of languages including Spanish, Korean, and Italian. For more information, click here.
ECTS Equivalent Points: 1
More Details About the Course Content
We’ll emphasize a streamlined, reproducible approach for social-science workflows in R. You’ll install and configure models through transforEmotion, then reproduce the full pipeline: load sample assets, score text with BERT encoders, classify images and video frames with CLIP, and optionally add short, structured rationales from small LLMs. You will design theory-aligned label sets (e.g., basic emotions, valence–arousal) and document settings to ensure transparent reporting.
The course balances method and practice. We’ll compare encoder vs. decoder architectures and explain how multimodal embeddings represent signals across text and vision. You will craft zero-shot label prompts, interpret similarity scores, and apply lightweight validation checks (accuracy, F1, confusion matrices). We’ll also cover practical issues: short texts, multilinguality, class imbalance, bias and privacy, face selection, and frame sampling strategies for video.
A capstone ties everything together on a curated subset of the MAFW dataset. You will combine text scores (BERT), image/video outputs (CLIP), and optional LLM rationales, then evaluate cross-modal alignment, run metrics, and iterate on labels and sampling choices to resolve disagreements.
Computing
This is a hands-on course with instructor-led software demonstrations and guided exercises. These guided exercises are designed for the R programming language, so you should use a computer with a recent version of R (version 4.3 or later) and RStudio, or VS Code with the R extension.
The course exercises will use the transforEmotion R package, which will install and manage the required dependencies. No manual Python or command-line installation is required. CPU-only workflows are supported, so specialized hardware is not needed. You should have approximately 20GB of free disk space available for models and cached embeddings. Curated, pre-extracted subsets of public datasets and sample text corpora will be provided for in-class use.
A basic familiarity with R, such as knowing how to use data frames, install packages, and run scripts, is sufficient to follow along with the course exercises.
If you’d like to take this course but are concerned that you don’t know enough R, there are excellent online resources for learning the basics. Here are our recommendations.
Who Should Register?
This seminar is designed for researchers in psychology, communication, digital humanities, and related fields who want a faster, clearer, and reproducible way to study emotion across media using R. No prior deep learning expertise is required.
You should be comfortable with:
-
- Basic measurement concepts (constructs, labels/taxonomies, validation logic).
- Foundational statistics (descriptives, classification metrics such as accuracy/precision/F1).
Outline
Module 1 — Transformers 101 and zero‑shot text analysis
-
- Encoder vs. decoder architectures; why encoder-style models suit sentiment workflows.
- Multimodal embeddings and how to interpret similarity scores and zero-shot outputs.
- Designing label sets and prompts; ethical and methodological caveats (bias, privacy).
- Hands-on: zero-shot BERT-style analyses on sample corpora; export tidy, analysis-ready tables.
Module 2 — Small LLMs for structured explanations
-
- Why compact models are useful for short, explainable outputs.
- Prompting and simple RAG-style approaches to produce JSON/table outputs.
- Hands-on: compare prompt-only vs. retrieval-augmented pipelines on a small corpus.
Module 3 — Vision with CLIP: images
-
- How CLIP differs from traditional neural network architectures; linking image and text cues via embeddings.
- Bias and coverage considerations for public training sets.
- Hands-on: apply CLIP models to facial-expression images; tackle face selection challenges.
Module 4 — Vision with CLIP: video
-
- Frame sampling strategies; consistency of face selection across frames.
- Integrating video frame outputs into the overall pipeline.
- Hands-on: run the CLIP video pipeline and compare sampling/label choices.
Module 5 — Capstone: end‑to‑end multimodal pipeline (MAFW subset)
-
- Combine BERT text scores, CLIP image/video outputs, and optional LLM rationales.
- Evaluate results with accuracy, F1, confusion matrices, calibration, and agreement checks.
- Iterate on label sets and sampling choices to improve alignment and clarity.
- Present and discuss model choices, evaluation strategies, and reporting.
Seminar Information
Wednesday, November 4 –
Friday, November 6, 2026
Schedule: All sessions are held live via Zoom. All times are ET (New York time).
Wednesday, November 4:
9:00am-12:30pm (convert to your local time)
1:30pm-3:30pm
Thursday, November 5:
10:00am-12:30pm
Friday, November 6:
9:00am-12:30pm
1:30pm-3:30pm
Payment Information
The fee of $995 USD includes all course materials.
PayPal and all major credit cards are accepted.
Our Tax ID number is 26-4576270.

Back to Public Seminars