Turn Documents into Data with AI - A Short Course
A 3-Day Livestream Seminar Taught by Karl Rohe, Ph.D.
Whether you call it data extraction, document coding, content analysis, structured abstracting or systematic review, the underlying work is the same: turning a body of literature into structured, analyzable, trustworthy data.
For a century, statistical methodology has systematized the final steps of inference—modeling and testing—but not the miles of reading and extraction that come before. That work has been done by humans, one paper at a time, with criteria written in a codebook and no measure of external validity. AI can change this, but not how you might think.
The stance of this course is that AI is a poor judge, but an excellent reader. Most researchers don’t trust AI for their research—and they shouldn’t, at least not blindly. But you can build your trust in AI the same way you can build it with a research assistant: give it a clear task, audit the first outputs, see where it stumbles, refine the instructions, and avoid assigning tasks it can’t do. Then, iterate. Instead of asking “Is the assistant trustworthy?” ask “Which tasks can I entrust to it? What should I ask so it understands? How do I know it did the job well?’
The craft lies in decomposing research judgment into reading tasks that AI can do reliably—and the codebook is the instrument for that decomposition. This course teaches you how to create codebooks with and for AI.
Starting September 10, this seminar will be presented as a 3-day synchronous, livestream workshop via Zoom. Each day will feature two lecture sessions with hands-on exercises, separated by a 1-hour break. Live attendance is recommended for the best experience. If you can’t join in real time, recordings will be available within 24 hours and can be accessed for four weeks after the seminar.
Closed captioning is available for all live and recorded sessions. Captions can be translated to a variety of languages, including Spanish, Korean, and Italian. For more information, click here.
ECTS Equivalent Points: 1
More details about the course content
Each session combines lecture and live demonstration: building codebooks, running extractions, diagnosing failures, and refining outputs. At times, you will watch the work unfold on the instructor’s screen; at others, you will make the same moves on your own, with supervision and feedback.
Along the way you will learn the transferable craft: how to spot where the AI is guessing rather than reading, how to refine a field to remove codebook ambiguity, and how to know when the answers are trustworthy enough to publish.
The course is built around four commitments that, together, provide an audit trail for trustworthy use of AI:
-
- The codebook is the durable artifact. It is the document that holds your operationalization—every choice you made about what counts and what doesn’t, written down so that reviewers (and your future self) can see them. A good codebook outlasts the model that ran it.
- AI is a reader, not a judge. Decomposing a research question into reading tasks an AI can do reliably is the core engineering move. In this course, you’ll learn decomposition as a craft, with concrete refinement techniques and a vocabulary for diagnosing failures.
- Disagreement is the diagnostic. When multiple AI readers run the same codebook and disagree, those are the cells worth auditing first. Multi-reader agreement is treated as publication-defensible evidence of reliability, but not a stamp of approval.
- Data are shared with codebook and reasoning. Your spreadsheet gets a URL that you can share. You can click any cell to see the codebook item, key quotes from your document, and the AI’s reasoning for the value in the cell.
The codebook can be reused on different documents by you or others and revised or refined as needed. The disagreements reveal where the AI readers interpret the codebook differently, or places where they make errors.
Computing
We will use Data Mint, a web-based platform built to do exactly what this course teaches. There is nothing to install, no command line, and no programming. All extraction work happens in a browser, and course accounts are provided so that no API keys or paid subscriptions are required.
To follow along with the hands-on exercises, you will need:
-
- A laptop or desktop computer (not a mobile phone), running Mac, Windows, or Linux, with a modern web browser and a reliable internet connection.
- A Data Mint account. Course accounts will be set up during the pre-course intake. Data Mint is open to anyone, so you can visit datamint.ing, request access with your academic email, and try it out before deciding whether to register.
- A spreadsheet program (Excel, Google Sheets, or similar) for inspecting exported data.
Recommended (but not required): you should arrive with a research question and a corpus of documents.
You will leave with three things:
- A publishable codebook for AI-assisted extraction, which you can apply to the rest of your corpus.
- A spreadsheet with the data you came for.
- Reliability metrics—based on multi-reader agreement—at publication-defensible levels.
No experience with Python, R, or any programming language is required (the course will not instruct on how to analyze your data, only how to generate the data).
Who should register?
If you’re a researcher who needs to turn documents into data, this course is for you. That includes—but is not limited to—anyone doing data extraction, document coding, content analysis, chart abstraction, structured literature review, or large-scale qualitative coding.
It fits especially well for researchers working on systematic reviews (screening, extraction, and risk-of-bias assessment), but the craft applies to any project where a body of text must become structured, analyzable data.
Outline
Day 1
First extraction
-
- Intro: How might we come to trust AI for research? What are key concerns? The idea of Minting Data.
- Build a draft codebook for your own research question and produce your first AI-extracted outputs
- Title-and-abstract screening as a low-stakes warm-up
- Introduction to the vocabulary of scaffolding—the structural moves that make an instruction reliably followable
- Hands-on exercise: a v0 codebook, a first round of extracted data, and at least one diagnosed failure
Day 2
The scaffolding deep dive
-
- The signature work: learning to refine your codebook
- Three legitimate fixes for a failing field:
- Reword the question—clarify decision boundaries, add worked examples
- Add scaffolding—upstream fields that break a multi-part judgment into steps the AI can handle
- Rebuild the field entirely
- Hands-on exercise: a cross-domain exercise that forces the meta-skill—diagnosing structural cracks—before you apply it to your own codebook
Day 3
Transfer and ship
-
- Apply the craft to risk-of-bias assessment—a new domain that uses the same skill
- Peer audits of each other’s codebooks
- Leave with: a shipped codebook, an exported dataset, and reliability metrics from multi-reader agreement
Seminar Information
Thursday, September 10 –
Saturday, September 12, 2026
Schedule: All sessions are held live via Zoom. All times are ET (New York time).
10:00am-12:30pm (convert to your local time)
1:30pm-3:30pm
Payment Information
The fee of $995 USD includes all course materials.
PayPal and all major credit cards are accepted.
Our Tax ID number is 26-4576270.

Back to Public Seminars