Overview:
Have you ever wanted to extract hidden themes, sentiment, or trends across thousands of historical records, social media posts, corporate documents, or research articles? Text analysis uses computational tools to examine and compare massive amounts of written content—revealing how word usage shifts over time, which topics cluster together, and how narrative tone varies across different sources.
This hands-on workshop provides a foundational introduction to text analysis using Python, boosted by generative AI assistance. To ground our learning in real-world application, we will analyze selected digital archive datasets from the Industry Documents Library (IDL). The IDL houses millions of internal records from the tobacco, pharmaceutical, chemical, and food industries. Working with these primary sources illustrates how computational text analysis can uncover critical evidence about corporate messaging, public health impacts, regulatory strategies, and historical accountability.
No prior programming experience is required. You will learn how to combine beginner-friendly Python libraries with generative AI tools (such as AI coding assistants) to write code, debug errors, and extract meaningful insights from complex archival text quickly and effectively.
Workshop Focus
- Core Concepts: Understand fundamental text analysis principles, key terminology, common analytical pipelines, and ethical considerations when working with public and archival records.
- Real-World Archival Datasets: Preprocess chosen Industry Documents Library (IDL) collections using core methods, setting the stage to uncover broader corporate strategies and their historical and societal effects.
- Python Basics for Text: Load real archival text datasets, calculate term frequencies, execute keyword analysis, and visualize key patterns over time.
- AI-Assisted Workflows: Leverage AI tools to write, debug, and explain Python code on the fly, making custom data science techniques accessible to complete beginners.
Contact information:
Instructor: Eileen Han
Please contact Sean Smith if you have questions about the Data@Rice workshop series.