Computational linguist 🧑🏫 · Visiting Assistant Professor, College of Humanities, Seoul National University
I model linguistic variation quantitatively and apply it to literary and historical texts. The same referent, meaning, or function is realized differently depending on conditions — time (diachronic variation) and the social relation between speaker and addressee, the medium, or the readership (synchronic variation). I describe these differences in corpora through frequency, co-occurrence, and distribution, at levels that reach beyond literal meaning into the pragmatic, stylistic, and social. I work on contemporary Korean, Classical Chinese sources, and Hangul manuscripts with the same methods, and I release the models, data, and code I use.
- The taste adjective 'dalda' (sweet) in Joseon cookbooks 🍯 — its distribution widens from fermented foods to a broad range of dishes (17th–19th c.), tracked through its co-occurrence with sweeteners.
- The character '鬼' in the Veritable Records of the Joseon Dynasty 👻 — per-reign Word2Vec models reveal how its nearest neighbors, and thus its semantic field, shift over time. (exploratory)
- Self-reference terms in the Standard Histories — how the same function is reassigned across forms (e.g. 臣 → 小人) as register and speaker–addressee relations change, traced across the Four Histories.
- Gendered references [X-男]/[X-女] in news headlines 📰 — a 124,857-token database (1990–2023) showing the rise and reversal of male/female suffix forms over time. (exploratory)
- Sentence representation — capturing not only semantic but also pragmatic and stylistic relations between texts; designing a multi-dimensional similarity dataset for literary and historical material.
- KR-SBERT — a Korean sentence-embedding model for clustering literary texts, comparing authorship and style, and surfacing relatedness that surface features miss.
- KOSAC — a Korean sentiment-analysis lexicon for quantifying sentiment and attitude in corpora.
- Gendered-reference database and analysis code · (more to come)
I am glad to reframe interpretive questions in the humanities into quantitatively testable form. I am especially keen to work with researchers in Classical Chinese philology, Joseon and early-medieval Chinese history, conceptual and religious history, media and gender studies, stylistics and narratology, and historical sociolinguistics. If you have a corpus and a question, let's talk. 📨



