Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

9 Commits
 
 

Repository files navigation

Sumin Park (박수민) 👋

Computational linguist 🧑‍🏫 · Visiting Assistant Professor, College of Humanities, Seoul National University

I model linguistic variation quantitatively and apply it to literary and historical texts. The same referent, meaning, or function is realized differently depending on conditions — time (diachronic variation) and the social relation between speaker and addressee, the medium, or the readership (synchronic variation). I describe these differences in corpora through frequency, co-occurrence, and distribution, at levels that reach beyond literal meaning into the pragmatic, stylistic, and social. I work on contemporary Korean, Classical Chinese sources, and Hangul manuscripts with the same methods, and I release the models, data, and code I use.

🔬 Research

🕰️ Variation across time

  • The taste adjective 'dalda' (sweet) in Joseon cookbooks 🍯 — its distribution widens from fermented foods to a broad range of dishes (17th–19th c.), tracked through its co-occurrence with sweeteners.
  • The character '鬼' in the Veritable Records of the Joseon Dynasty 👻 — per-reign Word2Vec models reveal how its nearest neighbors, and thus its semantic field, shift over time. (exploratory)

👥 Variation across social conditions

  • Self-reference terms in the Standard Histories — how the same function is reassigned across forms (e.g. 臣 → 小人) as register and speaker–addressee relations change, traced across the Four Histories.
  • Gendered references [X-男]/[X-女] in news headlines 📰 — a 124,857-token database (1990–2023) showing the rise and reversal of male/female suffix forms over time. (exploratory)

🧩 Meaning beyond similarity

  • Sentence representation — capturing not only semantic but also pragmatic and stylistic relations between texts; designing a multi-dimensional similarity dataset for literary and historical material.

🛠️ Resources

  • KR-SBERT — a Korean sentence-embedding model for clustering literary texts, comparing authorship and style, and surfacing relatedness that surface features miss.
  • KOSAC — a Korean sentiment-analysis lexicon for quantifying sentiment and attitude in corpora.
  • Gendered-reference database and analysis code · (more to come)

🤝 Collaboration

I am glad to reframe interpretive questions in the humanities into quantitatively testable form. I am especially keen to work with researchers in Classical Chinese philology, Joseon and early-medieval Chinese history, conceptual and religious history, media and gender studies, stylistics and narratology, and historical sociolinguistics. If you have a corpus and a question, let's talk. 📨

📬 Contact

About

No description, website, or topics provided.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors