Skip to content
View suparklingmin's full-sized avatar

Block or report suparklingmin

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
suparklingmin/README.md

Sumin Park (박수민) 👋

Computational linguist 🧑‍🏫 · Visiting Assistant Professor, College of Humanities, Seoul National University

I model linguistic variation quantitatively and apply it to literary and historical texts. The same referent, meaning, or function is realized differently depending on conditions — time (diachronic variation) and the social relation between speaker and addressee, the medium, or the readership (synchronic variation). I describe these differences in corpora through frequency, co-occurrence, and distribution, at levels that reach beyond literal meaning into the pragmatic, stylistic, and social. I work on contemporary Korean, Classical Chinese sources, and Hangul manuscripts with the same methods, and I release the models, data, and code I use.

🔬 Research

🕰️ Variation across time

  • The taste adjective 'dalda' (sweet) in Joseon cookbooks 🍯 — its distribution widens from fermented foods to a broad range of dishes (17th–19th c.), tracked through its co-occurrence with sweeteners.
  • The character '鬼' in the Veritable Records of the Joseon Dynasty 👻 — per-reign Word2Vec models reveal how its nearest neighbors, and thus its semantic field, shift over time. (exploratory)

👥 Variation across social conditions

  • Self-reference terms in the Standard Histories — how the same function is reassigned across forms (e.g. 臣 → 小人) as register and speaker–addressee relations change, traced across the Four Histories.
  • Gendered references [X-男]/[X-女] in news headlines 📰 — a 124,857-token database (1990–2023) showing the rise and reversal of male/female suffix forms over time. (exploratory)

🧩 Meaning beyond similarity

  • Sentence representation — capturing not only semantic but also pragmatic and stylistic relations between texts; designing a multi-dimensional similarity dataset for literary and historical material.

🛠️ Resources

  • KR-SBERT — a Korean sentence-embedding model for clustering literary texts, comparing authorship and style, and surfacing relatedness that surface features miss.
  • KOSAC — a Korean sentiment-analysis lexicon for quantifying sentiment and attitude in corpora.
  • Gendered-reference database and analysis code · (more to come)

🤝 Collaboration

I am glad to reframe interpretive questions in the humanities into quantitatively testable form. I am especially keen to work with researchers in Classical Chinese philology, Joseon and early-medieval Chinese history, conceptual and religious history, media and gender studies, stylistics and narratology, and historical sociolinguistics. If you have a corpus and a question, let's talk. 📨

📬 Contact

Pinned Loading

  1. LingDataSci2023 LingDataSci2023 Public

    언어데이터과학 (2023학년도 2학기, 서울대학교 언어학과)

    Jupyter Notebook 8 2

  2. LangComp2021 LangComp2021 Public

    언어와 컴퓨터 (2021학년도 2학기, 서울대학교 언어학과)

    13 2

  3. CompLing2022 CompLing2022 Public

    컴퓨터언어학 (2022학년도 1학기, 서울대학교 언어학과)

    22 3