Skip to content

Request to Review & Remove Personal Name Tokens (Muhamad Arki Munggaran) from IndoGPT Vocab #5

Description

@munggaranarki-lab

Hello IndoNLG team,

I am reaching out to you because I found that parts of my name—"arki" and "munggaran"—appear as separate tokens in the vocabulary file of the IndoGPT model (file link: https://huggingface.co/api/resolve-cache/models/indobenchmark/indogpt/cb7030394be8d42904aef27d3b24f0bc3f0a0811/IndoNLG_finals_indogpt_tokenizer.vocab?download=true&etag=%226a8c5f078e5f60ab72043030f17272f5af75ae3f%22).

My full name is Muhamad Arki Munggaran, and I have never given permission for my name to be used in the model's training data. I am concerned that the presence of my name could impact my privacy.

Under Indonesia's Personal Data Protection Law (Law No. 27 of 2022), I would like to submit the following requests:

1. To evaluate the source of how my name was included in the model's training corpus.

2. To remove or anonymize my name parts from the model data.

Please assist me in addressing this concern. I am happy to provide additional details if needed.

Thank you for your attention.

Sincerely,
[Muhamad Arki Munggaran]

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions