🌐 Large Language Models (LLMs) are reshaping global decision-making and societal systems through their ability to process vast and diverse datasets. However, their potential to homogenize human values—mirroring the risks of biodiversity loss in ecosystems—poses serious challenges to cultural diversity and ethical resilience.
🚀 This study introduces EthosGPT, an open-source framework designed to systematically map and evaluate LLMs against a globally representative spectrum of human values.
📊 By integrating cross-cultural survey data, interactive prompt-based evaluations with measurable outputs, and comparative statistical analyses, EthosGPT quantifies the cultural adaptability of LLMs while identifying critical disparities between AI-generated value indices and those of human populations across 100+ countries.
🧩 Findings reveal that while LLMs demonstrate partial alignment with localized cultural norms, significant gaps persist—especially in underrepresented regions. To mitigate these issues, EthosGPT offers actionable strategies for building inclusive AI systems:
- 📚 Diversifying training datasets
- 🏛️ Digitally preserving endangered cultural heritage
🎯 These efforts support the following UN Sustainable Development Goals (SDGs):
- 🔟 SDG 10: Reduced Inequalities
- 🏛️ SDG 11.4: Cultural Heritage Preservation
- ⚖️ SDG 16: Peace, Justice & Strong Institutions
🌱 EthosGPT positions value diversity as a cornerstone of societal innovation and sustainable prosperity. Through open-source tools and interdisciplinary collaboration, it advocates for AI systems that harmonize technical rigor with ethical pluralism. By bridging cultural gaps in machine learning, this work advances global AI governance frameworks toward a more inclusive and equitable future.
EthosGPT is an open-source framework for benchmarking cultural alignment in large language models (LLMs). It maps GPT-generated responses against global human values using survey data and cultural indices across countries and regions. Inspired by the philosophical concept of ēthos — moral character rooted in cultural identity — the project investigates how LLMs reflect or diverge from cultural plurality.
This repository supports the paper:
📝 "EthosGPT: Mapping Human Value Diversity to Advance Sustainable Development Goals (SDGs)"
-
ChatGPT/— Main datasets, notebooks, and high-resolution figures
📄 See full details inChatGPT/README.md -
figs/— Exported visualizations and plots -
data/— Cleaned cultural index files and regional metrics -
LICENSE— MIT License
- Cultural index simulation using GPT-4 (e.g., Traditional vs. Secular)
- Comparative visualizations between AI and survey-based responses
- Region-level error metrics (MSE, MAE)
- Tools for fairness, cross-cultural validation, and model benchmarking
EthosGPT lays the foundation for more inclusive, culturally aware AI. The roadmap below outlines key future directions to scale its impact and global relevance.
EthosGPT Roadmap
├── 🧩 Expand Cultural Indices
│ ├── Hofstede Dimensions
│ ├── ESS/EVS Distance Indices
│ ├── GLOBE Leadership Study
│ ├── D-PLACE (Ethnography & Ecology)
│ └── Ecology-Culture Dataset
├── 🧠 Evaluate Diverse LLMs
│ ├── GPT-4o (OpenAI)
│ ├── Gemini 1.5 Pro (Google)
│ ├── Claude 3 (Anthropic)
│ ├── GLM (Zhipu AI)
│ ├── DeepSeek V3
│ └── Qwen2.5 (Alibaba)
└── 🧬 Enhance Underrepresented Voices
├── 3D Cultural Digitization
├── Inclusive Training Filters
├── Socioeconomic Prompting
├── CCSV Self-Assessment
└── CultureLLM Fine-Tuning
| Dataset | Focus Area | Reference |
|---|---|---|
| Hofstede Dimensions | Cultural value scoring across 6 axes | [@hofstede2011dimensionalizing] |
| ESS/EVS Distance Indices | Regional value differences in Europe | [@kaasa2016dataset] |
| GLOBE Study | Leadership & society-wide values | [@house2004globe] |
| D-PLACE | Culture–language–environment links | [@kirby2016dplace] |
| Ecology-Culture Dataset | Environmental effects on culture | [@wormley2022ecology] |
Expanding EthosGPT's coverage with these datasets will strengthen its analytical power and global representation.
We aim to evaluate the cultural alignment of the most advanced LLMs available today:
This benchmark will reveal how architecture and data origin influence cultural representation.
| Strategy | Description | Reference |
|---|---|---|
| 🏛️ Cultural Digitization | 3D scanning and virtual records to preserve heritage | [@ocon2021digitalising] |
| 🌍 Inclusive Data Filtering | Cultural/socioeconomic-aware dataset curation | [@pouget2024filterculturalsocioeconomicdiversity] |
| 🧭 Prompt Personalization | Injecting regional & class-aware variables into prompts | [@nwatu2024upliftinglowerincomedatastrategies] |
| 🗳️ CCSV Self-Voting | Use model critiques & voting for demographic fairness | [@lahoti-etal-2023-improving] |
| 🧠 CultureLLM Fine-Tuning | Semantic augmentation & culture-specific alignment | [@li2024culturellmincorporatingculturaldifferences] |
These innovations aim to improve the fairness, inclusivity, and ethical sensitivity of LLM outputs.
✨ These roadmap pillars ensure that EthosGPT evolves as a globally inclusive, interdisciplinary, and ethically grounded framework for cultural intelligence in AI.
🥇 1st Prize – AI Governance Award
Awarded by AI Safety Fundamentals
For the project:
EthosGPT: Charting the Human Values Landscape on a Global Scale
This repository is released under the MIT License.