A comprehensive data analysis and visualization project based on the National Family Health Survey (NFHS-5) dataset. This project explores major health indicators across Indian states using Python, Pandas, Matplotlib, Seaborn, and Folium, providing statistical insights and interactive geographic visualizations.
The National Family Health Survey (NFHS-5) provides detailed information about the health and nutrition status of India's population. This project analyzes selected state-level health indicators and presents them through exploratory data analysis (EDA) and interactive maps.
The project focuses on identifying health trends, comparing states, and visualizing the distribution of obesity, anaemia, hypertension, and blood sugar levels.
- Load and clean the NFHS-5 dataset.
- Analyze important health indicators across Indian states.
- Compare male and female obesity levels.
- Identify states with high hypertension prevalence.
- Explore relationships between obesity and anaemia.
- Visualize state-wise health statistics using interactive maps.
- Build an end-to-end data analysis workflow suitable for beginners.
Source: Kaggle
Dataset Used:
- National Family Health Survey (NFHS-5) 2019–2020
- State-wise Health Indicators
The dataset contains health statistics such as:
- Male Obesity (%)
- Female Obesity (%)
- Child Anaemia (%)
- Women Anaemia (%)
- Male Hypertension (%)
- Female Hypertension (%)
- Male High Blood Sugar (%)
- Female High Blood Sugar (%)
| Technology | Purpose |
|---|---|
| Python | Programming Language |
| Google Colab | Development Environment |
| Pandas | Data Cleaning & Analysis |
| NumPy | Numerical Computations |
| Matplotlib | Data Visualization |
| Seaborn | Statistical Visualization |
| Folium | Interactive Maps |
| GeoJSON | Geographic Boundaries |
- Upload Kaggle dataset
- Extract ZIP file
- Load CSV into Pandas
- Clean missing values
- Convert numeric columns
- Create additional analysis columns
Performed analyses include:
- Distribution of obesity rates
- Top states with highest hypertension
- Female vs Male obesity comparison
- Scatter plot of obesity vs anaemia
- Correlation analysis
- Statistical summaries
Interactive maps built using Folium:
- Choropleth Map
- State-wise Circle Markers
- Health summary popups
- Color-coded health indicators
The notebook generates multiple visualizations including:
- Grouped Bar Charts
- Scatter Plots
- Correlation Heatmaps
- Choropleth Maps
- Interactive State Markers
The project examines:
- State-wise obesity distribution
- Hypertension prevalence
- Anaemia among women and children
- Gender obesity gap
- Geographic variation in health indicators
- Relationships between different health metrics
git clone https://github.com/yourusername/NFHS5-Health-Analysis.gitpip install pandas numpy matplotlib seaborn foliumjupyter notebook NFHS5_Kaggle_Analysis.ipynbor upload the notebook directly into Google Colab.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
import foliumThe project produces:
- Cleaned dataset
- Summary statistics
- Comparative visualizations
- Interactive state maps
- Health indicator insights
This project demonstrates practical skills in:
- Data Cleaning
- Exploratory Data Analysis (EDA)
- Data Visualization
- Geographic Data Visualization
- Interactive Mapping
- Feature Engineering
- Statistical Analysis
- Python for Data Science
- District-level analysis
- Time-series comparison with future NFHS surveys
- Machine Learning for health prediction
- Dashboard using Streamlit
- Power BI integration
- Automated report generation
N. Rithish Barath
Computer Science Engineering Student
Passionate about:
- Data Science
- Artificial Intelligence
- Machine Learning
- Data Visualization
- Python Development
This project is intended for educational and academic purposes.
- National Family Health Survey (NFHS-5)
- Ministry of Health and Family Welfare, Government of India
- Kaggle
- Pandas
- Matplotlib
- Seaborn
- Folium
"Transforming public health data into meaningful insights through analytics and visualization."