TelecoVision is an end-to-end Telecom Data Engineering Platform that combines both Batch and Streaming data processing to provide real-time customer monitoring, churn prediction, customer experience analytics, and business intelligence reporting.
The platform integrates telecom data from multiple operational systems and transforms it into actionable insights through a scalable modern data architecture.
Telecom companies lose a significant percentage of customers every year due to churn.
🔹 Customer data is scattered across multiple systems such as:
- CRM Systems
- Billing Systems
- Network Monitoring Systems
- Customer Service Systems
🔹 No real-time visibility into customer behavior and network events.
🔹 Lack of a centralized analytical layer for business reporting.
🔹 Difficulty identifying customers who are likely to churn before they leave.
TelecoVision unifies both streaming and batch data into a single platform, enabling:
✅ Real-Time Churn Monitoring
✅ Customer Analytics
✅ Usage Analytics
✅ Customer Experience Analytics
✅ Revenue Impact Analysis
✅ Interactive Dashboards
The platform combines both Batch and Streaming pipelines using a Medallion Architecture to support operational monitoring and analytical workloads.
The streaming layer captures telecom events from Kafka topics and processes them using Spark Structured Streaming for real-time analytics and churn prediction.
- Apache Kafka
- Apache Spark Structured Streaming
- Microsoft SQL Server
- Google Cloud Storage (GCS)
- Apache Spark
- Google Cloud Storage (GCS)
- BigQuery
- dbt Cloud
- Apache Airflow
- Grafana
- Looker Studio
Stores raw streaming and batch data exactly as received from source systems.
- Raw JSON Files
- Raw Telecom Events
- Google Cloud Storage
Processes raw data and applies business transformations.
✅ Data Cleaning
✅ Data Validation
✅ Standardization
✅ Mapping
✅ Feature Engineering
✅ Data Enrichment
- Parquet Files
- Google Cloud Storage
Business-ready analytical layer optimized for reporting and dashboarding.
- BigQuery
- Data Warehouse Tables
The warehouse follows Kimball Dimensional Modeling principles and implements a Fact Constellation (Galaxy Schema).
Type: Periodic Snapshot Fact
Grain: Customer + Date + Time Snapshot
Contains:
- Churn Score
- Risk Level
- Customer Metrics
- Network Metrics
- Usage Metrics
Type: Transaction Fact
Grain: One Customer Care Call
Contains:
- Call Duration
- Anger Rate
- Resolution Status
- Issue Type
Type: Transaction Fact
Grain: One Usage or Network Event
Contains:
- Internet Usage
- Voice Minutes
- SMS Usage
- Network Performance Metrics
Type: Transaction Fact
Grain: One Churn Scoring Event
Contains:
- Churn Probability
- Risk Category
- Churn Indicators
- Dim Customer
- Dim Date
- Dim Time
- Dim Device
- Dim Customer Loyalty
Telecom events are continuously generated and streamed into Kafka.
- 2 Partitions
Contains:
- Call Duration
- Anger Rate
- Issue Type
- Resolution Status
- 2 Partitions
Contains:
- Signal Strength
- Internet Speed
- Drop Calls
- Network State
- 3 Partitions
Contains:
- Internet Usage
- Voice Minutes
- SMS Count
- 3 Brokers
- Replication Factor = 2
✅ Scalability
✅ Parallel Processing
✅ High Availability
✅ Fault Tolerance
If a broker becomes unavailable, Kafka automatically switches to a replica and continues processing without interruption.
Spark consumes telecom events from Kafka and processes them using micro-batches.
✅ Read Events from Kafka
✅ Store Raw Data in SQL Server
✅ Store Raw Data in Bronze Layer
✅ Join Multiple Streams
✅ Calculate Churn Scores
✅ Generate Risk Levels
✅ Write Results to SQL Server
✅ Write Results to GCS
This enables near real-time churn monitoring and customer analytics.
Historical telecom datasets are uploaded to the Bronze Layer in Google Cloud Storage.
Spark processes data across multiple worker nodes.
✅ Cleaning
✅ Mapping
✅ Standardization
✅ Feature Engineering
✅ Data Enrichment
Processed data is stored in the Silver Layer as optimized Parquet files.
Silver Layer data is loaded into BigQuery and transformed using dbt.
✅ Data Modeling
✅ Fact Table Creation
✅ Dimension Table Creation
✅ Business Logic Implementation
✅ Data Warehouse Construction
Apache Airflow orchestrates the entire platform.

- Scheduling Pipelines
- Managing Dependencies
- Executing Spark Jobs
- Triggering dbt Runs
- Monitoring Workflow Execution
- Failure Recovery
Grafana provides real-time operational monitoring over streaming telecom events.
- Real-Time Churn Monitoring
- Churn Score Tracking
- Risk Distribution
- High-Risk Customer Detection
- Customer Service Monitoring
- Resolution Performance
- Anger Rate Tracking
- Quality Indicators
- Network Quality Monitoring
- Signal Strength Analytics
- Internet Performance
- Drop Call Monitoring
Looker Studio provides business-level analytics built on top of the Data Warehouse.
- Customer Segmentation
- Loyalty Analytics
- Customer Distribution
- Behavioral Analysis
- Churn Distribution
- Risk Categories
- Churn Drivers
- Retention Opportunities
- Revenue at Risk
- Churn Cost Analysis
- Revenue Trends
- Customer Value Analysis
- Network Performance Analytics
- Customer Experience Metrics
- Data Failure Analysis
- Quality Distribution
✅ Batch + Streaming Integration
✅ Real-Time Churn Monitoring
✅ Customer Analytics
✅ Usage Analytics
✅ Customer Experience Analytics
✅ Revenue Impact Analysis
✅ Enterprise Data Warehouse
✅ Automated Data Pipelines
✅ Interactive Dashboards
The project uses the public Telecom Customer Churn dataset from IBM Sample Data.
- Customer Demographics
- Service Usage Information
- Billing Information
- Customer Support Interactions
- Churn Indicators
- Reham Mohammed
- Sara Abuzeid
- Yasmin Shamakh
- Nermeen saad
⭐ If you found this project interesting, don't forget to give it a Star.









